Jobiba hiring network

Responsable Production Jobs

6,399 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current responsable production jobs. Use filters to narrow by work mode, employment type, experience and date posted.

M
Mongodb
📍 Gurugram• Full-time
1mo ago

The Infrastructure Engineering team is responsible for building and maintaining a self-service internal development platform that enables MongoDB engineering teams to reliably deploy and operate their own production services and products. We work with numerous engineering teams across the company to understand their infrastructure requirements and development workflows, develop broadly applicable self-service platform services and tooling, continuously monitor how platform services are being utilized, and look for ways to improve developer productivity through automation and education. We are big open source enthusiasts and use a number of open source tools in our stack (contributing upstream whenever possible). Some of the tools we use regularly include Go, AWS, Kubernetes, Crossplane, Terraform, Helm, Drone, Prometheus, and Grafana. However, technology is nothing without a stellar team of engineers that are focused on doing high quality work and working as a team to solve complex distributed computing and platform engineering problems. This is where you come in! We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Our ideal candidate 2+ years of experience managing and mentoring a team of 3+ engineers Has 5+ years of experience owning the design and implementation of large software/infrastructure projects Has built and operated large-scale distributed systems in cloud providers (AWS strongly preferred) Has a strong backend programming background. Fluency in Go is strongly preferred; deep experience with another compiled or strongly-typed backend language is acceptable Pragmatic, detail-oriented, self-motivated, and understands the benefits of collaboration Strong experience operating production Kubernetes clusters, not just deployed to it Has practical experience defining and operating against SLI/SLOs for services they owned Strong experience with observability tooling: metrics, logging, traces, Prometheus, Grafana, OpenTe

mongodbawsazure
View job →
S
1mo ago

About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team As Stripe’s user base and global footprint grows dramatically, we have distinctly unique support problems resulting from both our type of scale and the type of businesses we partner with. The Stripe Delivery Center (SDC) strategy will provide operational leverage and expand Stripe’s portfolio of operational capabilities to support the scaled needs for external users and internal Stripe teams. This role will be part of the Marketing Operations in SDC that is accountable for the production and delivery of all aspects of email demand and lead generation activities, lead nurturing, and conversion programs. What you’ll do We are looking for someone who is passionate about Marketing Automation and is naturally curious, and is comfortable operating in ambiguity and finding creative solutions. You will be responsible for the delivery of repeatable, measurable programs/campaigns that focus on highly qualified leads, improve the customer experience, and increase retention. Responsibilities Build, test and launch complex email and nurture campaigns in Marketo Perform test runs and quality control with strong attention to detail Conduct advanced A/B testing on email and landing page assets and, suggest testing opportunities Identify gaps in our marketing automation infrastructure and work cross functionally to address them (ex. Program structure inefficiencies, process flaws, etc) Scope, manage and deliver email campaigns from intake to execution Creat

awsgitai
View job →
O
1mo ago

About the Team The AI Deployment Engineering team is responsible for helping developers and enterprises safely and effectively deploy OpenAI technologies in production. We act as trusted technical advisors and thought partners for customers, working side by side with their teams to identify high-value use cases, design practical architectures, and move from prototype to durable deployment. Cybersecurity is one of the most urgent domains where AI can help. Security teams are under pressure to reason across code, logs, infrastructure, tickets, alerts, and vulnerability data faster than ever. As frontier models become more capable, organizations need deep technical guidance on how to evaluate, validate, and safely deploy AI systems in security-critical workflows. About the Role We are looking for a Cyber AI Deployment Engineer to partner with customers and help them apply OpenAI models, APIs, Codex, and agentic workflows to real cybersecurity use cases. You will work with CISOs, security executives, application security leaders, SOC teams, security engineering teams, and hands-on practitioners to identify where AI can create measurable security outcomes. This is a customer-facing technical role for someone who can move fluidly between executive strategy, practitioner-level cyber depth, and hands-on solution design. You will help customers evaluate and deploy workflows such as secure code review, vulnerability triage, threat modeling, remediation, SOC and incident response workflows, detection engineering, cloud security, GRC automation, and security validation. You will collaborate closely with Sales, Solutions Engineering, Product, Engineering, Research, and Security to turn customer needs into safe deployment patterns, reusable field assets, and product feedback. This role is based in our Dublin Office. We offer relocation support to new employees. In this role, you will: Deeply embed with strategic customers as the technical lead for AI-enabled cybersecurity workflo

javascriptpythonjava
View job →
O
1mo ago

About the Team The AI Deployment Engineering team is responsible for helping developers and enterprises safely and effectively deploy OpenAI technologies in production. We act as trusted technical advisors and thought partners for customers, working side by side with their teams to identify high-value use cases, design practical architectures, and move from prototype to durable deployment. Cybersecurity is one of the most urgent domains where AI can help. Security teams are under pressure to reason across code, logs, infrastructure, tickets, alerts, and vulnerability data faster than ever. As frontier models become more capable, organizations need deep technical guidance on how to evaluate, validate, and safely deploy AI systems in security-critical workflows. About the Role We are looking for a Cyber AI Deployment Engineer to partner with customers and help them apply OpenAI models, APIs, Codex, and agentic workflows to real cybersecurity use cases. You will work with CISOs, security executives, application security leaders, SOC teams, security engineering teams, and hands-on practitioners to identify where AI can create measurable security outcomes. This is a customer-facing technical role for someone who can move fluidly between executive strategy, practitioner-level cyber depth, and hands-on solution design. You will help customers evaluate and deploy workflows such as secure code review, vulnerability triage, threat modeling, remediation, SOC and incident response workflows, detection engineering, cloud security, GRC automation, and security validation. You will collaborate closely with Sales, Solutions Engineering, Product, Engineering, Research, and Security to turn customer needs into safe deployment patterns, reusable field assets, and product feedback. This role is based in our Tokyo office. We offer relocation support to new employees. In this role, you will: Deeply embed with strategic customers as the technical lead for AI-enabled cybersecurity workflow

javascriptpythonjava
View job →
O
1mo ago

About the Team The AI Deployment Engineering team is responsible for helping developers and enterprises safely and effectively deploy OpenAI technologies in production. We act as trusted technical advisors and thought partners for customers, working side by side with their teams to identify high-value use cases, design practical architectures, and move from prototype to durable deployment. Cybersecurity is one of the most urgent domains where AI can help. Security teams are under pressure to reason across code, logs, infrastructure, tickets, alerts, and vulnerability data faster than ever. As frontier models become more capable, organizations need deep technical guidance on how to evaluate, validate, and safely deploy AI systems in security-critical workflows. About the Role We are looking for a Cyber AI Deployment Engineer to partner with customers and help them apply OpenAI models, APIs, Codex, and agentic workflows to real cybersecurity use cases. You will work with CISOs, security executives, application security leaders, SOC teams, security engineering teams, and hands-on practitioners to identify where AI can create measurable security outcomes. This is a customer-facing technical role for someone who can move fluidly between executive strategy, practitioner-level cyber depth, and hands-on solution design. You will help customers evaluate and deploy workflows such as secure code review, vulnerability triage, threat modeling, remediation, SOC and incident response workflows, detection engineering, cloud security, GRC automation, and security validation. You will collaborate closely with Sales, Solutions Engineering, Product, Engineering, Research, and Security to turn customer needs into safe deployment patterns, reusable field assets, and product feedback. This role is based in our Singapore office. We offer relocation support to new employees. In this role, you will: Deeply embed with strategic customers as the technical lead for AI-enabled cybersecurity work

javascriptpythonjava
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the team The AI Deployment Engineering team is responsible for helping developers and enterprises safely and effectively deploy OpenAI technologies in production. We act as trusted technical advisors and thought partners for customers, working side by side with their teams to identify high-value use cases, design practical architectures, and move from prototype to durable deployment. Cybersecurity is one of the most urgent domains where AI can help. Security teams are under pressure to reason across code, logs, infrastructure, tickets, alerts, and vulnerability data faster than ever. As frontier models become more capable, organizations need deep technical guidance on how to evaluate, validate, and safely deploy AI systems in security-critical workflows. About the role We are looking for a Cyber AI Deployment Engineer to partner with customers and help them apply OpenAI models, APIs, Codex, and agentic workflows to real cybersecurity use cases. You will work with CISOs, security executives, application security leaders, SOC teams, security engineering teams, and hands-on practitioners to identify where AI can create measurable security outcomes. This is a customer-facing technical role for someone who can move fluidly between executive strategy, practitioner-level cyber depth, and hands-on solution design. You will help customers evaluate and deploy workflows such as secure code review, vulnerability triage, threat modeling, remediation, SOC and incident response workflows, detection engineering, cloud security, GRC automation, and security validation. You will collaborate closely with Sales, Solutions Engineering, Product, Engineering, Research, and Security to turn customer needs into safe deployment patterns, reusable field assets, and product feedback. This role is based in our San Francisco HQ. We offer relocation support to new employees. In this role, you will: Deeply embed with strategic customers as the technical lead for AI-enabled cybersecurity work

javascriptpythonjava
View job →
O
1mo ago

About the Team The Technical Success team is responsible for ensuring developers and enterprises are successful in building scalable production applications with the OpenAI API platform. We guide and support customers to achieve maximum benefits, value, and adoption from deploying our highly-capable models. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. About the Role We are looking for a technically savvy and business-minded AI Deployment Engineer to deeply partner with our most strategic and high-impact platform customers, guiding them through application ideation, development, delivery, and scale to accelerate and maximize the value of what they build with our platform. You will have the opportunity to work on the most novel and creative use cases being built on our API, serving as a critical partner for collecting and delivering high fidelity feedback to Product and Research teams. This role is based in Tokyo, Japan. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees In this role, you will: Deeply embed with our most strategic platform customers, serving as their technical thought partner in ideating and building novel applications on our API. Proactively provide guidance to our customers on how to maximize business impact from their applications, accelerating their time to value. Experiment and prototype solutions with and for your customers. Forge and manage relationships with our customers’ leadership and stakeholders to ensure their application’s successful deployment and scale. Contribute to our open-source developer and enterprise resources. Scale the AI Deployment Engineering function through sharing knowledge, codifying best practices, and publishing notebooks to our internal and external repositories. Validate, synthesize, and deliver high-signal feedback to the Product and Research teams. Use your expertise i

javascriptpythonjava
View job →
T
11 hrs ago

The QA Analyst will be responsible for testing and verifying software products and applications to ensure that they meet the company's quality standards. The successful candidate will work closely with developers, project managers, and other stakeholders to identify defects and ensure that they are fixed in a timely manner. Job Responsibilities Analyze the design of features and improvements to identify edge cases and balance issues (20%) Write comprehensive test plans derived from specs, emails, and conversations (20%) Execute manual tests on multiple projects simultaneously (20%) Log, track, and assign bugs in a bug database, and follow up to ensure that issues are reviewed and addressed in a timely manner (20%) Guide projects to completion, from the Development environment to Production (5%) Interface with management and the development teams (5%) Support Customer Service to help resolve field issues and maintain live products (5%) Support Operations to help resolve all stop-ship issues immediately (5%) Critical Skills & Experience Requirements Bachelor’s Degree in Computer Science or related field (Req) 3&#43; Years of experience in software quality assurance (Pref) 2&#43; Years of experience testing large-scale, consumer-facing web products (Pref) 1&#43; Year of experience with Agile development methodologies (Pref) Strong analytical and problem-solving skills Excellent communication and collaboration skills Knowledge of software development lifecycle and software testing methodologies Experience with test automation tools and frameworks is a plus Knowledge of bug tracking software, such as Jira Experience writing and executing test plans using software such as TestRail or Zephyr BENEFITS <

human resourcescustomer service
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! The integration team is responsible for developing and scaling machine learning algorithms and infrastructure for LLM post-training, with a focus on large-scale, distributed RL methods. We strive for excellence in both engineering and science by meticulously designing experiments and design docs. While tasks are assigned according to everyone’s expertise, there is a global team effort to write production code and support the team research efforts, depending on individual interests and organizational needs. In particular, this role aims to enhance the global quality of the post-training codebase by implementing new tools to ease and support research, optimizing post-training algorithms, and scaling distributed RL to unprecedented levels. Please Note: We have offices in London, Paris, Toronto, San Francisco, New York but we are also remote-friendly! Applicants for this role may work anywhere between UTC−06:00 and UTC+01:00. As a Member of Technical Staff, you will: Design and write high-performing and scalable software for training models. Develop new tools to support and accelerate research and LLM training. Coordinate with other

pythonkubernetesgit
View job →
O
1mo ago

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

pythonawsazure
View job →

About the team The Applied AI Engineering (AAE) team is responsible for helping developers and enterprises turn the potential of generative AI into real-world impact. We act as trusted advisors and technical partners to customers and ecosystem partners, helping identify high-impact AI use cases and bring them into production through strong architectural guidance and hands-on execution. The Partner Applied AI Engineering organization works closely with strategic cloud providers, systems integrators, consultancies, and implementation partners to scale successful adoption of OpenAI technologies. As the leader of the AWS Partner AAE pod, you will manage a team of Applied AI Engineers focused on enabling AWS-aligned partners and their customers to build, deploy, and operationalize AI applications on OpenAI’s platform. About the role We are seeking a Manager, Partner Applied AI Engineering – AWS to lead a team of Applied AI Engineers supporting strategic AWS ecosystem partnerships. In this role, you will own the technical success strategy for AWS-aligned partners and help build scalable, repeatable ways for partners and their customers to adopt OpenAI technologies. Your team will guide partners and customers across the full AI implementation lifecycle—from identifying and shaping high-value use cases to solution design, architecture, production deployment, optimization, and adoption growth. You will work cross-functionally with internal and external stakeholders across Sales, Partnerships, Product, Research, and Engineering to ensure the voice of partners and customers informs our platform roadmap and how we bring OpenAI technology into production at scale. This role requires a blend of technical depth, customer leadership, operational rigor, and people management. Success will be measured through production deployments, partner technical maturity, API adoption growth, team development, and the overall impact of the AWS partner ecosystem. This role is based in our San Fra

javascripttypescriptpython
View job →
O
OpenAI
📍 Washington• Full-time
1mo ago

About the team The AI Deployment Engineering team is responsible for ensuring the safe and effective deployment of Generative AI applications. We act as a trusted advisor and thought partner for our customers, working to build an effective backlog of GenAI use cases for their industry and drive them to production through strong technical guidance. As an AI Deployment Engineer (ADE) in the OpenAI for Government team, you’ll help government agencies transform their organization through solutions such as automated content generation, contextual search, and novel applications that make use of our newest, most exciting models and technology. About the Role We are looking for a driven solutions leader with a product mindset to partner with our public sector customers and ensure they achieve tangible value with GenAI. You will pair with government agencies (federal, state, and local), policymakers, and other public institutions to establish a GenAI strategy and identify the highest value applications. You’ll then partner with their technical teams, subject matter experts, systems integrators, and implementation partners to move from prototype through production. You’ll take a holistic view of their needs and design an architecture using the OpenAI API and other services to maximize customer value. You will collaborate closely with Sales, Solutions Engineering, Global Affairs, Applied Research, and Product teams. This role is based in Washington, DC. We offer relocation support to new employees. In this role, you will: Deeply embed with our most sophisticated public sector customers as the technical lead, serving as their technical thought partner to ideate and build novel applications on our API and other OpenAI products. Work with senior customer stakeholders to identify the best applications of AI in their industry and to build/qualify a comprehensive backlog to support their AI roadmap. Intervene directly to accelerate customer time to value through building hands-on pr

javascriptpythonjava
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing

pythonawskubernetes
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de

awsrestai
View job →
🔔

Get new responsable production jobs by email

Daily job updates · Unsubscribe anytime