About the team The Agent Enablement AI Deployment Engineering (ADE) team works across engineering, product, design, partnerships, and strategic customers to grow an open ecosystem of agent-enabled sites and services. We help partners adopt the OpenAI tech stack related to identity, permissioning, agent-auth primitives so users can safely connect ChatGPT and Codex to the tools, services, and workflows they already use. Our team also works with external partners on defining the standards for agent access, marketplace offerings as well as other agent enablement initiatives to ensure users of ChatGPT and Codex go from intent to task completion seamlessly. About the role We are looking for an AI Deployment Engineer to help strategic partners design, build, validate, launch, and operate agent enablement integrations across web applications, connectors, APIs, CLIs, MCP servers, and developer tools. This is a hands-on, partner-facing product engineering role for someone who can contribute to the platform itself, lead sophisticated technical engagements, and turn ambiguous identity and agent-workflow requirements into secure, production-ready integrations. You will work across partner product and engineering teams and OpenAI’s product, engineering, design, partnerships, legal, policy, security, support, and go-to-market teams. You will identify high-value user journeys, choose the right integration path, prototype and review architectures, write code, run evaluations and dogfood, trace failures end to end, guide launch and rollout, and support post-launch iteration. The best person for this role moves fluidly between full-stack code, OAuth/OIDC and identity systems, product judgment, project leadership, and clear communication with engineers and executives. This role is a fit for a product-minded engineer who wants to stay close to users and partners while going deep on authentication, permissions, reliability, safety, and developer experience. The principle objective is to
Jobiba hiring network
Production Operator Jobs
3,233 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current production operator jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team OpenAI's mission is to ensure that AGI benefits all of humanity. The Business Systems team helps make that mission possible by building the internal products and platforms that allow OpenAI to operate with speed, reliability, and care. We build internal applications and workflows for Finance and Supply Chain. Our work spans product discovery, React and TypeScript interfaces, Python services and APIs, data models, workflow orchestration, enterprise integrations, and the systems that connect people to systems of record. We work directly with the people who use these products and care about correctness, permissions, auditability, and production reliability. Examples of our work include building an integration platform for supply chain integrations, integrations with Oracle Fusion and Zip, contract intelligence applied to B2B revenue recognition, and Temporal-based agentic workflows for credit checks, duplicate bank detection, and invoice triaging. We turn these efforts into reusable patterns that can support many workflows, rather than one-off automations. About the Role We are looking for Product Engineers to build internal applications end to end. This role spans product discovery, user experience, frontend, backend services, data models, workflow orchestration, and integrations with order management, fulfillment, and supply chain systems. You will take a problem from a first conversation with a Finance or Supply Chain partner through design, implementation, rollout, and production support. Strong candidates combine product judgment with engineering depth. You should be comfortable moving between a React interface, a Python API, a durable workflow, and an integration with an enterprise system. You should be able to ship a useful first version quickly while building the foundations for reuse, security, and long-term maintainability. Direct AI experience is helpful, but the core requirement is strong product engineering judgment and reliable execution. I
About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. Our team builds the software, tooling, and operational systems that help manage this fleet at scale. We work across production engineering, distributed systems, capacity management, and operational automation to improve reliability, reduce manual work, and make better use of available compute. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will develop the systems that help manage the GPU fleet powering ChatGPT, including tooling for fleet health, capacity planning, operational automation, and incident response. You will work closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and compute utilization. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference. Develop systems for capacity planning, fleet health monitoring, and resource utilization. Automate operational workflows, including incident detection, diagnosis, and response. Identify and address bottlenecks affecting fleet reliability, scalability, and performance. Partner with infrastructure, research, and product engineering teams to improve the compute platform. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems. Have a background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering. Have built software that automates operational workflows and reduces manual work. Have worked with distributed infrastructure, cluster orchestration, or large-scale int
About the Team OpenAI's research training infrastructure powers how our frontier models are trained and evaluated. The Simulation team sits at the intersection between the agentic harness that powers OpenAI's products and the research infrastructure where GPT-next is trained, ensuring that our model's training environment is as realistic as possible. This team owns the integration layer that connects our production harness capabilities into the training stack. The work is highly cross-functional and high leverage: researchers depend on it to run experiments and evaluations reliably as well as to develop the next generation of harness capabilities. Failures in this surface can materially affect training velocity and correctness. About the Role We're looking for a Principal Software Engineer to lead the architecture and evolution of the Simulation Platform. You'll own a critical interface between research and engineering, building the systems, APIs, and operational patterns that let researchers use agentic coding infrastructure safely and effectively in training environments. This role is ideal for a senior backend or infrastructure engineer with strong technical judgment, product sense for highly technical users, and the ability to drive execution across multiple teams. The highest-leverage work is building robust infrastructure that supports and accelerates research without compromising engineering quality. In this role, you will Design, build, and evolve the integration between the Codex harness that powers OpenAI's products and research training infrastructure used for training GPT-next Build a platform for our LLMs to train and be evaluated in simulated environments that mimic their deployment setting as closely as possible, on every axis: agentic harness, compute substrate, timing, tools, data sources, humans in the loop, and more Own major integration surfaces end-to-end, from architecture and API design through rollout, operations, and long-term maintenance Bu
About the Team The AI Deployment Engineering team is responsible for ensuring the safe and effective deployment of Generative AI applications for developers and startups. We act as a trusted advisor and thought partner for our customers, working to build an effective backlog of GenAI use cases for their industry and drive them to production through strong technical guidance. As an AI Deployment Engineer (ADE) in the OpenAI for Global Affairs team, you’ll help government and non-profit agencies transform their organization through solutions such as automated content generation, contextual search, and novel applications that make use of our newest, most exciting models and technology. About the Role OpenAI’s mission is to ensure that general-purpose artificial intelligence benefits all of humanity. We believe that achieving our goal requires effective engagement with public policy stakeholders and the broader community impacted by AI. The Global Affairs team builds authentic, collaborative relationships with public officials and the broader AI policymaking community to inform and support our shared work in these domains. We ensure that insights from policymakers inform our work and - in collaboration with our colleagues and external stakeholders - help shape policy guardrails, industry standards, and safe and beneficial development of AI tools. We are looking for an AI Deployment Engineer to collaborate directly with our Global Affairs team to help public sector actors unlock the benefits of OpenAI tools and products, aiming to broadly benefit humanity. This includes supporting a workforce organization deploying Certifications programs or advising a government partner on responsible implementation practices. Your role will integrate technical expertise with our mission to ensure that artificial general intelligence benefits all of humanity. In this role, you will: Technical Enablement & Deployment Serve as the primary technical advisor for Global Affairs partnersh
About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. We work closely with hardware design teams, manufacturing partners, suppliers, and data center operations to deliver the compute platforms that power frontier AI. The Manufacturing Engineering team ensures our hardware can be built, tested, deployed, and supported at hyperscale. We bridge Hardware Engineering, Quality, Supply Chain, Contract Manufacturers, and Hardware Operations to continuously improve manufacturing performance throughout the product lifecycle. As our fleet grows globally, sustaining manufacturing engineering becomes increasingly important to maintain product quality, improve manufacturability, and rapidly resolve production issues. About the Role We are seeking a PCBA Manufacturing Engineer (Sustaining) to support production and continuous improvement of printed circuit board assemblies (PCBAs) used throughout Industrial Compute hardware platforms. This role focuses on sustaining engineering after product launch. You'll partner closely with Hardware Design, Quality, Manufacturing, Test Engineering, Supply Chain, and our contract manufacturers to resolve production issues, improve manufacturing yield, reduce failures, implement engineering changes, and ensure stable high-volume manufacturing. The ideal candidate has experience supporting complex server, networking, accelerator, storage, or high-performance electronics manufacturing environments. Key Responsibilities Own sustaining manufacturing engineering for PCBA production across multiple hardware platforms. Drive root cause investigations for manufacturing defects, field failures, and production escapes. Partner with Hardware Design Engineers to improve manufacturability (DFM/DFA/DFT). Support engineering change orders (ECOs) and manufacturing change implementation. Work directly with contract manufacturers to improve production yield, cycle time, and quality. Analyze manuf
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation
About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Manufacturing Test Engineer to own and drive manufacturing test strategy, development, and execution for complex AI hardware systems. This role will define and implement test coverage across the product lifecycle, including ICT, functional circuit test, tray-level functional test, and system manufacturing test. You will work closely with hardware design engineering, diagnostic/software teams, manufacturing engineering, quality, and external system integrators and suppliers to translate product requirements into robust, scalable, and production-ready test solutions. You will also play a key role in reviewing test data, debugging failures, improving yield, and ensuring manufacturing test readiness from early development through volume production. In This Role, You Will: Define and drive the manufacturing test strategy for boards, trays, and system-level hardware assemblies across EVT, DVT, PVT, and production ramp. Develop and manage test coverage for: ICT / structural test FCT / board-level functional test Tray-level functional and integration test System-level manufacturing and bring-up test Partner closely with electrical engineering, system engineering, and diagnostic/software teams to define test requirements, review manufacturing test scripts and diagnostics, validate failure isolation needs, debug hooks, logging, and production screening strategies. Translate engineering requirements into practical, scalable manufacturing test plans that
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for an experienced systems software engineer to help define and build the host software stack for our custom next-generation AI systems. You will work close to the hardware on performance-critical software, including Linux kernel drivers, high-throughput I/O paths, and system-scale networking and RDMA. This role spans architecture, implementation, platform bring-up, debugging, and performance optimization. You will work across hardware and software boundaries to make new systems usable end to end, from low-level device interfaces through userspace tooling and production validation. In this role you will: Design, implement, and debug host-side systems software for AI infrastructure, including Linux kernel drivers and supporting userspace components. Build and optimize software paths for high-throughput, low-latency communication, including RDMA and related networking functionality. Develop software around PCIe, DMA, NICs, accelerators, memory movement, and device interaction. Bring up new hardware platforms and diagnose complex issues across kernel, firmware, networking, and hardware boundaries. Build tooling for integration, testing, diagnostics, observability, qualification, and performance characterization. Collaborate with hardware, networking, and platform teams to define interfaces and integrate new capabilities. Work with external vendors where needed to integrate technologies and drive issues to resolution. Contribute across the systems sof
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are looking for a highly experienced RTL engineer to own critical on- and off-chip interconnect components for our custom AI accelerator platform. You will drive the microarchitecture and RTL implementation of scalable on-chip communication fabrics connecting high-bandwidth compute, memory, and I/O subsystems as well as purpose-built off-chip interfaces and protocols needed to enable custom computing at scale. This is a senior, hands-on engineering role with broad technical ownership. You will drive design from requirements through the full silicon lifecycle, from architecture definition and performance analysis through RTL implementation, verification closure, physical design convergence, bring-up, and production readiness. You will plan and oversee the work of junior engineers and help drive and develop productive engineering relationships with external partners and help manage partner execution. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the microarchitecture, RTL design, and delivery of major SoC interconnect components, including network-on-chip fabrics, switches, routers, bridges, protocol adapters, arbiters, and traffic-management logic as well as off-chip protocol bridges and interfaces. Drive third party engagements to develop novel networking and interface protocols and silicon IP while ensuring high quality and de
About the Team OpenAI’s API Multicloud team is responsible for extending OpenAI’s API platform into strategic cloud environments, starting with AWS . The team’s mission is to distribute OpenAI’s API broadly and safely by enabling key API technologies in cloud-native environments, in close partnership with Amazon and internal teams across Codex, Research, Safety Systems, and Applied. The team is focused on bringing core developer and enterprise capabilities into cloud-native environments, including cloud-hosted Codex, model customization / post-training as a service, and new stateful runtime environments for agentic workloads. This work sits at the intersection of production ML systems, developer platforms, model behavior, and large-scale infrastructure. About the Role We’re looking for a backend engineer who can quickly understand OpenAI’s models, products, and systems, then adapt first-party deployments for other cloud platforms. You’ll build backend services, APIs, SDK integrations, authentication flows, and cloud service infrastructure that let developers use OpenAI capabilities in the cloud environments where they already build. This role involves working across teams, sometimes embedded with partner product groups, to ship products quickly and across multiple platforms at the same time. It’s a strong fit for engineers who have built developer tools, especially AI-powered tools, communicate clearly across technical boundaries, and can shape architectures that support different deployment models; experience building cloud services is a strong plus. In this role, you will: Build backend and infrastructure systems that extend OpenAI’s API platform into cloud-native environments, like AWS. Design and ship cloud-contained products that allow customers to use OpenAI capabilities while keeping workloads and data within cloud environments. Help stand up cloud-hosted Codex experiences powered by the OpenAI Responses API. Build the infrastructure and runtime abstractions
About the Team The AI Architect team partners with organizations to turn OpenAI's most capable models into meaningful, real-world impact. We work with customers across industries and digital-native businesses to identify where AI can create value, design secure and scalable solutions, and help those solutions move from early exploration into sustained production adoption. The team brings together technical strategy, customer partnership, and practical deployment expertise, working closely with Sales, Product, Engineering, Research, and specialist delivery teams. About the Role As an AI Architect, you will be the senior technical owner for a named portfolio of customers and the primary technical counterpart to their leadership teams. You will act as the “CTO of your book of business”, shaping each customer's AI strategy and guiding their journey from pre-sales discovery and solution evaluation through deployment, adoption, and measurable business impact. You will own the technical account plan across ChatGPT Enterprise, the OpenAI API, Codex, and other agentic AI solutions. In partnership with the Account Director, you will translate business priorities into a focused use-case portfolio, an actionable adoption roadmap, and a clear path to durable customer value and growth. The Account Director owns commercial strategy; you own the technical strategy, customer journey, and path to production value. You will remain accountable for the technical outcome while bringing in the right specialists across deployment, implementation, enablement, security, product, and partners to provide deeper expertise and execute work where needed. This role calls for strong industry fluency, sound architectural judgment, and the ability to move confidently between executive strategy and hands-on technical conversations. In this role, you will: Serve as the primary technical advisor and long-term technical relationship owner for a named portfolio of existing customers and pre-sales prospect
About the Team OpenAI’s API Platform organization builds the products and infrastructure that help first-party and third-party developers build with OpenAI models. We ship the API primitives, tools, SDKs, documentation, playgrounds, and platform experiences that make OpenAI’s capabilities reliable, understandable, and useful in production. The API Experience team is focused on the end-to-end developer experience for the OpenAI API. We own the surfaces developers touch every day: docs, SDKs, the Playground, examples, onboarding flows, and the systems that help developers go from first request to production deployment quickly and confidently. About the Role We’re looking for full stack and frontend engineers to help define and build the next generation of OpenAI’s developer experience. In this role, you’ll work across frontend product surfaces, backend systems, SDK and documentation pipelines, and API workflows that serve millions of developers and companies. You’ll partner closely with product, design, research, API engineering, and developer-facing teams to make complex AI capabilities simple to understand, easy to test, and safe to launch in real-world applications. This is a highly cross-functional role for someone who cares deeply about craft, developer empathy, reliability, and product velocity. In this role, you will: Build and scale developer-facing products including the OpenAI API Playground, documentation experiences, onboarding flows, examples, and API workflow tools. Own full stack projects end to end, from product definition and UX collaboration through backend implementation, launch, measurement, and iteration. Improve the systems that generate, maintain, and publish SDKs, API references, docs, guides, and developer examples. Partner with API, research, design, and infrastructure teams to bring new model capabilities and API primitives to developers in a clear, usable way. Use developer feedback, product analytics, and direct customer insight to identif
About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. Role Overview We are seeking a Package Reliability Engineer to lead reliability engineering for advanced packages used in high-performance AI and computing systems. The primary focus of this role is to assess package level mechanical and thermal reliability risks and apply thermal and mechanical modeling to optimize package design, material selection, and assembly processes. The engineer will also develop reliability test plans with external partners, identify failure mechanisms, perform root-cause analysis, and recommend practical corrective actions. In this role, you will assess package reliability risks from early architecture development through product qualification and high-volume manufacturing. You will work closely with package design, silicon design, system engineering, manufacturing, and ASIC partners to predict package behavior, develop qualification strategies, resolve reliability issues, and improve overall package robustness and lifetime. In this role you will: Lead reliability test plan and assessments for advanced HPC packages, including risk identification, potential failure-mechanism analysis, root-cause investigation, mitigation planning, and corrective-action development. Drive reliability-focused package design optimization based on thermo-mechanical modeling to improve package reliability, power integrity, thermal performance, mechanical robustness, and platform scalability. Develop, validate, and apply package reliability models and lifetime-prediction
About the Team The Applied AI Engineer - Digital Natives team is responsible for ensuring the safe and effective deployment of Generative AI applications for developers and enterprises. We act as a trusted advisor and thought partner for our customers, working to build an effective backlog of frontier AI use cases for their industry and drive them to production through strong technical guidance. As an Applied AI Engineer in the Digital Native segment, you’ll help large and highly sophisticated companies transform their business through custom AI solutions applications such as customer service, automated content generation, contextual search, personalization, and other novel use cases leveraging OpenAI’s newest, most exciting models and latest capabilities. About the Role We are looking for a driven solutions leader with a product mindset to partner with our customers and ensure they achieve tangible business value with frontier AI. You will pair with senior customer leaders to establish AI strategic roadmaps and identify the highest value applications. You’ll then partner with their engineering and product teams to move from prototype through production. You’ll take a holistic view of their needs and design an enterprise architecture using OpenAI APIs and other services to maximize customer value. You will collaborate closely with Sales, Solutions Engineering, Applied Research, and Product. This role is based in our São Paulo, Brazil office. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Deeply embedd with our most sophisticated and technical platform customers, serving as their technical thought partner in ideating and building novel applications on our APIs. Proactively provide guidance to our customers on how to maximize business impact from their applications, accelerating their time to value. Experiment and prototype solutions with and for your customers. Forge and manage rel
Get new production operator jobs by email
Daily job updates · Unsubscribe anytime