About the team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a researcher working on Frontier Evals & Environments, you will help build north star model environments to drive progress towards safe AGI/ASI. Your work will directly guide the research programs of the most ambitious training runs happening at OpenAI. Some prior open-sourced evaluations built by researchers in this role include GDPval , SWE-bench Verified , MLE-bench , PaperBench , and SWE-Lancer . If you are interested in feeling firsthand the fast progress of our models, and steering them towards good outcomes, this is the role for you. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you'll: Create ambitious RL environments to push our models to their limits, and measure frontier
Jobiba hiring network
Codex Deployment Engineer Jobs
8,439 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current codex deployment engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Connectors, you will teach models how to interface with the top professional software using code. You will help train agents to use code, APIs, tools, and structured integrations to operate across applications like Slack, Google Workspace, GitHub, Notion, Linear, Salesforce, and other core systems of work. You will help enable models to take useful actions across a user’s digital context: finding information, updating systems, coordinating work, generating artifacts, and completing multi-step workflows through the tools teams already use. You will train models to be supercharged by the world’s most important productivity and enterprise software, turning connected tools into a powerful action surface for our agents. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people.
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, you will improve the capabilities, reliability, and product fit of OpenAI's agentic models. You might own a research direction, build the infrastructure that makes large training runs faster and more trustworthy, create evals that reveal where models fail, or drive a capability from an idea through experimentation, integration, and launch. This role is intentionally broad. The strongest candidates are not defined by one method or subfield; they are people who can take an ambiguous capability problem and make progress across research, engineering, data, evals, and product. You should be excited to work on models that act in the world: writing and debugging code, using tools, calling functions, operating computers, collaborating with other agents, and completing valuable work on behalf of users. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, meas
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a researcher working on Frontier Evals & Environments, you will help build north star model environments to drive progress towards safe AGI/ASI. Your work will directly guide the research programs of the most ambitious training runs happening at OpenAI. Some prior open-sourced evaluations built by researchers in this role include GDPval , SWE-bench Verified , MLE-bench , PaperBench , and SWE-Lancer . If you are interested in feeling firsthand the fast progress of our models, and steering them towards good outcomes, this is the role for you. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you might Create ambitious RL environments to push our models to their limits, and measure frontie
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team builds the data, environments, graders, training methods, and feedback loops that shape what OpenAI’s next agents can do and what they are like to work with, then carries those improvements through major training runs and into products used by people every day. About the Role As a member of the Agent Post-training Personality team, you will help make OpenAI’s agents exceptional collaborators. You will study what makes an agent thoughtful, clear, perceptive, appropriately proactive, and genuinely easy to work with, then translate those insights into evals, training data, reward signals, and model improvements. We use “personality” to mean much more than writing style or general likability. It includes whether an agent understands what the user is trying to accomplish, communicates with good judgment, adapts to context, asks useful questions, handles disagreement honestly and takes initiative at the right moments. The goal is to create a strong, tasteful default that can adapt to different people and situations. This work combines behavioral research, product thinking, research and communication taste. You will collaborate with product teams, human experts, and researchers across post-training and pretraining to ensure that improvements survive the full trai
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of this API & power-users team, you will improve the capabilities, reliability, and product fit of OpenAI’s agentic models for power users and API developers. You might design evals from real developer workflows, build training environments around production-like tool use, turn qualitative model failures into training data, evals, or post-training interventions, or drive a behavior improvement from discovery through post-training, integration, and launch. This role is intentionally broad. The strongest candidates are comfortable turning ambiguous model behavior problems into concrete progress, whether that means improving tool use, planning, instruction following, recovery from mistakes, or how models behave in API-based workflows. You should be excited to work across research, engineering, data, evals, and product to make models better at acting in real workflows. You will work closely with researchers, engineers, API/product teams, Codex, infrastructure, and safety/align
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Computer Use, you will teach models to operate computers. You will help train models that can navigate browsers and desktops, use tools and applications, reason through complex workflows, collaborate with users and other agents, and complete long-horizon tasks with reliability and judgment. This work sits at the intersection of frontier model training, product behavior, evaluation, and systems engineering, and will directly shape the computer-use capabilities shipped in OpenAI’s next generation of agents. Currently, our models are the best in the world at this behavior! You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you might Design and run experiments th
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Artifacts, you will train frontier models to create polished, useful work products: documents, spreadsheets, slide decks, dashboards, reports, analyses, and other interactive or editable artifacts. You will help teach our models to move from a vague user goal to a finished artifact with strong structure, visual taste, domain judgment, correctness, and low latency. This work will require owning improvements across our post-training stack, including RL, data pipelines, graders, reward signals, evals, and behavioral analysis. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you will: Design and run experiments that improve agentic model behavior for complex so
About the Team The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world. Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure. About the Role We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology. You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more. There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future. In this role, you will: Build, lead, and grow high-performing infrastructure engineering teams. Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search. Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach. Champion pragmatic use of agent technology to amplify execution velocity. Reduce operational toil and incident frequency through better abstractions, gua
About the Team API Frontiers turns OpenAI’s frontier models into production APIs that developers can use to build reliable products and agents. We own the core path connecting models to developers through the Responses API, with a focus on safety, reliability, and speed. Working closely with Research, Safety, Codex, and other API teams, we bring new model capabilities into production and improve them through developer feedback. About the Role We are looking for a backend software engineer to build and operate the services behind the Responses API. You will shape API behavior, bring new capabilities from research into production, and make long-running agent workflows dependable and fast. The work combines distributed systems engineering with product judgment: designing useful developer interfaces, managing staged rollouts, and following production issues through to durable fixes. In this role, you will: Design, build, and operate APIs and backend services that bring frontier model capabilities to developers. Partner with Research, Safety, Codex, and API teams to define API behavior and deliver safe, staged launches. Build API capabilities for agent workflows, including task delegation, context sharing, and parallel execution. Strengthen long-running request reliability across timeouts, cancellation, streaming, and background execution. Improve request-processing performance and tail latency through profiling, efficient systems code, and persistent connections. Turn developer feedback and production failures into better observability, diagnostics, and lasting product improvements. Your background might look something like: 5+ years of experience building and operating backend services or developer-facing APIs in production. Strong software engineering fundamentals, with practical knowledge of distributed systems, concurrency, and asynchronous execution. Ability to diagnose production failures and performance bottlenecks using observability data and profiling. Product
About the Team The Statsig team is responsible for the experimentation, feature rollout, dynamic configuration, and analytics systems that help OpenAI ship products with speed, safety, and evidence. Teams across ChatGPT, Codex, model measurement, monetization, business subscriptions, developer products, and shared infrastructure rely on Statsig to introduce capabilities safely, measure their impact, and make high-confidence product decisions. About the Role As a Product Lead on the Statsig team, you will define how experimentation, rollout, configuration, and analytics become a simple, reliable, and trusted part of how every OpenAI product team ships. You will set strategy across multiple product and platform workstreams, translate company-wide needs into durable capabilities, and help Statsig become a core part of OpenAI’s product development system. We’re looking for a product leader who combines strong product judgment, technical fluency, and deep analytical thinking. You should be comfortable navigating ambiguous customer needs, influencing teams across the company, and balancing rapid adoption with reliability, usability, and measurement quality. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Travel requirements should be confirmed with the recruiter before publishing. In this role, you will: Define the product vision, strategy, and roadmap for experimentation, feature management, dynamic configuration, rollout safety, and analytics. Partner with product, engineering, research, data, design, and infrastructure leaders to turn recurring launch and measurement needs into reusable platform capabilities. Develop a deep understanding of workflows across ChatGPT, Codex, model measurement, monetization, subscriptions, and developer products, then establish clear priorities across competing needs. Drive adoption by making sophisticated experimentation and analytics c
About the Team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. Our product teams work across research, engineering, product, data, and design to bring OpenAI’s technology to people and businesses around the world. Design plays a critical role in making powerful AI intuitive, useful, and trustworthy. We hold a high bar for quality and build experiences that earn users’ confidence, while moving quickly and learning from real-world feedback. We favor early validation and continuous refinement over waiting until every detail is perfect. About the Role We’re looking for a Product Design Leader to shape how startups, small businesses, and growing teams discover, adopt, and unlock lasting value from Codex and ChatGPT. You’ll lead design across the B2B growth journey, creating experiences that turn the potential of AI into meaningful, everyday impact for this user base. This role is a mix of team leadership and hands-on design execution. You’ll lead and develop a small team of designers, contribute directly to high-impact projects, and help set the standard for impactful, simple, user-centric experiences. You balance exceptional craft with momentum, creating a culture where the team launches, learns, and improves quickly. You already use tools such as Codex to extend what you can build. As a leader, you help designers strengthen their judgment, confidence, and influence through clear and actionable feedback. You’re also a highly effective cross-functional partner who brings people together, navigates ambiguity, and builds alignment through clear communication and collaborative problem-solving. Codex is on an extraordinary trajectory, and this role offers a rare opportunity to work across both sides of the product: the consumer experiences through which people first discover and adopt our technology, and the business experiences that help teams use it together. You’ll bring a high bar for craft, strong product judgment, and comfor
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. As a Staff Product Manager, AI Foundations at Vanta, you will build the core AI capabilities that power Vanta’s AI product experiences: the agent harness our agents are built on, and the quality and evaluation stack that makes them trustworthy. The AI pillar is making Vanta the intelligence layer for security and compliance work, wherever that work starts. We build the AI products and foundations that work across Vanta: a cohesive first-party Vanta Agent experience, headless surfaces like MCP and CLI that bring Vanta into tools like Claude Code and Codex, and the shared capabilities that let every team ship high-trust AI products. This role sits behind Vanta’s customer-facing AI experiences. Today, every team shipping AI features has to answer the same hard questions on their own: how to define quality, how to evaluate it, and how to know a change is safe to ship. You will turn those answers into shared foundations, so the AI experiences our customers rely on get better and more trustworthy over time. What you’ll do as a Staff Product Manager, AI Foundations at Vanta: Own the strategy and roadmap for Vanta’s core AI capabilities: the agent harness and the AI quality and evaluation stack Live in agent traces to build intuition for how customers actually use our AI products, and turn that into the highest-leverage improvements Make Vanta’s AI products win on quality: build golden datasets, automated evaluators aligned to expert judgment, and self-serve tooling so PMs, subject-matter experts, and designers can run evals without an engineer in the loop Scale the agent harness: shape the core capabilities Vanta’s agents are built on
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Customers increasingly use AI agents like Claude Code, Codex, Cursor, and internal Slack agents to run their business processes, and security and compliance work is no exception. For those agents to do that work, they need access to Vanta’s trust graph and intelligence layer. Vanta is building for that world: security and compliance products that customers can use through Vanta, through third-party agent platforms, and inside their own internal tools. As Senior Product Manager, Agent Ecosystem, you will own Vanta’s external agent surfaces: MCP, CLI, and plugins in marketplaces like Claude Code, Codex, and ChatGPT. Your job is to make Vanta the trusted compliance layer that agents rely on. You will define the strategy, ship the products, and own the outcomes. What you’ll do as a Senior Product Manager, Agent Ecosystem at Vanta: Own the strategy and roadmap for Vanta’s agent ecosystem products. Define how agents discover, access, and safely use Vanta capabilities, partnering with the platform teams that own APIs and authentication. Hold outcome parity as a product principle: work a customer can complete in Vanta should be possible through MCP, CLI, and plugin surfaces too. Lead marketplace and plugin launches, aligning product, engineering, legal, security, and go-to-market around each one. Partner with engineering and design to ship secure, reliable developer products. Build the developer-facing marketing motion: external agent documentation and customer stories that show what MCP and CLI users have unlocked. Partner with the product partnerships and alliances team on co-marketing with agent platforms like Anthropic and OpenAI,
We’re hiring a Senior Technical Product Marketing Manager to lead positioning, messaging, and product adoption across our AI and agentic context solutions. This is a high-impact role for a technical marketer who thinks and acts like a builder. Most AI systems don't fail because the model is bad—they fail because they lack a crucial, contextual foundation. MongoDB owns a piece of nearly every step in that chain. This role requires deep technical expertise across the context stack—including VoyageAI embedding and reranking models, MCP Server, and the coding agent ecosystem (Claude Code, Cursor, GitHub Copilot, Gemini CLI, Codex, etc.)—combined with product marketing instincts to deeply understand the market and translate technical depth into crisp, differentiated messaging for distinct user and buyer personas. The ideal candidate has experience building and shipping AI-enabled products and using developer tools in their day-to-day workflows. Individuals with prior experience in technical sales, developer relations, or software development are encouraged to apply. The ideal candidate is a voracious consumer of AI research and pays close attention to shifting patterns in how software and AI applications are built and consumed. This individual is confident in communicating with technical practitioners and non-technical decision makers in one-to-few and one-to-many engagements for internal and external audiences. You will leverage your technical depth to craft clear, compelling, and highly differentiated messaging by working backward from customer requirements. We are looking to speak to candidates who are based in Dublin for our hybrid working model. What You’ll Do Drive Strategy & Execution: Act as the strategic owner for MongoDB’s context engineering portfolio—MCP Server, Agent Skills, VoyageAI, and more— aligning roadmap and go-to-market with MongoDB’s long-term business goals, in collaboration with Product Management, Engineering, Developer Relations, Partn
Get new codex deployment engineer jobs by email
Daily job updates · Unsubscribe anytime