Jobs in United States

Production Planner in United States

1,337 active opportunities · Updated October 2026

Explore current production planner jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Ecosystem AI Deployment Engineering (ADE) team supports strategic partners as they build high-quality technical integrations into ChatGPT and Codex. Our goal is to create products users depend on, drive adoption and retention, and build an ecosystem where partners win when OpenAI wins. About the Role We are looking for an AI Deployment Engineer to help strategic partners design, build, evaluate, submit, launch, and maintain high-utility plugins for ChatGPT and Codex. This is a hands-on, partner-facing product engineering role for someone who can contribute to the platform itself, lead sophisticated partner engagements, and translate ambiguous product needs into production-ready integrations. You will work across partner product and engineering teams and OpenAI's product, engineering, partnerships, legal, policy, design, and go-to-market teams. You will identify the right use cases, prototype and review implementations, run evaluations, debug issues across systems, guide partners through submission and review, and support launch and post-launch iteration. The best person for this role moves fluidly between code, product judgment, project leadership, and clear communication with engineers and executives. This role is a fit for a product minded engineer who wants to stay close to users and partners while still going deep on code, reliability, evaluations, and developer experience. The goal is to help partners ship plugins that are not merely technically functional, but genuinely useful in ChatGPT and Codex. This role is based in our San Francisco office. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the technical partner journey for priority B2B plugins—from pitch and readiness assessment through architecture, build, evaluation, submission, launch, and ongoing maintenance. Identify strong plugin use cases, define crisp user journeys and expected behaviors, and

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

PythonAWSAzureGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Data Platform at OpenAI owns the foundational data stack powering critical product, research, and analytics workflows. We operate some of the largest Spark compute fleets in production; design, and build data lakes and metadata systems on Iceberg and Delta with a vision toward exabyte-scale architecture; run high throughput streaming platforms on Kafka and Flink; provide orchestration with Airflow; and support ML feature engineering tooling such as Chronon. Our mission is to deliver reliable, secure, and efficient data access at scale and accelerate intelligent, AI assisted data workflows. Join us to build and operate these core platforms that underpin OpenAI products, research, and analytics. We’re not just scaling infrastructure – we’re redefining how people interact with data. Our vision includes intelligent interfaces and AI-assisted workflows that make working with data faster, more reliable, and more intuitive. About the Role This role focuses on building and operating data infrastructure that supports massive compute fleets and storage systems, designed for high performance and scalability. You’ll help design, build, and operate the next generation of data infrastructure at OpenAI. You will scale and harden big data compute and storage platforms, build and support high-throughput streaming systems, build and operate low latency data ingestions, enable secure and governed data access for ML and analytics, and design for reliability and performance at extreme scale. You will take full lifecycle ownership: architecture, implementation, production operations, and on-call participation. You’ve supported Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms. You’re well-versed in infrastructure tooling like Terraform, experienced in debugging large-scale distributed systems, and excited about solving data infrastructure problems in the AI space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per wee

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Applied AI Engineering team helps organizations turn frontier AI capabilities into safe, reliable, and high-impact production systems. We work with customer executives, product and engineeriIng teams, security leaders, and transformation teams to identify valuable opportunities, accelerate technical implementation, and scale what works. Enterprise deployments are defined by complexity rather than any one industry: existing architectures, diverse data environments, security and governance requirements, multiple stakeholder groups, and organization-wide change. We turn lessons from these deployments into better products and reusable patterns for customers everywhere. About the Role As an Enterprise Applied AI Engineer you will partner directly with leading organizations to design, build, and deploy AI systems that deliver measurable business outcomes. You will combine deep technical judgment, hands-on engineering, and customer leadership to take ambitious ideas from use-case selection and architecture through prototyping, evaluation, production launch, and scale. You will write and debug code, build evaluation systems, resolve complex integrations, and guide decisions involving model behavior, reliability, latency, cost, safety, security, governance, and operational readiness. Success is measured by production systems, sustained adoption, and meaningful customer impact—not simply activity or successful demonstrations. This is a rare opportunity to work on consequential real-world deployments at the frontier of AI while directly influencing how OpenAI’s products evolve. This role is based in our SF or NYC office. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Partner directly with enterprise customers to identify high-value opportunities and translate them into technical architectures, implementation plans, evaluation strategies, and measurable success criteri

JavaScriptTypeScriptPythonJava
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Storage Infrastructure team builds and operates the storage foundation behind OpenAI’s most demanding workloads. We work directly with research to design storage systems for rapidly evolving experiments, while also powering production at scale. We own the platform end to end: backend systems, user-facing services and APIs, and the control planes that manage how data is placed, moved, and retained over time. Our stack spans cloud and in-house object stores across very different workload profiles, from GPU-attached systems to dedicated storage hardware. We also build the federation layer that unifies these backends behind a simple interface and routes each workload to the right storage solution. About the Role You will help build the storage platform that powers OpenAI’s research and production systems. This is a hands-on infrastructure role for engineers who want to work on deeply technical systems at scale and own them in production. You’ll work across object storage, cross-region data movement, lifecycle management, and the federation layer that provides a unified interface across multiple backends. Much of our stack runs on Kubernetes, and we primarily build services in Rust. In this role, you will: Build and operate storage services that underpin OpenAI’s research infrastructure Develop object storage systems across cloud and in-house environments Build systems for cross-region data movement, replication, and recovery Design lifecycle management capabilities that keep data durable, available, and cost-effective Evolve the federation layer that unifies multiple backend systems behind a simple interface Improve performance, reliability, and operational excellence across the platform Collaborate closely with researchers and infrastructure teams to support rapidly evolving workloads You might thrive in this role if you: Have experience building or operating distributed systems in production Have worked on storage infrastructure, object stores, dist

AWSKubernetesRestAI
O
📍 Washington, District of Columbia, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team The AI Deployment Engineering team is responsible for ensuring the safe and effective deployment of Generative AI applications. We act as a trusted advisor and thought partner for our customers, working to build an effective backlog of GenAI use cases for their industry and drive them to production through strong technical guidance. As an AI Deployment Engineer (ADE) in the OpenAI for Government team, you’ll help government agencies transform their organization through solutions such as automated content generation, contextual search, and novel applications that make use of our newest, most exciting models and technology. About the Role We are looking for a driven solutions leader with a product mindset to partner with our public sector customers and ensure they achieve tangible value with GenAI. You will pair with government agencies (federal, state, and local), policymakers, and other public institutions to establish a GenAI strategy and identify the highest value applications. You’ll then partner with their technical teams, subject matter experts, systems integrators, and implementation partners to move from prototype through production. You’ll take a holistic view of their needs and design an architecture using the OpenAI API and other services to maximize customer value. You will collaborate closely with Sales, Solutions Engineering, Global Affairs, Applied Research, and Product teams. This role is based in Washington, DC. We offer relocation support to new employees. In this role, you will: Deeply embed with our most sophisticated public sector customers as the technical lead, serving as their technical thought partner to ideate and build novel applications on our API and other OpenAI products. Work with senior customer stakeholders to identify the best applications of AI in their industry and to build/qualify a comprehensive backlog to support their AI roadmap. Intervene directly to accelerate customer time to value through building hands-on pr

JavaScriptPythonJavaAWS
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team ChatGPT is a rapidly evolving system: new capabilities ship continuously, product surfaces change quickly, and usage patterns shift week-to-week. Supporting that pace requires infrastructure that can handle real production constraints—high concurrency, unpredictable traffic patterns, complex dependency graphs, and frequent change. The ChatGPT Infrastructure team builds and operates the platforms that enable fast iteration without compromising performance or reliability. We design shared systems, data paths, rollout mechanisms, and reliability guardrails that teams rely on to ship changes to ChatGPT at scale. We focus on high-leverage infrastructure: primitives and “golden paths” that incorporate operational lessons as defaults, so engineers don’t need to rediscover failure modes, latency pitfalls, or integration issues each time they build something new. About the Role We’re hiring Senior and Staff Engineers to design and build infrastructure systems that underlie ChatGPT and multiply the effectiveness of teams building user experiences. This is not a support-only role. It’s a platform-building role: you’ll define interfaces, develop core abstractions, and create tooling to make safe, fast iteration the norm. Your work will reduce friction, prevent regressions, improve performance, and ensure systems scale gracefully as the product grows. Where You Can Have Impact You might work on one or more of the following areas (without being restricted to any single area): Platform foundations & frameworks: Core libraries, service frameworks, and shared components that standardize system building, integration, and evolution. Scalability & performance primitives: Patterns and infrastructure that reduce tail latency, improve throughput, and keep costs predictable as demand increases. Reliability guardrails: Mechanisms that prevent outages by design—rate limiting, load shedding, dependency isolation, backpressure, safe fallbacks, and robust regression contr

RedisAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the role We’re looking for an engineering manager to lead a team building software systems that detect and prevent harmful misuse of frontier AI models—before incidents occur. This is a builder’s role: you’ll lead engineers shipping production services, detection pipelines, and mitigation mechanisms that protect frontier model integrity and reduce high-severity misuse risk. While this work intersects with frontier model development, security and risk, we’re explicitly seeking someone with a software engineering foundation who is comfortable building reliable systems that can operate at billions of users scale. In this role you will: Lead a team of software engineers building detection + mitigation systems for frontier model misuse, with an emphasis on model IP protection / distillation detection and emerging risk surfaces from autonomous agents. Set the technical roadmap and execution strategy: prioritize, design, ship, iterate, measure impact. Build production systems: services, pipelines, tooling, instrumentation, and automation that scale with frontier model usage. Partner deeply with Research and Product to translate evolving model capabilities into concrete tests, signals, and mitigations that can be deployed at scale. Drive strong engineering fundamentals: architecture, reliability, monitoring, performance, and operational excellence. Hire and grow an exceptional team across backend, data systems, and applied ML engineering domains as needed. Anticipate what breaks at scale as agentic workflows become more capable. You might thrive in this role if you: Experience building systems in adversarial, fast-evolving environments Are comfortable with ambiguity and novelty Have experience adjacent to security (e.g., abuse prevention, fraud, integrity, platform defense, auth/identity, malware/spam, adversarial environments) Communicate clearly and build trust quickly with senior stakeholders—pragmatic, collaborative, and calm under scrutiny. Significant experience

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Codex Core Agent team builds the kernel of Codex. We own making the agent better, accelerating research, and making those improvements real in production for our users. That means working across the systems that make Codex actually function as an agent in the real world: the production performance envelope around tokens, latency, reliability, cost, and capacity; the core execution loop and interfaces that turn models into useful behavior; the shared infrastructure that enables other teams to build on Codex; and the feedback loops that turn real-world usage into better models and better agent behavior over time. About the Role We’re looking for engineers to build the infrastructure that powers Codex agents in production. This role focuses on the systems that let models safely execute code, interact with tools, complete long-running tasks, and operate reliably and efficiently at scale. You’ll design and operate the infrastructure behind sandboxed execution, orchestration, stateful workflows, app-server and SDK boundaries, and model rollouts. You’ll work at the intersection of distributed systems, developer tooling, and AI, building primitives that make Codex faster, safer, more reliable, and easier for the rest of the organization to build on. What You’ll Do Design and build execution environments for AI agents, including sandboxing, isolation, and reproducibility. Develop systems for agent orchestration across multi-step, tool-using workflows. Build infrastructure for running, testing, and debugging code generated by models. Create state and memory systems that allow agents to persist context across long-running tasks. Optimize tokens, latency, reliability, and cost across Codex’s production fleet. Support model rollouts, capacity planning, and the core tradeoffs between quality, speed, and economics to manage a fleet of frontier agents at scale. Build shared platform capabilities that unblock product teams, partner teams, and open source Codex. Yo

AWSCI/CDRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical. The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly. About the Role We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity. In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity. This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions. This role is based in San Francisco, CA, with travel as needed. Key Responsibilities Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint Drive readiness to convert contracted compute capacity into schedulable production clusters Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives Build integrated schedules spanning procurement, logistics, installation, st

AWSAzureRestAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%

From $230K/yr

Quick readStrong listing-quality and freshness signals

About the Role The Engineering Acceleration Delivery / Continuous Deployment team builds and operates the systems that safely ship OpenAI’s infrastructure and product code to production. We own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across OpenAI to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale infrastructure orchestration. In This Role, You Will Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces, and alerts to detect regressions and trigger safe rollbacks. Build systems that support agent-assisted or autonomous deployment workflows using modern AI tooling. Technologies commonly used in this environment include: Kubernetes for large-scale container orchestration and runtime infrastructure Python and FastAPI for internal services Terraform for infrastructure as code GitOps-based deployment workflows (e.g., ArgoCD, Flux, or similar systems) Buildkite for CI orchestration You may be a strong fit if you: Have worked with Kubernetes-based deployment systems at scale Have experience building or operating continuous deployment platforms Are familiar with GitOps tooling such as

PythonAWSKubernetesGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Codex Core Agent team builds the kernel of Codex. We own making the agent better, accelerating research, and making those improvements real in production for our users. That means working across the systems that make Codex actually function as an agent in the real world: the production performance envelope around tokens, latency, reliability, cost, and capacity; the core execution loop and interfaces that turn models into useful behavior; the shared infrastructure that enables other teams to build on Codex; and the feedback loops that turn real-world usage into better models and better agent behavior over time. About the Role We’re looking for applied AI engineers to help bring Codex agents from impressive demos to dependable tools. This role is about improving agent performance on real software engineering tasks and closing the gap between research capability and real-world usefulness. You’ll work closely with research, infrastructure, and product to ensure agents are not just powerful, but useful, steerable, and reliable in practice. The job is not only to improve model behavior in isolation, but to turn those improvements into measurable gains in solve rate, usefulness, and economic value for users. What You’ll Do Design and iterate on agent behaviors across real-world coding tasks and long-horizon workflows. Work closely with research to develop and run evals to measure agent performance, regressions, failure modes, and edge cases. Improve performance through prompting, tool-use strategies, context construction, and model-facing experimentation. Analyze failures in production and systematically improve robustness and reliability. Build feedback loops and data systems that get better real-task data into evaluation and research. Work with product teams to shape user-facing agent experiences and the interfaces the agent depends on. Help define what “good” looks like for agents completing complex tasks end-to-end. You Might Be a Good Fit If You Ha

PythonAWSRestMachine Learning
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -86%

From $270K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: Notion is looking for an experienced engineer who has designed and built secure software systems to help define the security foundations for our AI products. You’ll work with product and engineering teams on agent runtimes, tool permissions, retrieval, content writes, observability, and abuse-resistant design. You’ll turn product risks into clear architecture, reusable guardrails, automated tests, and production systems that help teams ship new features safely. This role is based in San Francisco. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: Define and build security architecture for product surfaces that operate across customer workspace content, including tool execution, content writes, retrieval, permission checks, provenance, and auditability. Make the secure path the easy path for product teams by shipping reusable libraries, review patterns, test fixtures, and guardrails that prevent classes of vulnerabilities.

RestAIGoRust
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Cortex Code is Snowflake’s coding agent for building with data. It ships inside the platform that thousands of the world’s largest enterprises — including a large share of the Forbes Global 2000 — run their data on, which means the quality of this agent is felt by the data teams behind a meaningful slice of the global economy. We are taking coding agents from impressive demos to tools Data Science and Engineering teams depend on every day, and we hold them to a rigorous, public bar: see our data engineering agent benchmark . About the Role This is a measurement-first role that owns the quality and efficiency of Cortex Code end to end: how good the agent is, how much it costs to run, and how reliably it behaves in production. You will take agents from research capability to real, measurable user value — turning fuzzy “the agent feels worse” signals into hard metrics, running the experiments that move them, and shipping the changes that stick. You will work on a small, high-powered modeling and infrastructure team where your work reaches every developer building on Snowflake. What you will do in this role Take agents from prototype to production: design and refine agent behaviors for real coding and data-engineering workflows, and make them reliable enough to depend on. Own a

TypeScriptPythonAIGo
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Data Clean Rooms team is Leading the market shift from traditional 2-party data sharing to multi-party collaboration hubs . Our vision is to provide a seamless, "safe-room" environment where enterprises can collaborate on shared datasets while maintaining absolute governance. We ensure that no party can exfiltrate another's underlying content, even while running complex joint workloads and getting high-value results. You will join a fast-paced, collaborative team of engineers on a journey to provide customers with an integrated set of innovative, AI-enabled capabilities to analyze data in a privacy-preserving way. You will have a real opportunity to impact and shape the future of secure data collaboration at Snowflake. AS A SOFTWARE ENGINEER IN DATA CLEAN ROOMS, YOU WILL: Architect and build highly scalable infrastructure that enables secure, multi-party collaboration. Design and implement core clean room features and services, intelligent agents, and robust developer APIs to expand platform capabilities and support custom AI/ML workflows. Partner closely with Product Management and cross-functional teams to drive complex projects from ideation and system design through to production deployment. Mentor peers and foster a warm, supportive culture of innovation, cross-tea

PythonJavaAIGo
🔔

Get new production planner jobs in United States by email

Daily job updates · Unsubscribe anytime