Jobs in United States

Production Operator in United States

1,337 active opportunities · Updated October 2026

Explore current production operator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Stargate organization is responsible for building and scaling the physical infrastructure systems that power OpenAI’s next generation of AI training and inference platforms. This includes the manufacturing, deployment, and operational execution required to bring large-scale compute infrastructure online globally. The team operates at the intersection of data center infrastructure, hardware manufacturing, supply chain, deployment operations, and systems planning. We partner closely across Infrastructure Strategy, Manufacturing Operations, Capacity Planning, Supply Chain, Deployment, and Engineering to execute one of the largest infrastructure scale-outs in the industry. About the Role We are seeking a Technical Program Manager, Rack Delivery to drive operational execution across rack manufacturing, site readiness, and deployment coordination for Stargate infrastructure programs. This role will serve as a key connective layer between manufacturing partners, deployment teams, and infrastructure readiness programs to ensure rack production and delivery timelines remain aligned with site availability and deployment sequencing. You will help manage operational execution across contract manufacturers (CMs), support build planning and RCCA processes, and coordinate deployment readiness across multiple concurrent infrastructure programs. You will also partner closely with Demand Planning teams to translate strategic planning inputs into actionable SKU-level manufacturing and delivery schedules. This role is ideal for someone who thrives operating across ambiguity, manufacturing operations, infrastructure deployment, and large-scale cross-functional execution. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation support. Key Responsibilities Drive cross-functional coordination between rack manufacturing, deployment operations, and site readiness programs. Manage operational execution acros

AWSRestAIGo
O
📍 Washington, District of Columbia, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role Our technologies support some of the most important and impactful work in the world, including our strategic and high-impact customers in the public sector. As a Forward Deployed Security Engineer (FDSecE) you will be responsible for securing these novel applications of OpenAI’s technology. We’re looking for motivated, tenacious, and curious people who will work closely with engineering teams to ensure our infrastructure deployments are highly secure against our adversaries. As an FDSecE, you will embed directly throughout the lifecycle, working on-site and being hands-on to ensure the overall security of these deployments from design to production and through ongoing operations. This role is preferred to be based in Washington DC but may consider remote work. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Travel to and working from customer sites is required for this role. In this role, you will: Deeply embed with our most strategic public sector customers to implement and maintain robust security controls. Be a design and technical thought partner by leveraging security expertise on protective controls including access controls, authentication, encryption, network, and system security. Collaborate closely with teammates, cross-functional teams, customers, and service providers to achieve security and compliance goals. Ensure continuity of critical security and monitoring c

PythonAWSAzureKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training. Our mission is to make some of the world's largest GPU clusters work reliably for frontier training. We bring new platforms and clusters online, safely maintain installed fleets, and partner with hardware, infrastructure, and research teams to resolve the system-level issues that keep jobs from running. That means building and maintaining the software closest to the machine: Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation. We make these components reproducible, compatible, and safe to operate across heterogeneous fleets. About the Role We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate. You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models

AWSLinuxRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Integrity team at OpenAI is dedicated to ensuring that our cutting-edge technology is not only revolutionary, but also secure from a myriad of adversarial threats. We strive to maintain the integrity of our platforms as they scale. The Integrity team is at the front lines of defending against misuse in all its forms: content abuse, scaled attacks, and other actions that could undermine the user experience or harm our operational stability. About the Role As a Machine Learning Engineer in OpenAI's Integrity team, you will have the opportunity to work with some of the brightest minds in AI. You’ll work on state-of-the-art models and classifiers, experiment with new architecture and approaches, and push forward our abilities in content and user understanding. You’ll help turn research breakthroughs into tangible solutions that improve the trust and safety of our platform. If you're excited about training LLMs and building ML models, this role is your chance to make a significant mark. In this role, you will: Innovate and Deploy: Design and deploy advanced machine learning models that solve real-world problems. Bring OpenAI's research from concept to implementation, creating AI-driven applications with a direct impact. Collaborate with the Best: Work closely with researchers, software engineers, and product managers to understand complex business challenges and deliver AI-powered solutions. Be part of a dynamic team where ideas flow freely and creativity thrives. Optimize and Scale: Implement scalable data pipelines, optimize models for performance and accuracy, and ensure they are production-ready. Contribute to projects that require cutting-edge technology and innovative approaches. Learn and Lead: Stay ahead of the curve by engaging with the latest developments in machine learning and AI. Take part in code reviews, share knowledge, and lead by example to maintain high-quality engineering practices. Make a Difference: Monitor and maintain deployed m

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

We’re looking for a Software Engineer to architect and build backend systems that enforce data privacy and automate compliance at scale. You’ll work closely with product, infrastructure, security, and legal teams to embed privacy-by-design into our data and access layers. This is a hands-on, high-impact role for an experienced engineer who is passionate about protecting user data while enabling innovation. What You’ll Do Design, build, and operate backend services that enforce policy-driven data access, lifecycle controls, and privacy protections. Develop distributed authorization and identity-aware enforcement mechanisms integrated directly into data services and control planes. Implement auditability, policy hooks, and enforcement observability to ensure compliance is continuously verifiable. Partner with Security, Legal, and Compliance to convert privacy requirements into scalable technical designs and developer-friendly APIs. Harden data platforms and backend services through schema-level controls and data handling constraints by default. Collaborate with infrastructure teams to ensure consistent enforcement across systems while minimizing duplicated implementations. Contribute patterns, libraries, and education that elevate trustworthy data access patterns across the organization. You Might Thrive in This Role If You Have 5+ years of industry experience building and operating backend or infrastructure systems in production. Strong software engineering fundamentals , with fluency in at least one major programming language (e.g., Python, Go, Rust, C++, Java). Experience with distributed authorization, RBAC/ACL systems, encryption-based access, or policy engines. Familiarity with global privacy regulations and their architectural implications. Ability to influence and collaborate with teams across legal, compliance, product, and engineering. A bias toward practical, impactful solutions that balance privacy protections with product needs. Nice to Have Experience wi

PythonJavaAWSAzure
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s User Operations team shepherds our customers’ adoption of AI and ensures that our customers' product experience is nothing short of exceptional. We are building the very first post-AGI support team. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others, to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. About the Role We’re looking for dedicated, experienced, and deeply curious individuals to help solve some of the most complex challenges faced by our customers while building the future of post-AGI support alongside us. In this role, you’ll work directly with customers through support tickets and live calls, troubleshooting high-impact issues and resolving novel, often ambiguous technical problems in one of the fastest-moving environments in technology. As AI adoption rapidly accelerates, the work you do will directly support mission-critical systems being built on OpenAI’s platform, serving as a critical line of defense for customers operating at massive scale. Beyond resolving technical issues, you’ll help define what world-class support looks like in an AGI-driven future. You’ll partner closely with Engineering, Product, and Operations to improve systems, reduce bugs, and elevate the customer experience, while leveraging automation, agents, and our own AI technology to transform how support operates at scale. You’ll be responsible for: Working directly with customers to troubleshoot and resolve their most complex technical issues, including API failures, integration challenges, authentication errors, and production incidents. Providing end-to-end ownership through debugging logs, analyzing system behavior, reproducing issues, and

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

AWSLinuxRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Foundations Research team works on high-risk, high-reward ideas that could shape the next decade of AI. Our goal is to advance the science and data that enable our training and scaling efforts, with a particular focus on future frontier models. Pushing the boundaries of data, scaling laws, optimization techniques, model architectures, and efficiency improvements to propel our science. The Search team sits within Foundations, building agentic search by co-designing model–system interfaces with the core search stack (serving, indexing, retrieval) to translate model intent into reliable, real-world actions. Operating at the frontier of AI and information retrieval, the team develops large-scale systems that transform and index vast corpora, enabling models to reason over global knowledge and act dependably. In close partnership with researchers, we rapidly bring modeling breakthroughs into production and redefine how intelligent systems discover, retrieve, and synthesize information at planetary scale. About the Role We’re looking for a researcher focused on our embedding retrieval efforts. You’ll work with a a team of world-class research scientists and engineers developing foundational technology that enables models to retrieve and condition on the right information, at the right time. This includes designing new embedding training objectives, scalable vector store architectures, and dynamic indexing methods. This work will support retrieval across many OpenAI products and internal research efforts, with opportunities for scientific publication and deep technical impact. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Responsibilities Tackle embedding models and retrieval systems optimized for grounding, relevance, and adaptive reasoning. Collaborate with a team of researchers and engineers building end-to-end infrastructure for training, evaluati

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Software Engineer, Distributed Data Systems, you will design, build, and operate some of the largest distributed data systems in the world. You will be responsible for the end-to-end stack to deliver and consume top-quality data for robotics training at exabyte-scale. You’ll manage distributed data pipelines, collaborate closely with researchers to translate requirements into robust systems, and harden pipelines that serve as the backbone for OpenAI’s rapid iteration cycles. We’re looking for engineers who are detail-oriented, have strong experience with distributed systems, and excel at building reliable, large-scale systems in high-stakes environments. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain data infrastructure such as exabyte-scale distributed data processing, data selection, automated labeling, and training data loaders. Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient. Partner with researchers to deeply understand requirements and translate them into production-ready systems. Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation. Deliver the best possible data for training robotics models. You might thrive in this role if you: Have strong experience with distributed systems and large-scale infrastructure with a strong interest in data. Are detail-oriented a

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team: The Database Systems team specializes in high-performance distributed databases. Our team built Rockset, the real-time search, analytics, and vector database that powers all vector search and retrieval augmented generation (RAG) at OpenAI. In addition to retrieval, as an online database, Rockset powers core functionality across all of OpenAI's product lines and many critical internal use cases. About the Role : We are looking for engineers passionate about distributed systems, close-to-the-metal performance optimization (our core engine is written in C++), and building scalable database infrastructure from the ground up. As an engineer on the Database Systems team, you'll contribute to the core database engine, driving improvements across ingestion, query execution, indexing, and storage. You'll partner with teams across OpenAI to unlock new product capabilities and help scale online database reliability and throughput as usage grows by orders of magnitude. In this role you will: Design, build, and operate high-performance distributed systems Identify and resolve performance bottlenecks to scale infrastructure to the next order of magnitude Define long-term technical direction and guide system evolution Collaborate with product, engineering, and research teams to deliver scalable and reliable infrastructure Dig deep into complex production issues across the stack Contribute to incident response, postmortems, and best practices for system reliability You might thrive in this role if you: Have significant experience building, scaling, and optimizing distributed systems at scale Are curious about database internals, storage engines, or low-latency query systems Enjoy debugging challenging performance issues in complex, high-throughput systems Have experience operating production clusters at scale (e.g., Kubernetes or other orchestration systems) Think rigorously about scalability, correctness, and reliability Thrive in fast-paced environments with high

AWSAzureGCPKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role OpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency of our systems. In this role, you'll apply deep technical expertise to optimize infrastructure and application-level performance across mission-critical products like ChatGPT and our developer API. You’ll work cross-functionally with teams building core services, training models, and developing real-time user experiences to push our latency, throughput, and cost-efficiency to the next level. We are looking for engineers who thrive in ambiguous environments, value deep systems understanding, and are motivated by delivering measurable impact. This is a highly technical, individual contributor role focused on root-cause analysis, profiling, instrumentation, and architecture-level performance improvements across our stack. In this role, you will: Analyze and optimize performance across application, middleware, runtime, and infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond. Develop tooling and metrics that provide deep observability into system performance. Collaborate closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements. Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale. Lead investigations into high-impact performance regressions or scalability issues in production. Drive performance testing strategies and help define SLAs/SLOs around latency and throughput for critical systems. You might thrive in this role if you: Have 7+ years of experience in software engineering with a strong tr

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Strategic Finance team at OpenAI plays a critical role in shaping the company’s long-term trajectory. We partner closely with Product, Engineering, and Go-To-Market teams to inform high-stakes decisions through rigorous data science and economic modeling. As part of our expanding Data Science function, we’re building a best-in-class Forecasting capability to drive real-time, data-driven decision-making across user growth, revenue, compute infrastructure, and more. We are developing scalable forecasting infrastructure to help us understand and anticipate business dynamics in an increasingly complex, usage-based world. Our models are foundational to planning, pricing, operational efficiency, and growth strategy - supporting key investment decisions and unlocking OpenAI’s full potential. About the Role We’re looking for a senior Machine Learning Data Scientist to lead our forecasting initiatives. You’ll be one of the founding members of the Forecasting pillar within Strategic Finance Data Science, responsible for building and scaling robust, interpretable, and production-ready forecasting systems. Your models will power critical business decisions by predicting core metrics such as DAU/WAU, revenue, LTV, compute consumption, and profitability. This is a highly cross-functional role, requiring technical excellence, strong product intuition, and business acumen. You’ll collaborate with product managers, researchers, engineers, and finance leaders to operationalize forecasting insights, influence company-wide strategy, and build foundational forecasting capabilities at OpenAI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build statistical and machine learning models to solve forecasting needs across product, finance, infrastructure, and GTM domains. Own the end-to-end modeling lifecycle , including scoping, feature engineerin

PythonSQLAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team At OpenAI, we’re building safe and beneficial artificial general intelligence. We deploy our models through ChatGPT, our APIs, and other cutting-edge products. Behind the scenes, making these systems fast, reliable, and cost-efficient requires world-class infrastructure. The Caching Infrastructure team is responsible for building a caching layer that powers many critical use cases at OpenAI. We aim to provide a high-availability, multi-tenant cache platform that scales automatically with workload, minimizes tail latency, and supports a diverse range of use cases. We’re looking for an experienced engineer to help design and scale this critical infrastructure. The ideal candidate has deep experience in distributed caching systems (e.g., Redis, Memcached), networking fundamentals, and Kubernetes-based service orchestration. In This Role, You Will: Design, build, and operate OpenAI’s multi-tenant caching platform used across inference, identity, quota, and product experiences. Define the long-term vision and roadmap for caching as a core infra capability, balancing performance, durability, and cost. Collaborate with other infra teams (e.g., networking, observability, databases) and product teams to ensure our caching platform meets their needs. You Might Thrive In This Role If You: Have 5+ years of experience building and scaling distributed systems, with a strong focus on caching, load balancing, or storage systems. Have deep expertise with Redis, Memcached, or similar solutions, including clustering, durability configurations, client-side connection patterns, and performance tuning. Have production experience with Kubernetes, service meshes (e.g., Envoy), and autoscaling systems. Think rigorously about latency, reliability, throughput, and cost in designing platform capabilities. Thrive in a fast-paced environment and enjoy balancing pragmatic engineering with long-term technical excellence. About OpenAI OpenAI is an AI research and deployment company d

RedisAWSKubernetesRest
O
📍 Seattle, Washington, United States· Full-time
✓ Quality checkedCompany trend -82%

About the team Online Data builds and operates Habitat, the single product surface of Online Data and the system of record for OpenAI’s online user data. As OpenAI’s scale and product requirements evolve, Habitat is becoming a full-stack, one-size-fits-most database platform with end-to-end ownership of: Provisioning and developer experience APIs and guardrails Scaling, performance, and reliability Data movement, caching, routing, and placement Privacy enforcement and access control Change Data Capture (CDC) as a first-class primitive The foundation for future storage backends You’ll work on the core online database platform behind OpenAI’s products, building and operating Habitat services that handle high-QPS, latency-sensitive workloads across regions. You’ll partner closely with internal platform and product teams to ship safe, reliable systems, then push them to be faster and more cost-efficient through better caching, routing, observability, and operational tooling. This is a critical role for engineers who like owning hard distributed-systems problems end to end and sweating the details from p99 latency to production operations at massive scale. In this role, you will Design and build core abstractions spanning storage, caching, routing, CDC, and privacy enforcement Own a major surface area end to end, from product and API design to operational excellence Improve latency, correctness, and cost efficiency for real production workloads at massive scale Build strong instrumentation, debugging workflows, and developer-first tooling Collaborate closely with internal product and infrastructure teams to understand requirements and ship pragmatic solutions Participate in an on-call rotation and raise the bar on reliability while aggressively improving performance and usability You might thrive in this role if you have A strong track record building and operating high-scale backend or data-intensive distributed systems in production Excellent systems judgment and the a

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We're seeking talented Robotics Software Engineers to expand our robotics data collection and evaluation program. This highly technical role involves designing, implementing, and optimizing software solutions across diverse robotics hardware. You'll work closely and collaboratively with multidisciplinary teams; including software, hardware, research, and operations - to drive advancements in our robotic systems. This role is based in San Francisco, CA, and requires in-person 4 days a week. In this role, you will: Help develop and grow our data collection labs, owning the entire integration lifecycle, from identifying and sourcing new hardware to collaborating with mechanical and electrical engineers on setup, software integration, and operational deployment. Develop innovative robot control interfaces suited to a variety of morphologies, environments, and tasks. Collaborate closely with research and engineering teams to develop automation tools and machinery that facilitate the evaluation of advanced robotic policies. Lead the design and implementation of data collection, visualization, and quality control processes. You might thrive in this role if you: Have 5+ years of professional software engineering experience developing and shipping production-quality systems in robotics or hardware-integrated environments. Have extensive experience integrating and deploying industrial automation systems, off-the-shelf robotics platforms, or custom hardware into production environments. Bring hands-on experience delivering production-quality software

AWSRestAIC++
🔔

Get new production operator jobs in United States by email

Daily job updates · Unsubscribe anytime