We’re looking for a Software Engineer to architect and build backend systems that enforce data privacy and automate compliance at scale. You’ll work closely with product, infrastructure, security, and legal teams to embed privacy-by-design into our data and access layers. This is a hands-on, high-impact role for an experienced engineer who is passionate about protecting user data while enabling innovation. What You’ll Do Design, build, and operate backend services that enforce policy-driven data access, lifecycle controls, and privacy protections. Develop distributed authorization and identity-aware enforcement mechanisms integrated directly into data services and control planes. Implement auditability, policy hooks, and enforcement observability to ensure compliance is continuously verifiable. Partner with Security, Legal, and Compliance to convert privacy requirements into scalable technical designs and developer-friendly APIs. Harden data platforms and backend services through schema-level controls and data handling constraints by default. Collaborate with infrastructure teams to ensure consistent enforcement across systems while minimizing duplicated implementations. Contribute patterns, libraries, and education that elevate trustworthy data access patterns across the organization. You Might Thrive in This Role If You Have 5+ years of industry experience building and operating backend or infrastructure systems in production. Strong software engineering fundamentals , with fluency in at least one major programming language (e.g., Python, Go, Rust, C++, Java). Experience with distributed authorization, RBAC/ACL systems, encryption-based access, or policy engines. Familiarity with global privacy regulations and their architectural implications. Ability to influence and collaborate with teams across legal, compliance, product, and engineering. A bias toward practical, impactful solutions that balance privacy protections with product needs. Nice to Have Experience wi
Jobiba hiring network
Production Operator Jobs
3,235 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current production operator jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team OpenAI’s User Operations team shepherds our customers’ adoption of AI and ensures that our customers' product experience is nothing short of exceptional. We are building the very first post-AGI support team. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others, to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. About the Role We’re looking for dedicated, experienced, and deeply curious individuals to help solve some of the most complex challenges faced by our customers while building the future of post-AGI support alongside us. In this role, you’ll work directly with customers through support tickets and live calls, troubleshooting high-impact issues and resolving novel, often ambiguous technical problems in one of the fastest-moving environments in technology. As AI adoption rapidly accelerates, the work you do will directly support mission-critical systems being built on OpenAI’s platform, serving as a critical line of defense for customers operating at massive scale. Beyond resolving technical issues, you’ll help define what world-class support looks like in an AGI-driven future. You’ll partner closely with Engineering, Product, and Operations to improve systems, reduce bugs, and elevate the customer experience, while leveraging automation, agents, and our own AI technology to transform how support operates at scale. You’ll be responsible for: Working directly with customers to troubleshoot and resolve their most complex technical issues, including API failures, integration challenges, authentication errors, and production incidents. Providing end-to-end ownership through debugging logs, analyzing system behavior, reproducing issues, and
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our
About the Team The Foundations Research team works on high-risk, high-reward ideas that could shape the next decade of AI. Our goal is to advance the science and data that enable our training and scaling efforts, with a particular focus on future frontier models. Pushing the boundaries of data, scaling laws, optimization techniques, model architectures, and efficiency improvements to propel our science. The Search team sits within Foundations, building agentic search by co-designing model–system interfaces with the core search stack (serving, indexing, retrieval) to translate model intent into reliable, real-world actions. Operating at the frontier of AI and information retrieval, the team develops large-scale systems that transform and index vast corpora, enabling models to reason over global knowledge and act dependably. In close partnership with researchers, we rapidly bring modeling breakthroughs into production and redefine how intelligent systems discover, retrieve, and synthesize information at planetary scale. About the Role We’re looking for a researcher focused on our embedding retrieval efforts. You’ll work with a a team of world-class research scientists and engineers developing foundational technology that enables models to retrieve and condition on the right information, at the right time. This includes designing new embedding training objectives, scalable vector store architectures, and dynamic indexing methods. This work will support retrieval across many OpenAI products and internal research efforts, with opportunities for scientific publication and deep technical impact. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Responsibilities Tackle embedding models and retrieval systems optimized for grounding, relevance, and adaptive reasoning. Collaborate with a team of researchers and engineers building end-to-end infrastructure for training, evaluati
About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Software Engineer, Distributed Data Systems, you will design, build, and operate some of the largest distributed data systems in the world. You will be responsible for the end-to-end stack to deliver and consume top-quality data for robotics training at exabyte-scale. You’ll manage distributed data pipelines, collaborate closely with researchers to translate requirements into robust systems, and harden pipelines that serve as the backbone for OpenAI’s rapid iteration cycles. We’re looking for engineers who are detail-oriented, have strong experience with distributed systems, and excel at building reliable, large-scale systems in high-stakes environments. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain data infrastructure such as exabyte-scale distributed data processing, data selection, automated labeling, and training data loaders. Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient. Partner with researchers to deeply understand requirements and translate them into production-ready systems. Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation. Deliver the best possible data for training robotics models. You might thrive in this role if you: Have strong experience with distributed systems and large-scale infrastructure with a strong interest in data. Are detail-oriented a
About the Team: The Database Systems team specializes in high-performance distributed databases. Our team built Rockset, the real-time search, analytics, and vector database that powers all vector search and retrieval augmented generation (RAG) at OpenAI. In addition to retrieval, as an online database, Rockset powers core functionality across all of OpenAI's product lines and many critical internal use cases. About the Role : We are looking for engineers passionate about distributed systems, close-to-the-metal performance optimization (our core engine is written in C++), and building scalable database infrastructure from the ground up. As an engineer on the Database Systems team, you'll contribute to the core database engine, driving improvements across ingestion, query execution, indexing, and storage. You'll partner with teams across OpenAI to unlock new product capabilities and help scale online database reliability and throughput as usage grows by orders of magnitude. In this role you will: Design, build, and operate high-performance distributed systems Identify and resolve performance bottlenecks to scale infrastructure to the next order of magnitude Define long-term technical direction and guide system evolution Collaborate with product, engineering, and research teams to deliver scalable and reliable infrastructure Dig deep into complex production issues across the stack Contribute to incident response, postmortems, and best practices for system reliability You might thrive in this role if you: Have significant experience building, scaling, and optimizing distributed systems at scale Are curious about database internals, storage engines, or low-latency query systems Enjoy debugging challenging performance issues in complex, high-throughput systems Have experience operating production clusters at scale (e.g., Kubernetes or other orchestration systems) Think rigorously about scalability, correctness, and reliability Thrive in fast-paced environments with high
About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role OpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency of our systems. In this role, you'll apply deep technical expertise to optimize infrastructure and application-level performance across mission-critical products like ChatGPT and our developer API. You’ll work cross-functionally with teams building core services, training models, and developing real-time user experiences to push our latency, throughput, and cost-efficiency to the next level. We are looking for engineers who thrive in ambiguous environments, value deep systems understanding, and are motivated by delivering measurable impact. This is a highly technical, individual contributor role focused on root-cause analysis, profiling, instrumentation, and architecture-level performance improvements across our stack. In this role, you will: Analyze and optimize performance across application, middleware, runtime, and infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond. Develop tooling and metrics that provide deep observability into system performance. Collaborate closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements. Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale. Lead investigations into high-impact performance regressions or scalability issues in production. Drive performance testing strategies and help define SLAs/SLOs around latency and throughput for critical systems. You might thrive in this role if you: Have 7+ years of experience in software engineering with a strong tr
About the Team The Strategic Finance team at OpenAI plays a critical role in shaping the company’s long-term trajectory. We partner closely with Product, Engineering, and Go-To-Market teams to inform high-stakes decisions through rigorous data science and economic modeling. As part of our expanding Data Science function, we’re building a best-in-class Forecasting capability to drive real-time, data-driven decision-making across user growth, revenue, compute infrastructure, and more. We are developing scalable forecasting infrastructure to help us understand and anticipate business dynamics in an increasingly complex, usage-based world. Our models are foundational to planning, pricing, operational efficiency, and growth strategy - supporting key investment decisions and unlocking OpenAI’s full potential. About the Role We’re looking for a senior Machine Learning Data Scientist to lead our forecasting initiatives. You’ll be one of the founding members of the Forecasting pillar within Strategic Finance Data Science, responsible for building and scaling robust, interpretable, and production-ready forecasting systems. Your models will power critical business decisions by predicting core metrics such as DAU/WAU, revenue, LTV, compute consumption, and profitability. This is a highly cross-functional role, requiring technical excellence, strong product intuition, and business acumen. You’ll collaborate with product managers, researchers, engineers, and finance leaders to operationalize forecasting insights, influence company-wide strategy, and build foundational forecasting capabilities at OpenAI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build statistical and machine learning models to solve forecasting needs across product, finance, infrastructure, and GTM domains. Own the end-to-end modeling lifecycle , including scoping, feature engineerin
About the Team At OpenAI, we’re building safe and beneficial artificial general intelligence. We deploy our models through ChatGPT, our APIs, and other cutting-edge products. Behind the scenes, making these systems fast, reliable, and cost-efficient requires world-class infrastructure. The Caching Infrastructure team is responsible for building a caching layer that powers many critical use cases at OpenAI. We aim to provide a high-availability, multi-tenant cache platform that scales automatically with workload, minimizes tail latency, and supports a diverse range of use cases. We’re looking for an experienced engineer to help design and scale this critical infrastructure. The ideal candidate has deep experience in distributed caching systems (e.g., Redis, Memcached), networking fundamentals, and Kubernetes-based service orchestration. In This Role, You Will: Design, build, and operate OpenAI’s multi-tenant caching platform used across inference, identity, quota, and product experiences. Define the long-term vision and roadmap for caching as a core infra capability, balancing performance, durability, and cost. Collaborate with other infra teams (e.g., networking, observability, databases) and product teams to ensure our caching platform meets their needs. You Might Thrive In This Role If You: Have 5+ years of experience building and scaling distributed systems, with a strong focus on caching, load balancing, or storage systems. Have deep expertise with Redis, Memcached, or similar solutions, including clustering, durability configurations, client-side connection patterns, and performance tuning. Have production experience with Kubernetes, service meshes (e.g., Envoy), and autoscaling systems. Think rigorously about latency, reliability, throughput, and cost in designing platform capabilities. Thrive in a fast-paced environment and enjoy balancing pragmatic engineering with long-term technical excellence. About OpenAI OpenAI is an AI research and deployment company d
About the team Online Data builds and operates Habitat, the single product surface of Online Data and the system of record for OpenAI’s online user data. As OpenAI’s scale and product requirements evolve, Habitat is becoming a full-stack, one-size-fits-most database platform with end-to-end ownership of: Provisioning and developer experience APIs and guardrails Scaling, performance, and reliability Data movement, caching, routing, and placement Privacy enforcement and access control Change Data Capture (CDC) as a first-class primitive The foundation for future storage backends You’ll work on the core online database platform behind OpenAI’s products, building and operating Habitat services that handle high-QPS, latency-sensitive workloads across regions. You’ll partner closely with internal platform and product teams to ship safe, reliable systems, then push them to be faster and more cost-efficient through better caching, routing, observability, and operational tooling. This is a critical role for engineers who like owning hard distributed-systems problems end to end and sweating the details from p99 latency to production operations at massive scale. In this role, you will Design and build core abstractions spanning storage, caching, routing, CDC, and privacy enforcement Own a major surface area end to end, from product and API design to operational excellence Improve latency, correctness, and cost efficiency for real production workloads at massive scale Build strong instrumentation, debugging workflows, and developer-first tooling Collaborate closely with internal product and infrastructure teams to understand requirements and ship pragmatic solutions Participate in an on-call rotation and raise the bar on reliability while aggressively improving performance and usability You might thrive in this role if you have A strong track record building and operating high-scale backend or data-intensive distributed systems in production Excellent systems judgment and the a
About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We're seeking talented Robotics Software Engineers to expand our robotics data collection and evaluation program. This highly technical role involves designing, implementing, and optimizing software solutions across diverse robotics hardware. You'll work closely and collaboratively with multidisciplinary teams; including software, hardware, research, and operations - to drive advancements in our robotic systems. This role is based in San Francisco, CA, and requires in-person 4 days a week. In this role, you will: Help develop and grow our data collection labs, owning the entire integration lifecycle, from identifying and sourcing new hardware to collaborating with mechanical and electrical engineers on setup, software integration, and operational deployment. Develop innovative robot control interfaces suited to a variety of morphologies, environments, and tasks. Collaborate closely with research and engineering teams to develop automation tools and machinery that facilitate the evaluation of advanced robotic policies. Lead the design and implementation of data collection, visualization, and quality control processes. You might thrive in this role if you: Have 5+ years of professional software engineering experience developing and shipping production-quality systems in robotics or hardware-integrated environments. Have extensive experience integrating and deploying industrial automation systems, off-the-shelf robotics platforms, or custom hardware into production environments. Bring hands-on experience delivering production-quality software
About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software is reliable, testable, and ready to ship. We design and maintain automated test frameworks, hardware-in-the-loop labs, and release pipelines that keep quality signals trustworthy and enable rapid, safe product launches. Our work spans developer tools, automation, systems integration, and cross-team collaboration to ensure every release meets the highest standards. About the Role As a Software Engineer, Quality and Developer Tools , you will build and own the systems that validate our device software—from test frameworks and regression infrastructure to hardware-in-the-loop labs and release gates. You’ll design the tooling and automation that keep quality signals trustworthy, integrate them into CI/CD, and make it easy for engineers and QA vendor technicians to execute reliable, repeatable workflows. We’re looking for engineers with deep experience in software quality, automation, developer tooling, and hardware-software integration who thrive on building scalable, reliable systems for validation and release readiness. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Test infrastructure & frameworks: Design, implement, and maintain a unified test framework for device software across unit, integration, system, and end-to-end testing, with reproducible runs and integrations with GitHub, Linear, and Slack. CI/CD integration & releases: Integrate test suites with Buildkite, enforce promotion criteria for staging and production, auto-file regressions, and publish traceable artifacts and release notes. Hardware-in-the-loop lab design & orchestration: Plan and bring up racks, power and networking systems, and orchestration for device testing; support automated flashing, provisioning
About the Team The Safety Systems team is dedicated to ensuring the safety, robustness, and reliability of AI models and their deployment in the real world. Learn more about OpenAI’s approach to safety. Building on the many years of our practical alignment work and applied safety efforts, Safety Systems addresses emerging safety issues and develops new fundamental solutions to enable the safe deployment of our most advanced models and future AGI, to make AI that is beneficial and trustworthy. About the Role At OpenAI, we're dedicated to advancing artificial intelligence, and we know that creating a secure and reliable platform is vital to our mission. That's why we're seeking a software engineer to help us build out our trust and safety capabilities. In this role, you'll work with our entire engineering team to design and implement systems that detect and prevent abuse, promote user safety, and reduce risk across our platform. You'll be at the forefront of our efforts to ensure that the immense potential of AI is harnessed in a responsible and sustainable manner. Your Responsibilities: Architect, build, and maintain anti-abuse and content moderation infrastructure designed to protect us and end users from unwanted behavior. Work closely with our other engineers and researchers to utilize both industry standard and novel AI techniques to measure, monitor and improve AI models’ alignment to human values. . Diagnose and remediate active incidents on the platform and build new tooling and infrastructure that address the root causes of system failure. You might thrive in this role if: You have built and run production services in a high growth, rapidly scaling environment. You can debug live issues and restore systems quickly. You have worked on content safety, fraud, or abuse, or are motivated and excited to work on present-day (“now-term”) AI safety. You have experience with Python or with modern languages such as C++, Rust, or Go, and are able to quickly ramp up on Py
About the Team With Codex we’re building an AI software engineer. One that you can pair with, delegate to, or even ask to take on future tasks proactively. Our team is a fast-moving group within OpenAI, bringing together research, engineering, design, and product. We iteratively build the Codex agent harness and product to get the most out of the model, and we iteratively train the model to be great at complex software engineering tasks. The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. We operate across research, engineering, product, and infrastructure; owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. Codex Enterprise builds the ecosystem building blocks, discovery surfaces, and enterprise capabilities that help Codex spread across developers, teams, and organizations worldwide. It is a cross-cutting team that works across the stack to build both delightful product experiences and fundamental platform capabilities. Its customers range from individual developers and small teams to large enterprises, and our mission is critical to achieving the vision of Codex as a proactive teammate. About the Role As we grow, we’re focused on turning Codex from a powerful individual tool into a production-grade teammate for entire organizations. You will work across internal OpenAI teams and external customers, from fast-moving startups to large enterprises, to make it possible to deploy, operate, and trust Codex in increasingly demanding real-world environments. As Codex’s consumer adoption accelerates, enterprise demand is growing just as quickly, and there is also increasing opportunity to expand Codex through ecosystem capabilities that unlock new workflows, integrations, and discovery. This team helps turn messy, real-world team requirements into robust, repeatable, and scalable product and
About the Team The Applied team at OpenAI safely brings cutting-edge technology to the world. We have released widely used products such as ChatGPT, Sora, and the OpenAI API, powering models including GPT-5 and a growing set of multimodal capabilities across text, image, audio, and video. Our team also manages large-scale inference and platform infrastructure that supports these experiences at global scale. With much more on the horizon, our impact continues to grow. Our customers build fast-growing businesses using our APIs, unlocking product capabilities that were previously unimaginable. ChatGPT and Sora exemplify the breadth of what’s now possible across text, image, audio, and video experiences. As these capabilities expand, we prioritize the responsible use of our technology, emphasizing safe and thoughtful deployment over unchecked growth. Within Applied Engineering, the Ads Monetization team in Financial Engineering builds the core systems dealing with all the money flows for ChatGPT Ads. These systems are a combination of low-latency, high scale, high reliability, while being built in a financially correct, accurate, auditable and explainable way. This role sits at the intersection of ads delivery, data engineering, and financial systems. In this role, you will: Architect and build the core monetization systems for ChatGPT Ads. Build and operate the core services and pipelines that power ads monetization end-to-end, from event capture and validation through aggregation, pricing, metering, and ultimately producing billable outputs. Define and implement the source of truth for ads monetization data, including schemas, data models, and invariants that ensure outputs are consistent, explainable, and auditable. Own correctness and reconciliation: align production outputs with downstream invoicing/finance requirements, build controls/monitors, and close gaps through investigations and backfills. Develop across the stack to create comprehensive billing integration
Get new production operator jobs by email
Daily job updates · Unsubscribe anytime