Jobs in United States

Software Reliability Engineer in United States

2,125 active opportunities · Updated October 2026

Explore current software reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team With Codex we’re building an AI software engineer. One that you can pair with, delegate to, or even ask to take on future tasks proactively. Our team is a fast-moving group within OpenAI, bringing together research, engineering, design, and product. We iteratively build the Codex agent harness and product to get the most out of the model, and we iteratively train the model to be great at complex software engineering tasks. The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. We operate across research, engineering, product, and infrastructure; owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. Codex Enterprise builds the ecosystem, governance, and enterprise capabilities that help Codex spread across developers, teams, and organizations worldwide. The Enterprise Controls team owns the systems that allow companies to safely deploy Codex across their organization while protecting their most sensitive code, data, and internal knowledge. About the Role As Codex adoption grows inside large organizations, customers are increasingly trusting Codex with their most valuable assets: proprietary codebases, internal documentation, customer data, and sensitive workflows. This role will help build the enterprise control plane that makes Codex secure, governable, and trustworthy at scale. You will design and operate backend systems that give enterprise administrators visibility and control over how Codex is used across their organization. You will work across identity, access, encryption, policy enforcement, auditability, and admin controls. This may include systems that let customers manage encryption keys, control which Codex capabilities are enabled, enforce organizational policies, and understand how data flows through Codex. This role owns systems end-to-end: from architecture and

PythonJavaAWSRest
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s mission is to ensure that artificial general intelligence (AGI) benefits all of humanity. A key part of achieving that mission is training models that deeply understand and reflect human preferences — the Human Data team is at the heart of that effort. The Human Data engineering team creates the systems that enable scalable, high-quality human feedback. These systems are essential to how OpenAI trains and improves its most advanced models. Engineers on this team collaborate closely with world-class researchers to bring alignment techniques to life — from experimental ideas to production-ready feedback loops. About the Role We’re looking for software engineers to join the Human Data team and build the platforms, prototypes, tools, and infrastructure that power how our AI models are trained, aligned, and evaluated. You’ll partner with researchers and cross-functional teams to bring alignment ideas to life, influence future model training, and shape how models interact with the real world. We’re looking for people who are excited by technical ownership, enjoy working across the stack, and are eager to solve ambiguous problems in a high-impact, fast-paced environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build and maintain robust full-stack systems for feedback collection, data labeling, and evaluation pipelines, while maintaining high levels of security. Translate experimental alignment research into scalable production infrastructure, including inference and model training stacks. Design and iterate on user-facing tools and backend services to support high-quality data workflows Partner with researchers, engineers, and program leads to shape feedback loops and model interaction paradigms Drive infrastructure improvements that enable faster iteration and scaling across OpenAI’s frontier models, from internal r

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Cooperative AI team is scaling OpenAI with OpenAI. We are building a model powered knowledge system that evolves and learns as our products, systems and customers evolve. We leverage our state of the art models, technologies, and products (some external, some still in the lab) to assist or completely automate robust operations supporting both internal and external customers. We support OpenAI customers and internal partners globally, powering systems from customer support to integrity to product insights. We are a self-contained multi-disciplinary team, who enjoy a lightning fast feedback loop with customers at scale, some of whom sit just a few pods away. We iterate fast, and engineer for reliable long-term impact. We're constantly looking for the similarities and patterns in different types of work, and focus on building simple primitives, to apply world class knowledge to many domains. The work of this team exemplifies use of OpenAI technologies. We build systems so everyone can see the leverage that is possible with well designed AI-based implementations. We do this by working through internal use cases focused on Customers (specifically knowledge systems, automation systems, and automated agent systems) to prove impact, then we scale. About the Role We’re looking for a Backend Software Engineer to help architect and scale the infrastructure that powers our knowledge systems. This is a deeply technical and highly cross-functional role where you’ll build robust systems and backend services that serve as the foundation for how knowledge is created, accessed, and applied across OpenAI. In this role, you will: Design, build, and maintain backend services and APIs to support intelligent automation and knowledge systems Integrate and structure data across internal platforms, transforming it into formats optimized for use by downstream systems and AI workflows. Collaborate closely with product, research, and engineering teams to integrate OpenAI mode

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, and scalable in production, and we partner closely with Research to bring the next generation of models into the world. We're a small, fast-moving team of engineers focused on delivering a world-class developer experience while pushing the boundaries of what AI can do. We’re expanding into multimodal inference, building the infrastructure needed to serve models that handle image, audio, and other non-text modalities. These workloads are inherently more heterogeneous and experimental, involving diverse model sizes and interactions, more complex input/output formats, and tighter coordination with product and research. About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal models at scale. You’ll be part of a small team responsible for building reliable, high-performance infrastructure for serving real-time audio, image, and other MM workloads in production. This work is inherently cross-functional: you’ll collaborate directly with researchers training these models and with product teams defining new modalities of interaction. You'll build and optimize the systems that let users generate speech, understand images, and interact with models in ways far beyond text. In this role, you will: Design and implement inference infrastructure for large-scale multimodal models. Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs. Enable experimental research workflows to transition into reliable production services. Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities. Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers. You might thrive in t

AWSRestAIRust
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software is reliable, testable, and ready to ship. We design and maintain automated test frameworks, hardware-in-the-loop labs, and release pipelines that keep quality signals trustworthy and enable rapid, safe product launches. Our work spans developer tools, automation, systems integration, and cross-team collaboration to ensure every release meets the highest standards. About the Role As a Software Engineer, Quality and Developer Tools , you will build and own the systems that validate our device software—from test frameworks and regression infrastructure to hardware-in-the-loop labs and release gates. You’ll design the tooling and automation that keep quality signals trustworthy, integrate them into CI/CD, and make it easy for engineers and QA vendor technicians to execute reliable, repeatable workflows. We’re looking for engineers with deep experience in software quality, automation, developer tooling, and hardware-software integration who thrive on building scalable, reliable systems for validation and release readiness. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Test infrastructure & frameworks: Design, implement, and maintain a unified test framework for device software across unit, integration, system, and end-to-end testing, with reproducible runs and integrations with GitHub, Linear, and Slack. CI/CD integration & releases: Integrate test suites with Buildkite, enforce promotion criteria for staging and production, auto-file regressions, and publish traceable artifacts and release notes. Hardware-in-the-loop lab design & orchestration: Plan and bring up racks, power and networking systems, and orchestration for device testing; support automated flashing, provisioning

PythonAWSCI/CDGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Cooperative AI team is scaling OpenAI with OpenAI. We are building an AI powered knowledge system that evolves and learns as our products, systems and customers evolve. We leverage our state of the art models, technologies, and products (some external, some still in the lab) to assist or completely automate robust operations supporting both internal and external customers. We support OpenAI customers and internal partners globally, powering systems from customer support to integrity to product insights. We are a self-contained multi-disciplinary team, who enjoy a lightning fast feedback loop with customers at scale, some of whom sit just a few pods away. We iterate fast, and engineer for reliable long-term impact. We're constantly looking for the similarities and patterns in different types of work, and focus on building simple primitives, to apply world class knowledge to many domains. The work of this team exemplifies use of OpenAI technologies. We build systems so everyone can see the leverage that is possible with well designed AI-based implementations. We do this by working through internal use cases focused on Customers (specifically knowledge systems, automation systems, and automated agent systems) to prove impact, then we scale. About the Role We’re looking for Software Engineers who're passionate about blending production-ready platform architecture with new tech and new paradigms. You’ll push the boundaries of OpenAI’s newest technologies to enable interactions and automations that are not only functional, but delightful. We value proactive, customer-centric engineers who can get the foundational details right (data models, architecture, security) in service of enabling great products. In this role, you will: Own the end-to-end development lifecycle for new platform capabilities and integrations with other systems Collaborate closely with engineers, data scientists, information systems architects, and internal customers to understand th

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Monetization team is a new cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products—including next-generation ads experiences—that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to help build the core monetization and ads systems at OpenAI. This is a foundational role responsible for designing and implementing the infrastructure, APIs, and user-facing experiences that will power OpenAI’s next-generation monetization products—including ads. You’ll work across the full technical stack to architect, build, and ship 0→1 systems that are robust, safe, and scalable. You will collaborate deeply with Product, Design, and Research to define the future of monetized AI experiences and ensure these systems meet OpenAI’s highest standards for safety, privacy, and policy alignment. This role is exclusively based across our San Francisco and Seattle sites. We offer relocation assistance to new employees. In this role, you will: Design, build, and scale the core infrastructure behind OpenAI’s monetization and ads products Develop advertiser-facing APIs and tools that enable the cre

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team GTM Growth Engineering builds AI-native products and systems that help OpenAI's go-to-market and B2B marketing organizations operate with greater speed, focus, and leverage. Our mandate is revenue leverage: products tied directly to pipeline quality, customer engagement, seller and marketer productivity, and the speed at which OpenAI can bring its technology to customers. We build the infrastructure and user experiences behind high-impact GTM workflows, including customer context, prioritization, routing, campaign execution, review surfaces, feedback loops, and measurement. Our work combines product craft, applied AI, reliable systems, and thoughtful operational design. About the Role We’re looking for a product-minded Software Engineer to build AI-powered products and full-stack experiences for GTM Growth Engineering. You will own meaningful product slices end to end, from user experience and frontend implementation to backend APIs, integrations, data models, instrumentation, and launch readiness. This is a role for engineers who want to build products that do real work in production. You will partner with Product, Design, Data Science, Sales, B2B Marketing, and operations teams to understand high-value workflows and ship systems that improve customer engagement, pipeline, conversion, and team productivity. The role is ideal for a strong product engineer who can move between product craft, systems engineering, applied AI, and measurable business outcomes. You should be excited to build from ambiguous problem statements, ship quickly, and improve products based on real user feedback. What You'll Do Build AI-powered products and workflows that help sales and B2B marketing teams identify opportunities, coordinate work, and engage customers more effectively. Own full-stack product experiences from prototype through launch, instrumentation, iteration, and production hardening. Design intuitive user journeys that combine polished interfaces, reliable servi

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team With Codex we’re building an AI software engineer. One that you can pair with, delegate to, or even ask to take on future tasks proactively. Our team is a fast-moving group within OpenAI, bringing together research, engineering, design, and product. We iteratively build the Codex agent harness and product to get the most out of the model, and we iteratively train the model to be great at complex software engineering tasks. The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. We operate across research, engineering, product, and infrastructure; owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. Codex Enterprise builds the ecosystem building blocks, discovery surfaces, and enterprise capabilities that help Codex spread across developers, teams, and organizations worldwide. It is a cross-cutting team that works across the stack to build both delightful product experiences and fundamental platform capabilities. Its customers range from individual developers and small teams to large enterprises, and our mission is critical to achieving the vision of Codex as a proactive teammate. About the Role As we grow, we’re focused on turning Codex from a powerful individual tool into a production-grade teammate for entire organizations. You will work across internal OpenAI teams and external customers, from fast-moving startups to large enterprises, to make it possible to deploy, operate, and trust Codex in increasingly demanding real-world environments. As Codex’s consumer adoption accelerates, enterprise demand is growing just as quickly, and there is also increasing opportunity to expand Codex through ecosystem capabilities that unlock new workflows, integrations, and discovery. This team helps turn messy, real-world team requirements into robust, repeatable, and scalable product and

PythonAWSCI/CDRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing

PythonAWSKubernetesRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. A majority of our users interact with our products in languages other than English, and our products must work seamlessly across languages, regions, and cultures. The Internationalization team builds the infrastructure that enables OpenAI products to ship globally by default. We develop the systems that power localization, international product launches, and high-quality global user experiences across all OpenAI products. About the Role As a Senior Software Engineer on the Internationalization team, you will build the systems that power localization and international product launches at OpenAI. You’ll work on the platform that manages product content, translation workflows, and localization infrastructure across our products. This role sits at the intersection of AI systems, developer platforms, and product infrastructure. In this role, you will Build and scale OpenAI’s localization, content, and experimentation platform used across OpenAI product teams, including open-source components: Develop AI-powered translation pipelines combined with human-in-the-loop review workflows. Design systems that reliably deliver localized product content across web and mobile apps. Build tools that enable linguists and localization teams to review and improve translations. Develop developer tooling that simplifies localization and internationalization workflows. Build and maintain internationalization libraries used across OpenAI products: Design systems that correctly handle numbers, currencies, dates, and pluralization across locales. Improve support for multilingual interfaces and right-to-left languages. Partner with product teams to improve the international readiness of new features. You might thrive in this role if you Have strong software engineering experience building backend or full-stack systems. Have familiarity with Java, React, MySQL, and cloud infrastructure p

JavaReactSQLMySQL
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Platform Analytics team builds the systems OpenAI researchers use to understand the quality and behavior of the models we train including what models are doing, why they behave in a particular way, and how that behavior changes across experiments. Neptune is a core part of this work. It ingests, stores, queries, and visualizes large volumes of metrics from pretraining, post-training, and reinforcement learning. Hundreds of researchers depend on these systems in their daily work to compare experiments, debug unexpected behavior, and decide what to try next. Our scope is broader than metrics. We also build platforms that help researchers analyze samples, traces, evaluation results, and other structured or unstructured data through dashboards, APIs, and increasingly agent-driven workflows. These systems need to remain fast, reliable, and understandable as the scale and complexity of research change quickly. We are not trying to become a consulting team that builds a separate solution for every research project. We work directly with researchers to understand recurring problems, then turn them into reusable infrastructure and platform capabilities that many teams can build on. About the Role We’re looking for a hands-on experienced software engineer who can take ownership of a critical system and drive it from problem definition through production adoption. This person should be able to own a platform such as CacheHouse end to end: define its technical direction, design its data model and storage architecture, integrate it with several research dashboards and workflows, guide one or two engineers, and ensure the system works reliably for its users. The right candidate should already bring the technical judgment, ownership, and execution expected at this level. The primary learning curve should be OpenAI’s stack and research problem space, not learning how to lead a complex engineering effort or deliver a production system. You will work directly with

AWSRestAIC++
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The GPT Infrastructure team builds systems that turn advances in model inference and optimization into reliable production capabilities. We enable OpenAI workloads to be qualified and optimized across new accelerator platforms without requiring a one-off port and tuning effort for every hardware target. Our work spans distributed systems, model execution, compilers and runtimes, performance engineering, secure partner integrations, evaluation systems, and developer tooling. We build the infrastructure that makes optimization workflows automated, reproducible, and trustworthy. About the Role We are seeking a software engineer to help build the platform that qualifies and optimizes inference workloads across heterogeneous compute environments. You will develop both OpenAI-hosted services and secure partner-side software for running long-lived optimization workflows. These workflows generate candidate kernels, runtime configurations, and serving-stack changes; compile and execute them on target hardware; verify their correctness; measure their performance; and use the results to guide further optimization. You will work across model architecture, distributed execution, compilers, runtimes, networking, and accelerator systems. A central part of the role is turning research prototypes and one-off hardware bring-up efforts into reliable, reusable infrastructure with clear contracts, reproducible results, strong observability, and well-defined security boundaries. Key Responsibilities Design, build, and operate APIs and control-plane services for long-running workload qualification and optimization campaigns, including scheduling, retries, checkpointing, resource budgets, and observability. Build secure partner-side execution and evaluation software that can compile, run, verify, profile, and benchmark candidate artifacts on accelerator hardware. Integrate model workloads, hardware profiles, compiler toolchains, runtimes, serving engines, and distributed-exe

PythonAWSLinuxRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, then turn that understanding into performance optimizations and models that project performance and capacity needs for future launches. About the Role In this role, you will model inference performance across application, model, and fleet layers with higher fidelity. You will build cost-to-serve estimates from microbenchmarks and create tools that help cross-functional teams reason about latency, capacity, utilization, and cost tradeoffs. In this role, you will Build and refine performance models that translate microbenchmark results into cost-to-serve estimates. Analyze inference workloads end to end across applications, models, and fleet infrastructure. Enhance tooling to identify bottlenecks across layers for latency and throughput. Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference. You might thrive in this role if you: Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency. Are comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling. Have deep expertise with performance profiling, benchmarking, analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve o

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are looking for a systems-minded engineer to help advance our kernel development, performance engineering, and hardware-software co-design capabilities, with a particular focus on AI-assisted workflows and tooling. This person will work at the intersection of kernel optimization, developer tooling, observability, and research infrastructure, helping us improve both how production kernels are built and optimized, and how future hardware-software systems are designed and evaluated. The role is ideal for someone who is excited by low-level performance work, but also sees AI and automation as powerful tools for accelerating engineering velocity. You will help define the future of kernel engineering in the era of AI-assisted development. In this role, you may: Build developer tooling and workflows that make kernel development and performance optimization faster, more scalable, and easier to debug, integrate, and deploy. Develop observability, diagnostics, and validation infrastructure that makes AI-assisted optimization systems more interpretable, reliable, and effective. Optimize production kernels end to end by formulating optimization problems, running search loops, analyzing bottlenecks, debugging generated implementations, and landing improvements into production. Design abstractions, interfaces, and automation systems that accelerate kernel optimization, correctness validation, and hardware-software co-design. Improve AI-assisted optimization systems for sp

AWSRestAIRust
🔔

Get new software reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime