About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Security Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments—GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Security Software Engineer, you will design and build critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Architect and implement production-grade security services (e.g., auth services, access brokers, secure proxies, key-management infrastructure) that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD. Partner with infrastructure and research engineers to embed security into high-performance compute clusters, enabling rapid model training and deployment without compromising protection. Develop automation and detection tooling to continuously identif
Jobiba hiring network
Software Reliability Engineer Jobs
6,326 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team At OpenAI, we’re building safe and beneficial artificial general intelligence. We deploy our models through ChatGPT, our APIs, and other cutting-edge products. Behind the scenes, making these systems fast, reliable, and cost-efficient requires world-class infrastructure. The Caching Infrastructure team is responsible for building a caching layer that powers many critical use cases at OpenAI. We aim to provide a high-availability, multi-tenant cache platform that scales automatically with workload, minimizes tail latency, and supports a diverse range of use cases. We’re looking for an experienced engineer to help design and scale this critical infrastructure. The ideal candidate has deep experience in distributed caching systems (e.g., Redis, Memcached), networking fundamentals, and Kubernetes-based service orchestration. In This Role, You Will: Design, build, and operate OpenAI’s multi-tenant caching platform used across inference, identity, quota, and product experiences. Define the long-term vision and roadmap for caching as a core infra capability, balancing performance, durability, and cost. Collaborate with other infra teams (e.g., networking, observability, databases) and product teams to ensure our caching platform meets their needs. You Might Thrive In This Role If You: Have 5+ years of experience building and scaling distributed systems, with a strong focus on caching, load balancing, or storage systems. Have deep expertise with Redis, Memcached, or similar solutions, including clustering, durability configurations, client-side connection patterns, and performance tuning. Have production experience with Kubernetes, service meshes (e.g., Envoy), and autoscaling systems. Think rigorously about latency, reliability, throughput, and cost in designing platform capabilities. Thrive in a fast-paced environment and enjoy balancing pragmatic engineering with long-term technical excellence. About OpenAI OpenAI is an AI research and deployment company d
About the Team The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. This role is based in San Francisco, CA. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change This role is exclusively based in our San Francisco HQ. We offer relocation assistance to new employees. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely depl
About the Team The Payments Team is working to make AGI financially sustainable by enabling OpenAI to accept payments across our existing products (ChatGPT, ChatGPT for Teams, API) and supporting future commercial experiences. We own the user-facing payments experience, direct integrations with payment service providers, and other supporting infrastructure that powers checkout and payment flows. Payments works closely with teams across Finance, Legal, Growth, Product, external vendors and more to ensure our payments infrastructure scales with the business and supports evolving commercial goals. About the Role As a Software Engineer on the Payments Team, you will play a pivotal role in shaping the infrastructure that powers OpenAI’s financial systems. You’ll design and build core payments systems, partner with cross-functional teams to deliver scalable solutions, and ensure that our financial systems support rapid product iteration and global reach. We’re looking for people who are pragmatic, collaborative, and energized by the opportunity to build systems that are foundational to OpenAI’s long-term mission and sustainability. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Architect, implement and scale core payment infrastructure. Build robust and extensible financial primitives to enable rapid iteration and support hyper-growth. Collaborate closely with internal and external stakeholders (e.g. Finance, GTM, Product, etc) to translate business requirements into clean, scalable technical systems. Demonstrate high levels of ownership to get changes shipped in the highly-regulated domain of payments. Create a seamless user experience for checkout around the world. Improve the reliability, observability and auditability of money movement at OpenAI. You might thrive in this role if you: Have 5+ years of professional experience in software engine
About the Team The Applied team at OpenAI safely brings cutting-edge technology to the world. We have released widely used products such as ChatGPT, Sora, and the OpenAI API, powering models including GPT-5 and a growing set of multimodal capabilities across text, image, audio, and video. Our team also manages large-scale inference and platform infrastructure that supports these experiences at global scale. With much more on the horizon, our impact continues to grow. Our customers build fast-growing businesses using our APIs, unlocking product capabilities that were previously unimaginable. ChatGPT and Sora exemplify the breadth of what’s now possible across text, image, audio, and video experiences. As these capabilities expand, we prioritize the responsible use of our technology, emphasizing safe and thoughtful deployment over unchecked growth. Within Applied Engineering, the Ads Monetization team in Financial Engineering builds the core systems dealing with all the money flows for ChatGPT Ads. These systems are a combination of low-latency, high scale, high reliability, while being built in a financially correct, accurate, auditable and explainable way. This role sits at the intersection of ads delivery, data engineering, and financial systems. In this role, you will: Architect and build the core monetization systems for ChatGPT Ads. Build and operate the core services and pipelines that power ads monetization end-to-end, from event capture and validation through aggregation, pricing, metering, and ultimately producing billable outputs. Define and implement the source of truth for ads monetization data, including schemas, data models, and invariants that ensure outputs are consistent, explainable, and auditable. Own correctness and reconciliation: align production outputs with downstream invoicing/finance requirements, build controls/monitors, and close gaps through investigations and backfills. Develop across the stack to create comprehensive billing integration
About the Team The Monetization team is a new cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products—including next-generation ads experiences—that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to help build the core infrastructure behind OpenAI’s monetization and ads systems. In this foundational role, you’ll architect and implement distributed systems that power OpenAI’s monetization stack—focusing on reliability, performance, privacy, and large-scale operation. You’ll work across backend, systems, and platform layers to define and implement 0→1 infrastructure, partnering closely with Product, Design, and Research to shape the future of monetized AI experiences. Your work will enable both internal and external teams to build on safe, scalable, and robust monetization primitives. This role is exclusively based across our San Francisco & Seattles sites. We offer relocation assistance to new employees. In this role, you will: Design and build the foundational backend and infrastructure powering OpenAI’s monetization and ads systems Architect large-scale distributed systems that
About the Team The Platform Systems team at OpenAI operates at the intersection of cutting-edge AI and large-scale distributed systems. We build the engineering and research infrastructure required to train OpenAI’s flagship models on some of the world’s largest, custom-built supercomputers. Our team develops core model training software and works deep in the stack - spanning collective communication, compute efficiency, parallelism strategies, fault tolerance, failure detection, and observability. The systems we build are foundational to OpenAI’s research velocity, enabling reliable, efficient training at frontier scale. We collaborate closely with researchers across the organization, continuously incorporating learnings from across OpenAI into the evolution of our training platform. About the Role As a Software Engineer, Platform Systems, you will design and build distributed systems that provide visibility into large-scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical to operating OpenAI’s training stack and is actively evolving to support new use cases and increasingly complex workloads. This role sits at the core of our training infrastructure, blending systems engineering, performance analysis, and large-scale debugging. In This Role, You Will Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high-throughput distributed systems Collaborate with systems, infrastructure, and research teams to evolve platform
About the Team The Platform Systems team at OpenAI operates at the intersection of cutting-edge AI and large-scale distributed systems. We build the engineering and research infrastructure required to train OpenAI’s flagship models on some of the world’s largest, custom-built supercomputers. Our team develops core model training software and works deep in the stack - spanning collective communication, compute efficiency, parallelism strategies, fault tolerance, failure detection, and observability. The systems we build are foundational to OpenAI’s research velocity, enabling reliable, efficient training at frontier scale. We collaborate closely with researchers across the organization, continuously incorporating learnings from across OpenAI into the evolution of our training platform. About the Role As a Software Engineer, Platform Systems, you will design and build distributed systems that provide visibility into large-scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical to operating OpenAI’s training stack and is actively evolving to support new use cases and increasingly complex workloads. This role sits at the core of our training infrastructure, blending systems engineering, performance analysis, and large-scale debugging. In This Role, You Will Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high-throughput distributed systems Collaborate with systems, infrastructure, and research teams to evolve platform
Location: San Francisco, CA (Hybrid: 4 days onsite/week). Relocation assistance available. About the Team: We build foundational platform software that enables reliable, secure, and performant products. The team works across system layers and partners closely with adjacent engineering groups to deliver robust capabilities from concept through launch. About the Role: We’re seeking a System Software Engineer to design, implement, and debug core platform components and the pipelines that build and update system images. You’ll work across operating system layers, focusing on performance, security, and deep system debugging to ship production‑grade systems. In this role, you will: Design, implement, and debug system‑level components and services across kernel and user space. Configure and maintain OS platform services (init, services, networking, security policies) and related tooling. Build and operate image and update pipelines, ensuring reliability, reproducibility, and rollback safety. Instrument and analyze performance using profiling and tracing; optimize CPU, memory, I/O, and power usage. Own platform observability and reliability: logging, crash capture, watchdogs, and diagnostics. Collaborate with cross‑functional teams to define interfaces and deliver end‑to‑end features. Establish strong engineering practices: code review, CI, reproducible builds, and release management. Partner with external suppliers to support builds and deployments. You might thrive in this role if you: Have shipped production systems software on modern operating systems. Are proficient in C/C++ and a scripting language, and comfortable with OS internals (concurrency, memory management, filesystems, networking, power management). Bring strong systems debugging skills using debuggers, tracers, profilers, and logs across kernel/user‑space boundaries. Understand configuration of platform services and interfaces, and can translate requirements into stable, well‑documented APIs. Are fluent in u
About the Team The Statsig team within OpenAI owns the experimentation, rollout, dynamic configuration, and analytics infrastructure that sits on the launch path for OpenAI products. Our systems help teams ship safely, evaluate product and model changes in production, and make high-confidence decisions from real-world usage. Statsig began as an independent company built around experimentation, feature management, and product analytics at scale. After Statsig joined OpenAI, the team began the next chapter: bringing that platform expertise and infrastructure into OpenAI as the experimentation and rollout foundation for every product we ship. This is infrastructure with a very direct product consequence. Teams working on ChatGPT, Codex, model measurement, consumer experiences including ads, business subscriptions, developer products, and shared platform systems depend on Statsig to evaluate configurations, move traffic safely, ingest experiment data, serve analytics, and roll changes forward or back when production reality demands it. We are at a critical point in the platform journey. Adoption is accelerating quickly across OpenAI, and the systems that were already important are becoming load-bearing for how the company launches. The infrastructure needs to stay fast under sharply increasing evaluation volume, reliable when more services depend on it, observable enough to debug quickly, and efficient enough to support OpenAI-wide scale. Recent SDK and server-side infrastructure work has already produced measurable wins in latency, reliability, memory usage, and compute efficiency across important services. The next phase is to make those gains systematic: a platform that can absorb rapidly growing product velocity while preserving low latency, data quality, operational safety, and developer trust. Based out of OpenAI's Bellevue office, we are a close-knit team that values in-person collaboration, technical depth, operational ownership, and building infrastructure that
About the Team We’re hiring software engineers to make OpenAI’s Model Performance teams more productive. These teams work on the systems, tooling, and infrastructure that help improve model performance across OpenAI’s training and inference workloads at frontier scale. About the Role We’re looking for an autonomous, high-ownership developer productivity engineer who cares deeply about helping other engineers move faster, safer, and with more confidence. This role will sit within OpenAI’s Model Performance organization, contributing to developer infrastructure, CI systems, testing workflows, tooling, and broader performance infrastructure efforts. There is also a strong opportunity to contribute to the Triton project and help improve the systems that support performance-critical engineering work across OpenAI. In this role you will: Improve development workflows for engineers working on model performance infrastructure Design and improve CI/CD, release, validation, and testing pipelines Build and maintain tools that improve reliability, iteration speed, and engineering confidence Partner closely with engineers to identify friction in testing, debugging, deployment, and development workflows Contribute to infrastructure efforts that support performance-critical training and inference systems Help improve developer experience across Python-heavy codebases and performance-oriented infrastructure Work in a high-context, ambiguous environment where ownership and good judgment matter You might thrive in this role if: You are motivated by enabling the people around you and helping engineers do their best work You have strong experience with CI/CD, developer infrastructure, testing systems, tooling, or build/release workflows You are highly collaborative, empathetic, and comfortable partnering deeply with technical teams You are strong in Python and enjoy building reliable, scalable developer tools and infrastructure You have experience improving large-scale engineering work
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches
About the Team We’re hiring software engineers to make OpenAI’s networking teams more productive. These teams build and operate the high-performance networking systems that support OpenAI’s training and inference infrastructure at frontier scale. About the Role We’re looking for someone who cares deeply about the developer experience of engineers working on complex infrastructure systems — especially around build systems, test architecture, release pipelines, and reliable development workflows. This role will be embedded with OpenAI’s networking team: making it faster, safer, and easier for engineers to build, test, validate, and ship changes across multi-server, networked, and hardware-adjacent environments. In this role you will: Improve development workflows for engineers building and operating OpenAI’s networking systems Design and improve continuous deployment, release, and validation pipelines Build and maintain test harnesses for multi-server, networked, and hardware-backed environments Improve iteration speed across C++, Python, and build-system-heavy codebases Partner with engineers to identify friction in CI, testing, debugging, and deployment workflows Drive testing and reliability strategy for infrastructure components that support large-scale training and inference workloads Work closely with centralized developer experience teams while staying deeply embedded with the networking engineers closest to the systems You might thrive in this role if: You are motivated by helping other engineers move faster and with more confidence You have experience with CI/CD, release pipelines, testing infrastructure, or build systems You are comfortable moving between C++, Python, and build systems such as CMake, Bazel, or Blaze You enjoy building test harnesses, automation, and workflow improvements for complex systems You do not need to be a networking expert, but you are excited to learn enough about the domain to make the team meaningfully more effective When you see
About the Team The Foundations Research team works on high-risk, high-reward ideas that could shape the next decade of AI. Our goal is to advance the science and data that enable our training and scaling efforts, with a particular focus on future frontier models. Pushing the boundaries of data, scaling laws, optimization techniques, model architectures, and efficiency improvements to propel our science. The Search team sits within Foundations, building agentic search by co-designing model–system interfaces with the core search stack (serving, indexing, retrieval) to translate model intent into reliable, real-world actions. Operating at the frontier of AI and information retrieval, the team develops large-scale systems that transform and index vast corpora, enabling models to reason over global knowledge and act dependably. In close partnership with researchers, we rapidly bring modeling breakthroughs into production and redefine how intelligent systems discover, retrieve, and synthesize information at planetary scale. About the Role We’re looking for a Software Engineer focused on building and scaling retrieval systems. You’ll work with a team of researchers and engineers to develop infrastructure that enables models to retrieve and act on the right information at the right time. This includes designing and operating indexing systems, retrieval pipelines, and serving layers. This work supports retrieval across OpenAI products and research, with direct impact on system performance, reliability, and scale. Responsibilities Build and scale retrieval infrastructure across indexing, serving, and query execution. Develop low-latency, high-throughput systems for real-time model interaction. Partner with research to productionize embedding and retrieval techniques. Support dense, sparse, and hybrid retrieval pipelines. Own system performance, reliability, and observability at scale. Collaborate across Pretraining, Inference, and Product teams to integrate retrieval end-to-e
About the Team The Statsig team within OpenAI builds the experimentation, feature rollout, dynamic configuration, and analytics systems that help OpenAI ship products with speed, safety, and evidence. Our work sits on the critical path for how product, engineering, research, and go-to-market teams learn from real-world usage and make high-confidence decisions. Statsig began as an independent company focused on helping builders move faster through trustworthy experimentation and feature management. After Statsig joined OpenAI, the team began the next chapter: bringing that deep product expertise, customer intuition, and mature platform infrastructure into OpenAI as the experimentation and rollout platform for every product we ship. Today, we support teams across ChatGPT, Codex, model measurement, consumer experiences including ads, business subscriptions, developer products, and the shared infrastructure that connects them. These teams rely on Statsig to safely introduce new capabilities, compare product and model behavior, measure impact, and roll changes forward or back with confidence. We are at a defining moment in the platform journey. OpenAI has the data, product surface area, and pace of innovation to learn faster than almost any organization in the world, but that potential only becomes real if teams can experiment responsibly, measure clearly, and roll out changes safely. Adoption of the platform is accelerating rapidly across the company, and recent SDK and server-side infrastructure work has already produced measurable wins in latency, reliability, memory usage, and compute efficiency for important services. Based out of OpenAI’s Bellevue office, we are a close-knit team that values in-person collaboration, urgency, craft, and impact. We build for other builders, and the best version of this team is one where every OpenAI product team can move faster because the experimentation and rollout layer is dependable, fast, and easy to use. About the Role We are l
Get new software reliability engineer jobs by email
Daily job updates · Unsubscribe anytime