Jobiba hiring network

Reliability Engineer Jobs

2,049 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About the Team The Privacy Engineering team builds secure, reliable systems that help OpenAI meet its legal obligations while protecting user data. We partner closely with Legal and engineering teams across OpenAI to support lawful data access requests and other critical legal workflows. Our work turns complex, high-stakes processes into auditable and dependable technical systems with clear human oversight and strong privacy and security controls. About the Role We’re looking for a full-stack Software Engineer to build the internal tools and data pipelines that power lawful data access request workflows and Legal Operations. You will work across product and data systems to make authorized retrieval and case handling accurate, efficient, and auditable. This role is well suited to someone who enjoys translating ambiguous operational requirements into durable systems, cares deeply about sensitive-data handling, and wants to improve both technical reliability and the day-to-day experience of the people operating these workflows. In this role, you will: Design, build, and operate backend systems and workflow tooling for the full lifecycle of lawful data access requests, from intake and scoping through authorized retrieval, review, preparation, and audit. Build reliable data pipelines and interfaces across products and data stores so authorized teams can locate and handle the right records accurately and reproducibly. Implement least-privilege access, approval gates, provenance, audit trails, data minimization, and safe failure modes for sensitive workflows. Partner with Legal and Legal Operations to translate legal and operational requirements into clear technical designs and intuitive operator experiences. Identify responsible automation opportunities that reduce repetitive work while preserving human review, judgment, and accountability. Own production systems through testing, observability, incident response, documentation, and continuous reliability improvements. Hel

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our Cyber team builds AI systems and products that help trusted defenders understand and respond to cyber threats while improving the safety and reliability of frontier models in security-sensitive settings. The team works across product engineering, model training, evaluations, safeguards, and deployment to make advanced cyber capabilities useful to defenders and responsibly managed. We collaborate closely with Safety/Preparedness, Research, Security, Legal, Communications, GTM, and external partners across OpenAI’s broader cyber work. About the Role We’re looking for research and software engineers to join Codex Cyber. You’ll help define and ship security products, work with trusted defenders and customers, shape model training and access patterns, and build research and evaluation systems for assessing cyber capabilities, validating safeguards, and improving training data. This role is hands-on and cross-functional, connecting product launches, model development, safety work, and real-world security use cases. In this role, you will: Help define and execute the technical roadmap for Codex Cyber’s security products, including evaluations, safeguards, trusted-defender workflows, and deployment decisions. Work with trusted defenders, customers, and partner teams to understand cyber use cases, evaluate risk, and turn feedback into product and research priorities. Shape cyber-specific model training and access patterns, including data, evaluations, validation, and deployment criteria. Build and validate systems for measuring cyber capabilities, monitoring misuse risk, and proving safeguards work in practice. Collaborate with Safety/Preparedness, Research, Security, Legal, Communications, Go-to-Market, and external partners on company-wide cyber priorities. Translate frontier cyber research into launch-ready tools, operational playbooks, and durable infrastructure for Codex and security products. You might thrive in this role if you: Enjoy 0 -> 1 envi

javascripttypescriptpython
View job →

About the Team The Privacy Engineering team builds secure, reliable systems that help OpenAI meet its legal obligations while protecting user data. We partner closely with Legal and Engineering teams across OpenAI to support lawful data access requests and other critical legal workflows. Our work turns complex, high-stakes processes into auditable and dependable technical systems with clear human oversight and strong privacy and security controls. About the Role We’re looking for a full-stack Software Engineer to build the internal tools and data pipelines that power lawful data access request workflows and Legal Operations. You will work across product and data systems to make authorized retrieval and case handling accurate, efficient, and auditable. This role is well suited to someone who enjoys translating ambiguous operational requirements into durable systems, cares deeply about sensitive-data handling, and wants to improve both technical reliability and the day-to-day experience of the people operating these workflows. This role is based in San Francisco, CA, with two additional locations under consideration: London, UK, and Dublin, Ireland. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and operate backend systems and workflow tooling for the full lifecycle of lawful data access requests, from intake and scoping through authorized retrieval, review, preparation, and audit. Build reliable data pipelines and interfaces across products and data stores so authorized teams can locate and handle the right records accurately and reproducibly. Implement least-privilege access, approval gates, provenance, audit trails, data minimization, and safe failure modes for sensitive workflows. Partner with Legal and Legal Operations to translate legal and operational requirements into clear technical designs and intuitive operator experiences. Identify responsible automation o

awsrestai
View job →

About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man

pythonsqlaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,

awskubernetesci/cd
View job →
O
1mo ago

About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p

pythonawsci/cd
View job →
O
1mo ago

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

pythonawslinux
View job →
O
1mo ago

About the Team Our London-based team builds the backend systems that help ChatGPT scale reliably. We work on infrastructure close to the product, partnering with engineering teams to improve the performance, resilience, and operability of critical user-facing systems. Our work combines backend software engineering with distributed systems and production reliability. We build shared capabilities, improve high-traffic workflows, and make it easier to introduce new product functionality without compromising performance or availability. About the Role This role is for software engineers who want to build and evolve backend systems operating at significant scale. You’ll write production code, design shared infrastructure, and solve technical challenges involving performance, distributed systems, and system reliability. You’ll also own how those systems behave in production: how changes are rolled out, how issues are detected and diagnosed, and how recurring operational problems can be addressed through better software and system design. This is a strong fit for backend engineers who enjoy complex systems problems and want a direct connection between the infrastructure they build and the experience of ChatGPT users. In this role, you will: Design, build, and maintain backend systems supporting high-traffic ChatGPT experiences. Develop shared services, APIs, and infrastructure that help product teams build and launch new capabilities safely. Improve the performance, scalability, and efficiency of production systems as usage and product complexity grow. Build and improve systems for asynchronous processing and other large-scale backend workloads. Lead architectural improvements and infrastructure migrations while maintaining correctness, compatibility, and safe rollout and rollback. Strengthen monitoring, alerting, and diagnostics to detect problems early and reduce customer impact. Participate in on-call, incident response, and root-cause analysis, and turn operational lea

awsrestai
View job →
O
1mo ago

About the Team The Applications Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and

awskubernetesrest
View job →
O
1mo ago

About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati

pythonsqlaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet Hardware team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automation to reduce manual work

pythonsqlaws
View job →

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa

awsrestai
View job →
O
1mo ago

About the Team We’re hiring software engineers to make the Workload team more productive. The Workload team maintains the core components of OpenAI’s training and inference frameworks and helps execute frontier experiments. About the Role We’re looking for someone who cares about the developer experience of working in and around OpenAI’s core training and inference frameworks. In this role you will: Be responsible for optimizing the development workflows of the engineers around you Work within various Workload teams to address their specific needs, but collaborate with the centralized teams that own various aspects of development experience Optimize iteration speed, both broadly, and in particular by optimizing specific teams’ CI Improve reliability, for instance, by driving testing strategy for particular components Work through the long tail of things that it takes to build libraries and systems that will delight researchers You might thrive in this role if: You are motivated by helping people. You believe a thing that separates great teams from good teams are the players willing to do whatever work it takes, without ego. You believe in the power of developer experience. Something magical happens when people can quickly and confidently iterate on a simple codebase, but this magic is fragile and must be fought for. When you see someone trip over something, no matter how small, your first instinct is asking yourself what it would take for that to not happen again. Your second instinct is clicking merge on the PR you’ve already written to make it so. You are pragmatic. You have the ability to see the world through a perfectionist’s eyes, but are not yourself a perfectionist. You know which problems to pick and when to switch to making progress on a different problem. You like going end-to-end on things. You love co-design — that feeling when you were only able to find the right solution because you both deeply understand the users that interact with a system and the

pythonawsrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role OpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency of our systems. In this role, you'll apply deep technical expertise to optimize infrastructure and application-level performance across mission-critical products like ChatGPT and our developer API. You’ll work cross-functionally with teams building core services, training models, and developing real-time user experiences to push our latency, throughput, and cost-efficiency to the next level. We are looking for engineers who thrive in ambiguous environments, value deep systems understanding, and are motivated by delivering measurable impact. This is a highly technical, individual contributor role focused on root-cause analysis, profiling, instrumentation, and architecture-level performance improvements across our stack. In this role, you will: Analyze and optimize performance across application, middleware, runtime, and infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond. Develop tooling and metrics that provide deep observability into system performance. Collaborate closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements. Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale. Lead investigations into high-impact performance regressions or scalability issues in production. Drive performance testing strategies and help define SLAs/SLOs around latency and throughput for critical systems. You might thrive in this role if you: Have 7+ years of experience in software engineering with a strong tr

pythonawsrest
View job →
O
1mo ago

About the team Online Data builds and operates Habitat, the single product surface of Online Data and the system of record for OpenAI’s online user data. As OpenAI’s scale and product requirements evolve, Habitat is becoming a full-stack, one-size-fits-most database platform with end-to-end ownership of: Provisioning and developer experience APIs and guardrails Scaling, performance, and reliability Data movement, caching, routing, and placement Privacy enforcement and access control Change Data Capture (CDC) as a first-class primitive The foundation for future storage backends You’ll work on the core online database platform behind OpenAI’s products, building and operating Habitat services that handle high-QPS, latency-sensitive workloads across regions. You’ll partner closely with internal platform and product teams to ship safe, reliable systems, then push them to be faster and more cost-efficient through better caching, routing, observability, and operational tooling. This is a critical role for engineers who like owning hard distributed-systems problems end to end and sweating the details from p99 latency to production operations at massive scale. In this role, you will Design and build core abstractions spanning storage, caching, routing, CDC, and privacy enforcement Own a major surface area end to end, from product and API design to operational excellence Improve latency, correctness, and cost efficiency for real production workloads at massive scale Build strong instrumentation, debugging workflows, and developer-first tooling Collaborate closely with internal product and infrastructure teams to understand requirements and ship pragmatic solutions Participate in an on-call rotation and raise the bar on reliability while aggressively improving performance and usability You might thrive in this role if you have A strong track record building and operating high-scale backend or data-intensive distributed systems in production Excellent systems judgment and the a

pythonawsrest
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.