Jobs in United States

Software Reliability Engineer in United States

2,007 active opportunities · Updated October 2026

Explore current software reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $295.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer on the Sharing team, you will lead large, multi-team initiatives with long-term technical vision and group-level impact. You will be expected to define and drive the platform-wide strategy powering how millions of users capture, share, and discover content on Roblox. In this role, you will set engineering standards, mentor senior engineers, and serve as a key architect of our long-term technical direction. You Will Drive Strategy & Execution: Own the outcome of complex, business-critical programs spanning several teams, often lasting years. Innovate at Scale: Develop and drive a multi-year technical vision for content creation and sharing, anticipating scale, technology, and business evolution. Elevate Reliability: Lead high-severity incident response across groups; drive durable systemic solutions that improve reliability and velocity. Architect Foundations: Regularly improve shared infrastructure and foundational systems, introducing frameworks that uplift development speed across the organization. Align Teams: Aligns multiple teams on shared technical direction, producing detailed design docs, phased roadmaps, and planning models that balance short an

JavaAWSGitAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $345K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. About the Team The CI and CD Foundation team is the backbone of our engineering velocity. Our mission is to engineer and deliver a CI/CD ecosystem that streamlines the integration and delivery of our core client code—including the Game Engine, Studio, and User Apps. We are laser-focused on three goals: enhancing overall build reliability, maximizing developer productivity, and ensuring consistent software quality. We’re building and improving pipelines and tools to transform SDLC in the age of AI-driven software development. The Role As a Principal Software Engineer on the CI/CD Foundation team, you will serve as the chief technical strategist and architect responsible for transforming how Roblox builds, tests, and ships code at massive scale. Our build and delivery ecosystem handles immense complexity: decades of C++ game engine code, extensive Lua applications, complex native platforms, and a sprawling monorepo with massive, growing test suites. You will lead the technical vision to modernize, decouple, and accelerate this ecosystem—driving radical improvements across our Local Dev Inner Loop (pre-commit/pre-push), CI (PR validation, speculative execution, merge queues), and CD (OTA

PythonAWSDockerKubernetes
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $345K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer on the Compute team, you will be the technical anchor for Roblox's GPU and AI accelerator capabilities. This is a battle-tested GPU expert role focused on the machine management layer and above: how GPU hosts are made production-ready, kept healthy, and turned into reliable compute for the workloads that depend on them. You will own the hard problems that show up only at scale, from driver and firmware management to GPU health, reliability, and performance across a rapidly growing fleet of accelerators spanning Roblox data centers and cloud environments. You will set the technical direction for GPU compute and up-level the entire organization's GPU expertise. You will: Serve as the GPU technical leader for the Compute team, partnering across Kubernetes, Machine Bootstrap, Networking, and Cloud to drive GPU strategy end to end. Own the GPU host lifecycle above raw fleet management: driver, firmware, and CUDA stack management, GPU health and telemetry, and remediation of GPU-specific failures (XID errors, ECC, thermal, NVLink and fabric faults). Architect how GPU capacity is exposed to compute platforms, including scheduling, isolation, and integration with Ku

AWSKubernetesGitAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Orchestration pod within Engine Productivity, you'll design and run the platform that executes large-scale end-to-end and integration tests, running the real, shipping client on real hardware, across Roblox's data centers, cloud, and our own device labs, so our engineering teams can ship the engine, clients, Studio, and more with speed and confidence. Every Roblox engine, client, and Studio change, along with the experiences built on top of them, should ship with confidence, and the Orchestration team is the layer that makes that possible. We build large-scale distributed services that turn thousands of test suites into a reliable, push-button pipeline: fanning work out across fleets of machines and real devices, moving artifacts to where they're needed, managing single- and multi-client test state, and giving test owners and maintainers a system to validate their own runs. It looks a lot like building a specialized cloud platform, with capacity-aware scheduling, isolation and sandboxing, and smart retry and backoff, plus the classic distributed systems problems (fairness, efficiency, failure handling, and reliability) at Roblox scale. You Will: Design a

PythonAWSGitAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $196.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Consumer Apps Consoles team owns the console app foundation that makes Roblox feel fast, fluid, and reliable for millions of players and is dedicated to delivering superior performance, reliability, and user experience on PlayStation and Xbox. As a Senior Software Engineer on the Consoles team, you'll be a driving technical force — leading complex client work end-to-end, raising the quality bar across our codebase, and partnering closely with product, design, and platform teams to ship experiences that are polished and performant at scale. You Will: Shape & improve the PlayStation and Xbox experience that millions of players see every day Translate ambiguous product ideas into well-scoped, high-quality engineering plans in close partnership with product, design, and data science. Raise the engineering bar through rigorous code review, proactive mentorship of junior engineers, and the standards you set in your own code. Drive platform-level investments that benefit the broader engineering organization. Work with our Xbox and PlayStation counterparts to ensure feature parity and platform-specific excellence. You Have: 5+ years of experience building and shipping client software, with

ReactAWSGitAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on our Release Engineering team, you will build the systems that hundreds of engineers across Roblox use every day to safely ship code to tens of millions of concurrent players. Your work will empower teams to ship bold, high-impact changes quickly and confidently on every device Roblox runs on. If you enjoy building developer-facing infrastructure where reliability and blast radius directly impact end-users, you will be right at home on our growing team. You Will: Design and develop backend services and automation that power our release and experimentation process across desktop, mobile, console, VR, and servers Work in C++ engine code to improve telemetry reporting, enhance feature rollout and automatic abort capabilities, and extend release functionality Build progressive client and server rollout, regression detection, and automated rollback systems that keep our weekly multi-platform releases safe at scale Work directly with engineering customers to turn pain points into durable, flexible, and safe-by-default tooling You Have: 5+ years building backend services with C#, Python, TypeScript, or similar Familiar with and comfortable working with C++ Familiar

TypeScriptPythonAWSGit
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's data infrastructure processes petabytes of data daily, powering analytics, ML, and product decisions. As a Senior Software Engineer in our Data Infra org, you will design, build, and scale the distributed data infrastructure platforms that power Roblox. You will own and drive the next-generation architecture of our core platforms, which span Kafka, Flink, Spark, Trino, Druid, Airflow and Data Catalog. This role combines high ambiguity and ownership to push the boundaries of what our infrastructure can handle at massive scale, giving you the unique opportunity to steer the evolution of the data landscape. You Will: Own and Scale Core Platform Components: Take responsibility for the design, architecture, and implementation of 1–2 key data platform frameworks within our stack Collaborate and Align: Partner with infra, data science, and product engineering teams to ensure your target platform's capabilities are directly guided by platform governance and product requirements. Optimize Performance at Scale: Dive deep into engine internals, query planning, state management, memory optimization, serialization efficiency to maximize throughput and reliability under heavy load. Drive Infrast

JavaAWSGCPKubernetes
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $192K/yr

Quick readStrong listing-quality and freshness signals

Senior Software Engineer - Streaming Platform Client Data streams are mission-critical at Datadog, powering near real-time communication across the vast majority of our services. Our Streaming Platform group builds the core infrastructure and abstractions that ensure Datadog remains a trusted partner for engineers worldwide. See our blog post . The Streaming Platform Client team sits at the heart of this ecosystem. We own the Rust client library (producers and consumers) with language bindings for Java, Go, and Python. We focus on building intuitive APIs and abstractions that make a powerful distributed system easy to adopt and operate for the hundreds of internal users of our library. Our library runs on critical data paths that handle hundreds of millions of messages per second making performance, observability, and reliability paramount. We also develop and operate the service that bridges the clients fleet with the platform's control plane, handling complex balancing, scaling, and static stability challenges. We are seeking a Senior Software Engineer to help us evolve these features. You will collaborate directly with our users, tackle performance-critical code, and solve complex distributed systems challenges across the control plane, client libraries, and data plane. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Work within a distributed, high-impact team spanning Europe and the US, building critical technologies that power data pipelines for dozens of internal teams and hundreds of services. Architect and implement resilient interactions between our client libraries and the control plane. Optimize our high-throughput, low-level streaming library to push the boundaries of performance and efficiency. Champion the developer experience by providing

PythonJavaGitAI
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $192K/yr

Quick readStrong listing-quality and freshness signals

Senior Software Engineer - Streaming Platform Client Data streams are mission-critical at Datadog, powering near real-time communication across the vast majority of our services. Our Streaming Platform group builds the core infrastructure and abstractions that ensure Datadog remains a trusted partner for engineers worldwide. See our blog post . The Streaming Platform Client team sits at the heart of this ecosystem. We own the Rust client library (producers and consumers) with language bindings for Java, Go, and Python. We focus on building intuitive APIs and abstractions that make a powerful distributed system easy to adopt and operate for the hundreds of internal users of our library. Our library runs on critical data paths that handle hundreds of millions of messages per second making performance, observability, and reliability paramount. We also develop and operate the service that bridges the clients fleet with the platform's control plane, handling complex balancing, scaling, and static stability challenges. We are seeking a Senior Software Engineer to help us evolve these features. You will collaborate directly with our users, tackle performance-critical code, and solve complex distributed systems challenges across the control plane, client libraries, and data plane. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Work within a distributed, high-impact team spanning Europe and the US, building critical technologies that power data pipelines for dozens of internal teams and hundreds of services. Architect and implement resilient interactions between our client libraries and the control plane. Optimize our high-throughput, low-level streaming library to push the boundaries of performance and efficiency. Champion the developer experience by pro

PythonJavaGitAI
A
📍 United States· Full-time
✓ High-confidence listingCompany trend -98.8%

From $212K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Passport & Commerce builds the trusted account and commerce foundation that lets anyone, anywhere, become part of the Airbnb community — from their first sign-up, through the profile and connections they build, to every booking and business transaction along the way. We're a newly formed org within Guest & Host that brings together identity, account, and commerce foundations under one roof. Our mission is to move Airbnb beyond the transaction — building a world where accounts create trust, value sticks, and what our community earns travels with them, and their businesses, wherever they go. You'll work closely with Payments and Wallet engineering, Identity & Privacy, Profile & Community, and Guest & Host product, design, and data science partners as we stand up this team's roadmap and technical foundations. The Difference You Will Make: As a Staff Software Engineer on the Passport team, you will be a key architect behind the next generation of our account platform, directly influencing how millions of guests and hosts experience Airbnb. You'll set technical direction for how account state, eligibility, and entitlements are computed, stored, and served at scale across every surface where a guest or host interacts with Airbnb. Success in this role looks like a Passport platform that is reliable and extensible enough to support new programs and partner integrations without re-architecture — measured through service reliability (uptime, latency, correctness of entitlement calculations), the speed at which new offerings can launch, and adoption of your platform by othe

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Privacy Engineering team builds secure, reliable systems that help OpenAI meet its legal obligations while protecting user data. We partner closely with Legal and Engineering teams across OpenAI to support lawful data access requests and other critical legal workflows. Our work turns complex, high-stakes processes into auditable and dependable technical systems with clear human oversight and strong privacy and security controls. About the Role We’re looking for a full-stack Software Engineer to build the internal tools and data pipelines that power lawful data access request workflows and Legal Operations. You will work across product and data systems to make authorized retrieval and case handling accurate, efficient, and auditable. This role is well suited to someone who enjoys translating ambiguous operational requirements into durable systems, cares deeply about sensitive-data handling, and wants to improve both technical reliability and the day-to-day experience of the people operating these workflows. This role is based in San Francisco, CA, with two additional locations under consideration: London, UK, and Dublin, Ireland. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and operate backend systems and workflow tooling for the full lifecycle of lawful data access requests, from intake and scoping through authorized retrieval, review, preparation, and audit. Build reliable data pipelines and interfaces across products and data stores so authorized teams can locate and handle the right records accurately and reproducibly. Implement least-privilege access, approval gates, provenance, audit trails, data minimization, and safe failure modes for sensitive workflows. Partner with Legal and Legal Operations to translate legal and operational requirements into clear technical designs and intuitive operator experiences. Identify responsible automation o

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man

PythonSQLAWSGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Applications Engineering organization builds and operates the products (such as ChatGPT & Codex) that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role We’re hiring Backend Software Engineers to design and implement safe services and infrastructure that power our core products. What You’ll Do Architect, build, and improve scalable backend systems and APIs. Drive performance, reliability, and safety across distributed services. Implement data storage, retrieval, compute, and integration solutions. Participate in long-term architectural planning and technical design reviews. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. You Might Thrive Here If You: Have strong experience with distributed systems, APIs, and backend languages (e.g., Go, Python, Rust, C++). Have experience setting up and maintaining production backend services and data pipelines. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Enjoy building resilient services that handle large scale and complexity. Are self-directed and enjoy figuring out the best way to solve a particular problem Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,

AWSKubernetesCI/CDLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p

PythonAWSCI/CDRest
🔔

Get new software reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime