Jobs in United States

Reliability Engineer in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

M
📍 United States· Full-time
✓ High-confidence listingCompany trend -93.7%

From $137K/yr

Quick readStrong listing-quality and freshness signals

The Code Gen team is tasked with building AI-powered code transformation tools that transform rigid, legacy applications that suffer from poor scalability and high operating costs into modern, microservices-based architectures that are built on top of MongoDB. Join our team and be at the forefront of innovation and creativity. We are looking for a Staff Engineer with domain expertise and years of experience in modernizing legacy applications that are based on traditional database systems. A significant advantage is profound prior experience in leveraging AI, particularly LLMs and GenAI capabilities, to enable reliable, self-driving automation of the code transformation, iterative build, and test processes. In this role, you will be instrumental in initiating technical strategies and ideas, lead the Code Gen team in designing, building, and optimizing our code transformation workflow and tools. You will work on critical components that ensure the scalability, efficiency, and reliability of our services. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation, build and test. This role will be based remotely in North America. A strong candidate for this position will have Extensive experience (8+ years) in software development and operations, with a proven track record of delivering high performance, correctness, and architectural excellence in fast-paced environments Experience using Relational Databases such as Oracle, MySQL, Microsoft SQL Server or PostgreSQL Experience with tools and methodologies for code analysis, refactoring, and automated testing Experience in designing and implementing complex software systems, collaborating effectively with engineers of all experience levels to achieve high reliability and performance Practical knowledge of integrating GenAI into large-scale, complex systems, including a clear unde

SQLPostgreSQLMySQLMongoDB
P
📍 United States· Full-time
✓ High-confidence listingCompany trend -86.3%

From $285.5K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . As a principal engineer on the Online Systems team, you’ll join a team that powers Pinterest’s most business-critical online systems at massive scale, driving the reliability, efficiency, and evolution behind every core Pinner and Advertiser experience. You'll lead major efforts like multi-region deployment and Kubernetes migration, set the standard for operational excellence, and define the long-term vision for our online serving infrastructure, supporting machine learning and product innovation across the company. This is an opportunity for high-impact technical leadership, broad visibility, and cross-functional influence at the heart of Pinterest’s platform. What you’ll do: Improve reliability, scalability and infra efficiency for Pinterest’s critical online systems across storage and caching, online service and realtime analytics syste

PythonJavaAWSKubernetes
P
📍 WA, United States· Full-time
✓ High-confidence listingCompany trend -86.3%

From $285.5K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is looking for a Principal Engineer to lead the technical strategy and execution for our Indexing & Retrieval Infrastructure across Core, Ads, and Shopping. This role sits at the center of our delivery infrastructure organization and will shape how fresh, relevant, and cost-efficient candidates are generated for the discovery experiences that power Pinterest at scale. What you’ll do: Define and drive the long-term technical vision for indexing and retrieval infrastructure across Core, Ads, and Shopping, aligning architecture investments to measurable improvements in freshness, quality, coverage, reliability, and cost efficiency. Lead cross-org initiatives to modernize Pinterest’s indexing architecture, including advancing unified real-time and incremental retrieval systems that support high-scale, high-quality candidat

P
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -86.3%
Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . What you’ll do: Lead a cross-platform team responsible for image rendering and video playback across Pinterest apps and web. Set the team’s technical direction and drive execution on quality, performance, and reliability for media experiences. Partner closely with Media Transcoding, Performance, Data Science, Product and Infrastructure teams to improve how media is delivered and experienced across Pinterest. Drive initiatives to reduce media load times, improve playback smoothness, and ensure high visual quality at scale. Guide the team through platform and architectural decisions, balancing product needs, technical investments, and long-term maintainability. Build and grow a high-performing team through hiring, coaching, and developing engineers. Raise the bar on engineering excellence through strong operational practices, performance mea

A
📍 United States· Full-time
✓ High-confidence listingCompany trend -98.8%

From $248K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Media Foundation team builds the core platforms and infrastructure that power photo and video capture, upload, processing, storage, and delivery across Airbnb's products, at a massive global scale. We partner closely with Product, Design, Data Science, Trust, and Infrastructure teams to ensure every image and video guests and Hosts see is fast, high-quality, and reliable, from the moment a Host uploads a photo to the moment a Guest views it in search results. The Difference You Will Make: As the Senior Engineering Manager for Media Foundation, you will lead a team of engineers to build and operate Airbnb's media infrastructure, the foundation powering every photo, video, and document that users interact with on the platform. Capabilities include uploads, AI/ML transformations, distribution, and presentation within the product. You will shape the team’s vision and help steer the organization towards our larger goals. This includes day-to-day operations, technical deep dives, community engagement with internal engineering teams, and strategic planning. As the technical and organizational leader for Media Foundation, you own the reliability, scalability, and evolution of the systems that power every media interaction on the Airbnb platform from the moment a host uploads a photo to the instant a guest streams a listing video. Your experience in both media technology and people management will be essential to the maintenance, modernization, and innovation of Airbnb's media platform. You set the bar for engineering excellence, define what great looks like for your team, and crea

A
📍 United States· Full-time
✓ High-confidence listingCompany trend -98.8%

From $212K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Passport & Commerce builds the trusted account and commerce foundation that lets anyone, anywhere, become part of the Airbnb community — from their first sign-up, through the profile and connections they build, to every booking and business transaction along the way. We're a newly formed org within Guest & Host that brings together identity, account, and commerce foundations under one roof. Our mission is to move Airbnb beyond the transaction — building a world where accounts create trust, value sticks, and what our community earns travels with them, and their businesses, wherever they go. You'll work closely with Payments and Wallet engineering, Identity & Privacy, Profile & Community, and Guest & Host product, design, and data science partners as we stand up this team's roadmap and technical foundations. The Difference You Will Make: As a Staff Software Engineer on the Passport team, you will be a key architect behind the next generation of our account platform, directly influencing how millions of guests and hosts experience Airbnb. You'll set technical direction for how account state, eligibility, and entitlements are computed, stored, and served at scale across every surface where a guest or host interacts with Airbnb. Success in this role looks like a Passport platform that is reliable and extensible enough to support new programs and partner integrations without re-architecture — measured through service reliability (uptime, latency, correctness of entitlement calculations), the speed at which new offerings can launch, and adoption of your platform by othe

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team API Agents builds the shared agent harness, tools, and infrastructure that turn OpenAI’s frontier models into systems that can reliably complete real work. We carry the capabilities behind Codex into a much broader set of products and workflows across software engineering, research, finance, healthcare, enterprise operations, and more. Our work spans search and connected context, computer use, memory, delegation and multi-agent coordination, and safe execution. Sitting at the intersection of Research, Codex, infrastructure, and applied product teams, we build reusable agent capabilities that compound across the ecosystem. About the Role We are looking for an experienced backend software engineer to build the core systems behind the next generation of agents. You will design reliable services and abstractions that help agents find the right context, use tools and computers, retain knowledge, coordinate over long-running workflows, and take action safely. The role combines deep backend and infrastructure work with strong product judgment, with opportunities to work across agent runtimes, orchestration, search, execution environments, identity and permissions, observability, and evaluations. This is software and systems engineering rather than model training: success comes from strong backend fundamentals, high agency, and the ability to turn fast-moving research capabilities into dependable production primitives. In this role, you will: Design, build, and operate the shared agent harness and backend infrastructure that power long-running, high-value workflows across OpenAI and third-party products. Build reusable capabilities across search and connected context, computer use, memory, tool execution, delegation, subagents, and multi-agent orchestration. Establish the foundations agents need to operate safely in production, including secure execution environments, identity and permissions, observability, evaluations, reliability, and cost and latency effi

TypeScriptPythonAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Data Platform at OpenAI owns the foundational data stack powering critical product, research, and analytics workflows. We operate some of the largest Spark compute fleets in production; design, and build data lakes and metadata systems on Iceberg and Delta with a vision toward exabyte-scale architecture; run high throughput streaming platforms on Kafka and Flink; provide orchestration with Airflow; and support ML feature engineering tooling such as Chronon. Our mission is to deliver reliable, secure, and efficient data access at scale and accelerate intelligent, AI assisted data workflows. Join us to build and operate these core platforms that underpin OpenAI products, research, and analytics. We’re not just scaling infrastructure – we’re redefining how people interact with data. Our vision includes intelligent interfaces and AI-assisted workflows that make working with data faster, more reliable, and more intuitive. About the Role This role focuses on building and operating data infrastructure that supports massive compute fleets and storage systems, designed for high performance and scalability. You’ll help design, build, and operate the next generation of data infrastructure at OpenAI. You will scale and harden big data compute and storage platforms, build and support high-throughput streaming systems, build and operate low latency data ingestions, enable secure and governed data access for ML and analytics, and design for reliability and performance at extreme scale. You will take full lifecycle ownership: architecture, implementation, production operations, and on-call participation. You’ve supported Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms. You’re well-versed in infrastructure tooling like Terraform, experienced in debugging large-scale distributed systems, and excited about solving data infrastructure problems in the AI space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per wee

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Applied AI team safely brings OpenAI's technology to the world. We released ChatGPT, Plugins, DALL·E, and the APIs for GPT-4, GPT-3, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate horizon. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. We serve end-users directly through ChatGPT, and serve developers through our APIs, which power product features that were never before possible. About the Role The Engineering Acceleration team designs, builds and maintains the foundational systems that engineers use to build ChatGPT and the API. This is a fast-growing team and you will get a chance to own and define the strategy, vision, and plan for how to increase developer productivity. In this role, you will: Drive the design, development, and implementation of tools, systems, and processes that accelerate engineering velocity, reduce manual effort, and increase the quality of output. Use our latest AI tools to re-think how we can be the most productive team in the industry. Work closely with various teams within OpenAI to understand their workflows, challenges, and needs, and ensure the tools and systems built by the Engineering Acceleration team address these requirements. Bring new features and research capabilities to the world by partnering with product engineers to lay the necessary technical foundations. Guide and advise product engineering teams on best practices for ensuring observable, scalable systems. Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years of experience in engineering, including 3+ years of experience in infrastructure building tooling for developers. Have experi

PythonAWSKubernetesRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Post-Training Frontiers team is responsible for training the frontier agents OpenAI ships to the world (GPT-Next). We train the flagship agentic models behind Codex, ChatGPT, and the API through large-scale reinforcement learning. The team’s work spans four areas. First, execution and science: working with teams across OpenAI to decide what can go into the final model and how, using scientific experiments and evals that are representative of the final pipeline so issues can be recognized early. Second, RL scaling: executing the final large-scale reinforcement learning run, making sure GPUs are used efficiently and training stays healthy. Third, research: improving horizontal capabilities like instruction following, factuality, memory, and multi-agent behavior, where the team’s broad visibility helps identify cross-cutting improvements across teams and domains. Fourth, engineering: maintaining the infrastructure stack and internal tools to ensure that both the final run and all integrations go as smoothly as possible and that the systems are easy to work with. About the Role This role focuses on keeping our frontier RL training runs fast, reliable, and unblocked. You will work across engineering and infrastructure problems as they emerge, from scaling and orchestration issues to inference bottlenecks, numerical problems, and hardware failures, as well as supporting large horizontal integrations in the big run, like multi-agent capabilities or memory. This is a role for a strong generalist who quickly learns anything needed for the task, has high attention to detail, debugs deeply, and is motivated by fixing the highest-impact problem in front of the team. In this role, you will: Keep large-scale async RL training runs moving by jumping into the most urgent engineering and infrastructure problems. Debug issues across training systems, inference, orchestration, scaling, and distributed infrastructure. Improve the reliability and efficiency of RL trai

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Private Computing team works across product, engineering, security, and safety to build advanced privacy products and infrastructure at OpenAI. Our mission is to provide world-class security features to users so their private data remains private, even from OpenAI. We use technologies like confidential computing, trusted execution environments, and end-to-end encryption to ship product features across ChatGPT, the API, and our future consumer devices. About the Role We’re looking for software engineers to design, build, and scale novel privacy features and infrastructure across ChatGPT, API, and future consumer devices. In this role, you will: Ship fast while balancing difficult trade-offs in complex domains Build core abstractions for trusted execution environments and end-to-end-encryption Build product features for private inference and storage across ChatGPT, API, and future consumer devices Update build systems to increase trust and verifiability Integrate with safety and integrity infrastructure Operate systems at scale with high reliability, including an on-call rotation Collaborate with a diverse set of cross-functional teams across product, engineering, security, safety, policy, and legal You might thrive in this role if you: Care deeply about user privacy and security Have 5+ years of experience in professional software engineering Have experience building and scaling confidential computing or encryption technologies in production environments Have experience with Kubernetes and cloud orchestration systems Take pride in building and operating scalable, reliable, secure systems Can collaborate well and drive alignment in the face of difficult trade-offs Are comfortable with ambiguity and rapid change Workplace & Location This role is based in San Francisco, CA. We follow a hybrid model with 4 days a week in the office and offer relocation assistance to new employees. About OpenAI OpenAI is an AI research and deployment company dedicat

AWSKubernetesRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role We are seeking a Software Engineer, Security Observability to join our Security team. In this role, you will be responsible for building secure, scalable systems that enhance our security observability infrastructure. Leveraging your strong engineering skills, you will collaborate with cross-functional teams to develop, deploy, and maintain robust software solutions that support our security and detection capabilities. This role is open to remote employees, or relocation assistance is available to one of our OpenAI offices in San Francisco, Seattle, or New York City. Due to requirements associated with work this role may support, applicants for this position must be U.S. citizens. In this role, you will: Design and develop scalable software systems that facilitate security observability across our infrastructure. Build and maintain data pipelines that centralize and store security-relevant data from diverse sources. Proactively improve the resilience and reliability of data systems to ensure high platform availability Collaborate closely with Detection & Response (D&R) and other security teams to reduce the company’s security risk. Contribute to data engineering in support of forensic investigations and compliance efforts. You might thrive in this role if you have: Strong software engineering experience, with proficiency in programming languages such as Python, Golang, or similar. A background in infrastructure as code, with exp

PythonAWSAzureRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI's mission is to ensure that artificial general intelligence benefits all of humanity. The Consumer Devices team is building a new generation of AI-powered products that seamlessly integrate hardware and software to create intuitive, transformative experiences. We bring together experts across embedded systems, machine learning, hardware, design, and product engineering to develop products at the intersection of AI and consumer technology. About the Role OpenAI is seeking a System Power Engineer to characterize, measure, and optimize power consumption across our embedded hardware products. In this role, you will work closely with Electrical Engineering and system software teams to build power test automation, measure subsystem-level power usage, and drive improvements that directly impact battery life, thermal behavior, charging performance, and system reliability. You will help establish the methodologies and metrics used to understand and improve power efficiency across real-world product experiences, from controlled lab environments to representative day-in-the-life usage scenarios. This role requires hands-on experience with embedded hardware platforms, power instrumentation, and the analysis of power profiles and system behavior. This role is based in San Francisco, CA. We use a hybrid work model of four days per week in the office and one day working remotely. Relocation assistance is available for new hires. In this role, you will: Define and develop power testing automation to evaluate system behavior across a range of workloads and operating conditions. Measure subsystem-level power consumption using power breakout probes and other lab instrumentation. Develop and execute power characterization tests spanning basic workloads, complex mixed-use scenarios, and representative day-of-use experiences. Partner closely with Electrical Engineers to identify opportunities to improve system power efficiency. Collaborate with software engineering

PythonAWSRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Applications Engineering organization builds and operates the products (such as ChatGPT & Codex) that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role We’re hiring Backend Software Engineers to design and implement safe services and infrastructure that power our core products. What You’ll Do Architect, build, and improve scalable backend systems and APIs. Drive performance, reliability, and safety across distributed services. Implement data storage, retrieval, compute, and integration solutions. Participate in long-term architectural planning and technical design reviews. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. You Might Thrive Here If You: Have strong experience with distributed systems, APIs, and backend languages (e.g., Go, Python, Rust, C++). Have experience setting up and maintaining production backend services and data pipelines. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Enjoy building resilient services that handle large scale and complexity. Are self-directed and enjoy figuring out the best way to solve a particular problem Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Business Systems / Enterprise Platform Technology builds the internal systems, data foundations, workflow infrastructure, and enterprise platforms that help OpenAI operate at scale. The EPT AI Pod builds AI-native internal apps, MCP connectors, multi-agent workflows, and reusable platform capabilities across Finance, People, and GTM. About the Role As an Enterprise Applied AI Engineer, you will build internal apps for enterprise operations and the shared platform components those apps run on. This includes MCP connectors, multi-agent orchestration, data architecture, evals, monitoring, auditability, and governance. We’re looking for a hands-on engineer who is strong in Python, system design, enterprise integrations, data architecture, and applied AI systems. You should be excited to turn ambiguous business workflows into reliable internal products and shared infrastructure. In this role, you will: • Build internal apps for enterprise operations across Finance, People, and GTM • Build MCP connectors and enterprise integrations with strong auth, permissions, idempotency, retries, and rate-limit handling • Design end-to-end multi-agent workflows with tool routing, human approvals, audit trails, and safe action boundaries • Design data architecture for operational AI systems, including ingestion, schemas, quality checks, lineage, and governance • Build evals, monitoring, metrics, and regression tests for agentic workflows • Create reusable infrastructure, patterns, and components that other enterprise teams can build on • Partner with system owners and business owners to turn messy enterprise workflows into reliable internal products You might thrive in this role if you: • Have strong Python engineering skills for backend services, MCP connectors, agent/tool workflows, eval harnesses, and data ingestion jobs • Have strong system design skills across shared infrastructure, app architecture, reliability, and scaling • Have experience building internal apps,

PythonAWSRestAI
🔔

Get new reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime