Jobs in United States

Reliability Engineer Iii in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer iii jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Applications Engineering organization builds and operates the products that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role You will be a Senior iOS engineer on OpenAI’s Applied Foundations team, building the core mobile experiences that power how users sign up, manage their account, family features, pay for services, stay safe, and interact with OpenAI’s products with confidence. This role is about creating high-quality products as well as reusable iOS foundations that product teams across different OpenAI apps depend on to ship quickly while meeting the highest standards for security, reliability, and user trust. You’ll own complex client-side systems spanning UI, networking, local state, payment integrations and Apple platform integrations, and work closely with backend, product, and safety partners to shape the architecture that supports OpenAI’s mobile ecosystem at global scale. In this role, you will: Build and ship new experiences on iOS that showcase the power of AI. Optimize app performance, reliability, and responsiveness at global scale. Design and maintain shared iOS frameworks and primitives for account, trust, and commerce flows that are used across OpenAI’s mobile apps. Establish robust testing frameworks and refine app architecture for long-term maintainability. Collaborate with product, design, research, and backend teams to deliver high-impact features. Provide technical leadership to shape the future of OpenAI’s iOS platform. You might thrive in this role if you: Have 4+ years of professional software engineering experience. Hav

AWSRestAISwift
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Consumer Products team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. Within Consumer Products, the camera stack is a critical sensing component. The team partners closely with electrical engineering, silicon vendors, systems, and higher-level perception and product teams to bring up new hardware, stabilize capture pipelines, and ensure camera systems are robust, debuggable, and ready for real-world deployment. This work spans early prototypes through production, with a strong emphasis on correctness, repeatability, and long-term reliability. About the Role As a Camera Firmware Engineer, you will own low-level camera enablement on custom hardware—from early board bring-up through stable production capture. You will develop and maintain the firmware and software that makes camera sensors reliable, controllable, and debuggable, forming the foundation for higher-level camera pipelines and product features. This role is highly hands-on and systems-oriented. You will work close to the hardware, diagnose real-world timing and integration issues, and build tooling that accelerates iteration across the entire camera stack. This role is based in San Francisco, CA. We follow a hybrid work model with four days per week in the office and offer relocation assistance to new employees. In This Role, You Will Bring up new camera sensors and modules on prototype and production boards, including link stability, sensor control, and correct power, reset, and clock sequencing. Develop and maintain low-level camera software, including sensor drivers, board configuration, and camera subsystem integration across hardware revisions. Enable and validate core capture paths for development and production, including RAW capture for debugging, still capture, and hardware-

AWSLinuxRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Platform Systems team at OpenAI operates at the intersection of cutting-edge AI and large-scale distributed systems. We build the engineering and research infrastructure required to train OpenAI’s flagship models on some of the world’s largest, custom-built supercomputers. Our team develops core model training software and works deep in the stack - spanning collective communication, compute efficiency, parallelism strategies, fault tolerance, failure detection, and observability. The systems we build are foundational to OpenAI’s research velocity, enabling reliable, efficient training at frontier scale. We collaborate closely with researchers across the organization, continuously incorporating learnings from across OpenAI into the evolution of our training platform. About the Role As a Software Engineer, Platform Systems, you will design and build distributed systems that provide visibility into large-scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical to operating OpenAI’s training stack and is actively evolving to support new use cases and increasingly complex workloads. This role sits at the core of our training infrastructure, blending systems engineering, performance analysis, and large-scale debugging. In This Role, You Will Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high-throughput distributed systems Collaborate with systems, infrastructure, and research teams to evolve platform

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Release Engineer team is responsible for building and maintaining the systems that power software delivery—from CI/CD pipelines and artifact management to release automation and fleet telemetry. We ensure software across bootloaders, firmware, operating systems, and cloud services is built reproducibly, validated rigorously, and released safely at scale. About the Role As a Release Engineer, you’ll design, build, and operate release infrastructure that enables reliable, secure, and traceable software delivery across complex multi-component systems. You’ll partner closely with embedded, cloud, and QA teams to ensure that every build—from development to OTA deployment—is fast, verifiable, and production-ready. We’re looking for engineers who take pride in automation, build reproducibility, and system reliability—and who enjoy building the connective tissue that allows hardware and software to ship together seamlessly. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and operate CI/CD pipelines for multi-component builds (bootloader, firmware, OS images, backend, companion apps) using hermetic toolchains. Define versioning and branching strategies; automate promotions, changelogs, and artifact retention. Integrate unit, integration, and hardware-in-the-loop (HIL) test results; quarantine flaky tests, auto-bisect failures, and block unsafe promotions. Build A/B OTA update flows with verity and health checks; run staged rollouts and canaries; implement safe rollback and roll-forward strategies. Implement code signing for binaries and firmware, generate SBOMs, run vulnerability scanning, and attach build attestations and provenance. Manage dashboards and alerts for build health, promotion latency, failure rates, and fleet update telemetry. You might thrive in this role if you: Have experience building and operating buil

PythonAWSCI/CDGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Connectivity Software Engineering team is responsible for enabling seamless, secure, and high-performance wireless connectivity across OpenAI’s products. We design and optimize Bluetooth, BLE, Wi-Fi, and emerging wireless technologies to ensure robust device pairing, network performance, and interoperability. Our work spans kernel drivers, system services, and user-level tools, with a focus on real-world performance, scalability, and reliability. About the Role OpenAI is seeking a Connectivity Software Engineer to design, implement, and optimize wireless connectivity features across our product ecosystem. You’ll work at the intersection of systems software, wireless standards, and hardware integration—building robust pairing and provisioning flows, debugging low-level protocols, and driving performance under real-world RF constraints. You will also support certification, field interoperability, and fleet-scale connectivity infrastructure. This role is based in San Francisco, CA . We use a hybrid work model of 4 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement, and debug Bluetooth/BLE and Wi-Fi features across kernel drivers, BlueZ/wpa_supplicant/hostapd, and systemd/D-Bus services Deliver robust pairing, bonding, and provisioning flows (GATT/GAP, LE Audio/LC3, WPA3/802.1X, captive portals, NAN) Optimize link performance: throughput, latency, jitter, roaming, coexistence (BT↔Wi-Fi), and power modes (TWT, WoWLAN) Build reliable network management using NetworkManager/nmcli, nl80211/cfg80211/mac80211, DNS/DHCP/mDNS, P2P/SoftAP Instrument and analyze with packet captures and tooling (btmon/hcidump, Wireshark, iperf, eBPF/perf, spectrum sniffers) Drive interoperability and certification readiness (Bluetooth SIG, Wi-Fi Alliance) and resolve field issues with root-cause fixes Contribute to OTA-safe configuration, telemetry, and diagnostics for fleet-scale operation You might thrive in

PythonAWSLinuxRest
O
📍 Seattle, Washington, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Statsig team within OpenAI owns the experimentation, rollout, dynamic configuration, and analytics infrastructure that sits on the launch path for OpenAI products. Our systems help teams ship safely, evaluate product and model changes in production, and make high-confidence decisions from real-world usage. Statsig began as an independent company built around experimentation, feature management, and product analytics at scale. After Statsig joined OpenAI, the team began the next chapter: bringing that platform expertise and infrastructure into OpenAI as the experimentation and rollout foundation for every product we ship. This is infrastructure with a very direct product consequence. Teams working on ChatGPT, Codex, model measurement, consumer experiences including ads, business subscriptions, developer products, and shared platform systems depend on Statsig to evaluate configurations, move traffic safely, ingest experiment data, serve analytics, and roll changes forward or back when production reality demands it. We are at a critical point in the platform journey. Adoption is accelerating quickly across OpenAI, and the systems that were already important are becoming load-bearing for how the company launches. The infrastructure needs to stay fast under sharply increasing evaluation volume, reliable when more services depend on it, observable enough to debug quickly, and efficient enough to support OpenAI-wide scale. Recent SDK and server-side infrastructure work has already produced measurable wins in latency, reliability, memory usage, and compute efficiency across important services. The next phase is to make those gains systematic: a platform that can absorb rapidly growing product velocity while preserving low latency, data quality, operational safety, and developer trust. Based out of OpenAI's Bellevue office, we are a close-knit team that values in-person collaboration, technical depth, operational ownership, and building infrastructure that

VueAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team We’re hiring software engineers to make OpenAI’s Model Performance teams more productive. These teams work on the systems, tooling, and infrastructure that help improve model performance across OpenAI’s training and inference workloads at frontier scale. About the Role We’re looking for an autonomous, high-ownership developer productivity engineer who cares deeply about helping other engineers move faster, safer, and with more confidence. This role will sit within OpenAI’s Model Performance organization, contributing to developer infrastructure, CI systems, testing workflows, tooling, and broader performance infrastructure efforts. There is also a strong opportunity to contribute to the Triton project and help improve the systems that support performance-critical engineering work across OpenAI. In this role you will: Improve development workflows for engineers working on model performance infrastructure Design and improve CI/CD, release, validation, and testing pipelines Build and maintain tools that improve reliability, iteration speed, and engineering confidence Partner closely with engineers to identify friction in testing, debugging, deployment, and development workflows Contribute to infrastructure efforts that support performance-critical training and inference systems Help improve developer experience across Python-heavy codebases and performance-oriented infrastructure Work in a high-context, ambiguous environment where ownership and good judgment matter You might thrive in this role if: You are motivated by enabling the people around you and helping engineers do their best work You have strong experience with CI/CD, developer infrastructure, testing systems, tooling, or build/release workflows You are highly collaborative, empathetic, and comfortable partnering deeply with technical teams You are strong in Python and enjoy building reliable, scalable developer tools and infrastructure You have experience improving large-scale engineering work

PythonAWSCI/CDRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal a dependable, scalable foundation for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between physical hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Security Engineer, Host Assurance to help build the trust foundations for bare-metal platforms across OpenAI’s global infrastructure. This is a deeply hands-on engineering role for a builder who can design, implement, and operate the core security infrastructure that establishes trust in hardware platforms before they are eligible to run workloads. Success in this role requires strong technical judgment, the ability to work comfortably at low levels of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—including novel hardware platforms and emerging deployment models– to make the secure path the easiest path. This role is well suited for engineers who enjo

AWSRestAgileAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Codex Core Agent team builds the kernel of Codex. We own making the agent better, accelerating research, and making those improvements real in production for our users. That means working across the systems that make Codex actually function as an agent in the real world: the production performance envelope around tokens, latency, reliability, cost, and capacity; the core execution loop and interfaces that turn models into useful behavior; the shared infrastructure that enables other teams to build on Codex; and the feedback loops that turn real-world usage into better models and better agent behavior over time. About the Role We’re looking for applied AI engineers to help bring Codex agents from impressive demos to dependable tools. This role is about improving agent performance on real software engineering tasks and closing the gap between research capability and real-world usefulness. You’ll work closely with research, infrastructure, and product to ensure agents are not just powerful, but useful, steerable, and reliable in practice. The job is not only to improve model behavior in isolation, but to turn those improvements into measurable gains in solve rate, usefulness, and economic value for users. What You’ll Do Design and iterate on agent behaviors across real-world coding tasks and long-horizon workflows. Work closely with research to develop and run evals to measure agent performance, regressions, failure modes, and edge cases. Improve performance through prompting, tool-use strategies, context construction, and model-facing experimentation. Analyze failures in production and systematically improve robustness and reliability. Build feedback loops and data systems that get better real-task data into evaluation and research. Work with product teams to shape user-facing agent experiences and the interfaces the agent depends on. Help define what “good” looks like for agents completing complex tasks end-to-end. You Might Be a Good Fit If You Ha

PythonAWSRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team We’re hiring software engineers to make OpenAI’s networking teams more productive. These teams build and operate the high-performance networking systems that support OpenAI’s training and inference infrastructure at frontier scale. About the Role We’re looking for someone who cares deeply about the developer experience of engineers working on complex infrastructure systems — especially around build systems, test architecture, release pipelines, and reliable development workflows. This role will be embedded with OpenAI’s networking team: making it faster, safer, and easier for engineers to build, test, validate, and ship changes across multi-server, networked, and hardware-adjacent environments. In this role you will: Improve development workflows for engineers building and operating OpenAI’s networking systems Design and improve continuous deployment, release, and validation pipelines Build and maintain test harnesses for multi-server, networked, and hardware-backed environments Improve iteration speed across C++, Python, and build-system-heavy codebases Partner with engineers to identify friction in CI, testing, debugging, and deployment workflows Drive testing and reliability strategy for infrastructure components that support large-scale training and inference workloads Work closely with centralized developer experience teams while staying deeply embedded with the networking engineers closest to the systems You might thrive in this role if: You are motivated by helping other engineers move faster and with more confidence You have experience with CI/CD, release pipelines, testing infrastructure, or build systems You are comfortable moving between C++, Python, and build systems such as CMake, Bazel, or Blaze You enjoy building test harnesses, automation, and workflow improvements for complex systems You do not need to be a networking expert, but you are excited to learn enough about the domain to make the team meaningfully more effective When you see

PythonAWSCI/CDRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Foundations Research team works on high-risk, high-reward ideas that could shape the next decade of AI. Our goal is to advance the science and data that enable our training and scaling efforts, with a particular focus on future frontier models. Pushing the boundaries of data, scaling laws, optimization techniques, model architectures, and efficiency improvements to propel our science. The Search team sits within Foundations, building agentic search by co-designing model–system interfaces with the core search stack (serving, indexing, retrieval) to translate model intent into reliable, real-world actions. Operating at the frontier of AI and information retrieval, the team develops large-scale systems that transform and index vast corpora, enabling models to reason over global knowledge and act dependably. In close partnership with researchers, we rapidly bring modeling breakthroughs into production and redefine how intelligent systems discover, retrieve, and synthesize information at planetary scale. About the Role We’re looking for a Software Engineer focused on building and scaling retrieval systems. You’ll work with a team of researchers and engineers to develop infrastructure that enables models to retrieve and act on the right information at the right time. This includes designing and operating indexing systems, retrieval pipelines, and serving layers. This work supports retrieval across OpenAI products and research, with direct impact on system performance, reliability, and scale. Responsibilities Build and scale retrieval infrastructure across indexing, serving, and query execution. Develop low-latency, high-throughput systems for real-time model interaction. Partner with research to productionize embedding and retrieval techniques. Support dense, sparse, and hybrid retrieval pipelines. Own system performance, reliability, and observability at scale. Collaborate across Pretraining, Inference, and Product teams to integrate retrieval end-to-e

AWSRestAIGo
O
📍 Seattle, Washington, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Statsig team within OpenAI builds the experimentation, feature rollout, dynamic configuration, and analytics systems that help OpenAI ship products with speed, safety, and evidence. Our work sits on the critical path for how product, engineering, research, and go-to-market teams learn from real-world usage and make high-confidence decisions. Statsig began as an independent company focused on helping builders move faster through trustworthy experimentation and feature management. After Statsig joined OpenAI, the team began the next chapter: bringing that deep product expertise, customer intuition, and mature platform infrastructure into OpenAI as the experimentation and rollout platform for every product we ship. Today, we support teams across ChatGPT, Codex, model measurement, consumer experiences including ads, business subscriptions, developer products, and the shared infrastructure that connects them. These teams rely on Statsig to safely introduce new capabilities, compare product and model behavior, measure impact, and roll changes forward or back with confidence. We are at a defining moment in the platform journey. OpenAI has the data, product surface area, and pace of innovation to learn faster than almost any organization in the world, but that potential only becomes real if teams can experiment responsibly, measure clearly, and roll out changes safely. Adoption of the platform is accelerating rapidly across the company, and recent SDK and server-side infrastructure work has already produced measurable wins in latency, reliability, memory usage, and compute efficiency for important services. Based out of OpenAI’s Bellevue office, we are a close-knit team that values in-person collaboration, urgency, craft, and impact. We build for other builders, and the best version of this team is one where every OpenAI product team can move faster because the experimentation and rollout layer is dependable, fast, and easy to use. About the Role We are l

VueAWSRestAI
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -86%

$272K – $320K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role Build the most advanced AI Meeting Notes product — and expand it into broader “AI data capture” features that help teams turn conversations into durable context, tasks, and knowledge. Our mission is to 10x the rate of business context & data that enters Notion — optimized for agents — so teams get superhuman memory across workstreams and customers. Notion workspaces that use AI Meeting Notes already enter 6x more data on a daily basis, so we’re well on our way. What You'll Achieve Ship end-to-end product experiences across capture → transcript → summary → follow-ups (full-stack ownership). Make meeting & data capture feel effortless and magical (e.g., speaker identification via audio waveforms, richer in-meeting UX, smarter organization). Improve summary quality that teams trust: structure, factuality, and citations that make downstream agents and humans more capable. Raise the bar on reliability & observability across the pipeline (SLOs, debugging workflows, incident response) for realtime systems. Build agentic meeting workflows that turn discussions into tasks, follow-ups, and organized knowledge — so “w

RestAIGoRust
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -86%

$135K – $155K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. To be eligible for this role, you need to graduate in December 2026 and be able to start full time by January/February 2027. What You'll Achieve: You'll work with others to plan, shape, and build new product features from start to finish: through conception, research, implementation, and maintenance. For example, you might work on driving adoption, by instrumenting key onboarding moments and iterating on activation flows. You'll help improve performance and reliability, or polish existing features. For example, you might update our spell checker to sync dictionaries across browsers or improve search to index file attachments. You'll build internal tools to support simplicity and productivity for the whole team. This might include writing a script to import user feedback into Notion. Qualifications: Pursuing a bachelor's or master’s degree in computer science, engineering, or another related field. To be eligible for this role, you need to graduate by Dec 2026 and be able to start FTE by January/February 2027. Previous internship experience. Working towards a proficiency of one or more programming languages such as TypeScript, Node.js

JavaScriptTypeScriptPythonJava
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -86%

From $130K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. This is NOT a new grad role! This role is for candidates with 0-2 YOE and can start full time right away. About the Role As an engineer at Notion, you’ll help shape core user experiences and accelerate how people discover value in Notion. You'll tackle meaningful challenges with increasing autonomy, crafting code that millions of users will experience. You'll take ownership of projects that matter, make critical technical decisions, and contribute your unique perspective to our product vision. Working alongside passionate experts across design, product, and data, you'll help shape the future of how people work. We're looking for an Early Career AI Engineer to join as a strategic partner in shaping Notion's AI vision. You'll work on cutting-edge AI-powered features, leveraging LLMs, embeddings, and other AI technologies to make Notion more intelligent and capable. This role may be aligned to one of multiple AI-focused teams at Notion. Depending on team match and business needs, you could work on: AI product engineering: building model-powered features end-to-end (UX, APIs, retrieval, orchestration, quality, and reliability) Model &amp

TypeScriptReactNode.jsSQL
🔔

Get new reliability engineer iii jobs in United States by email

Daily job updates · Unsubscribe anytime