Jobs in United States

Reliability Engineer in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

P
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -86.3%

From $158.8K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Are you passionate about designing and building enterprise finance platforms that power critical business operations at scale? Join Pinterest’s IT Enterprise Systems team and help shape the future of our Finance technology ecosystem. In this role, you will lead the design, development, and evolution of Oracle EBS-based solutions, modern integrations, and AI-enabled capabilities that support our fintech initiatives. This is a highly technical position offering the opportunity to work across Oracle EBS, AWS and emerging internal AI platforms to deliver secure, scalable, and business-critical systems. What You’ll Do: Design, build, and support scalable enterprise solutions across the Oracle EBS and Finance systems landscape, with a strong focus on technical architecture, integrations, and platform reliability. Serve as a senior technical expert for

SQLAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Enterprise Verticals builds role-specific ChatGPT Work experiences for high-value enterprise workflows. We combine product engineering, plugins and skills, connectors, data, evaluations, and customer evidence to turn useful demos into reliable daily work. This opening sits within the Technology vertical inside Enterprise Verticals. The group focuses on repeatable workflows for people at technology companies, beginning with functions such as data and analytics, sales, and design, and carries the shared platform needs—tool integration, permissions, quality measurement, and safe rollout—across those experiences. We work closely with Design, Research, GTM, Security, and platform teams, as well as with customers and design partners. Success means that people can reach a trustworthy first result, understand what the system did, and keep using the workflow—not merely that a prototype exists. About the Role We are looking for an exceptionally experienced, hands-on full-stack engineer to define and build the next generation of AI-powered enterprise workflows. You will take on the hardest and most ambiguous problems in the Technology vertical: translating real customer needs into product direction, designing the systems behind the experience, and personally writing and shipping production-quality code across the stack. You will own the technical direction and end-to-end delivery of products spanning ChatGPT Work surfaces, backend services, plugins, connectors, enterprise data, permissions, and evaluations. You will make foundational architecture and product tradeoffs; establish patterns other engineers can build on; and hold these experiences to a high bar for reliability, security, observability, and customer value. This is an individual-contributor role for an engineer who leads through technical judgment, direct execution, and influence—not people management. You should be equally comfortable working directly with customers, setting direction with senior cro

TypeScriptPythonReactNode.js
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking an Actuator Gear Design Engineer to lead the development of custom gears and gear stages for advanced robotic systems. You will own actuator development from early architecture and concept generation through prototype validation and system integration, partnering closely with mechanical, electrical, controls, firmware, and reliability teams. You will partner with external suppliers and internal manufacturing to create full gearbox assemblies. This role focuses on the design, integration, and validation of precision gearing, including broader knowledge around motor electromagnetics, transmission types, sensing, structural components, and thermal architectures. You will help drive actuator development across the full engineering lifecycle while establishing scalable design, test, and integration practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will: Lead the architecture, design, and integration of custom robotic actuator gearing. Define actuator requirements and system-level trade studies around torque density, bandwidth, efficiency, thermal performance, back drivability, inertia, reliability, manufacturability, and cost. Design precision electromechanical assemblies with strong attention to tolerances, alignment, load paths, thermal expansion, sealing, wear, and serviceability. Drive actuator integration into robotic systems, partnering closely with controls, firmware, electrical, and robotics software teams to optimize closed-lo

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a senior Actuator Electromagnetic Design Engineer to lead the development of custom electromechanical actuators for advanced robotic systems. You will own actuator development from early architecture and concept generation through prototype validation and system integration, partnering closely with mechanical, electrical, controls, firmware, reliability, and manufacturing teams. This role focuses on the design, integration, and validation of precision electromechanical systems, including motors, transmissions, sensing, structural components, and thermal architectures. You will help drive actuator development across the full engineering lifecycle while establishing scalable design, test, and integration practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will: Lead the architecture, design, and integration of custom robotic actuators, including the design, simulation, integration and sourcing of custom electromagnetic components. Define actuator requirements and system-level trade studies around torque density, bandwidth, efficiency, thermal performance, inertia, reliability, manufacturability, and cost. Design precision electromechanical assemblies with strong attention to tolerances, alignment, load paths, thermal expansion, sealing, wear, and serviceability. Drive actuator integration into robotic systems, partnering closely with controls, firmware, electrical, and robotics software teams to optimize closed-loop performance. Devel

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We're looking for an Optical Interconnect System Engineer to design, qualify, and deploy scalable optical connectivity for large-scale AI infrastructure. This role spans fiber-system architecture, optical-mechanical integration, validation, reliability, deployment, and serviceability. You will work with optical, mechanical, electrical, networking, manufacturing, reliability, and data-center teams to translate system needs into practical interconnect solutions. This is a hands-on role for someone who can connect design decisions with installation, qualification, troubleshooting, and long-term operational performance. In this role, you will: Define optical interconnect architectures and requirements across hardware platforms and rack-level systems. Design high-density fiber systems for performance, density, reliability, installation, and serviceability. Lead optical-mechanical integration and cross-functional design reviews. Develop test and qualification plans for optical components, modules, switching platforms, and integrated systems. Own optical loss budgets, routing guidelines, handling requirements, and serviceability criteria. Support system bring-up, deployment, troubleshooting, failure analysis, and reliability improvement. Create reusable design guidelines, interface requirements, and qualification methods. You might thrive in this role if you have: Core experience Experience desi

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the team The Agent Enablement AI Deployment Engineering (ADE) team works across engineering, product, design, partnerships, and strategic customers to grow an open ecosystem of agent-enabled sites and services. We help partners adopt the OpenAI tech stack related to identity, permissioning, agent-auth primitives so users can safely connect ChatGPT and Codex to the tools, services, and workflows they already use. Our team also works with external partners on defining the standards for agent access, marketplace offerings as well as other agent enablement initiatives to ensure users of ChatGPT and Codex go from intent to task completion seamlessly. About the role We are looking for an AI Deployment Engineer to help strategic partners design, build, validate, launch, and operate agent enablement integrations across web applications, connectors, APIs, CLIs, MCP servers, and developer tools. This is a hands-on, partner-facing product engineering role for someone who can contribute to the platform itself, lead sophisticated technical engagements, and turn ambiguous identity and agent-workflow requirements into secure, production-ready integrations. You will work across partner product and engineering teams and OpenAI’s product, engineering, design, partnerships, legal, policy, security, support, and go-to-market teams. You will identify high-value user journeys, choose the right integration path, prototype and review architectures, write code, run evaluations and dogfood, trace failures end to end, guide launch and rollout, and support post-launch iteration. The best person for this role moves fluidly between full-stack code, OAuth/OIDC and identity systems, product judgment, project leadership, and clear communication with engineers and executives. This role is a fit for a product-minded engineer who wants to stay close to users and partners while going deep on authentication, permissions, reliability, safety, and developer experience. The principle objective is to

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b

TypeScriptPythonAWSDocker
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Ecosystem AI Deployment Engineering (ADE) team supports strategic partners as they build high-quality technical integrations into ChatGPT and Codex. Our goal is to create products users depend on, drive adoption and retention, and build an ecosystem where partners win when OpenAI wins. About the Role We are looking for an AI Deployment Engineer to help strategic partners design, build, evaluate, submit, launch, and maintain high-utility plugins for ChatGPT and Codex. This is a hands-on, partner-facing product engineering role for someone who can contribute to the platform itself, lead sophisticated partner engagements, and translate ambiguous product needs into production-ready integrations. You will work across partner product and engineering teams and OpenAI's product, engineering, partnerships, legal, policy, design, and go-to-market teams. You will identify the right use cases, prototype and review implementations, run evaluations, debug issues across systems, guide partners through submission and review, and support launch and post-launch iteration. The best person for this role moves fluidly between code, product judgment, project leadership, and clear communication with engineers and executives. This role is a fit for a product minded engineer who wants to stay close to users and partners while still going deep on code, reliability, evaluations, and developer experience. The goal is to help partners ship plugins that are not merely technically functional, but genuinely useful in ChatGPT and Codex. This role is based in our San Francisco office. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the technical partner journey for priority B2B plugins—from pitch and readiness assessment through architecture, build, evaluation, submission, launch, and ongoing maintenance. Identify strong plugin use cases, define crisp user journeys and expected behaviors, and

AWSRestAIGo
O
📍 Seattle, Washington, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Statsig team within OpenAI builds the experimentation, feature rollout, dynamic configuration, and analytics systems that help OpenAI ship products with speed, safety, and evidence. Our work sits on the critical path for how product, engineering, research, and go-to-market teams learn from real-world usage and make high-confidence decisions. Statsig began as an independent company focused on helping builders move faster through trustworthy experimentation and feature management. After joining OpenAI, the team began its next chapter: bringing deep product expertise, customer intuition, and mature platform infrastructure into the product development system used by every OpenAI team. Today, teams across ChatGPT, Codex, model measurement, consumer monetization, business subscriptions, developer products, and shared infrastructure rely on Statsig to safely introduce new capabilities, measure impact, and roll changes forward or back with confidence. We are at a defining moment as adoption accelerates and the platform becomes a company-wide standard. About the Role We are looking for an Engineering Manager, Statsig Product to lead the product engineering organization responsible for Statsig’s post-acquisition journey at OpenAI. You will define how experimentation, rollout, configuration, and analytics become a simple, reliable, and trusted part of how every OpenAI product team ships. You will set strategy across multiple product and platform workstreams, build the organization and leadership structure needed for the next phase, and establish the operating model for a platform that serves teams across the company. The right leader can operate across product strategy, technical architecture, organizational design, developer experience, reliability, and executive alignment. You will help preserve what made Statsig strong while integrating it deeply into how OpenAI launches, measures, learns, and makes product decisions. In this role, you will: Build, lead,

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Applied AI Engineering team helps organizations turn frontier AI capabilities into safe, reliable, and high-impact production systems. We work with customer executives, product and engineeriIng teams, security leaders, and transformation teams to identify valuable opportunities, accelerate technical implementation, and scale what works. Enterprise deployments are defined by complexity rather than any one industry: existing architectures, diverse data environments, security and governance requirements, multiple stakeholder groups, and organization-wide change. We turn lessons from these deployments into better products and reusable patterns for customers everywhere. About the Role As an Enterprise Applied AI Engineer you will partner directly with leading organizations to design, build, and deploy AI systems that deliver measurable business outcomes. You will combine deep technical judgment, hands-on engineering, and customer leadership to take ambitious ideas from use-case selection and architecture through prototyping, evaluation, production launch, and scale. You will write and debug code, build evaluation systems, resolve complex integrations, and guide decisions involving model behavior, reliability, latency, cost, safety, security, governance, and operational readiness. Success is measured by production systems, sustained adoption, and meaningful customer impact—not simply activity or successful demonstrations. This is a rare opportunity to work on consequential real-world deployments at the frontier of AI while directly influencing how OpenAI’s products evolve. This role is based in our SF or NYC office. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Partner directly with enterprise customers to identify high-value opportunities and translate them into technical architectures, implementation plans, evaluation strategies, and measurable success criteri

JavaScriptTypeScriptPythonJava
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i

PythonReactNode.jsAngular
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role OpenAI's Hardware organization builds supercompute platforms from silicon and boards to full rack-scale systems to power advanced AI workloads. This role owns end-to-end quality for high-speed interconnect hardware across the product lifecycle: early design influence, supplier/contract manufacturer readiness, qualification, ramp, and fleet quality in lab and data center environments. You will be the quality lead for advanced interconnect components and assemblies, including high-speed copper cables, cable cartridges, patch panels, backplane/cable-backplane solutions, high-speed connectors, and related electro-mechanical interfaces. You will partner closely with electrical, mechanical, SI/PI, systems, reliability, operations, and external vendors to prevent escapes and drive rapid, data-driven containment and corrective action. In this role you will: Own quality for advanced interconnect components and assemblies: high-speed connectors, high-speed copper cables, cable cartridges (e.g., cable cassette style assemblies), patch panels & optics, and backplane/cable-backplane interconnect solutions. Drive quality-by-design: participate in design reviews, DFM/DFx, tolerance stacks, material and plating selections, connector mating strategy, strain relief, and assembly methods to reduce variation and field failures. Define and track quality and reliability metrics (DPPM, yield, escapes, RMA/FRACAS trends, Cpk/Ppk where applicable) for interconnects across NPI and m

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will develop and evolve the tooling ecosystem that hardware engineers rely on every day — from hardware compilers and IR transformations to simulation, debugging, and automation infrastructure. The work spans software engineering, compiler concepts, and practical hardware workflows, with direct impact on how quickly and effectively we design next-generation AI systems. You’ll collaborate closely with architects, RTL designers, and verification engineers to translate real engineering friction into durable, scalable tooling solutions. In this role you will: Build and improve the software tooling that makes hardware teams faster: compilation, IR transforms, RTL generation, simulation, debug, and automation. Extend and integrate hardware compiler stacks (frontends, IR passes, lowering, scheduling, codegen to Verilog/SystemVerilog) and connect them to real design workflows. Improve developer experience and reliability: reproducible builds, better error messages, faster iteration loops, and dependable CI and regression infrastructure. Work closely with designers and verification engineers to turn real pain points into durable tools. Dive into RTL when needed: read and reason about Verilog/SystemVerilog to debug issues, validate tool output, and improve debuggability. Be willing to go all the way down the stack when necessary, including gate-level views, synthesis results, and implementation artifacts. Help enable PPA optimization loops by building analysis and au

PythonAWSGitRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

Location: San Francisco, CA (Hybrid: 4 days onsite/week). Relocation assistance available. About the Team: We build foundational platform software that enables reliable, secure, and performant products. The team works across system layers and partners closely with adjacent engineering groups to deliver robust capabilities from concept through launch. About the Role: We’re seeking a System Software Engineer to design, implement, and debug core platform components and the pipelines that build and update system images. You’ll work across operating system layers, focusing on performance, security, and deep system debugging to ship production‑grade systems. In this role, you will: Design, implement, and debug system‑level components and services across kernel and user space. Configure and maintain OS platform services (init, services, networking, security policies) and related tooling. Build and operate image and update pipelines, ensuring reliability, reproducibility, and rollback safety. Instrument and analyze performance using profiling and tracing; optimize CPU, memory, I/O, and power usage. Own platform observability and reliability: logging, crash capture, watchdogs, and diagnostics. Collaborate with cross‑functional teams to define interfaces and deliver end‑to‑end features. Establish strong engineering practices: code review, CI, reproducible builds, and release management. Partner with external suppliers to support builds and deployments. You might thrive in this role if you: Have shipped production systems software on modern operating systems. Are proficient in C/C++ and a scripting language, and comfortable with OS internals (concurrency, memory management, filesystems, networking, power management). Bring strong systems debugging skills using debuggers, tracers, profilers, and logs across kernel/user‑space boundaries. Understand configuration of platform services and interfaces, and can translate requirements into stable, well‑documented APIs. Are fluent in u

AWSRestAIC++
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the role We’re looking for an engineering manager to lead a team building software systems that detect and prevent harmful misuse of frontier AI models—before incidents occur. This is a builder’s role: you’ll lead engineers shipping production services, detection pipelines, and mitigation mechanisms that protect frontier model integrity and reduce high-severity misuse risk. While this work intersects with frontier model development, security and risk, we’re explicitly seeking someone with a software engineering foundation who is comfortable building reliable systems that can operate at billions of users scale. In this role you will: Lead a team of software engineers building detection + mitigation systems for frontier model misuse, with an emphasis on model IP protection / distillation detection and emerging risk surfaces from autonomous agents. Set the technical roadmap and execution strategy: prioritize, design, ship, iterate, measure impact. Build production systems: services, pipelines, tooling, instrumentation, and automation that scale with frontier model usage. Partner deeply with Research and Product to translate evolving model capabilities into concrete tests, signals, and mitigations that can be deployed at scale. Drive strong engineering fundamentals: architecture, reliability, monitoring, performance, and operational excellence. Hire and grow an exceptional team across backend, data systems, and applied ML engineering domains as needed. Anticipate what breaks at scale as agentic workflows become more capable. You might thrive in this role if you: Experience building systems in adversarial, fast-evolving environments Are comfortable with ambiguity and novelty Have experience adjacent to security (e.g., abuse prevention, fraud, integrity, platform defense, auth/identity, malware/spam, adversarial environments) Communicate clearly and build trust quickly with senior stakeholders—pragmatic, collaborative, and calm under scrutiny. Significant experience

AWSRestAIRust
🔔

Get new reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime