Jobs in United States

Reliability Engineer Iii in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer iii jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team The Applied organization brings OpenAI’s most advanced technology to the world through products like ChatGPT and the APIs that power a growing ecosystem of developer and enterprise applications. Data Engineering builds and operates the trustworthy, secure, and reliable data systems that power decisions across OpenAI. About the Role We’re looking for a Data Engineering Manager to lead the Growth & Revenue data engineering team. This leader will own the data strategy and execution for the data subject areas spanning growth accounting across all product surfaces, product partnerships, checkout, billing, payments, revenue, and monetization, helping OpenAI understand how people adopt, engage with, and pay for our products. You will partner closely with several Data Science, Business, and Engineering partners to connect product behavior to trustworthy subscriber, payment, and revenue measurement. In this role, you will: Build, manage, and grow a high-performing, inclusive team across the Growth & Revenue data subject areas. Define the data strategy for all the data subject areas you own. Deliver durable, well-modeled data products that connect product behavior, subscription state, checkout events, payment outcomes, and revenue. Establish trusted metric definitions and data quality standards so product, growth, finance, and executive leaders can make fast, consistent decisions. Partner with Data Science and Product teams to support experimentation, causal measurement, funnel analysis, and scalable self-serve analytics. Partner with Finance and Financial Engineering to ensure analytical revenue views reconcile to financial truth and production billing systems. Raise operational excellence for critical pipelines, including reliability, observability, privacy, governance, and incident response. Set a clear roadmap, make principled tradeoffs, and communicate progress and risk across technical and business stakeholders. You might thrive in this role if yo

PythonSQLAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team builds next-generation AI-native silicon and systems while working closely with software, research, and manufacturing partners to co-design hardware tightly integrated with AI models. In addition to delivering systems for OpenAI’s supercomputing infrastructure, the team develops the tools, methodologies, and strategic partnerships needed to accelerate hardware innovation. About the Role We’re seeking an experienced Hardware Strategic Sourcing Manager to own sourcing strategy and supplier partnerships for fiber and optical interconnect components across OpenAI’s next-generation AI infrastructure. Reporting to the Head of Partnerships & Strategic Sourcing, you will lead sourcing across fiber cable assemblies, internal optical harnesses, fiber shuffles, optical backplane assemblies, connectorized and standalone passive optical assemblies, fiber-array units (FAUs), fiber-to-chip and coupling interfaces, detachable connectors, optical routing, and assigned optical packaging, assembly, and test services. You will work closely with electrical engineering, optical engineering, systems engineering, mechanical and packaging engineering, quality, rack integration, data-center deployment,manufacturing, supply chain, finance, legal, and program management teams to translate demanding bandwidth, signal integrity, reliability, and scale requirements into resilient supplier partnerships and scalable commercial strategies. Your work will directly support the performance, reliability, manufacturability, and scale of the high-speed optical connectivity required for OpenAI’s next-generation AI systems. In this role, you will: Develop and execute a comprehensive sourcing strategy for fiber and optical interconnect components supporting high-bandwidth AI systems and infrastructure. Own sourcing across optical fiber cable assembli

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team The Statsig team is responsible for the experimentation, feature rollout, dynamic configuration, and analytics systems that help OpenAI ship products with speed, safety, and evidence. Teams across ChatGPT, Codex, model measurement, monetization, business subscriptions, developer products, and shared infrastructure rely on Statsig to introduce capabilities safely, measure their impact, and make high-confidence product decisions. About the Role As a Product Lead on the Statsig team, you will define how experimentation, rollout, configuration, and analytics become a simple, reliable, and trusted part of how every OpenAI product team ships. You will set strategy across multiple product and platform workstreams, translate company-wide needs into durable capabilities, and help Statsig become a core part of OpenAI’s product development system. We’re looking for a product leader who combines strong product judgment, technical fluency, and deep analytical thinking. You should be comfortable navigating ambiguous customer needs, influencing teams across the company, and balancing rapid adoption with reliability, usability, and measurement quality. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Travel requirements should be confirmed with the recruiter before publishing. In this role, you will: Define the product vision, strategy, and roadmap for experimentation, feature management, dynamic configuration, rollout safety, and analytics. Partner with product, engineering, research, data, design, and infrastructure leaders to turn recurring launch and measurement needs into reusable platform capabilities. Develop a deep understanding of workflows across ChatGPT, Codex, model measurement, monetization, subscriptions, and developer products, then establish clear priorities across competing needs. Drive adoption by making sophisticated experimentation and analytics c

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. The Power Execution team owns the strategy and execution required to secure reliable, scalable, and economically resilient power for OpenAI’s global data center portfolio. The team sits at the intersection of commercial, technical, policy, legal, and operational work, partnering across OpenAI and with utilities, grid operators, regulators, counterparties, and public-sector stakeholders. About the Role The Energy Regulatory Lead will own energy regulatory strategy and execution for OpenAI’s infrastructure growth. This role will be the primary bridge between the Power Execution team and Public Policy and Government Affairs on energy regulatory matters, ensuring that OpenAI’s external engagement is grounded in project realities and that changing policy and regulatory conditions are translated into actionable infrastructure decisions. This is an individual contributor lead role and does not have direct reports initially. The role combines portfolio-level regulatory positioning with transactional regulatory work: evaluating jurisdictional pathways, supporting utility and energy transactions, coordinating approvals and filings, and helping project teams navigate tariffs, interconnection, load-service requirements, market rules, and regulatory risk from diligence through execution. In this role, you will: Develop and maintain OpenAI’s energy regulatory strategy across priority U.S. markets and, as needed, emerging geographies for infrastructure expansion. Coordinate closely with Public Policy and Government Affairs to shape energy regulatory priorities, engagement plans, messaging, and positions before utilities, public utility commissions, grid operators, state energy offices, and other relevant policymakers. Translate project requirements—load size, timing, reliability, cost, carbon, and expansion needs—into clear regulatory objectiv

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s People team hires, engages, and retains world-class talent to safely build and deploy AGI that benefits all of humanity. The People Analytics team helps leaders make rigorous, evidence-based talent decisions and ensures that the systems supporting those decisions are valid, reliable, fair, and accountable. About the Role As a People Data Scientist focused on AI fairness and bias testing, you will help establish how OpenAI evaluates AI-assisted People systems and high-impact talent processes. You will design and conduct rigorous assessments to identify, measure, and mitigate potential bias across the lifecycle of models, agents, decision-support tools, and automated workflows. Your work will span the entire employee life-cycle, such as hiring, performance, promotion, employee development, workforce planning, etc. You will evaluate both technical systems and the broader human-AI decision processes in which they operate, examining not only model performance but also data quality, measurement validity, differential outcomes, human oversight, and unintended consequences. We’re looking for an experienced data scientist or applied researcher who can translate complex fairness questions into defensible evaluation strategies, scalable testing infrastructure, and clear recommendations for technical teams and senior leaders. This role is preferred to be based in San Francisco, CA. In this role, you will: Define and lead fairness and bias-testing strategies for AI-assisted People processes, models, agents, and decision-support systems from development through deployment and ongoing monitoring. Design rigorous algorithmic audits and validation studies, including adverse-impact analysis, subgroup and intersectional evaluation, error-rate analysis, calibration, measurement invariance, reliability, criterion-related validity, and sensitivity testing. Identify the appropriate fairness criteria for each use case, evaluate tradeoffs among competing definitions

PythonSQLAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Industrial Compute team is building and productizing infrastructure capabilities that help organizations deploy and operate advanced AI systems at scale. The team works across AI hardware, systems engineering, physical infrastructure, and customer delivery to turn emerging technologies into reliable, repeatable infrastructure solutions. Our work sits at the intersection of technical strategy, product development, engineering, and deployment. We partner closely with customers and internal engineering teams to solve complex infrastructure challenges spanning compute, power, cooling, controls, and facility efficiency. About the Role We are seeking a senior, hands-on Data Center Infrastructure Architect to develop and optimize the physical infrastructure required for large-scale AI deployments. This is a broad technical role spanning data center architecture, electrical and mechanical systems, high-density compute, controls, telemetry, and digital modeling. You will use simulation, operational data, and digital-twin approaches to evaluate infrastructure designs, identify system-level constraints, and improve efficiency, reliability, cost, and speed of deployment. The ideal candidate can move fluidly between first-principles analysis, facility and equipment design, computational modeling, engineering review, and real-world implementation. You should be comfortable working across disciplines rather than operating solely within electrical, mechanical, or software boundaries. Key Responsibilities Define system-level architectures for high-density AI data centers across power, cooling, IT equipment, controls, and facility infrastructure. Develop digital twins and other computational models that represent the behavior of data center systems under changing workloads, environmental conditions, equipment configurations, and failure scenarios. Use design and operational data to identify constraints, improve PUE and related efficiency metrics, and optimize

PythonAWSGitRest
G
📍 United States· Full-time
✓ High-confidence listingCompany trend -100%

From $128K/yr

Quick readStrong listing-quality and freshness signals

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This position may be a hybrid or fully remote position, as decided by your manager. If designated as hybrid, you’ll divide your time between working remotely from your home and an office location, so you should live within commuting distance. If designated as remote, you’ll be working remotely from your home and may occasionally visit a GoDaddy office to meet with your team for events or meetings. Your hiring manager can share more about this role’s hybrid or remote designation. This position is not eligible to be performed in Alaska, Mississippi, North Dakota, or the Virgin Islands. GoDaddy is not currently considering candidates for this role in California, Seattle, or NYC. Join Our Team Join a team powering secure, scalable email services for millions of customers worldwide! As part of GoDaddy's Professional Email team, you'll solve complex challenges in distributed systems, cloud infrastructure, security, and AI while modernizing critical platforms that businesses rely on every day. If you enjoy owning impactful systems, working across a diverse technology stack, and building innovative solutions at scale, you'll feel right at home here. What you'll get to do... Design, build, and maintain highly available, scalable APIs and services used by millions of customers Deploy, manage, and optimize cloud infrastructure in AWS Architect and implement modern solutions that improve performance, reliability, and security Leverage AI technologies to enhance development workflows and create innovative customer experiences Monitor, troubleshoot, and resolve complex production issues using modern observability and monitoring tools Drive continuous improvement through automation, modernization, and operational excellence C

PythonAWSCI/CDGit
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA builds the silicon behind AI, accelerated computing, and graphics. Every watt of performance and every degree of thermal headroom traces back to decisions made in power, performance, and thermal architecture. We are the Silicon Co-Design Group (SCG). We identify, own, and drive system-level co-design ideas. We start with initial concepts and advance to product differentiation across NVIDIA's roadmap. We are hiring a Principal System Power Management and Performance Architect who operates at the ambiguous boundary where workload behavior, silicon capabilities, firmware policies, and platform constraints collide, and who turns that ambiguity into architecture that survives across multiple silicon generations. SCG scope spans architecture, design, software, operations, platforms, and productization. This role shapes system, platform, and data center features and behavior, and partners with teams across NVIDIA. What You'll Be Doing: The work here is rarely well-defined when it arrives. You will be given problems that appear to be performance gaps or power anomalies and encouraged to build a framework for solving them, not just tackle a single instance. Define the multi-generation roadmap for system-level power and performance features, grounded in prototyping, use-case analysis, and cost/benefit trade-offs across segments. You will decide what problems are worth solving and why. Own the architecture and integration strategy for HSIO power management, DVFS, P-states, and low-power features. Your decisions improve product performance, power, and reliability across product lines — not just the current program. Lead system-level boot and IST architecture defining how power and clock domains initialize, sequence, and recover across complex multi-IP systems where the interaction space is large and the failure modes matter. Drive power management strategy at data

C
📍 Fredericksburg, United States
✓ Quality checkedCompany trend +340.2%

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary Starting Pay: Based on experience Scheduled - Monday-Thursday Hours - 3pm-1am Position Summary The Maintenance Technician is responsible for maintaining, troubleshooting, repairing, and improving equipment within a fast-paced, mid-to-large volume distribution or manufacturing facility. This role supports operational efficiency by performing preventative maintenance, diagnosing equipment issues, reading schematics, conducting voltage testing, and repairing conveyor systems and other facility equipment. The ideal candidate is mechanically inclined, safety-focused, and skilled at identifying and resolving equipment failures to minimize downtime. What You'll Do • Perform preventative maintenance on conveyors, machinery, and facility equipment to ensure optimal performance and reliability. • Troubleshoot mechanical, electrical, pneumatic, and conveyor system issues to quickly identify root causes and implement repairs. • Read and interpret electrical schematics, blueprints, wiring diagrams, and technical manuals. • Utilize electrical testing equipment to conduct voltage readings and diagnose electrical system issues. • Inspect equipment regularly for signs of wear, damage, or potential failures and recommend corrective actions. • Execute repairs on conveyor systems, motors, gearboxes, drives, sensors, and other product

B
📍 Berkeley, United States
✓ Quality checkedCompany trend +515.8%

Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations organization is currently seeking a Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking an experienced cloud infrastructure professional to help design, build, and sustain the foundational cloud environment supporting critical program needs. In this role, the selected candidate will help establish and operate secure, scalable, and resilient cloud infrastructure environments in Microsoft Azure to enable enterprise applications, software toolchains, and digital engineering workloads. As both an individual contributor and technical leader, this position will work across network, computer, storage, identity, security, and automation domains to deliver repeatable cloud infrastructure patterns and operational excellence. This role is focused on infrastructure operations, sustainment, automation, and reliability, rather than application software development. Position Responsibilities: Design, implement, and maintain Microsoft Azure-based infrastructure solutions including networking, compute, storage, identity integration, and supporting services Develop and maintain Infrastructure as Code (IaC) and configuration automation solutions using Terraform, Ansible, PowerShell, and Bash Implement cloud policies to enforce security, ensure regulatory compliance, and manage user access Build repeatable landing zones and cloud infrastructure patterns that support mul

AzureTerraformAnsibleSap
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $263.7K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. WHY DATA SCIENCE & ANALYTICS? The Data Science & Analytics organization's mission is to increase our speed, frequency and acumen of making decisions at scale by instilling a data-influenced approach to building products. We cover a wide area of the data spectrum including analytical data engineering, product analytics, experimentation, causal inference, statistical modeling and machine learning. Aligned and partnering with product groups, we use this vast tool belt to discover new opportunities and unmet use cases, influence and shape the product roadmap and prioritization, build data products and measure the impact on our community of players and developers. WHY Consumer Apps? Roblox is used by tens of millions of people every day across a wide range of devices, including mobile, desktop, console, TV, and VR. The Consumer Apps team owns the app foundation that makes Roblox feel fast, fluid, and reliable for players and is dedicated to delivering superior performance, reliability, and user experience across all platforms Roblox supports. This team ensures a seamless and engaging user interface that facilitates intuitive interactions while enabling efficient, high-quality experiences

PythonSQLAWSGit
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as &#34;the AI computing company.&#34; We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent Senior Engineering Manager to lead a large firmware engineering organization delivering end-to-end manageability firmware for NVIDIA's next generation Data Center Compute Systems. This role owns HGX product line and OpenBMC-based management firmware and MCU firmware components in data center platforms, including architecture, execution, quality, reliability, telemetry, and customer readiness. We are seeking an experienced senior leader with strong technical depth, broad system perspective, and a proven ability to lead large teams through complex product cycles. This role is onsite in Santa Clara, CA, USA. If you're creative and autonomous, we want to hear from you! What you'll be doing: Lead a large firmware engineering organization delivering OpenBMC based firmware and MCU firmware for next-generation Data Center Compute Systems. Own HGX platform as a lead for Firmware and System software readiness working across the organization. Define and drive the long-term firmware roadmap, balancing architectural innovation with product execution and delivery milestones. Drive architecture strategy across BMC, MCU, platform software, manageability, health management, and data center firmware interfaces. <spa

PythonGitLinuxArtificial Intelligence
D
📍 Boston, Massachusetts, United States
✓ Quality checkedCompany trend -83.5%

As an Enterprise Customer Success Manager, you will proactively drive new product attachment and effective strong relationships across our largest and most strategic customers. You’ll advocate for the customer internally and focus on a positive customer experience. Interactions are rooted in relationship-management, first and foremost, while also advocating for growth opportunities. Enterprise Customer Success Managers follow a well-defined methodology that helps them identify the customer's unique needs and clearly convey the value of the Datadog product. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Act as a strategic partner to customers, orchestrating cross-functional internal teams and engaging executive, technical, and business stakeholders to understand customer goals and translate them into a clear, deliverable Datadog value narrative. Proactively build and maintain executive relationships to deliver clear, outcome-driven value stories that connect Datadog technical use cases to measurable business results. Lead QBRs and strategic reviews as a forum to demonstrate impact, align on priorities, and define next-step initiatives. Analyze adoption and usage trends to quantify value delivered, extract insights from large datasets, identify gaps, and drive financially grounded commercial recommendations and strategic opportunities. Position Datadog as a critical observability platform that enables reliability, efficiency, and informed decision-making. Own and project manage the on-boarding process for new customers Collaborate cross-functionally with AEs, SEs, TAM, Product, Support, Enablemen and other technical teams to ensure consistent value delivery and messaging. Who You Are: Customer-centric with 3+ years in a Customer Success

S
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -91.7%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is a high-growth SaaS observability platform built on the Snowflake AI Data Cloud, enabling businesses to troubleshoot modern distributed applications 10x faster. Now, as a core part of Snowflake, we’ve reached a major milestone in the evolution of the Snowflake platform. By bringing AI-powered observability directly into the Snowflake ecosystem, we’ve created the first truly unified platform for telemetry and business data. We’re looking for a Technical Account Manager to partner with our most strategic enterprise customers and ensure they derive sustained operational value from Observe. This is a hands-on, post-sales technical role focused on long-term platform adoption, optimization, and technical partnership. You will work directly with SRE, DevOps, platform, and engineering teams to embed Observe into daily workflows, evolve telemetry strategy over time, and continuously improve reliability, performance, and cost efficiency. This role is ideal for an experienced observability practitioner who enjoys being deeply embedded with customer teams, solving real production challenges, and acting as a trusted technical advisor in complex enterprise environments. What You’ll Do Serve as the primary technical owner and trusted advisor for assigned strategic a

AWSAzureGCPKubernetes
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -79.2%

About Team Our Robotics team is focused on unlocking general-purpose robotics and advancing toward AGI-level intelligence in dynamic, real-world environments. Working across the full model and systems stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the physical constraints of real-world systems to improve people’s lives. About the Role We are looking for a TPM to drive development and integration of a range of sensor systems for robotics. This role will drive cross-functional alignment across requirements, engineering design, integration, validation, manufacturing, supply chain, and release processes, helping turn complex sensor-system needs into clear plans, decisions, and milestones. Location and in-person expectations: This role is based in San Francisco, CA and requires in-person presence 4 days a week. In this role you will: Drive requirements alignment across engineering design, integration, testing, and validation for camera modules, LiDAR, IMUs, RADAR, proximity sensors, audio components and the systems they interact with. Coordinate the integration of modules including electrical, mechanical, harnessing, and software interfaces with the full robotic system with deep understanding of timelines to drive the respective PCBAs, enclosures, build and test fixtures, connectors and cables. Establish effective cadences for technical reviews, BOM readiness, change management, production releases, approvals, and decision tracking. Align harnesses, fasteners, assembly fixtures, test fixtures, and documentation so cross-functional teams can execute against a clear plan. Lead validation planning around functional, reliability, NVH failure modes, including testing needs, schedules, and exit criteria. Partner with manufacturing and supply chain to manage handoffs, lead times, dependencies, and production readiness. Drive tradeoff decisions across cost, qua

AWSRestAIRust
🔔

Get new reliability engineer iii jobs in United States by email

Daily job updates · Unsubscribe anytime