Jobs in United States

Reliability Engineer Iii in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer iii jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Platform and Infrastructure Engineering organization advances the mission of deploying artificial general intelligence (AGI) for the benefit of all by delivering secure, scalable, and resilient technology solutions. Our team builds and maintains robust infrastructure that safeguards OpenAI’s data and systems while ensuring employees are well-equipped and seamlessly connected. By prioritizing security, reliability, and user-centric solutions, we empower OpenAI employees to drive impactful AI research, corporate operations, and product innovation. About the Role As a Software Engineer: Internal Applications, Enterprise, you will build internal products that make technology support and administration safer, faster, and less dependent on manual intervention. You will help reduce reliance on broadly privileged human actions, turn recurring technology problems into paved paths, and build agentic systems that can help resolve tickets end to end. A core part of the role is building the interfaces that bring employees, AI agents, and human responders together in a shared ITSM experience, with the right context, controls, and handoffs at each step. We are seeking engineers who enjoy working across frontend and backend layers on ambiguous, high-leverage enterprise problems. You should bring strong product judgment, solid backend engineering fundamentals, and an interest in building software that changes how technology support, system administration, and agent-assisted operations are delivered. The best fit will care as much about the quality of the operator and employee experience as the correctness of the backend systems behind it. In this role, you will: Build frontend experiences that let employees request help, let agents gather context and take safe actions, and let human responders review, approve, or take over without losing the thread. Reduce reliance on broadly privileged manual actions by replacing them with narrow, auditable, policy-aware aut

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%

From $230K/yr

Quick readStrong listing-quality and freshness signals

About the Role The Engineering Acceleration Delivery / Continuous Deployment team builds and operates the systems that safely ship OpenAI’s infrastructure and product code to production. We own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across OpenAI to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale infrastructure orchestration. In This Role, You Will Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces, and alerts to detect regressions and trigger safe rollbacks. Build systems that support agent-assisted or autonomous deployment workflows using modern AI tooling. Technologies commonly used in this environment include: Kubernetes for large-scale container orchestration and runtime infrastructure Python and FastAPI for internal services Terraform for infrastructure as code GitOps-based deployment workflows (e.g., ArgoCD, Flux, or similar systems) Buildkite for CI orchestration You may be a strong fit if you: Have worked with Kubernetes-based deployment systems at scale Have experience building or operating continuous deployment platforms Are familiar with GitOps tooling such as

PythonAWSKubernetesGit
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Platform Engineer An individual contributor who will serve as a hands-on technical member of the SMAI Platform Engineering team. In this role, you will help build, operate, and maintain the platforms that power Micron's analytics and AI workloads. You will contribute to platform reliability and scalability through day-to-day engineering, collaborative problem solving, and close partnership with solution architects, project teams, and multi-functional partners. Responsibilities: Collaborate with global platform teams, customers, partners, and vendors to deliver effective technical solutions. Know the latest platform roadmaps, emerging technologies, and new service offerings; evaluate and recommend adoption opportunities. Partner with solution architects to design, implement, and optimize solutions across IaaS, PaaS, SaaS, and Infrastructure as Code (Terraform). Document findings, operational procedures, guidelines, and reusable patterns while providing feedback to vendors and internal teams. Deliver high-quality platform support by managing customer requests, maintaining service standards, and implementing controlled platform adjustments. Monitor platform performance, observability, costs, and resource utilization to identify optimization opportunities and improve reliability. Collaborate with multi-functional teams to ensure seamless operations, scalable architectures, and automation, including AI-driven business solutions. Implement and m

KubernetesMachine LearningAITerraform
B
📍 Raleigh, North Carolina, United States
✓ High-confidence listingCompany trend +350%
Quick readStrong listing-quality and freshness signals

This is where your work makes a difference. At Baxter, we believe every person—regardless of who they are or where they are from—deserves a chance to live a healthy life. It was our founding belief in 1931 and continues to be our guiding principle. We are redefining healthcare delivery to make a greater impact today, tomorrow, and beyond. Our Baxter colleagues are united by our Mission to Save and Sustain Lives. Together, our community is driven by a culture of courage, trust, and collaboration. Every individual is empowered to take ownership and make a meaningful impact. We strive for efficient and effective operations, and we hold each other accountable for delivering exceptional results. Here, you will find more than just a job—you will find purpose and pride. Your Role at Baxter As a Principal DevOps Engineer, you will provide technical leadership in the design, implementation, deployment, and support of cloud and platform engineering solutions that enable product development teams. You will partner across engineering functions to build scalable, reliable, and secure infrastructure while helping drive operational excellence through automation, observability, and continuous improvement. What You'll Do: Design, implement, and support cloud infrastructure solutions that enable the development and operation of Baxter products and applications. Manage infrastructure, platform, and deployment processes for product teams to ensure reliability, scalability, and performance. Deploy, manage, and troubleshoot containerized applications using Kubernetes and cloud-native technologies. Develop and maintain Infrastructure as Code (IaC) solutions using Terraform to automate infrastructure provisioning and management. Implement and support observability, monitoring, and alert

AzureKubernetesTerraformRecruitment
P
📍 New York, NY, United States· Full-time· Hybrid
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Backend Software Engineers at Palantir build software at scale to transform how organisations use data. Our Software Engineers are involved throughout the product lifecycle, from idea generation, design, prototyping, and production delivery. You will collaborate closely with technical and non-technical teammates to understand our customers' problems and build products that solve them. We encourage movement across teams to share context, skills, and experience, so you'll learn about many different technologies and aspects of each product. Engineers work autonomously and make decisions independently, within a community that will support and challenge you as you grow and develop, becoming a strong technical contributor and engineering leader. Your day-to-day workflow will vary, adapting to the requirements of our users and the technical challenges that arise. One day, you may find yourself collaborating with other engineers to architect a new system that enables a novel workflow, the next you could be fine-tuning performance to enable low-latency operational outcomes. Our Product Development organisation is made up of small teams of Software Engineers. Each team focuses on a specific aspect of a product and work collaboratively to build cross functional capabilities, streamline user workflows and continuously improve our software's efficiency and reliability. We’re hiring engineers who are passionate about solving real-world problems and empowering both developers and end-users to work optimally. If you’re motivated to develop reliable, performant, and scalable systems, and to design robust APIs and primitives, this role offers the opportunity to make a

PythonJavaC++Supply Chain
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

PythonArtificial IntelligenceAI
O
📍 Atlanta, Georgia, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge As a Senior Staff Software Engineer, you will serve as a technical leader for OneTrust’s AI Governance (AIG) platform, driving the design, scalability, and reliability of systems that enable enterprises to deploy and govern AI and LLM-powered applications responsibly. You will deeply understand how customers build, deploy, and operate AI systems, and translate those needs into secure, compliant, and observable platform capabilities. Your Mission Development Lead the design and development of Java/Python microservices and shared libraries integrating with AI platforms for OneTrust’s AI Governance product. Design, build, and test cloud-native applications deployed on Microsoft Azure using Core Java, REST, and the Spring ecosystem. Lead the architecture and development of reusable AIG reporting and dashboard capabilities that integrate governance data from SQL databases and analytical platforms with runtime observability signals. Design reusable semantic-layer and metric-abstraction capabilities, including dataset contracts, metric defini

PythonJavaSQLAWS
O
📍 San Francisco, California, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge As a Senior Staff Software Engineer, you will serve as a technical leader for OneTrust’s AI Governance (AIG) platform, driving the design, scalability, and reliability of systems that enable enterprises to deploy and govern AI and LLM-powered applications responsibly. You will deeply understand how customers build, deploy, and operate AI systems, and translate those needs into secure, compliant, and observable platform capabilities. Your Mission Development Lead the design and development of Java/Python microservices and shared libraries integrating with AI platforms for OneTrust’s AI Governance product. Design, build, and test cloud-native applications deployed on Microsoft Azure using Core Java, REST, and the Spring ecosystem. Lead the architecture and development of reusable AIG reporting and dashboard capabilities that integrate governance data from SQL databases and analytical platforms with runtime observability signals. Design reusable semantic-layer and metric-abstraction capabilities, including dataset contracts, metric defini

PythonJavaSQLAWS
M
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend -75%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Department Intro: The Advanced Packaging Technology Development (APTD) department at Micron Technology is at the forefront of innovation, driving the advancement of memory and storage Interconnects and Packaging solutions that transform how the world uses information. Our team is dedicated to developing innovative processes and technologies that enable the creation of next-generation semiconductor products which drive the AI revolution. Position Overview: As a Staff Process Engineer in Advanced Packaging CMP team, you will be primarily responsible for starting up, developing and optimizing processes to improve product quality and reliability, working on process yield improvement, cost reduction, productivity improvement and risk management as well as resolving manufacturing line problems. You will also be required to identify, diagnose and resolve assembly process related problems by applying failure analysis, FMEA, 8D or SPC/FDC methodology. This position requires communication and collaboration with associates both within the region and internationally. Strict adherence to Micron Intellectual Property Protection policy is vital. This role will also extensively collaborate with vendors to develop processes that meet integration requirements. Responsibilities: Perform fundamental research, consumables development, and hardware evaluation, as well as test processes for novel applications Initiate and manage experiments to widen process margins and evaluate man

AIRecruitment
A
📍 California, United States
✓ High-confidence listingCompany trend +365.2%
Quick readStrong listing-quality and freshness signals

Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: At Lingo, we’re building a groundbreaking health platform that combines continuous biosensor data, real-time analytics, and personalized insights to help people live fuller, longer, and healthier lives. Our systems ingest millions of sensor readings daily, powering experiences for consumers and partners worldwide, with the reliability and scalability of cloud-native, enterprise-grade platforms. At Abbott, you can do work that matters, grow, and learn, care for yourself and family, be your true self and live a full life. You’ll also have access to: Career development with an international company where you can grow the career you dream of. Employees can qualify for free medical coverage in our Health Investment Plan (HIP) PPO medical plan in the next calendar year An excellent retirement savings plan with high employer contribution Tuition reimbursement, the Freedom 2 Save student debt program and FreeU education benefit - an affordable and convenient path to getting a bachelor’s degree. A company recognized as a great place to work in dozens of countries around the world and named one of the most admired companies in the world by Fortune. A company that is recognized as one of the best big companies to wor

ReactGraphqlAIKotlin
M
📍 Minnesota, United States of America, United States
✓ High-confidence listingCompany trend +1850%
Quick readStrong listing-quality and freshness signals

We anticipate the application window for this opening will close on - 9 Oct 2026 Careers that change lives start here. Medtronic is a global leader in healthcare technology with a Mission to alleviate pain, restore health, and extend life. Our 95,000 employees work across more than 150 countries to put patients first — developing innovative medical technologies that improve the lives of 72+ million patients each year. Your unique talents will help shape the future of healthcare while building a career grounded in purpose, growth, and impact. A Day in the Life The Process Engineer II will support the development of manufacturing processes and battery designs for New Product Development at the Medtronic Energy and Component Center (MECC) for Cardiac Rhythmn Management (CRM) business. MECC supports the design, development, and production of power components used in implantable devices. At Medtronic, we bring bold ideas forward with speed and decisiveness to put patients first in everything we do. In-person exchanges are invaluable to our work. We are working onsite a minimum of 4 days per week as part of our commitment to fostering a culture of professional growth and cross-functional collaboration as we work together to engineer the extraordinary. This role will require less than 10% travel to enhance collaboration and ensure successful completion of projects. Responsibilities may include the following and other duties may be assigned: Work with cross-functional teams including Design, Operations, Quality, and Reliability to develop processes for batteries in a regulated industry; ensure processes and designs are compatible and in compliance with regulations Contribute to the innovation, development and/or optimization of new manufacturing concepts, processes and procedur

RecruitmentHRCRM
I
📍 Arizona, Phoenix, United States
✓ High-confidence listingCompany trend +315.4%
Quick readStrong listing-quality and freshness signals

Job Details: Job Description: As a Facilities Mechanical Engineer, you will play a pivotal role in ensuring the reliable operation and maintenance of Intel's advanced mechanical systems, supporting our cutting-edge manufacturing, cleanroom, and research and development activities. Your work will directly impact Intel's ability to innovate and deliver world-class technologies by maintaining the critical infrastructure that enables peak operational performance. This role offers a unique opportunity to drive engineering excellence, collaborate with cross-functional stakeholders, and shape the future of Intel's facilities globally. Key Responsibilities: Ensure the availability, reliability, and maintenance of mechanical systems such as oil-free air systems, HVAC systems, chilled water plants, boilers, environmental abatement systems, and compliance exhaust systems. Support daily tactical efforts to meet safety, reliability, and environmental targets for mechanical systems. Design and analyze mechanical systems and equipment, troubleshoot systematic issues, and provide evaluations, recommendations, and solutions to resolve problems. Develop engineering scopes of work for mechanical projects and conduct design reviews to ensure system integrity and performance. Conduct feasibility studies and testing on new and modified designs, while overseeing prototype fabrication and design testing. Partner with internal business units and operations teams to deliver engineering solutions tailored to their specific needs.

Project ManagementRecruitment
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team Consumer Monetization builds the experiences and systems that power how customers purchase and pay for OpenAI products. Our scope spans purchasing flows, payments, subscriptions, and billing, along with the shared capabilities that support new products, offers, and distribution channels. We own both customer-facing experiences and the underlying platforms that power them. We partner closely with Product, Design, Growth, Data Science, and engineering teams across OpenAI to make purchasing effective and reliable, and new offerings easier to launch and monetize. About the Role We’re looking for experienced Staff+ engineers to evolve the payments, billing, and subscription capabilities that support OpenAI’s growing product portfolio. You’ll tackle problems where correctness, reliability, and flexibility are essential: supporting new billing requirements, managing the billing lifecycle, synchronizing state across internal systems and external providers, and enabling new products and commercial models. Your work may span several areas based on your expertise and team priorities: Billing and monetization capabilities: Extend billing capabilities and improve integrations and state consistency across systems to support new products and business models. Subscriptions: Orchestrate purchases, renewals, plan changes, cancellations, and recovery, ensuring customers are charged correctly and receive the right benefits. Payments: Expand payment capabilities through processor integrations, routing, and broader payment-method coverage. Risk and integrity: Partner with risk and integrity teams to integrate controls into purchasing flows, reducing abuse while protecting legitimate customer experiences. You’ll help set technical direction while remaining hands-on in implementation and delivery. This is an opportunity to solve complex engineering problems at scale, connect architecture decisions to customer and business outcomes, and help other engineers take on broader ow

Artificial IntelligenceAIFinance
H
📍 Texas, United States of America, United States
✓ High-confidence listingCompany trend +103.7%

$105.1K – $161.8K/yr

Quick readStrong listing-quality and freshness signals

NPI Mechanical Quality Engineer Description - As part of the Personal Systems Quality organization, the NPI Mechanical Quality Engineer drives product quality and reliability throughout the NPI product development lifecycle. This role partners with cross-functional teams to evaluate designs, identify and mitigate reliability risks, drive failure analysis and root cause investigations, and implement corrective actions that improve product quality and customer experience. This role provides technical leadership in reliability test development, mechanical design assessment, and Design for Quality and Reliability (DFX) initiatives to ensure high-quality products from concept through production. Strong mechanical engineering expertise, problem-solving skills, and team collaboration are essential for success in this role. *Onsite in Spring office is required 4-days a week Responsibilities: • Provides technical analysis of compute products during NPI Phase through RFQ design proposal reviews, CAD reviews, and DFX product tear downs. • Provides advanced failure analysis, root cause identification, and corrective/preventive actions for reliability concerns and/or issues during development, manufacturing, and field operations. • Proactively addresses customer requests or issues related to mechanical reliability, ensuring timely resolution while liaising with the relevant stakeholders. • Provides consultation to product development architecture teams regarding material component selection, design risk assessment, and associated reliability test plans. • Provides NUD (new, unique, difficult) risk mitigation and test plan development and develops issue prevention. • Actively mentors less-experienced engineers in the field of mechanical reliability and contributes to their growth in the organization. Education & Expe

MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. The TD Electrical Device Characterization team pushes the boundaries of Micron’s memory technology. We explore device behavior, uncover insights through data, and collaborate across engineering teams to enable next-generation products. We are a hands-on team that values curiosity, precision, technical rigor, and collaboration. In this role, you will drive electrical characterization for advanced memory technologies and partner with device, process, and product teams to turn measurements into actionable technical insights. Your work will help guide technology development and shape the performance and reliability of future NAND, DRAM, and emerging memory products. Responsibilities: Lead electrical characterization for advanced memory devices Develop measurement methods, algorithms, and scalable characterization standards Automate data pipelines and analysis using Python, JSL, C/C&#43;&#43;, and dashboard tools Architect test setups and manage wafer probe and ATE test flows Investigate device behavior, identify root causes, and communicate technical findings Minimum Qualifications: <span style="color:#6

PythonLinuxAIRecruitment
🔔

Get new reliability engineer iii jobs in United States by email

Daily job updates · Unsubscribe anytime