Jobs in United States

Reliability Engineer in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

MT
📍 Richardson, TX, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. The Staff Design Engineer is responsible for designing, simulating, and optimizing DRAM circuits while supporting the development of digital and analog circuitry used in advanced memory products. This role collaborates globally across design, verification, product engineering, test, probe, process integration, assembly, and marketing to ensure manufacturable, high‑quality, cost‑optimized solutions. The position drives innovation in future memory generations and contributes to silicon‑to‑system development in a dynamic engineering environment. Responsibilities Design digital, analog, and memory core circuits using CMOS logic and transistors, implementing device specifications from concept to solution. Utilize AI-Enabled tools to assist with schematic editing, data collection and layout floorplan. Create optimized floorplans for circuit placement, routing, power delivery, sense margins, array timing, and die size, including layout leadership. Conduct circuit simulations using FINESIM, HSPICE, and VERILOG; analyze power, performance, reliability, and parasitic impacts. Validate builds through reticle experiments, tape‑out revisions, and simulation‑to‑silicon correlation, identifying required schematic edits. Prepare and maintain design documentation while contributing to best‑known practices and departmental training. Collaborate with global build, verification, product engineering, test, probe, process integration, assembly, and marketing teams to ensure manufacturability and quality. <l

AIRecruitment
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

Are you ready to contribute to world-class innovation and push the boundaries of what's possible? At NVIDIA, you'll have the opportunity to be part of a team that is driving groundbreaking impacts across various markets. As a Thermal Solutions Development Engineer, you will play a pivotal role in our Silicon Codesign Group, transforming thermal solution concepts into lab-ready builds and beyond. What you will be doing: Build thermal solutions for engineering characterization and validation of next-gen GPU/SOC products, ensuring flawless delivery from concept to lab. Drive end-to-end development and deployment of thermal solutions, collaborating with internal teams and external vendors on build requirements, prototype evaluation, test system integration, and software automation. Improve thermal design processes by incorporating feedback and findings, developing workflow and maintaining our world-class standards. Work closely with system architects, chip and board designers, and software/firmware engineers in a dynamic and high-energy environment to bring industry-defining products to market. Apply AI-enabled approaches and AI tools to accelerate design iteration, test planning, and characterization/validation triage (e.g., requirements/spec summarization, experiment prioritization, log/telemetry summarization, anomaly/outlier detection), improving cycle time, coverage, and traceability while validating outputs against physics, specs, and lab measurements. Partner with AI/tooling teams as the thermal domain SME to define use-cases, success criteria, and evaluation methods; provide feedback to improve tool reliability and usability. What we need to see:

C
📍 United States· Remote
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary CVS Health's Adjudication & Client Experience Engineering organization is seeking a motivated and highly skilled Senior Analyst - Software Development Engineering to join our Application Production Support team. This role will support critical business applications by providing production support, troubleshooting technical issues, and contributing to ongoing application enhancements and stability improvements. As a Sr. Analyst, you will work closely with Lead Engineers, Software Development Engineers, Product Owners, QA teams, and business stakeholders to investigate and resolve production incidents, perform root cause analysis, and implement code fixes. You will be responsible for supporting enterprise applications built on Java, Angular, APIs, and Cloud platforms while ensuring the reliability and performance of systems that serve our PBM (Pharmacy Benefit Management) business. This role is ideal for a hands-on engineer who enjoys solving production challenges, developing software solutions, and collaborating within a fast-paced environment. The successful candidate will contribute to application support activities, system enhancements, and continuous improvement initiatives while growing their technical and business domain expertise. Required Qualifications 5-8 years of experience in software development, application s

TypeScriptJavaAngularSQL
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

We are seeking a highly skilled and hard-working Senior Test Developer / test engineer to join our multifaceted Enterprise Software QA team. This role offers an outstanding opportunity to leave your mark on the design, construction, optimization and testing of large-scale infrastructure for various foundational NVIDIA unified cloud services and data center offerings. If you are a dedicated engineer with strong expertise in cloud infrastructure and distributed systems and want to apply your skills with AI tools, this role could fit you perfectly. You will thrive in an exciting, innovative environment. What you'll be doing: Work with development teams on test plans for all layers of SW stack for cloud infrastructure, execution, reviews, failure analysis and assessing overall quality and risk. Work with customer PMs on software issues including technical feedback from OEMs and CSPs. Develop key benchmarks to track execution and deploy process improvements to improve efficiency Leverage AI skills to expedite the test scope, test plan, execution and automation workflows. Lead NVIDIA Cloud and Data Center bring up activities which will involve validation, reporting, working with engineering to debug issues, providing design input at times, adding coverage in different areas. Design, develop and maintain CI/CD pipelines for continuous testing in cloud environments when needed. Perform performance, scalability, and reliability testing of cloud services. Implement and maintain test environments in cloud platforms such as AWS, Azure, or Google Cloud. Supervise the infrastructure to alert on significant events, ensuring the highest level of system performance and reliability. Work with various different partner teams to ensure availability of clusters to test on and take the lead in resolve all issues. Working with tea

AWSAzureDockerKubernetes
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable &#34;flight recorder&#34; for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as &#34;Hardware Fault,&#34; &#34;Software Bug,&#34; or &#34;Environment Issue.&#34; This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define &#34;signals of impending failure,&#34; enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6&#43; years in systems programming. Experience building automated

PythonKubernetesLinuxMachine Learning
M
📍 Arizona, United States of America, United States
✓ High-confidence listingCompany trend +1850%
Quick readStrong listing-quality and freshness signals

We anticipate the application window for this opening will close on - 6 Oct 2026 Careers that change lives start here. Medtronic is a global leader in healthcare technology with a Mission to alleviate pain, restore health, and extend life. Our 95,000 employees work across more than 150 countries to put patients first — developing innovative medical technologies that improve the lives of 72&#43; million patients each year. Your unique talents will help shape the future of healthcare while building a career grounded in purpose, growth, and impact. A Day in the Life At Medtronic, we bring bold ideas forward with speed and decisiveness to put patients first in everything we do. In-person exchanges are invaluable to our work. We’re working onsite 5 days a week as part of our commitment to fostering a culture of professional growth and cross-functional collaboration as we work together to engineer the extraordinary. This Principal Process Development Engineer will be responsible for the development of laser based processes through release to manufacturing. The Engineer will lead process development and process improvement projects in laser joining and laser cutting of metal and glass materials. Process development work scope may span across multiple process areas with a primary focus in laser equipment. They will coordinate risk burn down and problem-solving experiments by utilizing DRM/DFSS (Design and Reliability for Manufacturing / Design For Six Sigma) and data driven methods. They will also be responsible for managing the validation activities for the development work including requirements flow down, effective control and monitoring strategies and ensuring compliance to regulations and safety for our patients. Responsibilities may include the following and other duties may be assigned.

RecruitmentHR
M
📍 Arizona, United States of America, United States
✓ High-confidence listingCompany trend +1850%
Quick readStrong listing-quality and freshness signals

We anticipate the application window for this opening will close on - 6 Oct 2026 Careers that change lives start here. Medtronic is a global leader in healthcare technology with a Mission to alleviate pain, restore health, and extend life. Our 95,000 employees work across more than 150 countries to put patients first — developing innovative medical technologies that improve the lives of 72&#43; million patients each year. Your unique talents will help shape the future of healthcare while building a career grounded in purpose, growth, and impact. A Day in the Life At Medtronic, we bring bold ideas forward with speed and decisiveness to put patients first in everything we do. In-person exchanges are invaluable to our work. We’re working onsite 5 days a week as part of our commitment to fostering a culture of professional growth and cross-functional collaboration as we work together to engineer the extraordinary. This Principal Process Development Engineer will be responsible for the development of material finishing, material handling and automated handling systems and processes through release to manufacturing. The Engineer will lead equipment and system development and process improvement projects to support various component and device handling, finishing, and assembly needs in new and current manufacturing lines. Process development work scope often spans across multiple process areas including wafer processing functional areas, component assembly, laser processes, and device test and finishing processes. They will coordinate risk burn down and problem-solving experiments by utilizing DRM/DFSS (Design and Reliability for Manufacturing / Design For Six Sigma) and data driven methods. They will also be responsible for managing the validation activities for the development work including requirements flow down, effective control

RecruitmentHR
DC
📍 New York, New York, United States· Full-time
✓ High-confidence listing

From $131K/yr

Quick readStrong listing-quality and freshness signals

Help shape the technology that enables a global organisation to do its best work. As Senior Manager, Platform Engineering, you’ll lead the team responsible for Diligent’s Atlassian and Microsoft platforms while setting the architectural direction for the wider internal IT estate. You’ll combine people leadership, enterprise platform strategy and hands-on technical judgement to create secure, reliable and scalable experiences for employees worldwide. From modernising service management and automating joiner, mover and leaver processes to enabling AI safely through Microsoft Copilot and Atlassian Rovo, your work will reduce friction, strengthen governance and deliver measurable business impact. Working across IT, Security, HR, Finance, Legal, Compliance and business teams, you’ll turn complex requirements into well-governed platforms that are easy to use, resilient and ready for the future. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead, coach and grow a global team of platform engineers and systems administrators, building a high-performing and inclusive culture. Own the strategy, architecture, governance and roadmap for Atlassian Cloud, including Jira, Jira Service Management, Confluence, Atlassian Guard and Rovo. Set the direction for Diligent’s Microsoft 365 E5 estate, including Teams, SharePoint, Exchange Online, Intune, Defender, Purview, Power Platform and Copilot. Design scalable integration and automation patterns across identity, HRIS, ITSM and business systems using APIs, event-driven automation, Okta Workflows, Power Platform and scripting. Partner with IT Support to improve self-service, automate repetitive work and reduce ticket volume, escalation effort and time to resolution. Establish strong standards for security, access governance, AI adoption, reliability, compliance and business continuity across the internal technology estate. These are the essentials you’ll need to get an interview Significant experience in i

PythonAWSGitAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxAIC++
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxAIC++
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debug compl

PythonLinuxAIC++
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxAIC++
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu

PythonLinuxAIC++
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team The Emerging Products team is a lean, high-output product lab group that builds products at the forefront of model capabilities. We collaborate across all teams within the company, from research and infrastructure to consumer products. The team is responsible for identifying new product opportunities, building them quickly, dogfooding them internally, and then launching the successful products to users. We use data, user research, and analytics to inform our ideas, and make decisions on what experiments are worth iterating, stopping, or scaling. About the Role We’re looking for a senior, product-minded software engineer to own ambiguous 0-to-1 work from idea through prototype, validation, and handoff. This is a full-stack role with a strong frontend and product emphasis: you will build the interfaces and supporting backend systems needed to test new experiences quickly, while making sound architectural choices that enable successful concepts to scale. This role is based in our Mission Bay office in San Francisco. In this role, you will: Build and ship high-quality, product experiments across the full stack. Turn ambiguous user needs and emerging technical capabilities into testable product concepts, using research and metrics to guide iteration. Own technical direction for 0-to-1 projects, balancing speed, reliability, and a clear path from prototype to scalable product. Partner closely with design, product, research, and engineering teams to dogfood, evaluate, launch, and transition successful experiments. You might thrive in this role if you: Have a track record of building and shipping end-to-end products in fast-moving, startup, founder-led, growth, or other high-ownership environments. Bring strong frontend engineering skills and enough backend and systems depth to make sound full-stack architectural decisions. Pair product intuition with evidence, using user research and product data to identify opportunities and make pragmatic tradeoffs. Operat

AWSRestAIRust
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA is seeking a Senior Software Engineer to help us develop distributed storage services for AI/ML. In this role you will work closely with the broader NVIDIA team to design and build a reliable, scalable, and efficient storage-as-a-service tailored to AI applications that can be deployed anywhere and scale without limitations. This service supports the whole NVIDIA critical business from graphics drivers to autonomous vehicles to deep learning frameworks. To achieve this goal, we are looking for an engineer with a deep understanding of distributed systems, outstanding design skills, and a track record in building and delivering large-scale distributed services. What you will be doing: Leading the overall architecture and design of our distributed storage service optimized for AI/ML Develop and maintain distributed, robust and scalable Go programs deployed to state of the art open-source ecosystems, including Kubernetes. Develop and maintain user-space applications, containers, Go-bindings, and CLI tools. Building features for a distributed storage service to enhance availability and reliability for large-scale deployments Engaging and collaborating with NVIDIA Research, Computing, Product teams, cross-functional teams, and external customers to deliver Cloud services. Automating distributed storage service end-to-end, including deployment, management, and monitoring What we need to see: Bachelor’s of Science in Computer Science, or related field (or equivalent experience) with 8&#43; years of industry experience Strong background in developing distributed systems involving Golang, Kubernetes, and Cloud Service Provider integrations Strong track record of delivering distributed services in a variety of distributed computing environments Experience in i

KubernetesArtificial IntelligenceAIGolang
🔔

Get new reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime