Jobs in United States

Reliability Engineer in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

M
📍 Boise, ID - SIG Building, United States
✓ Quality checkedCompany trend -75%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Department Introduction The Systems Integration Group (SIG) develops advanced test solutions that support semiconductor manufacturing and product quality. The team partners across engineering, operations, procurement, and external suppliers to design, build, and deploy innovative hardware platforms that enable reliable testing and validation of semiconductor devices. Join the team building semiconductor memory test platforms for tomorrow’s AI infrastructure. Position Overview As an Electrical Design Engineer, you will design and develop electronic hardware used in advanced semiconductor test systems. You will apply electrical engineering principles to create analog, digital, and mixed-signal solutions that support manufacturing and product development. This role spans the full hardware development lifecycle, from concept and design through validation, deployment, and sustaining support. You will collaborate with cross-functional teams around the world to deliver reliable, high-quality test solutions. Responsibilities Design analog, digital, and mixed-signal electronic circuits from concept through implementation and validation Evaluate and select electronic components to meet system performance, reliability, and cost objectives Develop schematics and collaborate with PCB designers to create optimized board layouts Perform circuit simulation, analysis, and design verification using industry-standard tools Debug, test, and val

AIProcurementRecruitment
H
📍 Texas, United States of America, United States
✓ High-confidence listingCompany trend +103.7%

$123.1K – $150.4K/yr

Quick readStrong listing-quality and freshness signals

Software Product Security Engineer Description - This role supports the development and maintenance of secure software products under the guidance of senior engineers. The position focuses on learning software engineering and security best practices while contributing to the design, implementation, testing, and maintenance of desktop, web, and cloud-based applications and services. Key Responsibilities Assist in developing, testing, and maintaining software applications and security solutions. Participate in software development activities including coding, debugging, testing, and integration. Support the development and maintenance of Windows desktop applications and services. Assist in developing and maintaining web applications, APIs, and cloud-connected services. Troubleshoot software issues with guidance from senior team members. Write clean, maintainable, and well-documented code. Create and execute unit tests to verify software functionality and reliability. Participate in code reviews and learn software development best practices. Contribute to Agile ceremonies, sprint planning, and team activities. Learn and apply secure coding and software security principles. Support the deployment, monitoring, and maintenance of cloud-based applications and services. Collaborate with cross-functional teams to deliver end-to-end software solutions. Support product release and maintenance activities. Education & Experience Bachelor's or Master's Degree in Computer Science, Software Engineering, or a related discipline. 0-2 years of software development experience. <li

JavaScriptTypeScriptPythonReact
MT
📍 Manassas, VA - Fab 6, United States
✓ Quality checkedCompany trend +1266.7%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Fab Engineer in the RDA & Metrology Team , you will collaborate with other equipment engineers, technicians, and application owners. Your responsibilities include overseeing the installation, modification, upgrade, and maintenance of manufacturing equipment, providing technical support to manufacturing equipment repair and process engineering organizations. You will define and write preventative maintenance schedules, maintain records on technical notices, upgrades, and safety issues, and study equipment performance to establish solutions that improve tool uptime. Responsibilities: Drive performance standards, critical measures, and accountability for operational decisions and results. Lead equipment issue resolution, troubleshooting efforts, and root cause investigations. Manage project priorities and recommend changes, continuation, or cancellation to meet fab objectives. Collaborate daily with Equipment Leads to align priorities, resolve issues, and optimize resource utilization. Plan for capacity and product mix requirements while enhancing equipment flexibility and performance. Maintain and improve equipment reliability through maintenance procedures, best-known methods (BKMs), program monitoring, and documentation updates. Partner with vendors and multi-functional teams to implement continuous improvement initiatives that enhance throughput, productivity, and equipment health. Leverage AI, analytics, and

PythonSQLAIPower Bi
C
📍 Dallas 8000 Frankford Road, United States
✓ Quality checkedCompany trend +340.2%

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary We are seeking a highly experienced Principal / Director-level Full-Stack Software Development Engineer to join our Digital Caremark organization and lead the architecture, design, and delivery of next-generation digital solutions. This role spans AI-enabled applications, scalable digital platforms, and enterprise integrations that power critical healthcare and pharmacy experiences. This is a senior technical leadership role for a hands-on engineer who can operate across the full stack—from intuitive front-end applications to resilient backend services—while setting architectural direction, influencing engineering standards, and mentoring teams. The ideal candidate combines deep technical expertise, platform thinking, and strong collaboration skills to build secure, scalable, API-first solutions in a highly regulated environment. Key Responsibilities: Architecture & System Design • Define and drive architecture for large-scale distributed systems and digital platforms • Lead design reviews and set architecture standards and best practices • Champion API-first, microservices, and event-driven architecture patterns • Ensure systems meet scalability, reliability, security, and compliance requirements • Balance performance, cost, and speed in technical decision-making Fu

TypeScriptPythonJavaReact
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build cloud-native data and storage services for hybrid and multi-cloud infrastructure, including dataset discovery, ingestion, governance, checkpointing, observability, and low-latency access. Develop scalable cloud-native services and APIs that support exabyte-scale, high-performance GPU training and inference workflows. Work closely with product managers, internal AI teams, platform teams, and partner engineering teams to understand requirements and turn them into reliable production systems. Collaborate with SRE, operations, and support teams to improve service reliability, performance, observability, on-call readiness, and operational scale. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, and verification. What we need to see: BS in Computer Science, Information Systems, Computer Engineering, or equivalent experience, with 5&#43; years of software engineering experience. Strong foundation in algorithms, data structures, distributed systems, and practi

PythonJavaAWSAzure
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys

PythonJavaKubernetesLinux
P
📍 New York City, New York, United States· Full-time
✓ Quality checkedCompany trend -85.7%

About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep

PythonJavaAWSKubernetes
P
📍 United States· Full-time
✓ Quality checkedCompany trend -85.7%

About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep

PythonJavaAWSKubernetes
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -93.7%

From $168K/yr

Quick readStrong listing-quality and freshness signals

MongoDB’s Developer Productivity organization exists to help engineers build and deliver high-quality software through a highly effective software development process and a strong foundation of shared tools and services. We are looking for a Senior Director to lead our Pipeline team. This role is tasked with bringing together the major systems and experiences that power software delivery at MongoDB. The team’s mission is to provide a reliable, scalable, secure, and effective platform for ensuring fast software deployability, leveraging AI native approaches. We are open to in-office, flexible or remote hiring across the US. The Team The Pipeline organization sits within Developer Productivity and is responsible for the systems, services, and user experiences that define MongoDB’s software delivery ecosystem. This is mission-critical infrastructure operating at substantial scale and supports a variety of software product delivery needs. Success in this role requires excellent product judgment for developer-facing experiences, strong systems and platform leadership, and the ability to align multiple teams around a cohesive strategy. Candidate Profile We’re looking for a senior engineering leader who can unify product-minded developer tooling with deep platform and operational excellence. The right candidate is passionate about developer productivity and has a track record of leading managers and teams through organizational growth, technical complexity, and cross-functional change. They should be comfortable owning a broad portfolio that spans developer experience, reliability and scale, release systems, telemetry, and operational health. They should also be able to work effectively with senior leaders and partners across engineering and product to set direction, allocate resources, and make trade-offs that balance near-term delivery with long-term platform function. The right candidate for this role will 12+ years of hands-on software engineering experience bui

MongoDBAWSAzureKubernetes
M
📍 Boise, ID - Main Site, United States· Full-time
✓ Quality checkedCompany trend -75%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Position Overview: As a Staff/Principal Equipment Engineer within PROCT ADT, you will provide technical leadership to improve equipment performance, capability, and global manufacturing outcomes. You will lead complex, multi-site initiatives to resolve equipment challenges, enable technology node readiness, and drive yield and defectivity improvements. This role requires deep expertise, strong ownership, and the ability to influence across teams and suppliers to deliver measurable impact in cost, performance, and scalability. Responsibilities Lead global ownership of equipment performance across toolsets, driving improvements in stability, availability, matching, and efficiency. Direct complex, multi-fab root cause analysis for equipment-driven yield, defectivity, and reliability issues using data-driven and physics-based approaches. Provide technical leadership in tool hardware and chamber behavior to expand process capability and improve performance limits. Partner with Technology Development, integration, and process teams to enable node readiness, volume ramp, and disciplined change control. Influence equipment supplier strategy by leading technical engagements, driving design improvements, and aligning roadmaps to business needs. Drive global standardization through development, validation, and deployment of Best Known Methods (BKMs) across sites. Lead initiatives to improve cost of ownership, reduce variation,

AIRecruitment
O
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s API Platform organization builds the products and infrastructure that help first-party and third-party developers build with OpenAI models. We ship the API primitives, tools, SDKs, documentation, playgrounds, and platform experiences that make OpenAI’s capabilities reliable, understandable, and useful in production. The API Experience team is focused on the end-to-end developer experience for the OpenAI API. We own the surfaces developers touch every day: docs, SDKs, the Playground, examples, onboarding flows, and the systems that help developers go from first request to production deployment quickly and confidently. About the Role We’re looking for full stack and frontend engineers to help define and build the next generation of OpenAI’s developer experience. In this role, you’ll work across frontend product surfaces, backend systems, SDK and documentation pipelines, and API workflows that serve millions of developers and companies. You’ll partner closely with product, design, research, API engineering, and developer-facing teams to make complex AI capabilities simple to understand, easy to test, and safe to launch in real-world applications. This is a highly cross-functional role for someone who cares deeply about craft, developer empathy, reliability, and product velocity. In this role, you will: Build and scale developer-facing products including the OpenAI API Playground, documentation experiences, onboarding flows, examples, and API workflow tools. Own full stack projects end to end, from product definition and UX collaboration through backend implementation, launch, measurement, and iteration. Improve the systems that generate, maintain, and publish SDKs, API references, docs, guides, and developer examples. Partner with API, research, design, and infrastructure teams to bring new model capabilities and API primitives to developers in a clear, usable way. Use developer feedback, product analytics, and direct customer insight to identif

TypeScriptPythonReactAWS
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Cybersecurity Products team builds products at the frontier of AI and cybersecurity. Our work includes Codex Security and related cyber products that turn advances in model capability into dependable tools for defenders. We help teams find, validate, and remediate vulnerabilities, continuously improve the security of software, and test AI-powered applications before they reach production. About the Role As a Full Stack Software Engineer, you will build the product experiences and systems that make AI-powered security useful in real engineering environments. You will work across web surfaces, APIs, orchestration, data models, and integrations to help security and engineering teams move from a codebase or application to evidence-backed findings, prioritized remediation, and revalidation. You will collaborate closely with product engineers, security researchers, and customer-facing teams. The work spans fast-moving product development and hard systems problems: long-running workflows, large repositories, sensitive data, reliability, observability, and a high bar for earning user trust. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build end-to-end workflows for vulnerability discovery, security scanning, red teaming, findings review, remediation, and reruns. Design and operate backend services for long-running security work, including APIs, asynchronous orchestration, durable state, and integrations with developer workflows. Make complex security results actionable through clear product surfaces, strong evidence, thoughtful prioritization, and reliable reporting. Partner with security researchers, product teams, and users to evaluate quality, reduce noise, improve coverage, and ship safely. You might thrive in this role if you: Have experience shipping production full-stack products across modern web frontends and backend s

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. This team is central to this mission, setting the core infra strategy and implementing this vision. From site selection to the buildout process, this team sits at the intersection of commercial, technical, strategy, and operations, interacting with teams and executives inside and outside of OpenAI. About the Role We are seeking experienced Data Center Mechanical and Electrical/Power Design Engineers with expertise in designing, operating, and maintaining large-scale data center campuses. The ideal candidate for this role will have extensive background and experience in design and managing critical equipment and facilities, design and operation of MEP (Mechanical, Electrical, Plumbing) systems, and overseeing operational activities from initial phases of Data Center build through delivery and ongoing maintenance. The ideal candidate will have a strong technical background, operational leadership experience, and a proven ability to collaborate with external vendors on critical infrastructure. This role offers the opportunity to lead transformative data center projects with high visibility and impact. If you are passionate about delivering cutting-edge infrastructure solutions, we encourage you to apply. Key Responsibilities Oversee building and MEP design, operation, and maintenance, including reviewing building and MEP drawings and proposals across all project phases. Lead operational activities for large-scale data center campuses, from early design phases through delivery and daily operation. Operate and maintain critical data center facilities and equipment, ensuring reliability and performance. Collaborate with external vendors to select, procure, and manage critical equipment, such as generators, UPS, chillers, and CDUs. Provide technical expertise on all aspects of data center building, equipment, and

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des

AWSRestAIGo
🔔

Get new reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime