Jobs in United States

Reliability Engineer in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake's Data Engineering organization builds the platform that ingests, transforms, and stores data for modern lakehouse architectures — powering billions of queries, DML, and DDL operations with industry-leading price-performance. We lead the industry's shift to open data lakes through our work on Iceberg and Polaris, and we deliver capabilities like Snowpark, Dynamic Tables, cross-region replication, time travel, and zero-copy cloning at enterprise scale. We are investing in a new line of applied research — building toward verified data infrastructure and trustworthy data systems — that brings formal methods, automated reasoning, and modern AI techniques to bear on the hardest problems in our distributed systems and developer tooling. The goal is to improve correctness, reliability, and engineering velocity at a scale very few platforms operate at. We're hiring at both the Staff and Principal level; we'll calibrate the offer to the candidate's experience and scope of impact. What you'll do Lead research projects that apply formal methods, program analysis, automated reasoning, and AI-driven techniques (including code generation and modeling) to real problems in our cloud data platform. Translate research ideas into prototypes, then into shipped capabilities that move

S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. The Team On the AI Backend team at Observe by Snowflake you'll be at the forefront of how AI is reshaping the way engineering teams operate; building the platform that powers intelligent, automated workflows across observability and beyond. Our team is small, and moves fast, with real ownership over hard problems that span agentic APIs, real-time pipelines, and AI quality. You'll work alongside talented engineers across ML, product, and

TypeScriptPythonAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team builds next-generation AI-native silicon and systems while working closely with software, research, and manufacturing partners to co-design hardware tightly integrated with AI models. In addition to delivering systems for OpenAI’s supercomputing infrastructure, the team develops the tools, methodologies, and strategic partnerships needed to accelerate hardware innovation. About the Role We’re seeking an experienced Hardware Strategic Sourcing Manager to own sourcing strategy and supplier partnerships for fiber and optical interconnect components across OpenAI’s next-generation AI infrastructure. Reporting to the Head of Partnerships & Strategic Sourcing, you will lead sourcing across fiber cable assemblies, internal optical harnesses, fiber shuffles, optical backplane assemblies, connectorized and standalone passive optical assemblies, fiber-array units (FAUs), fiber-to-chip and coupling interfaces, detachable connectors, optical routing, and assigned optical packaging, assembly, and test services. You will work closely with electrical engineering, optical engineering, systems engineering, mechanical and packaging engineering, quality, rack integration, data-center deployment,manufacturing, supply chain, finance, legal, and program management teams to translate demanding bandwidth, signal integrity, reliability, and scale requirements into resilient supplier partnerships and scalable commercial strategies. Your work will directly support the performance, reliability, manufacturability, and scale of the high-speed optical connectivity required for OpenAI’s next-generation AI systems. In this role, you will: Develop and execute a comprehensive sourcing strategy for fiber and optical interconnect components supporting high-bandwidth AI systems and infrastructure. Own sourcing across optical fiber cable assembli

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. The Systems Integration team is critical in this mission, turning complex hardware-software development into reliable product signals. We build the shared infrastructure, tooling, and lab environments that let teams test quickly, understand failures, and ship with confidence. About the Role As a Systems Integration Manager , you will lead the team responsible for device validation infrastructure, test automation, developer tooling, and lab operations. This is a player-coach leadership role: you’ll set technical and operational direction, build and develop a team of engineers and lab operations professionals, and stay close to the architecture and hardest systems problems. You will partner closely with device software, OS, firmware, hardware, reliability, QA, and release infrastructure teams to define validation strategy, improve release readiness, and ensure our test environments and quality signals scale with the product. Because this is a new category of devices, you’ll have the rare opportunity to build the validation foundation early—shaping the systems, standards, and operating model that will support products from prototype through launch. We’re looking for a leader who combines strong technical judgment with people leadership, operational rigor, and experience building reliable systems for complex hardware-software products. This role is b

Artificial IntelligenceAI
MT
📍 Richardson, TX, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. In this position, you will be joining a brilliant team to collaborate with Micron’s various design and verification teams all over the world, support the efforts of design verification flow development, optimization, and implementation and work with cross-functional team to assure the best cost, quality, reliability, time-to-market, and customer satisfaction As Memory Design Group CAD Engineer at Micron Technology, Inc., you will be working in a collaborative, production support role evaluating, developing enhancing and debugging both in-house and commercial Electronic Design Automation (EDA) tools and flows for the physical layout and design of CMOS integrated circuits. You will work closely with the Design, Layout teams, supporting their use of electronic design automation tools and methodologies to increase design and layout productivity and work efficiency. Responsibilities : Develop tools, flows, and methodologies to increase the productivity and reliability of our Memory designs Integrate commercial EDA (Electronic Design Automation) tools into design flows Provide training, documentation, and support to end users on new tools and method. Work closely with design, layout, verification and process teams. Undertake independent research and development of ne

PythonAIRecruitment
I
📍 Arizona, Phoenix, United States
✓ High-confidence listingCompany trend +315.4%
Quick readStrong listing-quality and freshness signals

Job Details: Job Description: The Advanced Packaging Technology Development (APTD) Substrate and Wafer Assembly organization is responsible for developing cutting edge substrate packaging and wafer assembly solutions for our customers. As a member of the Substrate Technology and Manufacturing Group, you will play a critical role in enabling new manufacturing capacity across advanced packaging suppliers. In this position, you will lead substrate supplier development and qualification activities for advanced technologies such as high density FCBGA, EMIB-T and glass core substrate, ensuring suppliers achieve the required readiness to support growing customer demand. This includes driving technical enablement, managing readiness milestones, and collaborating closely with cross functional engineering teams and supplier partners. A combination of detail-oriented project management, strong technical risk assessment / problem solving skills, and effective supplier / stakeholder management will be needed to be successful. Please note: This role requires periodic evening meetings with suppliers in East Asia, as well as occasional travel to supplier sites to support essential program objectives. Key Responsibilities Define process integration roadmaps and establish manufacturing procedures to meet technology transfer milestones. Extract insights from structured and unstructured data, applying statistical analysis and coding techniques for yield improvement. Lead low yield analysis and corrective actions to address quality excursions in high-volume production environments. Drive continuous improvement initiatives to minimize defects and optimize process reliability. Assess quality and reliability risks for manufacturing material and process changes, ensuring sustained product performance. C

Project ManagementRecruitment
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world's most transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the future of AI computing. The Opportunity As Technical Services Director, you will lead the teams that operate and evolve Graphcore's engineering labs, high-performance computing (HPC) platforms, and data center environments globally. You will be accountable for reliable, secure, cost-effective infrastructure that supports demanding engineering, AI, silicon-development, and validation workloads. This role combines people leadership, infrastructure strategy, operational excellence, capacity and financial planning, procurement, and program delivery. You will partner with Engineering, Information Technology, Security, Finance, Facilities, Supply Chain, customers, and external suppliers. The position is based onsite in Austin and requires travel to company facilities, data centers, and supplier locations, including international travel. What You'll Do Lead, recruit, mentor, and develop the systems administration, lab operations, and technical services teams responsible for the facility supporting global Engineering and Research and Development. Own the reliability, efficiency, protection, safety, supportability, and continuous improvement of engineering labs, HPC systems, and infrastructure facilities. Establish service levels, operating standards, escalation paths, performance measures, monitoring, observability, automation, ticketing, and configuration-management practices. Translate engineering and customer requirements into infrastructure roadmaps, capacity p

LinuxAIGoExcel
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Industrial Compute team is building and productizing infrastructure capabilities that help organizations deploy and operate advanced AI systems at scale. The team works across AI hardware, systems engineering, physical infrastructure, and customer delivery to turn emerging technologies into reliable, repeatable infrastructure solutions. Our work sits at the intersection of technical strategy, product development, engineering, and deployment. We partner closely with customers and internal engineering teams to solve complex infrastructure challenges spanning compute, power, cooling, controls, and facility efficiency. About the Role We are seeking a senior, hands-on Data Center Infrastructure Architect to develop and optimize the physical infrastructure required for large-scale AI deployments. This is a broad technical role spanning data center architecture, electrical and mechanical systems, high-density compute, controls, telemetry, and digital modeling. You will use simulation, operational data, and digital-twin approaches to evaluate infrastructure designs, identify system-level constraints, and improve efficiency, reliability, cost, and speed of deployment. The ideal candidate can move fluidly between first-principles analysis, facility and equipment design, computational modeling, engineering review, and real-world implementation. You should be comfortable working across disciplines rather than operating solely within electrical, mechanical, or software boundaries. Key Responsibilities Define system-level architectures for high-density AI data centers across power, cooling, IT equipment, controls, and facility infrastructure. Develop digital twins and other computational models that represent the behavior of data center systems under changing workloads, environmental conditions, equipment configurations, and failure scenarios. Use design and operational data to identify constraints, improve PUE and related efficiency metrics, and optimize

PythonAWSGitRest
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $204K/yr

Quick readStrong listing-quality and freshness signals

The opportunity Datadog’s Infrastructure products help engineers understand and operate the systems their applications depend on. Our customers work in complex environments like Kubernetes and serverless, where infrastructure changes constantly, information is dense, and decisions about reliability, performance, and cost are closely connected. We’re looking for a Staff Product Designer to join Modern Compute, with an initial focus on Containers Autoscaling. Autoscaling helps engineering teams make better decisions about how their applications and infrastructure use resources. Designing these experiences requires making deeply technical systems understandable, helping customers act with confidence, and fitting into the tools and workflows they already use. The team is rethinking how workload and cluster autoscaling come together as a more coherent product experience. This includes how customers get started, understand recommendations, evaluate value, and safely apply changes across their environments. The work also connects to other parts of Datadog, including observability, Cloud Cost Management, permissions, and AI-assisted workflows. As a Staff Product Designer, you will help define that direction and lead the work from early problem framing through shipped product. You will partner closely with product and engineering, bring a high level of interaction and visual craft to complex workflows, and help raise the quality of design across Modern Compute. At Datadog, we place value in our office culture, the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to help our Datadogs find a work-life rhythm that works for them. What you’ll do Lead end-to-end product design for Modern Compute, initially focused on our Autoscaling product. Help define the product direction for an area that is still evolving, from early framing and exploration through detailed design and delivery. Design clear, trustwort

KubernetesGitAIGo
H
📍 Colorado, United States of America, United States
✓ High-confidence listingCompany trend +103.7%

$83K – $127.8K/yr

Quick readStrong listing-quality and freshness signals

Software Developer in Test Description - This role is responsible for ensuring quality, reliability and performance of software applications throughout the development lifecycle primary through software test automation. The role designs, codes, and implements software test automation using appropriate programming languages, frameworks, and tools. The role works closely with cross-functional teams to gather requirements, provide technical insights, and ensure the successful execution of test automation with the main goal of improve quality of the solution. The role also creates and executes comprehensive test plans, test cases, and test scripts based on project specifications. The role sets and provides design guidance to other developers and SQA engineers regarding test automation and the test framework. *Onsite in Ft Collins 4-days a week Responsibilities • Designs quality assurance and test processes for portions of end-user video conferencing/collaboration application, systems software running on android hardware, local, networked, and Internet-based platforms. • Analyzes design and determines test scripts, coding, automation, and integration activities required based on specific objectives and established project guidelines. • Designs and maintains Test automation framework • Executes and writes portions of testing plans, protocols, and documentation for assigned portion of application; identifies and debugs issues with code and suggests changes or improvements. • Identifies opportunities for performance improvements and optimizes code and application performance. • Utilizes latest AI tools and technologies in speeding up test automation • Executes test cases depending on the needs of the project • Provides valuable input into the development of user stories and acceptance criteria, shaping a quality-oriented d

PythonJavaDockerAI
M
📍 Boise, ID - Main Site, United States
✓ Quality checkedCompany trend -75%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As an Intern Product Yield Enhancement (PYE) Engineer at Micron Technology Inc., you will assist in taking next generation DRAM devices from design to mass production! PYE Engineers are crucial partners between Research and Development, Manufacturing, Product Engineering and Global Quality teams. Our goal is to offer hands-on and substantial engineering projects that allow interns to gain valuable experience with memory systems and the product life cycle while adding impact to Micron’s business. Our intern programs help prepare engineers for future roles and offer an extraordinary experience that supplements your academic experience! The PYE Engineering role is critical to Micron as it serves the company's charter to be the most cost-effective producer of DRAM memory while delivering the quality and reliability that customers expect from the Micron brand. This position is for a 3-month internship. Responsibilities: Your job duties will include detailed DRAM electrical and physical failure analysis. Identify defects introduced by the semiconductor fabrication process. You will data mine using statistical analysis, AI assisted engineering tools, and automation workflows. Find the root cause of failure mechanisms, communicating results with Management, Product, Process and Test Engineering teams. Complete 3-month internship. Minimum Qualifications: Must be pursuing a minimum of a bachelor’s degree in Electrical Engineering, Computer Engineering,

PythonAIRecruitment
D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $152K/yr

Quick readStrong listing-quality and freshness signals

The Manager of Networking at Datadog leads the global management of network services across all Datadog offices worldwide. This role is responsible for ensuring seamless, high-performance Wi-Fi and direct internet access in our global offices and conference room technology, supporting a rapidly growing global enterprise. As Datadog continues to grow rapidly, this role plays a critical part in scaling both the team and network infrastructure to meet increasing demand. This is a hybrid role that sits in global headquarters in New York city and requires three days in the office each week with occasional travel to our offices around the world. What You’ll Do: Lead the global network engineering teams, managing both full-time employees and third-party vendors to ensure consistent and high-quality service delivery for Datadog offices worldwide. Oversee the design, deployment and scaling of office network infrastructure, including Wi-Fi (Cisco Meraki) and edge networking devices (Cisco, Palo Alto, Juniper, and others), ensuring these services operate according to Datadog's defined service level objectives. Define and implement standards, policies, and processes for network infrastructure to ensure security, reliability and scalability. Collaborate with IT Security, Enterprise Technology, and Workplace teams to align network services with broader IT and business objectives. Develop and track operational metrics for service availability, network performance, driving continuous improvement and optimization. Who You Are: An experienced people manager with at least 5+ years of leadership experience managing teams of network engineers. Proven expertise in Wi-Fi network engineering, including deep knowledge of Cisco Meraki and edge networking solutions from vendors like Cisco, Palo Alto and Juniper. Experience managing large-scale office technology projects in global enterprises with more than 10 offices and 7,000+ employees, ensuring infrastructure keeps pace with ra

I
📍 Arizona, Phoenix, United States
✓ High-confidence listingCompany trend +315.4%
Quick readStrong listing-quality and freshness signals

Job Details: Job Description: As a Module Equipment Technician, you will play a vital role in ensuring the smooth operation of advanced manufacturing equipment used in semiconductor production. You will perform hands-on troubleshooting, maintenance, and calibration of electromechanical systems while supporting experiments and equipment modifications that drive cutting-edge advancements. Your work will directly contribute to enhancing equipment reliability and optimizing production efficiency, making an essential impact on Intel's manufacturing operations. Business Group Join Intel's Advanced Packaging Technology Development - Substrate and Wafer Assembly (ATPD: SWA) organization, a leader in delivering innovative and cost-effective substrate packaging solutions. This business group is focused on advancing Intel's capabilities in manufacturing technology to maintain competitive excellence across global markets. As part of this dynamic team, you will support critical equipment and collaborate on continuous improvement initiatives that align with Intel's broader mission of innovation and growth. Key Responsibilities Perform electrical and mechanical troubleshooting to diagnose issues with manufacturing equipment. Execute setup, calibration, corrective and preventative maintenance on production equipment, including wet chemistry, plating, and dry high-vacuum toolsets. Monitor tool performance and analyze data to identify and address equipment-related issues. Collaborate with engineering and support teams to improve equipment reliability and reduce downtime. Document maintenance activities, repair findings, and parts usage using established systems and procedures. Lead or contribute to continuous improvement projects, focusing on equipment optimization and yie

Recruitment
MT
📍 Boise, ID - ID1, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron is seeking a Facilities Water Services UPW and Wastewater Coordinator to support wafer manufacturing by coordinating daily work, maintenance, vendors, contractors, and projects for ultrapure water, process water, reclaim, and wastewater systems. This role partners with Operations, Engineering, Maintenance, Construction, Procurement, EHS, vendors, and leadership to plan, implement, document, and align work with production needs. Success requires strong organization, technical understanding, communication, attention to detail, and the ability to lead priorities across operations, maintenance, projects, and production schedules. Ideal candidates are proficient with SAP or other CMMS tools, Microsoft Office, trackers, dashboards, drawings, and documentation systems. They coordinate schedules, track action items, communicate status, support scope development, identify gaps, and improve safety, reliability, documentation, cost control, and execution quality. This role helps maintain critical facility systems while demonstrating Micron’s core values of People, Innovation, Tenacity, Collaboration, and Customer Focus. Responsibilities: Coo

AISapProcurementRecruitment
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team The Growth Platforms team builds the systems and operating foundations that help OpenAI grow responsibly. We partner across the product portfolio to connect customer signals, identity and consent, campaign workflows, measurement, and product experiences into an AI-enabled growth engine. Our work helps teams launch, learn, and scale with a high bar for data quality, privacy, reliability, and customer trust. About the Role We’re looking for an experienced marketing technology and operations leader to drive cross-functional work at the intersection of growth, measurement, data, and automation. Your mission will be to turn fragmented tools, signals, and workflows into reliable, measurable, AI-enabled capabilities that teams can use safely at scale. You’ll work across Growth, Marketing Operations, Product, Engineering, Data Engineering, Data Science, Security, Privacy, Legal, and Revenue Operations, as well as external advertising platforms, measurement providers, and implementation partners. You’ll translate business requirements and privacy constraints into data contracts, integration designs, rollout plans, and reliable first-party data systems. This is a hands-on, high-impact role for someone who brings structure to ambiguity and moves from event schemas, APIs, and data quality assurance to operating cadences, partner enablement, and executive updates. This role is based in San Francisco or New York City with a hybrid office expectation. In this role, you will: Own the operating model for Growth’s marketing technology stack across identity, consent, audiences, activation, measurement, and experimentation. Own and operate the complete paid-media tracking and measurement system, including website pixels, server-to-server conversion events, mobile measurement integrations, identity and consent controls, attribution methods, and timely signal delivery to advertising platforms. Design and implement event schemas, data mappings, APIs, and integrations; valid

SQLAWSRestAI
🔔

Get new reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime