Jobs in United States

Reliability Engineer in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse using open formats like Apache Iceberg — at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world's leading data platforms. We are hiring a Senior Software Engineer for the Observe Data Management team. This team owns the core pipelines that ingest and process over 1 petabyte of telemetry data per day — the foundational infrastructure powering Observe's entire observability stack. You'll be working at the intersection of massive scale, open-source innovation, and real-world reliability challenges for enterprise customers around the globe. AS A SENIOR SOFTWARE ENGINEER - OBSERVE

AWSAzureAIC++
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. About the team The Apps & Experiences Platform team is powering systems and services for all our user facing apps including Snowsight , Snowflake Intelligence , and new mobile apps . Our mission is to craft innovative backend services, features, tools, infrastructure, and AI tooling that bring such products to life with delightful user experiences. As part of our team, you'll dive into a mix of building & managing platform infrastructure, and building AI self-serve tools to support the platform and its developer’s needs. We're passionate about building a platform that is highly reliable, available, maintainable, and scalable. We are a high growth AI data cloud company and we are looking for exceptional talent like you to help build and grow our infrastructure to scale us to the next level. AS MANAGER FOR APPS & EXPERIENCES PLATFORM TEAM, YOU WILL: Own the technical strategy and execution for the team, driving projects from initial idea formulation and detailed system design to high-quality implementation and successful deployment. Provide deep technical oversight by actively participating in design reviews, architecture discussions, and drilling into complex system implementations to ensure reliability and scalability. Serve as a subject matter expert , setting

AWSAzureGCPAI
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We’re hiring a talented Software Engineering Manager to lead the Snowtrail infrastructure team at Snowflake. Snowtrail is the infrastructure that enables Snowflake to deliver dedicated coverage for customer-specific workloads. Its innovative approach allows Snowflake to precisely test and measure the impact of changes on individual customers, making it essential for ensuring the platform’s reliability, correctness, and performance. Through query replay, Snowtrail helps us catch regressions early. By leveraging machine learning models to intelligently sample queries and workloads, we continuously optimize for both cost and performance. Evolving Snowtrail to incorporate new engine features while improving scalability, efficiency, and reliability is central to our continued success OUR IDEAL MANAGER WILL HAVE : Strong passion and proven track record for shipping quality software in high code velocity environments 10+ years industry experience designing and building distributed data systems. Excellent problem solving skills, and strong CS fundamentals including data structures, algorithms, and distributed systems. Fluency in SQL, Java, C++, Python or Go. Ability to collaborate well across teams, build high-performing teams and mentor junior engineers. Excellent interpersonal co

PythonJavaVueSQL
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - Container Platform (Menlo Park) About the Role We build the foundational container platform that runs Snowflake's production, AI/ML, and CI workloads across AWS, Azure, and GCP, including a rapidly growing AI/ML footprint. Hundreds of large Kubernetes clusters under management and growing. The work is to make that fleet reliable, automated, and invisible to the thousands of engineers building on top of it. This is a staff-level role on a senior, high-performing platform team. You'll own hard problems end to end, drive technical direction across teams, and build the automation and platform abstractions that make operating at this scale sustainable. There is significant unsolved work ahead: improving the developer experience for thousands of internal engineers and continuing to scale the platform to meet Snowflake's growth. What You'll Do Own the design and delivery of large, complex platform initiatives spanning cluster lifecycle management, multi-cloud automation, and internal developer tooling. Identify and drive cross-team technical improvements across the platform, from architecture through adoption. Make and defend architectural trade-offs grounded in reliability, scalability, and operational reality. Act as a technical anchor for the team, dev

AWSAzureGCPKubernetes
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. In this role you will: Develop interactive, data-rich user interfaces using React, TypeScript, and Vega, with a focus on integrating LLM-driven features (e.g., natural language querying, generative UI, and AI-assisted data storytelling). Lead the end-to-end delivery of substantial product features, ensuring AI outputs are presented with high reliability and low latency. Work closely with PMs, UX designers, and AI/ML engineers to bridge t

JavaScriptTypeScriptJavaReact
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Our vision is to transform how the world uses information to enrich life for all. Join an inclusive team passionate about one thing: using their expertise in the relentless pursuit of innovation for customers and partners. The solutions we build help make everything from virtual reality experiences to breakthroughs in artificial intelligence possible! We do it all while committing to integrity, sustainability, and giving back to our communities, because doing so can fuel the very innovation we are pursuing. As a process integration engineer in the DRAM organization, you will be part of a team of world-class engineers who are working in an industry leading 300mm R&D facility on technology which enables future memory scaling. You will focus on the performance and manufacturability of Micron’s cutting-edge next-generation memories. Areas of concentration will include integration and circuit issues specific to DRAM technology. In addition to a broad basis of process integration knowledge, this position will require a strong background in semiconductor device physics. Role interacts significantly with the process development, simulation, modeling, characterization, and reliability groups. Responsibilities include, but not limited to: Lead technology integration activities in enabling manufacturability of state-of-the-art DRAM technology Engage with numerous cross-functional teams including process development, product engineering, design, yield improvement, advanced mask and probe to arrive at solutions to

ReactArtificial IntelligenceAIProject Management
I
📍 Oregon, Hillsboro, United States
✓ High-confidence listingCompany trend +315.4%
Quick readStrong listing-quality and freshness signals

Job Details: Job Description: Join an enthusiastic team of engineers in Intel's Networking Solutions Group (NSG) focused on enabling next generation of programmable Infrastructure Processing Units (IPUs) with our lead customers as part of the Customer Experience Support (CES) organization. Intel brings decades of leadership in networking, virtualization, packet processing, storage, and security to a new class of IPU products that accelerate host networking functions and support emerging use cases such as security, virtualization, storage, load balancing, and data path optimization. Working closely with major cloud service providers and Intel development teams, you will help deliver customized IPU based solutions that enhance isolation, security, performance, storage and system management for our customers. A big part of the day-to-day job is to help customers manage feature request processes, enable solutions, and debug issues. Projects and responsibilities include but are not limited to: • Gain our customers' trust, understand their needs, and build POCs to meet them. Work closely with internal and external partners to understand use cases and requirements. • Be the go-to technical resource for customers building complex Datacenters, AI infrastructure as well as helping them understand performance characteristics for solutions. • Prepare and deliver technical content to customers including presentations, workshops, etc. • Contribute across the full IPU lifecycle, including board and platform bring up, low-level device initialization, OS driver and kernel configuration, system management, feature enablement, use case testing, debugging, and verification. • Defines systems implementation and integration solutions and plans to ensure optimum performance and reliability across hardware, firmware and software w

DockerGitLinuxAI
I
📍 Oregon, Hillsboro, United States
✓ High-confidence listingCompany trend +315.4%
Quick readStrong listing-quality and freshness signals

Job Details: Job Description: Embark with us on a journey of growth and transformation as we create exceptionally engineered technology and bring AI everywhere. As a valued team member, your adaptability and attention to detail will contribute to our drive for results and relentless pursuit of quality, ensuring we meet our customers' needs with precision. Join us and build on our legacy of innovation and collaboration as we deliver world‑changing technology that improves the life of every person on the planet. Life at Intel: https://jobs.intel.com/en/life-at-intel This position is in the Intel Mask Operations within the Logic Technology Development team working in one of the most advanced semiconductor process technologies in the world. In this position the engineer will be an integral contributor to the ongoing production of photolithography masks for Intel's leading-edge silicon manufacturing solutions. In this position, you will work on-site, in a dynamic and collaborative environment solving complex and challenging technical problems on sophisticated manufacturing processes and equipment. As an IMO Module Development Engineer, the responsibilities may include but not limited to: Process equipment installation and development. Equipment maintenance, management of troubleshooting activities, regular monitoring of process performance, defect analysis and reduction. Drives improvements on quality, reliability, cost, yield, process stability/capability, productivity, and safety/ergonomics. Working with cross functional teams to solve technical process and defect issues. Plans and conducts experiments to fully characterize the process throughout the development cycle. Establishes control systems to optimize and sustain production performance. D

SQLAIRecruitment
MT
📍 Boise, ID - ID1, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Our Vision Our vision is to transform how the world uses information to enrich life for all. Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate, and advance faster than ever. Micron is advancing a historic $15 billion investment in semiconductor manufacturing in Boise, Idaho, with DRAM production planned for the second half of the decade. As a global leader in memory and storage solutions, Micron is investing more than $150 billion worldwide over the next decade to expand leading-edge manufacturing capabilities and drive innovation across the semiconductor industry. Join the Boise expansion team and help build the future of semiconductor manufacturing. You will work alongside experienced engineers to develop advanced processes, solve complex manufacturing challenges, and support the production of technologies that power everything from artificial intelligence to next-generation computing systems. Position Overview As a PCVD Process Engineer , you will contribute to the development, optimization, and sustainment of semiconductor manufacturing processes. You will gain hands-on experience in a state-of-the-art fabrication environment, partnering with teams to improve yield, quality, reliability, and productivity. This role offers an outstanding opportunity to apply engineering fundamentals, build technical expertise, and develop a career in advanced semi

PythonSQLArtificial IntelligenceAI
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxArtificial IntelligenceAI
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxArtificial IntelligenceAI
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxArtificial IntelligenceAI
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team Consumer Monetization builds the experiences and systems that power how customers purchase and pay for OpenAI products. Our scope spans purchasing flows, payments, subscriptions, and billing, along with the shared capabilities that support new products, offers, and distribution channels. We own both customer-facing experiences and the underlying platforms that power them. We partner closely with Product, Design, Growth, Data Science, and engineering teams across OpenAI to make purchasing effective and reliable, and new offerings easier to launch and monetize. About the Role We’re looking for experienced Staff+ full-stack engineers to shape purchasing experiences and monetization capabilities across OpenAI products. You’ll work across web experiences, product APIs, and shared components, combining strong product judgment with technical depth. The work ranges from improving checkout conversion and performance to enabling new pricing models, offers, and ways for customers to purchase our products. You’ll help identify opportunities, turn ambiguous goals into concrete technical plans, and lead initiatives from exploration through launch. As patterns emerge across products, you’ll develop reusable capabilities that make future launches faster and more consistent. This is a hands-on technical leadership role with substantial ownership over architecture, implementation, and product outcomes. This role is based in San Francisco, with three days per week in the office. In this role, you will: Architect and build purchasing experiences end to end, from frontend interactions through the product APIs and backend integrations that support them. Improve checkout conversion, performance, and reliability through experimentation, product analytics, and customer insights. Develop shared checkout components and monetization capabilities that support new products, pricing models, offers, and distribution channels. Partner with Product, Design, Growth, and Data Science to

Artificial IntelligenceAI
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team Consumer Monetization builds the experiences and systems that power how customers purchase and pay for OpenAI products. Our scope spans purchasing flows, payments, subscriptions, and billing, along with the shared capabilities that support new products, offers, and distribution channels. We own both customer-facing experiences and the underlying platforms that power them. We partner closely with Product, Design, Growth, Data Science, and engineering teams across OpenAI to make purchasing effective and reliable, and new offerings easier to launch and monetize. About the Role We’re looking for experienced Staff+ engineers with deep iOS or Android expertise and demonstrated experience contributing beyond mobile to frontend web or backend development. You’ll shape and build purchasing and subscription experiences across mobile applications, web, and supporting product APIs. The work spans improving conversion, performance, and reliability, building reusable components, and enabling new product launches. You’ll combine product judgment with technical depth to set direction, lead initiatives across teams, and remain hands-on through implementation and delivery. Approximately 50% of the work will initially be native mobile development, with the remainder across web and product APIs. We’re looking for engineers who enjoy working across the stack and have concrete examples of doing so professionally. Depth in either iOS or Android is required; experience in both is not required. This role is based in San Francisco, with three days per week in the office. In this role, you will: Architect and build purchasing and subscription experiences across native mobile, mobile web, and supporting product APIs. Partner with Product, Design, and Growth to identify customer needs, prioritize improvements, and shape technical direction for purchasing experiences. Improve conversion, performance, and reliability through experimentation, product analytics, and customer insights

Artificial IntelligenceAI
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron’s DRAM Product Architecture Group partners across fabs, test, module, and system teams worldwide to deliver high-performance memory products. We work across geographies and functions to solve complex technical challenges, drive product innovation, and enable industry-leading DRAM solutions. Our team values collaboration, continuous learning, and making a measurable impact through engineering excellence. As a New College Graduate DRAM Yield & Quality Engineer, you will play a key role in monitoring and improving the health of Micron’s server DRAM products. Working with experienced engineers and cross-functional teams around the globe, you will analyze product and manufacturing data, investigate technical issues, and contribute to strategies that improve yield, quality, and performance throughout the product lifecycle. This is a highly visible role with direct impact on customer success and product reliability. Responsibilities Analyze silicon parametric, test, and manufacturing data to identify trends, resolve product issues, and support root cause investigations. Contribute to yield and quality improvement initiatives across new product introduction (NPI), qualification, ramp, and high-volume manufacturing. Collaborate with fab, test, module, quality, and operations teams to drive execution of product health objectives. Develop and maintain automation solutions to improve engineering efficiency, data analysis, and reporting. Apply AI, Agentic AI, and Generative AI tools to en

AIRecruitment
🔔

Get new reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime