Jobs in United States

Reliability Engineer in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We lead complex technical programs that help Plaid scale its engineering platform. We partner across Engineering, Infrastructure, Data, Security, ML, Legal, and Product to deliver company-wide technical initiatives that improve reliability, scalability, and developer productivity. You'll lead strategic technical programs from planning through execution. You'll partner with engineering leaders to align stakeholders, manage dependencies, drive decisions, and ensure successful delivery of complex initiatives. You'll work across a variety of technical domains, adapting quickly to new challenges and helping teams execute effectively. As a Technical Program Manager, you will lead high-impact, cross-functional initiatives. As a generalist, you may work on a variety of programs. An example is one that strengthens Plaid's data and machine learning platforms. You will partner with engineering, product, data, legal, privacy, and business stakeholders to drive complex technical programs from planning through execution. Your work will help improve data governance, modernize machine learning infrastructure, and accelerate the adoption of trusted, high-quality datasets that power analytics, artificial intelligence

AWSMachine LearningAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Through strategic partnerships and self-built campuses, we are scaling one of the world's fastest-growing AI infrastructure platforms. The Supply Chain organization ensures critical infrastructure components—from compute systems and networking equipment to integrated rack solutions—are sourced, manufactured, qualified, and delivered with the speed and reliability required to support frontier AI development. We partner closely with Hardware Engineering, Manufacturing Quality Engineering, Infrastructure Delivery, Hardware Operations, Finance, and suppliers worldwide to build a resilient, scalable supply chain capable of supporting rapid infrastructure expansion. As Industrial Compute continues to grow, Supply Chain serves as the operational bridge between engineering innovation and large-scale infrastructure deployment. About the Role We are seeking a Supply Chain Manager to lead strategic execution across sourcing, supplier operations, manufacturing quality, and infrastructure delivery for OpenAI's AI infrastructure portfolio. This role will oversee a multidisciplinary team responsible for strategic sourcing, manufacturing quality engineering, and technical program management while partnering closely with engineering, finance, hardware operations, and deployment teams. You will drive supplier strategy, manufacturing readiness, production planning, quality performance, and operational execution across the full hardware lifecycle. Success requires balancing long-term supplier strategy with day-to-day execution. You'll establish scalable operating mechanisms, strengthen supplier partnerships, manage complex cross-functional programs, and ensure OpenAI can rapidly deploy AI infrastructure without compromising quality, cost, or reliability. This is a people leadership role responsible for developing a high-performing organization while driving operati

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Platform team powers how millions of developers and enterprises build with our models. We provide APIs and agentic solutions used by global startups and fortune 500s. We work closely with product, engineering, design, and go-to-market to build a world-class platform that pushes the frontier of AI capabilities. About the Role As a Data Scientist on the Platform team, you will drive a data-driven culture for OpenAI’s API and B2B solutions. You’ll define the metrics that matter for developer success and enterprise value, measure the impact of new models and features, and partner with PMs and engineers to improve model quality, reliability, latency, and cost. Your work will shape how thousands of products adopt agentic AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will Embed with the Platform product team as a trusted partner, uncovering ways to improve developer experience, reliability, and usage growth Define north-star metrics across the developer funnel (activation, retention, growth), as well as latency/cost guardrails for new features and models Design and interpret A/B tests and controlled rollouts (e.g., new model versions, pricing/limits, new API features, new B2B products) Build source-of-truth dashboards and self-serve data tools for product, engineering, and go-to-market teams Translate product learnings into actionable feedback for Research (e.g., failure modes, eval gaps, model response quality) You might thrive in this role if you have 5+ years in a quantitative role in ambiguous, high-growth environments (platforms, APIs, or B2B products a plus) Depth in SQL and Python, with a track record proposing, designing, and running rigorous experiments Experience defining and operationalizing metrics from scratch (including reliability/latency/cost and safety) Strong cross-functional communication with PMs, enginee

PythonSQLAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Human Data team at OpenAI is responsible for identifying and mitigating risks in advanced AI systems by designing evaluations, surfacing vulnerabilities, and collaborating closely with researchers to strengthen model reliability and public trust. About the Role As a Research Program Manager, you will lead initiatives that test the safety and robustness of OpenAI’s models through creative experimentation and structured evaluation. You’ll coordinate efforts across research and engineering teams to transform ambiguous risks into concrete research programs and influence future model development and deployment. We’re looking for people who are technically savvy, comfortable with ambiguity, and excited about shaping the future of safe AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead programs that explore unexpected model behaviors and identify failure modes. Translate vague or emergent risk signals into clear priorities and actionable research plans. Design and run creative evaluations, experiments, and red-teaming campaigns. Collaborate with research, product, and deployment teams to integrate findings into model training and deployment cycles. Develop repeatable systems for tracking model performance and understanding emerging behavior patterns. You might thrive in this role if you: Have strong experience in technical program management, with excellent organizational and communication skills. Are familiar with large language models, prompt engineering, or model evaluation techniques. Are comfortable managing fast-paced, high-uncertainty projects and shaping them from the ground up. Are creative and resourceful in devising new methods for testing model behavior and performance. Can effectively coordinate across technical and non-technical stakeholders to drive alignment and execution. About OpenAI OpenAI is an AI resear

AWSRestAIRust
MT
📍 Richardson, TX, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. At Micron Technology, we transform how the world uses information to enrich life for all. The Heterogeneous Integration Group (HIG) HBM Architecture team develops next-generation High-Bandwidth Memory (HBM) solutions that power AI, high-performance computing, cloud infrastructure, and advanced networking systems. The team works across architecture, design, verification, packaging, product engineering, and technology development to evaluate innovative memory architectures and deliver scalable, high-performance semiconductor solutions. As an HBM Design Architect, New College Graduate, you will contribute to the evaluation and development of future HBM and DRAM architectures. Working with experienced architects and engineering teams, you will analyze system and block-level design tradeoffs related to performance, power, area, thermal behavior, reliability, and manufacturability. This role provides an opportunity to leverage AI, Large Language Models (LLMs), and data-driven engineering methodologies to accelerate architecture exploration and improve decision-making. Responsibilities Analyze HBM and DRAM architectures using analytic

AIRecruitment
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the hardware platforms that support the next generation of AI systems. The Opportunity As a Systems Engineering Intern, you will contribute to projects that combine hardware, firmware, and software engineering for advanced AI compute platforms. You will work with experienced engineers on subsystem design, laboratory testing, system validation, automation, and performance analysis. The internship provides hands-on experience with modern hardware development and system-level engineering. You will own clearly defined technical tasks with guidance from the team and document your methods, results, and conclusions. What You Will Do Support the design and testing of CPU and high-speed input and output subsystems for advanced compute platforms. Run laboratory tests and measurements to help evaluate performance, power, signal behavior, and reliability. Contribute to system-level validation by creating scripts and tools that streamline testing, data collection, and analysis. Explore emerging input and output technologies, including PCIe 6.0 and 800G Ethernet, and learn how they support advanced computing workloads. Assist with investigations into platform power, cooling, and energy efficiency, including liquid-cooling systems for high-performance processors. Use power meters, oscilloscopes, logic analyzers, or comparable lab equipment under appropriate supervision. Analyze test results, identify unexpected behavior, and work with engineers to reproduce and investigate issues. Collaborate across hardware, firmware, software, m

PythonArtificial IntelligenceAIC++
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA’s Silicon Co-Design Group sits at the crossroads of architecture, silicon, systems, and manufacturing, where first-principles thinking and engineering judgment at the highest level translate directly into product outcomes at scale. We are looking for a Principal Performance and Manufacturing Architect who has built the models, defined the specs, and seen them validated through silicon. You have owned the connection between design intent and manufacturing reality, not as a reviewer or a contributor, but as the person who set the methodology and proved it worked. You turn ambiguous physical phenomena into quantified, defensible margin terms. You do not wait for data to confirm your hypothesis; you design the experiment that gets it. You improve how the organization ships products after every program. The exceptional hire also uses AI deliberately — with proven workflow impact and the judgment to know where it compresses real work and where it introduces risk. What you'll be doing: Own the physics, from mechanism to margin. Build first-principles models connecting AVF, defect mechanisms, and DVFS transients to field FIT, system-level yield, and DPPM vs. coverage — calibrated per node and population shift — so every margin term in the V/F curve and P-state table is named, sourced, and defensible. Set the screen that resolves escapes. Specify ATE and SLT voltage, frequency, and timing conditions that capture worst-case transient VF windows — making it unambiguous whether a marginal defect or timing violation is detected or escapes at every manufacturing stage. Make the POR the authoritative source. Author the methodology document for each program and drive alignment across build, product definition, reliability, and test engineering — so every team is making decisions from the same model. Prove the model before produc

I
📍 Arizona, Phoenix, United States
✓ High-confidence listingCompany trend +315.4%
Quick readStrong listing-quality and freshness signals

Job Details: Job Description: Join Intel's Advanced Packaging Technology Development - Substrate and Wafer Assembly (ATPD: SWA) team and help us fulfill our mission to be the supplier of choice for leading-edge, cost-effective substrate packaging. As a Manufacturing Equipment Technician, you will support critical manufacturing equipment and drive continuous improvement in a highly technical, cleanroom production environment. Shift 5: Shift 5: (days frontend) AM-PM Sun-Tues and alternating Wednesday days / Hours: 6:00 a.m.-6:00 p.m. Manufacturing Equipment Technicians in Advanced Packaging Technology Development: Substrate and Wafer Assembly are responsible for: Performing electrical and/or mechanical troubleshooting to diagnose issues with non-functioning electromechanical equipment used in semiconductor manufacturing. Dismantling, adjusting, repairing, and assembling equipment according to layout plans, blueprints, operating or repair manuals, rough sketches, or engineering drawings. Performing setup, calibration, and preventive maintenance on production equipment, with a focus on wet chemistry, plating, and dry high-vacuum toolsets. Documenting maintenance activities, repairs, parts usage, and findings in accordance with established procedures and systems. Collaborating with engineering, process, and support teams to improve tool performance, reduce equipment downtime, and enhance process capability and reliability. Leading or contributing to continuous improvement projects, including root cause analysis, yield and uptime improvement initiatives, and optimization of maintenance practices. Assignments are semi-routine in nature and performed within generally defined parameters. Employees normally receive general instructions on most work and are expected to exer

Recruitment
MT
📍 Richardson, TX, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. In this position, you will work and support the efforts of groups such as Design Engineering, Product Engineering, Process Integration, Probe, and Assembly to proactively design products that optimize all manufacturing functions and assure the best cost, quality, reliability, time-to-market, and customer happiness. How To Qualify: B.S.EE degree with + 2 years of experience M.S.EE degree with 1+ years of experience. Coursework in CMOS Circuit Design and VLSI Coursework that included crafting the physical layout of circuits Knowledge and experience using Cadence Virtuoso Excellent problem-solving and analytical skills A self-motivated, hard-working team player who enjoys working with others. Good communication and documentation skills Drive innovation into the future Memory generation with a dynamic work environment. What Sets You Apart: Knowledge of layout synthesis and place and route techniques Experience with LVS and DRC tools such as Calibre Utilize AI-Enabled tools to assist in layout floor planning, place and route Standard cell construction and analog layout techniques Having an innovative approach that is open to improving upon any of our processes or products. As a world leader in the semiconductor industry, Micron is dedicated to your persona

AIRecruitment
MT
📍 Manassas, VA - Fab 6, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron is seeking motivated Process Technicians to support wafer manufacturing by monitoring operations, reviewing data, and working closely with engineering and production teams to enhance yield, quality, and efficiency. This role requires strong analytical thinking, attention to detail, and disciplined execution. Process Technicians must have the ability to recognize abnormal conditions and apply structured problem-solving to troubleshoot issues in a fast-paced manufacturing environment while adhering to established procedures. Ideal candidates are comfortable working with data, dashboards, and computer-based systems and are capable of extracting, analyzing, and interpreting trends and clearly communicating insights. Success depends on effective communication, strong organization and time management skills, the ability to work independently, and flexibility with responsibilities and shift schedules. As part of a fabrication team driving innovation for a global leader in memory and storage solutions, Process Technicians directly support manufacturing excellence while demonstrating reliability, engagement, and a growth mindset aligned with Micron’s core values: People, Innovation, Tenacity, Collaboration, and Customer Focus. Responsibilities: • Manage semiconductor equipment, perform technical evaluations, and partner with cross functional teams to monitor, stabilize, and improve wafer fabrication processes. • Work with engineering teams to enhance workstations and apply structured troubleshooting and escalati

AIRecruitment
C
📍 Tampa Florida United States, United States
✓ High-confidence listingCompany trend +800%
Quick readStrong listing-quality and freshness signals

We are seeking an experienced Senior Generative AI Developer to help drive the design, development, and integration of state-of-the-art Generative AI and agentic AI solutions across our enterprise Controls Technology platform. You will collaborate with cross-functional teams, contribute deep technical expertise in context engineering, retrieval systems, knowledge graphs, and multi-agent orchestration, and play a key role in delivering scalable, grounded AI solutions to enhance automation and operational efficiency. This role centers on architecting robust applications and agent systems on top of pre-trained and hosted foundation models — not on training or fine-tuning models. Key Responsibilities Collaborate with AI architects, leads, and stakeholders to design and implement generative and agentic AI solutions that address business challenges. Architect advanced context engineering strategies — context layering, chaining, compression, pruning/offloading, and memory management — to maximize reliability, provenance, and token efficiency in production. Design and implement advanced generative AI methods, including sophisticated prompt engineering and Retrieval-Augmented Generation (RAG) . Build and optimize RAG systems , including hybrid search, multi-vector retrieval, and re-ranking pipelines. Design and implement knowledge graphs and Graph RAG architectures to enable multi-hop reasoning, explainability, and traceable, grounded responses for high-value business domains. Architect agentic workflows and multi-agent systems using Google Agent Development Kit (ADK) and comparable frameworks (LangGraph, Microsoft Agent Framework, CrewAI), applying orchestration patterns such as supervisor/worker, hierarchical, and peer-to-peer. Design robust agent harnesses — governance, constraints, feedback loops, state/session management, and

PythonJavaSQLAWS
DR
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Senior Product Manager, Robotics & Autonomy What we're doing isn't easy, but nothing worth doing ever is. At Diligent Robotics, we envision a future powered by robots that work seamlessly with human teams. We build artificial intelligence that enables service robots to collaborate with people and adapt to dynamic human environments. Our robots operate every day in hospitals, helping healthcare staff spend less time on routine work and more time caring for patients. Operating a real-world fleet gives us something few robotics companies have: continuous customer feedback and operational data that directly shapes the next generation of Physical AI. We're looking for a Senior Product Manager, Robotics & Autonomy to define and execute the product strategy for some of the most critical capabilities in our robotics platform. You'll work at the intersection of robotics, autonomy, AI, and software engineering to translate business priorities, customer needs, and technical opportunities into a clear product roadmap that drives measurable outcomes. This role is ideal for someone who understands complex autonomous systems and enjoys working alongside world-class engineers to bring ambitious technology from concept into production. Responsibilities Own the product strategy and roadmap for key Robotics and Autonomy initiatives, balancing customer impact, technical feasibility, and long-term platform investments. Define product requirements for autonomy, navigation, perception, fleet intelligence, simulation, and robotics platform capabilities. Partner closely with Engineering, AI, Robotics, Customer Success, Operations, and Leadership to align priorities across the organization. Translate customer feedback, fleet telemetry, and operational insights into product decisions that improve robot performance, reliability, and user experience. Prioritize investments using data, customer value, technical complexity, and business impact. Drive cross-functional execution from concep

AgileMachine LearningAIExcel
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI's data and storage infrastructure spans data platforms, online databases, and file/object storage. These systems underpin data ingestion and processing, durable persistence, indexing and retrieval, and product file experiences. As frontier models and agents evolve how they use memory, history and snapshots, the underlying architecture increasingly shapes the capabilities products can deliver—and their latency, reliability, cost and efficiency. About the Role We are looking for a technically deep TPM to independently define and lead multiple programs across data platforms, online databases and storage infrastructure. You will connect model, product and data-consumer requirements to architecture, and work with the relevant engineering teams to take new capabilities through production adoption and repeatable expansion. The design scope is exabyte-scale storage and infrastructure spanning multiple millions of CPU cores. The challenge is not simply forecasting more resources: it is making complete, workload-ready capacity repeatable, with a clear path from product requirements through architecture, deployment and validation. A data pipeline, database query, file operation or execution snapshot can affect whether a product or agent succeeds; you will connect those outcomes to the systems underneath. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate model, product and data-platform needs into precise access patterns, consistency, durability, freshness, availability and scalability requirements. Connect memory, history, retrieval and resumable work to capability and end-to-end latency. Partner with engineering to transform data and storage architecture into repeatable scale units: standardized provisioning, placement, routing, data movement and readiness checks that bring storage, compute and networking online together.

AWSAzureRestAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT engineering org builds and operates the systems that bring product improvements to users across backend services, web, mobile, and desktop platforms. The Developer Velocity team partners with product engineering, platform, infrastructure, reliability, engineering acceleration, and observability teams to make everyday development faster and releases safer, more predictable, and easier to operate. About the Role We are seeking a Technical Program Manager to improve developer velocity and deployment excellence across ChatGPT. You will lead durable improvements to local development, CI, testing, build systems, release trains, progressive rollout, and post-deployment validation. You will identify the highest-leverage sources of engineering friction, align teams around shared standards and metrics, and drive adoption of tooling and workflows that improve both speed and reliability. This role combines systems thinking, technical program leadership, and hands-on operating rigor across a broad engineering surface. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the cross-functional roadmap for improving local development, CI, testing, build workflows, and release infrastructure. Create a durable intake and prioritization mechanism for developer friction, using evidence to focus teams on the highest-impact improvements. Lead programs that improve deployment speed and safety, including pre-merge confidence, progressive rollout, release guardrails, rollback readiness, and post-deploy valida

AWSCI/CDRestAI
O
📍 Washington, District of Columbia, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%

From $2M/yr

Quick readStrong listing-quality and freshness signals

About the team The OpenAI for Government team is a dynamic, mission-driven group leveraging frontier AI to transform how governments achieve their missions. Our team works to empower public servants with secure, compliant AI tools (e.g., ChatGPT Enterprise, ChatGPT Gov) and mission-aligned deployments that meet government technical requirements with strong reliability and safety. About the role Our Federal Sales team has a unique mission to help government customers understand the transformative impact that highly capable AI models can bring to their agencies and missions. This role combines technical understanding, strategic vision, partnership management, and value-driven strategy tailored specifically to federal customers. You’ll drive key opportunities through the entire federal sales cycle, from pipeline generation to closure. You’ll collaborate closely with researchers, engineers, and solution strategists to help government customers advance their missions through AI. This role is based in Washington DC. We use a hybrid work model of 3 days in the office per week. We offer relocation assistance. In this role, you'll: Manage a focused set of key federal accounts, developing and executing comprehensive federal account plans. Lead federal customers through their AI adoption journey, from consideration to successful deployment. Partner with solutions and research engineering to build and execute complex government customer programs and projects. Own and manage a federal consumption revenue target. Oversee consumption revenue forecasting and reporting. Analyze key federal account metrics and provide insights to internal and external stakeholders. Closely monitor the federal landscape (agencies, policies, competitors, partners, etc.) to inform product roadmaps and corporate strategies. Collaborate cross-functionally with solutions, marketing, communications, business operations, people operations, finance, product management, and engineering. Support the recruitment

AWSRestAIGo
🔔

Get new reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime