What we’re doing isn’t easy, but nothing worth doing ever is. Diligent builds helpful robots that work safely and autonomously in real world environments. We move quickly, solve messy problems, and care deeply about reliability at scale. We’re hiring a Manufacturing Reliability Engineer to own production test for our robots at our contract manufacturer: you’ll design and run robust end-to-end test protocols, provision fleets of robots for production, and own the KPIs that define production quality. This role is based in Austin, TX. However, the position will require 50% travel to the Milwaukee, WI area and requires close collaboration across software, hardware, operations, and product engineering teams. Key Responsibilities End-to-end test process ownership. Create, validate, and maintain production test protocols and gating criteria from incoming inspection through final test and shipment. Provisioning of bots. Design and operate provisioning flows (imaging, firmware deployment, configuration, validation) and the tooling/fixtures needed to provision and handoff robots for production. KPIs and continuous improvement. Own key production metrics — First Pass Yield (FPY), cycle time, and test coverage — and drive continuous improvements to meet throughput and quality targets. Test automation & infrastructure. Architect, implement, and maintain automated test frameworks, harnesses, and test rigs used at the CM site. Ensure tests are stable, fast, and provide actionable failure data. Cross-functional escalation & RCA. Lead root-cause analysis for field and production failures; coordinate corrective actions with design, firmware, and CM engineering to close quality loops. On-site production leadership. Be the onsite technical authority at the contract manufacturer: train operators, debug failures on the line, and continuously refine processes with CM partners. What Success Looks Like Improved FPY and reduced rework rates across production builds. Reduced per
Jobs in United States
Manufacturing Reliability Engineer in United States
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current manufacturing reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. Role Overview We are seeking a Package Reliability Engineer to lead reliability engineering for advanced packages used in high-performance AI and computing systems. The primary focus of this role is to assess package level mechanical and thermal reliability risks and apply thermal and mechanical modeling to optimize package design, material selection, and assembly processes. The engineer will also develop reliability test plans with external partners, identify failure mechanisms, perform root-cause analysis, and recommend practical corrective actions. In this role, you will assess package reliability risks from early architecture development through product qualification and high-volume manufacturing. You will work closely with package design, silicon design, system engineering, manufacturing, and ASIC partners to predict package behavior, develop qualification strategies, resolve reliability issues, and improve overall package robustness and lifetime. In this role you will: Lead reliability test plan and assessments for advanced HPC packages, including risk identification, potential failure-mechanism analysis, root-cause investigation, mitigation planning, and corrective-action development. Drive reliability-focused package design optimization based on thermo-mechanical modeling to improve package reliability, power integrity, thermal performance, mechanical robustness, and platform scalability. Develop, validate, and apply package reliability models and lifetime-prediction
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Infrastructure Quality team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, general contractors, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. Our work spans from vendor qualification through commissioning, ensuring operational readiness across our global portfolio. About the Role We are seeking an experienced Manufacturing Quality Engineer (MQE) to establish, implement, and manage a manufacturing-focused quality program for datacenter infrastructure. This role will be responsible for vendor oversight, quality assurance, process improvement, and issue resolution for all critical systems. You will lead vendor audits, monitor performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced risks, and operational reliability. By partnering with vendors, construction teams, and internal stakeholders, you will help ensure OpenAI’s datacenters are delivered on time and built to the highest operational standards. Travel Domestic and international travel as needed (estimated 40–60%) to manufacturing sites, datacenter locations, and partner facilities. Key Responsibilities Vendor Oversight & Performance Management Conduct manufacturing evaluation, audits, and improve vendor performance across production, inspection, testing, and delivery phases. Develop and track quality metrics to assess manufacturing performance and identify trends. Partner with vendors to refine processes, training, and quality controls to mitigate risks before shipment. Program Development & Execution Develop and maintain a datacenter-focused m
Manufacturing Test Engineering Manager Position Summary We are seeking an experienced Manufacturing Test Engineering Manager to lead the development and execution of the end-to-end manufacturing test strategy for next-generation AI server platforms and datacenter infrastructure. This role is responsible for defining and driving the manufacturing test architecture from L6 board assembly through L11 rack-level integration and final system validation , ensuring world-class product quality, manufacturability, and production scalability. This leader will manage a team of 3–5 Manufacturing Test Engineers while partnering closely with Hardware, Firmware, Platform, Validation, Quality, Operations, and Joint Design Manufacturing (JDM) partners. The role owns the manufacturing test strategy, test coverage, factory test infrastructure, manufacturing capacity planning, and continuous improvement of manufacturing quality. Key Responsibilities Manufacturing Test Strategy Define and own the end-to-end manufacturing test strategy from L6 board assembly through L11 rack integration . Develop standardized manufacturing test methodologies that optimize quality, throughput, cost of test, and scalability across multiple products and JDM sites. Establish manufacturing test standards, best practices, and engineering processes that support high-volume server manufacturing. Technical Leadership & People Management Lead, mentor, and develop a team of 3–5 Manufacturing Test Engineers supporting multiple hardware programs. Establish team priorities, allocate resources, and ensure successful execution of manufacturing test deliverables. Foster a culture of technical excellence, accountability, collaboration, and continuous improvement. Serve as the primary technical escalation point for manufacturing test and production issues. Cross-Functional Engineering Collaboration Partner with Hardware, Platform, Firmware, Validation, Reliability, Quality, and Operations teams to ensure manufac
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a product manufacturing & quality engineer, who will be responsible for driving technical initiatives related to the manufacturing, quality and reliability of our AI supercomputer hardware systems to ensure product success from concept to launch and through mass production. You’ll have the opportunity to coordinate with functional SMEs and work with a wide range of stakeholders, from design engineering and operations teams, TPMs, external industry vendors and partners to ensure that all products are developed and delivered on time and to the highest quality standards. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees In this role, you will: Own the integrated manufacturing and quality readiness for a product across L6, L10, and L11, with clear gates, milestones, deliverables, owners, and closure criteria. Lead readiness of process flows, tooling, fixtures, assembly operations, test interfaces, and production controls. Review and contribute to work instructions. Translate product requirements into qualification plans, process controls, test requirements and acceptance criteria with design engineering and Area SMEs Coordinate and drive execution of product and process qualification, reliability testing, and validation with the relevant SMEs. Maintain traceable evidence that assigned products and processes meet agreed performance, reliability,
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Quality leadership within Manufacturing Operations, the Senior Reliability Scientist is responsible for leading reliability activities across complex, high-performance systems. Working closely with established reliability experts and cross-functional teams, this role uses experimental data and advanced modelling to inform design decisions, validate product reliability and optimise serviceability strategies, including spares provisioning. The Team The Quality team within Manufacturing Operations is responsible for ensuring product robustness, reliability and lifecycle performance across Graphcore’s hardware portfolio. The team includes experienced reliability specialists and works closely with technology research, chip, board, system design, platform and operations teams to translate reliability insights into actionable improvements across the product lifecycle. Responsibilities and Duties: · Define and refine reliability requirements across silicon, board and system levels, working in partnership with research and design teams · Apply ad
About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a Senior Actuator Design and Integration Engineer to lead the development of custom electromechanical actuators for advanced robotic systems. You will own actuator development from early architecture and concept generation through prototype validation and system integration, partnering closely with mechanical, electrical, controls, firmware, reliability, and manufacturing teams. This role focuses on the design, integration, and validation of precision electromechanical systems, including motors, transmissions, sensing, structural components, and thermal architectures. You will help drive actuator development across the full engineering lifecycle while establishing scalable design, test, and integration practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will Lead the architecture, design, and integration of custom robotic actuators, including motors, transmissions, sensing, thermal systems, structural components, and packaging. Define actuator requirements and system-level trade studies around torque density, bandwidth, efficiency, thermal performance, backdrivability, inertia, reliability, manufacturability, and cost. Design precision electromechanical assemblies with strong attention to tolerances, alignment, load paths, thermal expansion, sealing, wear, and serviceability. Drive actuator integration into robotic systems, partnering closely with controls, firmware, electrical, and robotics software teams to optimize closed-loop pe
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our
What we’re doing isn’t easy, but nothing worth doing ever is. Diligent builds helpful robots that work safely and autonomously in real world environments. We move quickly, solve messy problems, and care deeply about reliability at scale. As a Fleet Engineer, you'll own the reliability and continuous improvement of our deployed robotic fleet — leading hands-on investigations into how and why robots fail in the field, across the mobile base, charging/docking, motion and power, connectivity (modem), and sensor hardware. You'll combine remote data analysis with bench/lab failure analysis at our Austin HQ, turning field-technician reports and fleet data into clear problem statements, validated root causes, and corrective actions driven to closure with engineering, operations, manufacturing, and vendors. We are hiring a Lead Engineer, Issue Management & Triage to lead the systems, tooling, and team at the intersection of our Customers, Remote Operations Center (ROC), and Engineering. This is a highly technical, hands-on role focused on building the infrastructure that powers how we detect, triage, diagnose, and resolve issues across a deployed robotic fleet. You will work deeply with Engineering teams to design classification frameworks, build internal tools, and develop automation pipelines that improve reliability at scale. Location: Austin preferred, Remote possible (U.S.) Travel: if remote up to ~50% travel to Austin, TX (especially in the your first 90 days) What You’ll Do: Own Issue Management & Triage Systems Design and own end-to-end systems for issue intake, triage, and escalation. Define severity frameworks, SLAs, and ensure issues are consistently structured for engineering prioritization. Build Tools & Automation (Hands-On) Develop automation and pipelines to ingest, process, and classify operational data, reducing manual triage effort. Contribute directly to codebases (Python, backend services) and partner with Engineer
1743 - This position is in Austin, Texas. Position Summary We are seeking an experienced Board-Level Hardware Validation Engineer to define and execute the validation and verification of complex electronic systems throughout the product lifecycle. This role is responsible for defining validation strategies, developing test plans, executing hands-on testing, analyzing failures, and working directly with ODM partners to ensure products meet performance, reliability, quality, and compliance requirements before mass production. The ideal candidate combines strong electrical engineering fundamentals with practical lab expertise and is comfortable personally performing validation activities while coordinating with cross-functional teams and manufacturing partners. Key Responsibilities Validation Strategy & Planning Define comprehensive board-level and inter-board validation plans based on product requirements, design specifications, and customer use cases. Develop validation methodologies covering functional, electrical, thermal, power, signal integrity, reliability, and stress testing. Establish test coverage, acceptance criteria, qualification requirements, and release gates. Review hardware architecture, schematics, component specifications, and interface topologies to identify validation risks early in the design cycle. Define incremental validation and regression coverage for component substitutions, design changes, and firmware updates. Hands-On Validation Execution Develop, automate, and execute validation tests on prototype and production-intent hardware. Perform board bring-up, functional verification, electrical characterization, and system-level integration testing. Validate communication interfaces, control signals, and timing requirements. Verify power sequencing, reset behavior, leakage current, and recovery across operating states. Execute temperature and voltage corner testing against approved operating limits. Use oscilloscopes, logic an
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We're looking for an Optical Interconnect System Engineer to design, qualify, and deploy scalable optical connectivity for large-scale AI infrastructure. This role spans fiber-system architecture, optical-mechanical integration, validation, reliability, deployment, and serviceability. You will work with optical, mechanical, electrical, networking, manufacturing, reliability, and data-center teams to translate system needs into practical interconnect solutions. This is a hands-on role for someone who can connect design decisions with installation, qualification, troubleshooting, and long-term operational performance. In this role, you will: Define optical interconnect architectures and requirements across hardware platforms and rack-level systems. Design high-density fiber systems for performance, density, reliability, installation, and serviceability. Lead optical-mechanical integration and cross-functional design reviews. Develop test and qualification plans for optical components, modules, switching platforms, and integrated systems. Own optical loss budgets, routing guidelines, handling requirements, and serviceability criteria. Support system bring-up, deployment, troubleshooting, failure analysis, and reliability improvement. Create reusable design guidelines, interface requirements, and qualification methods. You might thrive in this role if you have: Core experience Experience desi
About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking an Actuator Gear Design Engineer to lead the development of custom gears and gear stages for advanced robotic systems. You will own actuator development from early architecture and concept generation through prototype validation and system integration, partnering closely with mechanical, electrical, controls, firmware, and reliability teams. You will partner with external suppliers and internal manufacturing to create full gearbox assemblies. This role focuses on the design, integration, and validation of precision gearing, including broader knowledge around motor electromagnetics, transmission types, sensing, structural components, and thermal architectures. You will help drive actuator development across the full engineering lifecycle while establishing scalable design, test, and integration practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will: Lead the architecture, design, and integration of custom robotic actuator gearing. Define actuator requirements and system-level trade studies around torque density, bandwidth, efficiency, thermal performance, back drivability, inertia, reliability, manufacturability, and cost. Design precision electromechanical assemblies with strong attention to tolerances, alignment, load paths, thermal expansion, sealing, wear, and serviceability. Drive actuator integration into robotic systems, partnering closely with controls, firmware, electrical, and robotics software teams to optimize closed-lo
About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a senior Actuator Electromagnetic Design Engineer to lead the development of custom electromechanical actuators for advanced robotic systems. You will own actuator development from early architecture and concept generation through prototype validation and system integration, partnering closely with mechanical, electrical, controls, firmware, reliability, and manufacturing teams. This role focuses on the design, integration, and validation of precision electromechanical systems, including motors, transmissions, sensing, structural components, and thermal architectures. You will help drive actuator development across the full engineering lifecycle while establishing scalable design, test, and integration practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will: Lead the architecture, design, and integration of custom robotic actuators, including the design, simulation, integration and sourcing of custom electromagnetic components. Define actuator requirements and system-level trade studies around torque density, bandwidth, efficiency, thermal performance, inertia, reliability, manufacturability, and cost. Design precision electromechanical assemblies with strong attention to tolerances, alignment, load paths, thermal expansion, sealing, wear, and serviceability. Drive actuator integration into robotic systems, partnering closely with controls, firmware, electrical, and robotics software teams to optimize closed-loop performance. Devel
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
Other cities to consider
More places hiring for this role
Get new manufacturing reliability engineer jobs in United States by email
Daily job updates · Unsubscribe anytime