About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. In this role you will: As a Hardware Test Engineer, you will work on Machine Learning/AI hardware system projects to craft the solutions for current and future data center deployments. You will bring a strong understanding of hardware system testing, excellent project management skills, and the ability to collaborate across multiple teams to ensure efficient lab operations. You will be responsible for designing, implementing, and executing comprehensive test plans that ensure the reliability, performance, and scalability of our supercomputing hardware systems. You will develop detailed test plans and methodologies tailored to hardware components, including processors, memory modules, custom accelerators and interconnects. You will collaborate with hardware design, manufacturing, firmware teams and vendors to identify, analyze, and resolve issues affecting hardware, power, thermal and high-speed interconnects. You will perform in-depth debugging on the hardware system Excellent analytical skills to diagnose hardware issues, troubleshoot problems, and propose solutions. Ability to interpret complex test data, identify trends, and draw meaningful conclusions. High-speed links, with a focus on SerDes (Serializer/Deserializer) technology to assess signal integrity, error rates, and overall link performance. You will collaborate with the lab manager to maintain the equipment and hardware systems, including oscilloscopes, thermal test chambers, liquid cooling systems, and other mea
Jobs in United States
Hardware Engineer in San Francisco
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current hardware engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will develop and evolve the tooling ecosystem that hardware engineers rely on every day — from hardware compilers and IR transformations to simulation, debugging, and automation infrastructure. The work spans software engineering, compiler concepts, and practical hardware workflows, with direct impact on how quickly and effectively we design next-generation AI systems. You’ll collaborate closely with architects, RTL designers, and verification engineers to translate real engineering friction into durable, scalable tooling solutions. In this role you will: Build and improve the software tooling that makes hardware teams faster: compilation, IR transforms, RTL generation, simulation, debug, and automation. Extend and integrate hardware compiler stacks (frontends, IR passes, lowering, scheduling, codegen to Verilog/SystemVerilog) and connect them to real design workflows. Improve developer experience and reliability: reproducible builds, better error messages, faster iteration loops, and dependable CI and regression infrastructure. Work closely with designers and verification engineers to turn real pain points into durable tools. Dive into RTL when needed: read and reason about Verilog/SystemVerilog to debug issues, validate tool output, and improve debuggability. Be willing to go all the way down the stack when necessary, including gate-level views, synthesis results, and implementation artifacts. Help enable PPA optimization loops by building analysis and au
About the Team OpenAI’s 1P Hardware Systems team operates at the intersection of Hardware Engineering, Infrastructure, Supply Chain, Manufacturing, Deployment, and Finance to translate infrastructure demand into an executable systems plan. The Planning Lead connects system demand and deployment timing to site and configuration requirements, power availability, XPU needs, hardware supply commitments, manufacturing capacity, and site readiness. About the Role OpenAI is seeking a 1P Hardware Systems Planning Lead to own the integrated demand, supply, and deployment plan for our 1P AI infrastructure systems. This role will connect infrastructure demand and site power availability to system configurations, XPU requirements, and hardware supply commitments—creating a single, actionable view of whether our deployment plan can be met. You will establish the planning mechanisms that allow teams to see what changed, understand the impact, and act quickly when demand, supply, configuration, site readiness, or power timing moves. You will work across Hardware Engineering, Infrastructure, Supply Chain, Manufacturing, Deployment, Finance, and external partners to identify gaps early and drive recovery plans to closure. This is a highly cross-functional role for someone who combines systems-level thinking, strong analytical judgment, and rigorous program execution. The ideal candidate can turn complex and changing inputs into a clear operating plan, surface the decisions that matter, and drive accountability across teams without relying on formal authority. Success also requires strong communication judgment: the ability to align cross-functional teams, brief leadership at a concise and decision-oriented level, and go deep into the underlying assumptions, dependencies, risks, and recovery plans when needed. In this role, you will: Own the integrated planning process for 1P hardware systems, connecting system demand, deployment timing, site and configuration requirements, power ava
About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, cloud services, mechanical engineering, electrical engineering, and product design to deliver reliable, production-ready devices at scale. Within Consumer Devices, Hardware Engineering eXperience, or HEX, is a new bootstrapped team building the environments, applications, compute, product-data systems, and workflows that let hardware engineers do their work without needing to troubleshoot the machinery underneath. HEX owns virtual engineering environments, HPC/GPU compute, storage, networking, licensing, MCAD/ECAD/CAE applications, PLM, product data, automation, validation, and support as one connected system. About the Role As a Staff PLM & Engineering Applications Engineer, you will be one of the first technical builders of HEX and the primary counterpart to the HEX lead. You will own the engineering-application and product-data side of the hardware engineering experience, with an initial focus on NX, Teamcenter, licensing, parts import, integrations, packaging, validation, and user workflows. This is not a traditional Teamcenter administration role and not a Corporate IT application-support role. You will take complex, fragile workflows and turn them into reliable engineering systems. This role is highly hands-on and systems-oriented. You will not inherit a mature environment and support queue. You will help build a fresh one, replacing manual setup guides, tribal knowledge, repeated support issues, and team handoffs with tested automation and reliable workflows. In This Role, You Will Own the technical architecture, deployment, configuration, integration, validation, and long-term operation of NX and Teamcenter. Build reliable workflows for parts import, product-data migration, metadata quality, BOMs, revisions, lifecycle states, and releas
About the Team The Software Engineering Firmware team builds reliable, high-performance systems on custom hardware. We work closely with hardware engineers to design, optimize, and ship software that bridges cutting-edge devices and real-world constraints like memory, power, and latency. Our work spans early prototyping through product launch, ensuring that our embedded platforms are robust, efficient, and production-ready. About the Role As a Firmware Engineer , you will design, implement, and debug software for embedded devices. You’ll own low-level bring-up, write production C/C++ code, and partner closely with hardware teams to deliver reliable, high-performance systems. We’re looking for engineers with deep embedded expertise, strong debugging skills, and a passion for building systems that perform under real-world conditions. This role is based in San Francisco, CA . We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement, and debug software for embedded devices. Contribute to defining software requirements, interfaces, and test plans. Bring up and debug new boards. Analyze performance, memory, and power profiles and implement optimizations. Investigate field issues, perform root-cause analysis, and deliver robust fixes. Foster good software engineering practices. You might thrive in this role if you: Have deep experience shipping embedded systems (around 10+ years). Are proficient in C and C++. Are familiar with embedded toolchains, operating systems, and debugging tools. Have experience with both rapid prototyping and scalable product development. (Nice to have) Have experience with Zephyr RTOS. (Nice to have) Have worked with networking/wireless stacks (BLE, Wi-Fi). (Nice to have) Have experience with robotic system bring-up or Linux kernel development. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose arti
About the Team The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads. As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI’s rapidly expanding compute environment. About the Role We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI’s third-party infrastructure partners, including major cloud service providers and strategic compute vendors. In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership. Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure. Key Responsibilities Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations. Develop operationa
About the Team Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training. Our mission is to make some of the world's largest GPU clusters work reliably for frontier training. We bring new platforms and clusters online, safely maintain installed fleets, and partner with hardware, infrastructure, and research teams to resolve the system-level issues that keep jobs from running. That means building and maintaining the software closest to the machine: Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation. We make these components reproducible, compatible, and safe to operate across heterogeneous fleets. About the Role We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate. You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Engineering Program Manager, you will help turn complex infrastructure strategy into executable programs across electrical, mechanical, controls, network, hardware, construction, commissioning, deployment, and operations workstreams. You will partner with research, hardware engineering, data center engineering, site development, supply chain, security, EHS, finance, legal, operations, and external delivery partners to bring OpenAI's infrastructure vision to life. About the Role We are looking for an Engineering Program Manager (EPM) to lead assigned infrastructure programs focused on production and non-production network integration, controls coordination, and the design and deployment of data hall or whitespace facilities. The EPM will support functional Directly Responsible Individuals (DRIs) across network, controls, structural, electrical, and mechanical disciplines. Key responsibilities include coordinating assigned workstreams and program controls, maintaining risks and interfaces, and supporting readiness within the network and data hall deployment track. The ideal candidate thrives on bringing structure to complex environments characterized by ambiguous technical requirements, large partner ecosystems, tight deadlines, and high operational stakes. This individual must be adept at keeping teams aligned on decisions, risks, dependencies, schedules, and readiness criteria, and escalating gaps or decision points when needed. Candidates should have a proven track record of managing technically challenging engineering programs across major lifecycle phases, including design, validation, procurement, construction, c
About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de
About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. In partnership with leading cloud providers, hardware manufacturers, utilities, construction partners, and internal engineering organizations, we are delivering hyperscale AI campuses that power the next generation of frontier AI models. Infrastructure Delivery Operations sits at the center of this effort. Our team develops the operating model that connects infrastructure strategy, supply planning, manufacturing operations, and delivery into a single, integrated system that enables OpenAI to deploy AI infrastructure predictably at scale. We partner across Hardware Engineering, Network Engineering, Capacity Delivery, Hardware Operations, Security, Finance, Strategic Sourcing, and external infrastructure partners to create a single, integrated view of program health. Through governance, operational analytics, executive reporting, and scalable operating mechanisms, we enable leaders to proactively manage risk, optimize capacity, and deliver infrastructure predictably at Industrial Compute speed. About the Role We are seeking a Technical Program Manager, Infrastructure Delivery Operations to drive integrated strategy and delivery across OpenAI's rapidly expanding AI infrastructure portfolio. This role sits at the intersection of infrastructure strategy, New Product Introduction (NPI), supply planning, manufacturing operations, and infrastructure delivery. You will lead highly cross-functional programs spanning engineering, supply planning, manufacturing, logistics, construction, commissioning, and operations, ensuring technical and operational dependencies remain synchronized from planning through production readiness. Beyond driving program execution, you will leverage operational insights to improve capacity planning, infrastructure strategy, and deployment readiness. You will also help operationalize new technologies and suppliers by partnering w
About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Through strategic partnerships and self-built campuses, we are scaling one of the world's fastest-growing AI infrastructure platforms. The Supply Chain organization ensures critical infrastructure components—from compute systems and networking equipment to integrated rack solutions—are sourced, manufactured, qualified, and delivered with the speed and reliability required to support frontier AI development. We partner closely with Hardware Engineering, Manufacturing Quality Engineering, Infrastructure Delivery, Hardware Operations, Finance, and suppliers worldwide to build a resilient, scalable supply chain capable of supporting rapid infrastructure expansion. As Industrial Compute continues to grow, Supply Chain serves as the operational bridge between engineering innovation and large-scale infrastructure deployment. About the Role We are seeking a Supply Chain Manager to lead strategic execution across sourcing, supplier operations, manufacturing quality, and infrastructure delivery for OpenAI's AI infrastructure portfolio. This role will oversee a multidisciplinary team responsible for strategic sourcing, manufacturing quality engineering, and technical program management while partnering closely with engineering, finance, hardware operations, and deployment teams. You will drive supplier strategy, manufacturing readiness, production planning, quality performance, and operational execution across the full hardware lifecycle. Success requires balancing long-term supplier strategy with day-to-day execution. You'll establish scalable operating mechanisms, strengthen supplier partnerships, manage complex cross-functional programs, and ensure OpenAI can rapidly deploy AI infrastructure without compromising quality, cost, or reliability. This is a people leadership role responsible for developing a high-performing organization while driving operati
About the Team Our Robotics team is focused on unlocking general-purpose robotics and advancing toward AGI-level intelligence in dynamic, real-world environments. Working across the full model and systems stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the physical constraints of real-world systems to improve people’s lives. About the Role We are seeking an experienced commercial attorney to serve as the primary legal partner supporting Robotics. In this role, you will provide practical, business-oriented legal counsel across hardware development, manufacturing, supply chain, procurement, and strategic commercial initiatives. You will work closely with Robotics leadership and cross-functional partners to help build scalable legal frameworks that enable innovation while thoughtfully managing risk. This is a unique opportunity to help shape the legal foundation of a rapidly growing robotics organization developing cutting-edge technologies. You will negotiate high-impact commercial agreements, advise on complex operational matters, and partner closely with technical and business teams to support the development and commercialization of next-generation robotics systems. This role is based in San Francisco, CA and requires in-person presence 4 days a week. In this role, you will: Serve as the primary legal partner supporting Robotics leadership and cross-functional teams. Provide practical, business-oriented legal advice to teams across hardware engineering, manufacturing operations, supply chain, procurement, finance, and operations. Draft, review, and negotiate complex commercial agreements, including supplier, manufacturing, development, consulting, licensing, procurement, and strategic partnership agreements. Advise on legal issues arising throughout the hardware development lifecycle, including manufacturing, supply chain operations, vendor relatio
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. ABOUT THE ROLE This role owns Baseten's relationships and market intelligence across the hardware and chip layer of the compute stack: NVIDIA directly and key OEM partners such as Dell, Lenovo, Pegatron, and Supermicro. As Baseten's compute strategy increasingly depends on hardware access and terms, this role is central to keeping Baseten ahead of the market. WHAT YOU'LL DO Build and maintain relationships across NVIDIA and key OEM partners (e.g. Dell, Supermicro) Track market intelligence on hardware availability, roadmaps, and terms to keep Baseten informed and strategically well-positioned Support deal structuring and negotiation in partnership with Baseten's deal-making function Work closely with Infrastructure and Hardware Platform engineering teams to ensure consistent, high-quality provider relationships and engineering partnerships Represent Baseten credibly across senior relationships in the hardware ecosystem, escalating to company leadership when strategically valuable WHAT WE'RE LOOKING FOR Existing relationships and credibility within the NVIDIA, OEM, and HPC ecosystem Strong relationship-management instincts, with the judgment to know when to bring in senior leadership for maximum impact Comfort operating in a fast-moving, high-stakes market where hardware access can be a major competitive differentiator Collaborative style — this role depends on close coordination with engineering counterparts, not just ex
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role As an Engineer on our hardware optimization and co-design team, you will co-design future hardware from different vendors for programmability and performance. You will work with our kernel, compiler and machine learning engineers to understand their unique needs related to ML techniques, algorithms, numerical approximations, programming expressivity, and compiler optimizations. You will evangelize these constraints with various vendors to develop and influence future hardware architectures towards efficient training and inference on our models. If you are excited about efficiently distributing a large language model across devices, dealing with and optimizing system-wide/rack-wide networking bottlenecks and eventually tailoring the compute pipe and memory hierarchy of the hardware platform, simulating workloads at different abstractions and working closely with our partners, this is the perfect opportunity! This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Key Responsibilities Co-design future hardware for programmability and performance with our hardware vendors Assist hardware vendors in developing optimal kernels and add support for it in our compiler Develop performance estimates for critical kernels for different hardware configurations and drive decisions on compute core and memory h
Other cities to consider
More places hiring for this role
Get new hardware engineer jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime