About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui
Jobs in United States
Hardware Systems Planning Lead in San Francisco
153 active opportunities · Updated September 2026
Showing
15 jobs
Explore current hardware systems planning lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Industrial Compute team is responsible for building and scaling large-scale compute capacity across first-party data centers, strategic partners, and industrial infrastructure environments. We focus on converting power, land, hardware, and operational execution into reliable compute capacity that can support frontier AI training and inference workloads. This team operates at the intersection of infrastructure delivery, hardware systems, utilities, supply chain, and capacity strategy—ensuring OpenAI can scale compute faster than traditional models allow. About the Role We are seeking a Tokens-as-a-Service (TaaS) Lead to drive the end-to-end conversion of industrial-scale infrastructure investments into usable token capacity for OpenAI workloads. In this role, you will own execution across complex compute programs where raw infrastructure capacity must be transformed into operational GPU throughput. You will coordinate across data center delivery, power, networking, hardware deployment, workload enablement, finance, and external partners to ensure capacity becomes productive tokens as quickly and efficiently as possible. This role is ideal for someone who can bridge physical infrastructure delivery with compute utilization outcomes. Success requires strong systems thinking, elite program leadership, and the ability to drive accountability across internal teams and strategic partners. In this role, you will Lead Tokens-as-a-Service programs across industrial compute environments, including first-party and partner-owned capacity. Convert delivered power, space, and hardware capacity into production-ready token throughput. Build integrated execution plans spanning construction, power energization, rack deployment, networking, cluster readiness, and workload onboarding. Partner with infrastructure engineering, hardware, networking, finance, supply chain, and operations teams. Drive external providers, EPCs, OEMs, utilities, and strategic partners t
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are looking for a systems-minded engineer to help advance our kernel development, performance engineering, and hardware-software co-design capabilities, with a particular focus on AI-assisted workflows and tooling. This person will work at the intersection of kernel optimization, developer tooling, observability, and research infrastructure, helping us improve both how production kernels are built and optimized, and how future hardware-software systems are designed and evaluated. The role is ideal for someone who is excited by low-level performance work, but also sees AI and automation as powerful tools for accelerating engineering velocity. You will help define the future of kernel engineering in the era of AI-assisted development. In this role, you may: Build developer tooling and workflows that make kernel development and performance optimization faster, more scalable, and easier to debug, integrate, and deploy. Develop observability, diagnostics, and validation infrastructure that makes AI-assisted optimization systems more interpretable, reliable, and effective. Optimize production kernels end to end by formulating optimization problems, running search loops, analyzing bottlenecks, debugging generated implementations, and landing improvements into production. Design abstractions, interfaces, and automation systems that accelerate kernel optimization, correctness validation, and hardware-software co-design. Improve AI-assisted optimization systems for sp
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. About the role We are looking for an Operating Systems Engineer to build and harden the OS foundations for OpenAI products. We are especially interested in experienced, passionate, and innovative operating systems developers who thrive on building foundational platform software and solving hard problems in security, privacy, performance, power, and reliability. You will work across the OS kernel, core OS services, security and privacy primitives, performance and power, and the frameworks that connect applications and UI to the system. This role emphasizes deep debugging and systems ownership from development through production. You will collaborate closely with embedded, firmware, hardware, application, and product engineering teams. Experience with hardware bring-up is a plus, but not required. What you will do Work on end-to-end OS capabilities spanning the OS kernel, userspace services, application frameworks, UI toolkits, and application-facing APIs. Develop, integrate, and maintain OS components, both kernel-bound and in userspace, including scheduling, memory management, filesystems, drivers, IPC/RPC mechanisms, and security-relevant subsystems. Build and maintain core OS services and daemons (init, service management, device discovery, networking primitives, time, logging, update hooks, crash handling, and so on). Design and implement security and privacy mechanisms: Secure boot and measured boot integration points (where applicable). Mandatory access control and sandboxing. Secrets management, secure storage, key handling, and least-privilege service design. Privacy-preserving telemetry, data minimization, and user-consent oriented system behaviors. Establish a perfo
About the Team The Consumer Products team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. Within Consumer Products, the camera stack is a critical sensing component. The team partners closely with electrical engineering, silicon vendors, systems, and higher-level perception and product teams to bring up new hardware, stabilize capture pipelines, and ensure camera systems are robust, debuggable, and ready for real-world deployment. This work spans early prototypes through production, with a strong emphasis on correctness, repeatability, and long-term reliability. About the Role As a Camera Firmware Engineer, you will own low-level camera enablement on custom hardware—from early board bring-up through stable production capture. You will develop and maintain the firmware and software that makes camera sensors reliable, controllable, and debuggable, forming the foundation for higher-level camera pipelines and product features. This role is highly hands-on and systems-oriented. You will work close to the hardware, diagnose real-world timing and integration issues, and build tooling that accelerates iteration across the entire camera stack. This role is based in San Francisco, CA. We follow a hybrid work model with four days per week in the office and offer relocation assistance to new employees. In This Role, You Will Bring up new camera sensors and modules on prototype and production boards, including link stability, sensor control, and correct power, reset, and clock sequencing. Develop and maintain low-level camera software, including sensor drivers, board configuration, and camera subsystem integration across hardware revisions. Enable and validate core capture paths for development and production, including RAW capture for debugging, still capture, and hardware-
About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples' lives. About the Role You will work at the cutting edge of OpenAI's robotics hardware, partnering with hardware and software teams to develop, build, test, and iterate on robotic systems. Working closely with engineers, you will build prototypes, execute experiments, fabricate fixtures, troubleshoot hardware, and help drive rapid iteration on new robotic technologies. Your practical problem-solving and technical judgment will help move projects from concept to reality. We are looking for versatile, hands-on generalists who enjoy solving problems across mechanical, electrical, fabrication, and experimental domains. The ideal candidate will have strong experience in high-velocity and early-stage hardware development environments, provide thoughtful feedback to engineering teams, and help improve both hardware and development processes. This role will help accelerate robotics development by enabling fast, high-quality iteration on new hardware concepts and systems. This role is based in San Francisco, CA, and is in-person 5 days a week. In this role, you will: Partner closely with engineers, researchers, and other members of the robotics team to support prototype development, experimentation, and hardware iteration. Build, modify, troubleshoot, and repair robotic systems and electromechanical assemblies spanning structural, electrical, sensing, and actuation subsystems. Design and fabricate fixtures, adapters, test equipment, and other prototype hardware that accelerate development efforts. Execute low-volume builds, engineering changes, and rework activities acro
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for an experienced systems software engineer to help define and build the host software stack for our custom next-generation AI systems. You will work close to the hardware on performance-critical software, including Linux kernel drivers, high-throughput I/O paths, and system-scale networking and RDMA. This role spans architecture, implementation, platform bring-up, debugging, and performance optimization. You will work across hardware and software boundaries to make new systems usable end to end, from low-level device interfaces through userspace tooling and production validation. In this role you will: Design, implement, and debug host-side systems software for AI infrastructure, including Linux kernel drivers and supporting userspace components. Build and optimize software paths for high-throughput, low-latency communication, including RDMA and related networking functionality. Develop software around PCIe, DMA, NICs, accelerators, memory movement, and device interaction. Bring up new hardware platforms and diagnose complex issues across kernel, firmware, networking, and hardware boundaries. Build tooling for integration, testing, diagnostics, observability, qualification, and performance characterization. Collaborate with hardware, networking, and platform teams to define interfaces and integrate new capabilities. Work with external vendors where needed to integrate technologies and drive issues to resolution. Contribute across the systems sof
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions tailored to the demands of advanced AI workloads. We work across the full stack—from silicon to system integration—partnering closely with internal teams and external vendors to define and deliver next-generation AI infrastructure. Our team focuses on defining scalable, high-performance system architectures and reference designs that balance performance, cost, and operational efficiency across rapidly evolving technologies. About the Role We are seeking a 3P Architect to define and drive rack- and cluster-level reference designs in collaboration with external partners. This role is responsible for translating workload requirements and system-level goals into concrete architectures, aligning partners on critical design attributes, and ensuring vendor roadmaps meet our infrastructure needs. You will work closely with performance modeling and internal architecture teams to evaluate tradeoffs, while owning the end-to-end definition and execution of third-party system designs. This includes identifying gaps in current technologies, driving vendor development, and shaping future infrastructure capabilities. This role requires strong system intuition, cross-functional leadership, and the ability to operate effectively across internal teams and external ecosystems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Define rack- and cluster-level reference architectures for AI infrastructure deployments. Translate workload requirements into clear system design specifications and partner deliverables. Collaborate with performance modeling teams to evaluate architectural tradeoffs and system behaviors. Align internal stakeholders and external partners on critical system attributes (performance, cost, power, reliability, scalability). Identify gaps in current technology offerings and dr
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will work on the systems software strategy and execution that brings new AI silicon from first power-on to a fully integrated system running production-representative models at expected functionality and performance. You will define how software exercises and validates compute, memory, interconnect, and I/O subsystems, then build the diagnostics, automation, and observability needed to find issues quickly. This role sits at the center of silicon, firmware, platform, systems, and workload teams. You will turn hardware specifications and performance targets into an end-to-end bringup plan, drive cross-functional debug, and establish the stress and regression infrastructure that makes each new platform reliable across operating environments. In this role, you will: Contribute to the end-to-end software bringup and validation strategy for new silicon and first-party systems. Define software-driven test coverage across compute, memory, interconnect, I/O, and their system-level interactions. Build diagnostics, test automation, telemetry, and regression infrastructure that accelerate first-silicon learning and issue isolation. Lead bringup from initial silicon arrival through board and system integration, docking, runtime enablement, and model execution. Design stress tests that characterize reliability, performance, and stability across workloads and operating conditions. Translate architecture specifications and performance models into measurable acceptance crit
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will develop and evolve the tooling ecosystem that hardware engineers rely on every day — from hardware compilers and IR transformations to simulation, debugging, and automation infrastructure. The work spans software engineering, compiler concepts, and practical hardware workflows, with direct impact on how quickly and effectively we design next-generation AI systems. You’ll collaborate closely with architects, RTL designers, and verification engineers to translate real engineering friction into durable, scalable tooling solutions. In this role you will: Build and improve the software tooling that makes hardware teams faster: compilation, IR transforms, RTL generation, simulation, debug, and automation. Extend and integrate hardware compiler stacks (frontends, IR passes, lowering, scheduling, codegen to Verilog/SystemVerilog) and connect them to real design workflows. Improve developer experience and reliability: reproducible builds, better error messages, faster iteration loops, and dependable CI and regression infrastructure. Work closely with designers and verification engineers to turn real pain points into durable tools. Dive into RTL when needed: read and reason about Verilog/SystemVerilog to debug issues, validate tool output, and improve debuggability. Be willing to go all the way down the stack when necessary, including gate-level views, synthesis results, and implementation artifacts. Help enable PPA optimization loops by building analysis and au
About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. The Systems Integration team is critical in this mission, turning complex hardware-software development into reliable product signals. We validate complete device experiences across software, cloud services, connectivity, accessories, and real-world operating environments, combining hands-on system testing, structured test development, hardware-in-the-loop environments, diagnostics, and automation to uncover issues that component-level testing alone cannot reveal. About the Role As a Systems Test Engineer, End-to-End Validation , you will design and execute end-to-end testing for complex device experiences spanning hardware, software, connectivity, cloud services, and accessories. You’ll translate product behavior and real-world use cases into structured, reproducible test procedures and build test environments that allow failures to be reliably reproduced and diagnosed. You’ll also identify opportunities to automate repetitive or high-value scenarios, working with engineers to turn complex manual workflows into scalable validation systems. Because this is a new category of devices, you’ll have the opportunity to build the end-to-end validation foundation early—shaping test coverage, environments, and workflows from prototype through launch. We’re looking for someone who combines strong systems thinking, hands-on testing skills, technical curiosi
About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an OS / K8s Systems Engineer at Baseten, you’ll build the automation and systems that turn raw GPU hardware into production-ready compute. From provisioning to orchestration, you’ll own the software layer that makes our infrastructure reproducible, scalable, and reliable across data centers. This is a senior, hands-on role focused on building systems not operating them. You’ll work close to the metal designing OS images, building provisioning pipelines, and automating cluster bring-up from scratch. Your work will define how quickly we can turn new capacity into usable compute. EXAMPLE INITIATIVES Zero-to-cluster automation Build workflows that take new hardware from unprovisioned to fully operational cluster. Provisioning systems Design PXE-based or equivalent systems for imaging and lifecycle management. Reproducible infrastructure — Ensure clusters deploy consistently across data centers. RESPONSIBILITIES Own the end-to-end automation of cluster bring-up and lifecycle management. Build and maintain OS images, provisioning systems, and configuration pipelines. Deploy and operate cluster orchestration platforms (Kubernetes, Slurm, or similar). Design systems for reproducibility across sites and hardware generations. Automate upgrades, rollouts, and failure recovery. Optimize system performance, including GPU utilization and networking. Partner with hardware and network teams to validate and improve system b
Other cities to consider
More places hiring for this role
Get new hardware systems planning lead jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime