About the Team OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical. The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly. About the Role We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity. In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity. This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions. This role is based in San Francisco, CA, with travel as needed. Key Responsibilities Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint Drive readiness to convert contracted compute capacity into schedulable production clusters Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives Build integrated schedules spanning procurement, logistics, installation, st
Jobs in United States
Server Cpu Hardware Systems Lead in San Francisco
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current server cpu hardware systems lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de
Technical Program Manager – Applied Infrastructure About the Team The Applied team safely brings OpenAI’s technology to the world, powering products like ChatGPT, and the APIs for GPT and more. Behind these products is a complex and rapidly evolving infrastructure platform that enables scale, performance, and safety. The Applied Infrastructure TPM team partners across engineering to lead foundational programs that ensure OpenAI’s infrastructure can meet current and future demand. About the Role We’re looking for a seasoned Technical Program Manager to drive critical infrastructure programs across the Applied organization. This TPM will focus on cross-cutting initiatives such as general compute capacity planning, process transformation, cost and quota attribution and optimization, and coordination across infrastructure and product stakeholders. There will also be focus on evolving OpenAI’s infrastructure to support growth, scale and new products. This work is core to how OpenAI manages and grows its infrastructure footprint in a disciplined, scalable way. Location: San Francisco, CA (Hybrid – 3 days/week in-office) In this role, you will: Serve as the DRI for complex infrastructure programs spanning CPU planning, orchestration, and other resource management domains (e.g. networking, storage). Build and operationalize systems to capture demand signals, model future capacity needs, and align infrastructure planning across internal teams and partners external to the company. Partner closely with Infrastructure, Product and Finance teams to forecast infrastructure usage patterns and ensure supply/demand alignment. Lead cost attribution and quota enforcement programs to promote stability and ensure equitable access to resources across teams. Drive simplification and standardization of infrastructure tooling and processes across Applied and Infra organizations. Drive cross functional programs to evolve our infrastructure to support new growth and scale Work with external v
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI compute infrastructure ecosystem. The InfraDev team is central to this mission, setting the strategy and executing the roadmap to scale our supercomputing footprint globally. From site planning to system integration, this team operates at the intersection of commercial, technical, and operational excellence, partnering with leaders across OpenAI and the industry. About the Role We are seeking a Strategic Sourcing Manager who is ready to take on global-scale challenges in AI compute supply and manufacturing — an opportunity to shape the future of supercomputing. As part of the Infrastructure Strategy & Delivery organization supporting Industrial Compute, OpenAI’s next-generation supercomputing platform, you will lead sourcing and strategic supplier engagements for compute infrastructure at hyperscale. This role will drive commercial strategy and supplier accountability across server platforms, accelerators, rack systems, and associated thermal & power delivery components. This is not traditional procurement — it is foundational work enabling OpenAI to deploy compute faster and more efficiently than anyone in the world, while building deeply-integrated partnerships with the global compute supply chain. Key Responsibilities Develop and execute sourcing strategies for the next generation of AI compute infrastructure in partnership with engineering and program leadership. Stay at the leading edge of industry and supply-chain trends to inform category strategy and long-range planning. Initiate, negotiate, and manage commercial agreements across OEM/ODM/JDMs for accelerators, GPUs/CPUs, server platforms, rack systems, liquid-cooling components, PSUs, and supporting mechanical/electrical subsystems. Secure capacity and optionality across a rapidly scaling compute supply chain while mitigating risk and ensuring manufacturing readiness. Partner cross-funct
About the Team We’re hiring software engineers to make OpenAI’s networking teams more productive. These teams build and operate the high-performance networking systems that support OpenAI’s training and inference infrastructure at frontier scale. About the Role We’re looking for someone who cares deeply about the developer experience of engineers working on complex infrastructure systems — especially around build systems, test architecture, release pipelines, and reliable development workflows. This role will be embedded with OpenAI’s networking team: making it faster, safer, and easier for engineers to build, test, validate, and ship changes across multi-server, networked, and hardware-adjacent environments. In this role you will: Improve development workflows for engineers building and operating OpenAI’s networking systems Design and improve continuous deployment, release, and validation pipelines Build and maintain test harnesses for multi-server, networked, and hardware-backed environments Improve iteration speed across C++, Python, and build-system-heavy codebases Partner with engineers to identify friction in CI, testing, debugging, and deployment workflows Drive testing and reliability strategy for infrastructure components that support large-scale training and inference workloads Work closely with centralized developer experience teams while staying deeply embedded with the networking engineers closest to the systems You might thrive in this role if: You are motivated by helping other engineers move faster and with more confidence You have experience with CI/CD, release pipelines, testing infrastructure, or build systems You are comfortable moving between C++, Python, and build systems such as CMake, Bazel, or Blaze You enjoy building test harnesses, automation, and workflow improvements for complex systems You do not need to be a networking expert, but you are excited to learn enough about the domain to make the team meaningfully more effective When you see
About the Team OpenAI’s mission is to ensure the responsible and widespread adoption of artificial intelligence. In support of that mission, the Ads Solutions team partners closely with advertisers to deeply understand their businesses and needs, helping inform the development of products and solutions that drive meaningful revenue growth and long-term success on the platform—while maintaining strong standards for user trust and platform integrity. About the Role We’re looking for an Ads Solutions Engineer to partner with advertisers and agencies throughout the pre-sales and growth lifecycle, helping them successfully evaluate, launch, and scale on OpenAI’s advertising platform. You will serve as the technical expert in the sales process, translating advertiser objectives into scalable technical solutions across measurement, integrations, data activation, and campaign execution. This role sits at the intersection of sales, product, and engineering. You’ll work closely with Client Partners, Customer Success Managers, Product, Engineering, Policy, and Operations teams to remove technical blockers, accelerate revenue growth, and shape the future of OpenAI’s ads platform. This role is based in our San Francisco HQ and we offer generous relocation support for new hires. In this role, you will: Partner with Client Partners during the sales cycle to provide technical expertise that unlocks new business and accelerates deal closure. Lead technical discovery with advertisers and agencies to understand data flows, martech stack, measurement requirements, and activation goals. Design and recommend implementation approaches for pixels, APIs, server-to-server integrations, identity solutions, offline conversions, and measurement frameworks. Guide advertisers through onboarding and launch readiness, ensuring successful setup of tracking, attribution, audience signals, and reporting. Troubleshoot technical issues related to implementation, signal quality, delivery, and measurement
About the Team OpenAI’s mission is to ensure the responsible and widespread adoption of artificial intelligence. In support of that mission, the Ads Solutions team partners closely with advertisers to deeply understand their businesses and needs, helping inform the development of products and solutions that drive meaningful revenue growth and long-term success on the platform—while maintaining strong standards for user trust and platform integrity. About the Role We’re looking for an Ads Solutions Engineer to partner with advertisers and agencies throughout the pre-sales and growth lifecycle, helping them successfully evaluate, launch, and scale on OpenAI’s advertising platform. You will serve as the technical expert in the sales process, translating advertiser objectives into scalable technical solutions across measurement, integrations, data activation, and campaign execution. This role sits at the intersection of sales, product, and engineering. You’ll work closely with Client Partners, Customer Success Managers, Product, Engineering, Policy, and Operations teams to remove technical blockers, accelerate revenue growth, and shape the future of OpenAI’s ads platform. This role is based in our San Francisco HQ and we offer generous relocation support for new hires. In this role, you will: Partner with Client Partners during the sales cycle to provide technical expertise that unlocks new business and accelerates deal closure. Lead technical discovery with advertisers and agencies to understand data flows, martech stack, measurement requirements, and activation goals. Design and recommend implementation approaches for pixels, APIs, server-to-server integrations, identity solutions, offline conversions, and measurement frameworks. Guide advertisers through onboarding and launch readiness, ensuring successful setup of tracking, attribution, audience signals, and reporting. Troubleshoot technical issues related to implementation, signal quality, delivery, and measurement
About the Team The Growth Platforms team builds the systems and operating foundations that help OpenAI grow responsibly. We partner across the product portfolio to connect customer signals, identity and consent, campaign workflows, measurement, and product experiences into an AI-enabled growth engine. Our work helps teams launch, learn, and scale with a high bar for data quality, privacy, reliability, and customer trust. About the Role We’re looking for an experienced marketing technology and operations leader to drive cross-functional work at the intersection of growth, measurement, data, and automation. Your mission will be to turn fragmented tools, signals, and workflows into reliable, measurable, AI-enabled capabilities that teams can use safely at scale. You’ll work across Growth, Marketing Operations, Product, Engineering, Data Engineering, Data Science, Security, Privacy, Legal, and Revenue Operations, as well as external advertising platforms, measurement providers, and implementation partners. You’ll translate business requirements and privacy constraints into data contracts, integration designs, rollout plans, and reliable first-party data systems. This is a hands-on, high-impact role for someone who brings structure to ambiguity and moves from event schemas, APIs, and data quality assurance to operating cadences, partner enablement, and executive updates. This role is based in San Francisco or New York City with a hybrid office expectation. In this role, you will: Own the operating model for Growth’s marketing technology stack across identity, consent, audiences, activation, measurement, and experimentation. Own and operate the complete paid-media tracking and measurement system, including website pixels, server-to-server conversion events, mobile measurement integrations, identity and consent controls, attribution methods, and timely signal delivery to advertising platforms. Design and implement event schemas, data mappings, APIs, and integrations; valid
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Voice is becoming the internet’s next interface, but a production-grade Voice AI system is "hard to build" . You’ll join a small founding team of Baseten Voice AI, focused on bringing state-of-the-art open source models into production for Voice AI customers across productivity, customer service, clinical conversation, creator tools, education, and more. You’ll make a meaningful impact on people’s daily lives and help reshape these industries. This is a high-impact, high-ownership role. You will be the primary owner of Baseten Voice AI - our in-house inference stack to power Voice AI models - from product roadmap through engineering implementation. You’ll partner closely with Forward Deployed Engineers, Model Performance Engineers, and sister engineering teams to push the boundaries of Voice AI. EXAMPLE INITIATIVES: Develop world-class model serving stack for state-of-the-art open-source voice models - reduce end-to-end and tail latency (p95/p99), increase throughput, and improve GPU efficiency via profiling, runtime tuning, and server-level optimizations. Build large-scale, real-time infrastructure for multi-model voice agents - orchestrate STT, TTS, and agent components with streaming I/O to meet customer SLOs. Design tight training and inference iteration loops for voice model customization - enable fast evaluation, safe rollout, and rapid experimentation for custom voice model development. Past projects:
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About this role Are you ready to redefine the future of JavaScript development? Are you a seasoned JavaScript expert who gets a thrill from tackling complex challenges across the entire JavaScript ecosystem? Do you believe that AI can be a powerful partner in crafting elegant, high-impact code? If you're ready to leave the mundane behind and join a team that's shaping the tools used by millions of developers globally, we've got an opportunity for you at Sentry. This isn't your typical Senior Software Engineer position. As a key member of our growing JavaScript SDK team, you'll be at the forefront of innovation, working on everything from our cutting-edge SDKs for Node.js, Bun, Deno, Cloudflare Workers, and other modern server runtimes. You won't just be maintaining code; you'll be pushing the boundaries of what's possible in developer tooling across the rapidly evolving server-side JavaScript landscape. In this role you will Join our JavaScript SDK team and get ready to build the future. You'll be at the forefront, working on: A Universe of JavaScript Challenges: Dive deep into our extensive suite of JavaScript SDKs, with a sharp focus on server-side and edge runtimes — from the battle-tested Node.js ecosystem to cutting-edge alternatives like Bun and Deno, and distributed edge environments like Cloudflare Workers. Your work will directly empower millions of developers to build better, more reliable software, no matter which runtime powers their stack End-to-End Ownership: We believe in giving our engineers the autonomy to see their vision through. You'll have the freedom to plan, implement, and ship your code, from writing
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About this role Are you ready to redefine the future of JavaScript development? Are you a seasoned JavaScript expert who gets a thrill from tackling complex challenges across the entire JavaScript ecosystem? Do you believe that AI can be a powerful partner in crafting elegant, high-impact code? If you're ready to leave the mundane behind and join a team that's shaping the tools used by millions of developers globally, we've got an opportunity for you at Sentry. This isn't your typical Senior Software Engineer position. As a key member of our growing JavaScript SDK team, you'll be at the forefront of innovation, working on everything from our cutting-edge SDKs for the React, Next.js, Vue, Nuxt, Hono, NestJS, and beyond. You won't just be maintaining code; you'll be pushing the boundaries of what's possible in developer tooling across the full spectrum of the modern JavaScript framework landscape. In this role you will Join our JavaScript SDK team and get ready to build the future. You'll be at the forefront, working on: A Universe of JavaScript Challenges: Dive deep into our extensive suite of JavaScript SDKs, with a broad focus on framework support spanning the modern JS ecosystem — from frontend frameworks like React, Vue, and their meta-frameworks Next.js and Nuxt, to server-side and edge runtimes like NestJS and Hono. Your work will directly empower millions of developers to build better, more reliable software, regardless of their framework of choice End-to-End Ownership: We believe in giving our engineers the autonomy to see their vision through. You'll have the freedom to plan, implement, and ship your code, from writi
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As an AI Accelerator Systems Software Technical Program manager at OpenAI, you will help bring our chips/system hardware roadmap to life, navigating an array of technical and partnership challenges. We’re looking for people excited to push the frontiers of computing by navigating technical explorations and are passionate about building. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Manage the end-to-end software development from design to implementation for our AI acceleration systems, working across technical, cross-functional and external stakeholders Lead planning and scheduling of AI system software designs with our strategic partners and vendors Coordinate and lead internal resources and communication for efficient interaction with partners and vendors. You might thrive in this role if you: Have experience as a software technical program manager for data center system products (server, GPU, TPU, networking, storage and so on) taking products from concept to volume in a data center environment ensuring the systems scale with high quality Know end-to-end software development program management techniques from concept, design, production, deployment into the data center Want to help design some of the world’s largest supercomputing systems, working at the edge of complex hardware challenges Enjoy working with and enabling world-clas
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet Hardware team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automation to reduce manual work
Other cities to consider
More places hiring for this role
Get new server cpu hardware systems lead jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime