About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside leading cloud providers, engineering firms, construction partners, utilities, and equipment manufacturers, we are delivering hyperscale AI campuses that enable the next generation of frontier AI models. The Strategic Sourcing team develops and executes the commercial strategies that ensure our infrastructure programs have reliable access to the equipment, materials, and strategic partners needed to deliver at unprecedented scale. We partner closely with Infrastructure Delivery, Capacity Planning, Design Engineering, Hardware Operations, Finance, Legal, and our external suppliers to build a resilient global supply network capable of supporting Industrial Compute's long-term growth. As we continue expanding globally, strategic sourcing becomes a critical competitive advantage, ensuring our infrastructure programs remain cost-effective, resilient, and capable of executing against aggressive deployment timelines. About the Role We are seeking a Strategic Sourcing Manager, Data Center Infrastructure to lead sourcing strategy for the critical infrastructure systems that power Industrial Compute campuses. This role will develop commercial strategies, negotiate strategic supplier agreements, and manage relationships across engineering, construction, manufacturing, and infrastructure partners responsible for delivering mission-critical facilities. You will work closely with Infrastructure Delivery, Capacity Planning, Engineering, Finance, Construction, and external suppliers to ensure Industrial Compute has the capacity, supplier relationships, and commercial frameworks required to support rapid global expansion. The ideal candidate has experience sourcing major infrastructure systems for hyperscale data centers, mission-critical facilities, industrial construction, semiconductor manufacturing, energy infrastructure, or similarly comple
Jobs in United States
Compute Strategy And Transactions in San Francisco
145 active opportunities · Updated September 2026
Showing
15 jobs
Explore current compute strategy and transactions jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Research Program Management team partners with researchers and engineers to advance the development of increasingly capable, safe, and beneficial AI systems. We work alongside teams developing our core models, helping turn ambitious research goals into coordinated execution across model training, alignment and safety, and research infrastructure. The team also regularly collaborates with our closest cross-functional partners such as Security, Applied product and engineering, Strategy, and Scaling. About the Role As a Research Program Manager, you will embed with research teams and help drive some of the most technically complex and consequential work behind OpenAI’s model development. Depending on your focus, your work may span training, reasoning, evaluations, compute, research infrastructure, safety, model launch readiness, and governance. You will translate evolving research priorities into actionable programs, help teams navigate technical and operational tradeoffs, and keep important work moving as new issues emerge. This is a hands-on technical role: you will engage directly with research workflows, experimental results, technical systems, and engineering constraints; not simply coordinate from the sidelines. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. You might thrive in this role if you: Have 5+ years of experience in research program management, technical program management, or related roles in fast-moving environments. Can engage substantively with researchers and engineers on topics such as model training, experimental design, model safety, evaluation methods, data workflows, compute infrastructure, or distributed systems. Are comfortable working directly with technical tools, research data, experimental results, or operational workflows to understand problems and develop practical solutions. Have a strong track record of movi
About the Team The Strategic Finance team partners across OpenAI to turn company strategy into financial decisions, helping allocate resources and steer the business toward its highest-impact long-term outcomes. Within Strategic Finance, the B2B Product Finance team owns the financial perspective on growth across OpenAI’s B2B portfolio. We connect monetization and customer economics with compute demand, capacity, margin, and resource planning, partnering closely with Product, GTM, Strategic Deal Desk, Compute Finance, Capacity Planning, Data, Accounting, and other Finance teams. About the Role We are hiring a Strategic Finance leader for B2B Product to build the 0→1 financial foundations and decision-making frameworks that will help OpenAI scale its B2B business sustainably. You will own high-impact work across B2B monetization, compute demand modeling, enterprise deal economics, contribution margin, and more. This is a portfolio-wide individual contributor role spanning our API Platform and other B2B products. You will connect customer demand and commercial terms to revenue, compute consumption, and margin outcomes; influence some of our largest enterprise deals; and help leadership make sound growth investment, compute capacity, and resource allocation decisions. This role is based in San Francisco, CA. In this role, you will: Own an integrated view of B2B monetization and product economics across the API Platform and other B2B products – including pricing, packaging, channel and usage mix, discounts, credits, commitments, revenue, and customer profitability Build, maintain, and improve complex driver-based financial models that connect customer usage, product and model mix, pricing, discounting, credits, commitments, and contribution margin Build financial views for B2B compute demand and margin planning; identify risks and opportunities, explain key drivers, and recommend capacity and resourcing decisions that improve portfolio economics Evaluate some of OpenAI’
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with research, software, and external hardware partners to shape the next generation of AI systems, from silicon through full-scale deployments. Our team focuses on understanding and optimizing performance across the full system stack—ensuring that architectural decisions are grounded in rigorous, quantitative analysis of real-world workloads. About the Role We are seeking a Performance Modeling Lead to build and lead a small, high-impact team responsible for answering forward-looking architectural questions across AI infrastructure systems. You will develop modeling frameworks and methodologies to evaluate system-level tradeoffs and guide key design decisions. Your work will directly influence reference architectures, vendor designs, and long-term infrastructure strategy. This role sits at the intersection of AI workloads, system architecture, and quantitative modeling, and requires strong technical judgment, ownership, and the ability to translate complex analysis into clear, actionable guidance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Build and own a performance modeling framework/toolchain to evaluate AI systems across multiple levels of abstraction. Analyze and quantify architectural tradeoffs across compute, memory, networking, storage, and system topology. Develop performance models to guide decisions on: scale-up vs. scale-out architectures interconnect and network design memory hierarchy and system balance. Translate modeling outputs into clear recommendations for internal teams and external hardware vendors. Influence reference designs and vendor roadmaps through data-driven insights. Partner closely with machine learning, systems, and hardware teams to understand workload characte
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des
About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de
About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability Team Focus Areas Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely. Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads. Databases: Building high performance, distributed database systems that power all of OpenAI's product stack. In this
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Controls Network Engineer, you will design, validate, and scale the controls and OT network architectures that support high-density AI data centers. You will work across controls systems, OT infrastructure, telemetry, commissioning, deployment, and operations, partnering with mechanical, electrical, IT/networking, security, and external delivery teams. About the Role We are seeking a mid to senior OT Network Engineer with a strong controls systems background to lead the design and operation of resilient, secure, and scalable OT network architectures for high-density AI data centers. This role translates compute, power, cooling, and operational requirements into practical OT network designs, evaluates vendor solutions, and drives technical decisions across controls infrastructure, telemetry, commissioning, and operations. The ideal candidate has strong hands-on experience in mission-critical OT environments, including industrial networking, virtualized infrastructure, and OT network operations, with expertise in routing, switching, segmentation, firewall policy, time synchronization, monitoring, and network lifecycle support. Key Responsibilities Define controls, automation, and OT network requirements for AI data center campuses. Develop reference architectures, engineering standards, and reusable design templates. Review and develop basis-of-design and functional design documents, including OT network diagrams, IP/VLAN schemes, telemetry architectures, data flow diagrams, and commissioning requirements. Design OT and infrastructure network architectures, including physical topology, logical topology, IP addressing, subnetting, VLA
From $232K/yr
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Relevance & Personalization (R&P) team is Airbnb's matching intelligence engine — a talented team of ML engineers, applied researchers, and technical program managers who connect guests to the right listings across every surface and help hosts compete and thrive in our marketplace. We work at the intersection of search ranking, recommendations, personalization, and generative AI, and we're building toward a future where Airbnb feels less like a search engine and more like a knowledgeable travel companion that understands your needs across your entire trip journey. The Difference You Will Make: As Product Manager for Relevance & Personalization, you will help set the strategy and drive execution for some of Airbnb's highest-leverage AI systems. You'll own the roadmap and shape how personalization works across the guest journey, and help define how we close the feedback loop for hosts. You'll partner with engineers, researchers, designers, and cross-functional teams to ship systems that directly drive bookings, guest satisfaction, and host success — at global scale. A Typical Day: Define and drive the roadmap for Airbnb's relevance and personalization platform — from natural language query understanding to multi-turn, context-aware discovery experiences Make prioritization calls that balance multiple competing objectives: guest experience, host success, revenue, fairness, and marketplace health Partner with ML engineers and applied researchers to shape model strategy, evaluation frameworks, and experimentation design Align cross-functional partners — Guest, Host,
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. ROLE We're looking for a Solutions Product Marketing Manager to define and own Baseten's up-market, enterprise, and industries go-to-market strategy. You'll build the playbook that helps Baseten win large, complex deals and the vertical go-to-market strategies that let us compete in Financial Services, and Health & Life Sciences, and the use cases that scale across all of them. This is a highly cross-functional, largely greenfield role. You'll partner closely with the Industries sales team, product marketing, demand gen, and executive leadership to build on our enterprise go to market motion. RESPONSIBILITIES Enterprise Deal Playbook & Buyer Enablement Build the enterprise deal playbook and full bill of materials (BOM) — pricing guides, one-pagers, TCO/ROI guides, landing pages, and event/field collateral — mapped to every stage of the buyer's journey Build the system for keeping every sales asset current — clear ownership, last-updated dates, and a refresh cadence Develop industry-specific event concepts featuring anchor customers and AI-native accounts to build credibility with target executive audiences Vertical GTM & Industry Go-to-Market Kits Build vertical sales kits and ICP/positioning frameworks for priority industries Own industry-specific landing pages, messaging, and TCO guides, extrapolating from early AI-native accounts to larger enterprise targets Design and help launch an ABM experiments in p
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation
About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui
About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t
£215K – £260K/yr
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the Role Cohere is seeking a Global Public Policy Manager to lead policy engagement on compute infrastructure, export controls, AI competitiveness, and sovereign AI strategies. Governments increasingly view AI as critical national infrastructure and are investing heavily in compute capacity, energy resources, and domestic AI ecosystems. This role will help position Cohere as a trusted partner in emerging discussions around AI infrastructure, national competitiveness, and sovereign AI deployment. Key Responsibilities Monitor and analyze developments related to AI infrastructure, data centers, energy policy, semiconductor policy, export controls, and national AI strategies. Develop policy positions on sovereign AI, compute access, digital sovereignty, and AI competitiveness. Support engagement with governments developing AI infrastructure investment programs and national AI initiatives. Collaborate with commercial, product, and corporate development teams on strategic opportunities involving public-private partnerships. Represent Cohere in policy discussions related to AI infrastructure, energy requirements, and technology c
About the Team Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training. Our mission is to make some of the world's largest GPU clusters work reliably for frontier training. We bring new platforms and clusters online, safely maintain installed fleets, and partner with hardware, infrastructure, and research teams to resolve the system-level issues that keep jobs from running. That means building and maintaining the software closest to the machine: Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation. We make these components reproducible, compatible, and safe to operate across heterogeneous fleets. About the Role We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate. You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models
Other cities to consider
More places hiring for this role
Get new compute strategy and transactions jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime