About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role On the Accelerators team, you will help OpenAI evaluate and bring up new compute platforms that can support large-scale AI training and inference. Your work will range from prototyping system software on new accelerators to enabling performance optimizations across our AI workloads. You’ll work across the stack, collaborating with both hardware and software aspects - working on kernels, sharding strategies, scaling across distributed systems, and performance modeling. You'll help adapt OpenAI's software stack to non-traditional hardware and drive efficiency improvements in core AI workloads. This is not a compiler-focused role, rather bridging ML algorithms with system performance - especially at scale. In this role, you will: Prototype and enable OpenAI's AI software stack on new, exploratory accelerator platforms. Optimize large-scale model performance (LLMs, recommender systems, distributed AI workloads) for diverse hardware environments. Develop kernels, sharding mechanisms, and system scaling strategies tailored to emerging accelerators. Collaborate on optimizations at the model code level (e.g. PyTorch) and below to enhance performance on non-traditional hardware. Perform system-level performance modeling, debug bottlenecks, and drive end-to-end optimization. Work with hardware teams and vendors to evaluate alternatives to existing platforms and adapt the software stack to their architectures. Contribute to runtime improvements, compute/communication over
Jobs in United States
Compute Strategy And Transactions in San Francisco
147 active opportunities · Updated September 2026
Showing
15 jobs
Explore current compute strategy and transactions jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About OpenAI OpenAI is dedicated to ensuring that artificial general intelligence (AGI) benefits all of humanity. Our mission requires building not only world-class AI models, but also the infrastructure that enables those models to be deployed reliably, efficiently, and at global scale. As demand for AI continues to grow, we are expanding the ways OpenAI can bring high-performance inference capacity online across a diverse hardware ecosystem. About the Team The GPT Infrastructure team builds software that turns advanced inference and optimization research into production products. One focus is enabling strategic infrastructure partners and accelerator vendors to qualify and onboard new compute without a bespoke porting and optimization effort for every hardware platform. We build the control planes, APIs, secure partner-side execution environments, evaluation systems, artifact pipelines, and operational tooling that make these workflows repeatable and trustworthy. The work sits at the intersection of distributed systems, AI inference, compilers and runtimes, performance engineering, security, and external partnerships. About the Role We are seeking an experienced systems generalist who can work comfortably across the stack to help build an automated inference optimization platform. Given a workload, target hardware profile, compiler and runtime context, and a trusted verifier, the system runs durable optimization campaigns that generate, compile, execute, grade, and improve candidate kernels, runtime configurations, and serving-stack changes. You will design both the OpenAI-hosted control plane and the partner-side software that evaluates candidates on real accelerator hardware. The product must keep long-running workflows reliable, make performance results reproducible, and maintain clear trust boundaries around sensitive model and hardware information. This is a deeply cross-stack role, combining strong software engineering fundamentals with systems thinking and
About the Team The Platform Systems team at OpenAI operates at the intersection of cutting-edge AI and large-scale distributed systems. We build the engineering and research infrastructure required to train OpenAI’s flagship models on some of the world’s largest, custom-built supercomputers. Our team develops core model training software and works deep in the stack - spanning collective communication, compute efficiency, parallelism strategies, fault tolerance, failure detection, and observability. The systems we build are foundational to OpenAI’s research velocity, enabling reliable, efficient training at frontier scale. We collaborate closely with researchers across the organization, continuously incorporating learnings from across OpenAI into the evolution of our training platform. About the Role As a Software Engineer, Platform Systems, you will design and build distributed systems that provide visibility into large-scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical to operating OpenAI’s training stack and is actively evolving to support new use cases and increasingly complex workloads. This role sits at the core of our training infrastructure, blending systems engineering, performance analysis, and large-scale debugging. In This Role, You Will Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high-throughput distributed systems Collaborate with systems, infrastructure, and research teams to evolve platform
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
About the Team The GPT Infrastructure team builds systems that turn advances in model inference and optimization into reliable production capabilities. We enable OpenAI workloads to be qualified and optimized across new accelerator platforms without requiring a one-off port and tuning effort for every hardware target. Our work spans distributed systems, model execution, compilers and runtimes, performance engineering, secure partner integrations, evaluation systems, and developer tooling. We build the infrastructure that makes optimization workflows automated, reproducible, and trustworthy. About the Role We are seeking a software engineer to help build the platform that qualifies and optimizes inference workloads across heterogeneous compute environments. You will develop both OpenAI-hosted services and secure partner-side software for running long-lived optimization workflows. These workflows generate candidate kernels, runtime configurations, and serving-stack changes; compile and execute them on target hardware; verify their correctness; measure their performance; and use the results to guide further optimization. You will work across model architecture, distributed execution, compilers, runtimes, networking, and accelerator systems. A central part of the role is turning research prototypes and one-off hardware bring-up efforts into reliable, reusable infrastructure with clear contracts, reproducible results, strong observability, and well-defined security boundaries. Key Responsibilities Design, build, and operate APIs and control-plane services for long-running workload qualification and optimization campaigns, including scheduling, retries, checkpointing, resource budgets, and observability. Build secure partner-side execution and evaluation software that can compile, run, verify, profile, and benchmark candidate artifacts on accelerator hardware. Integrate model workloads, hardware profiles, compiler toolchains, runtimes, serving engines, and distributed-exe
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions optimized for advanced AI workloads. We collaborate across research, software, and external hardware partners to design and deploy next-generation AI systems at scale. Our team works closely with silicon vendors and system partners to evaluate emerging technologies, validate performance characteristics, and ensure that hardware capabilities translate effectively to real-world AI workloads. About the Role We are seeking a 3P Hardware Architecture Expert with deep expertise in GPU and accelerator architectures to engage directly with silicon vendors and guide hardware decisions for AI infrastructure. In this role, you will evaluate architectural tradeoffs across compute, memory, and interconnect systems, translating vendor specifications into real-world workload impact. You will play a critical role in early silicon evaluation, benchmarking, and performance validation, helping ensure that next-generation hardware meets the needs of our workloads. This role is highly hands-on and requires both deep technical understanding and the ability to engage at a high level with partners such as NVIDIA and AMD on architectural direction and design tradeoffs. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Engage deeply with silicon vendors (e.g NVIDIA & AMD) on GPU and accelerator architecture tradeoffs. Analyze and interpret performance, power, and efficiency characteristics of next-generation hardware. Translate vendor specifications into expected real-world performance for AI workloads. Evaluate architectural aspects including: compute throughput and utilization memory systems (HBM, cache hierarchies, bandwidth constraints) data types and precision tradeoffs (FP16, BF16, FP8, etc.) interconnect and scaling behavior. Run benchmarks and profiling to validate hardware performance a
About the team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the role: We are seeking an experienced Optical Network Engineer to lead Laser related work within our optical interconnect efforts for large-scale compute systems. The role also requires broad, hands-on optical validation experience across IM/DD-based interconnects, working from lab characterization through production readiness and scaled deployment. In this role you will: Drive laser-focused requirements and technical direction within the broader optical interconnect roadmap. Lead evaluation and validation of optical components and subsystems, including laser-based elements, in lab and production-representative environments. Support end-to-end optical testing for IM/DD interconnects (e.g., module/system bring-up, characterization, debug, and readiness for scale). Work with external partners to align on development milestones, performance targets, and quality expectations. Own technical issue triage and resolution across performance, reliability, and manufacturability topics. Collaborate across internal teams to support integration, rollout, and operational success at scale. You might thrive in this role if you have: Strong experience in laser-focused optical engineering (development, validation, manufacturing readiness, or field support). Broad hands-on background with IM/DD optical technologies and optical test/debug workflows. Experience working with external suppliers/manufacturing partners and production-oriented execution. Demonstrated ability to debug complex t
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and inform next-generation infrastructure design. About the Role We are seeking Performance Modeling Engineers to develop and apply modeling tools that evaluate AI system performance and inform architectural decisions. In this role, you will work closely with the Performance Modeling Lead and partner teams to analyze system behavior, run simulations or analytical models, and help quantify tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks and applying them to real-world questions that impact system design and vendor decisions. This role is well-suited for engineers with strong software or modeling backgrounds who are interested in developing deeper expertise in system architecture and AI infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Develop and maintain performance modeling tools and frameworks. Build models to evaluate system behavior across: compute, memory, and interconnect subsystems distributed system scaling and bottlenecks. Run simulations and analytical models to support architectural tradeoff analysis. Collaborate with performance modeling lead and system architects to answer forward-looking design questions. Analyze and interpret modeling outputs, translating results into actionable insights. Validate models against real system measurements and workload behavior. Contribute to improving modeling fidelity, usability, and scalability. Qualifications Strong software engineeri
About the Team OpenAI’s Infrastructure organization builds and evaluates the systems that power advanced AI workloads. We work closely with hardware, modeling, and architecture teams to ensure that new platforms deliver real-world performance aligned with workload needs. Our team focuses on understanding workload behavior across evolving hardware platforms—bridging the gap between theoretical capability and observed system performance. About the Role We are seeking a Workload Porting & Performance Engineer to evaluate new hardware platforms by porting benchmarks and real-world workloads, analyzing performance, and identifying system bottlenecks. In this role, you will bring up workloads on new systems, characterize performance behavior, and adapt workloads to better utilize hardware capabilities. You will play a critical role in validating new platforms and ensuring that performance aligns with expectations across compute, memory, and networking subsystems. This role requires strong hands-on experience with performance analysis, workload optimization, and system-level debugging across hardware and software boundaries. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Port and enable benchmarks and real-world workloads on new hardware platforms. Evaluate system performance across compute, memory, storage, and networking subsystems. Identify and analyze performance bottlenecks and inefficiencies. Adapt and optimize workloads to better utilize hardware capabilities. Develop and run performance experiments and profiling workflows. Compare expected vs. observed performance and provide feedback to: hardware architecture teams performance modeling teams system and software engineers. Debug issues across the stack, including software, runtime, and hardware interactions. Provide actionable insights to guide platform readiness and deployment decisions. Qualifications E
About the Team OpenAI's data and storage infrastructure spans data platforms, online databases, and file/object storage. These systems underpin data ingestion and processing, durable persistence, indexing and retrieval, and product file experiences. As frontier models and agents evolve how they use memory, history and snapshots, the underlying architecture increasingly shapes the capabilities products can deliver—and their latency, reliability, cost and efficiency. About the Role We are looking for a technically deep TPM to independently define and lead multiple programs across data platforms, online databases and storage infrastructure. You will connect model, product and data-consumer requirements to architecture, and work with the relevant engineering teams to take new capabilities through production adoption and repeatable expansion. The design scope is exabyte-scale storage and infrastructure spanning multiple millions of CPU cores. The challenge is not simply forecasting more resources: it is making complete, workload-ready capacity repeatable, with a clear path from product requirements through architecture, deployment and validation. A data pipeline, database query, file operation or execution snapshot can affect whether a product or agent succeeds; you will connect those outcomes to the systems underneath. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate model, product and data-platform needs into precise access patterns, consistency, durability, freshness, availability and scalability requirements. Connect memory, history, retrieval and resumable work to capability and end-to-end latency. Partner with engineering to transform data and storage architecture into repeatable scale units: standardized provisioning, placement, routing, data movement and readiness checks that bring storage, compute and networking online together.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments. We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment. This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment.. RESPONSIBILITIES Core Responsibilities: Drive suppliers to keep the maximum amount of the GPU fleet online and healthy. Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs. Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier. SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short. Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability. Scope and
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a Delivery Director, Capacity programs for our on-premises data center builds and neo cloud (GPU cloud) delivery programs. This is a high-visibility, execution-critical role sitting at the intersection of infrastructure engineering, capacity planning, vendor/partner management, and customer delivery. You will own the end-to-end delivery lifecycle for large-scale compute infrastructure — from initial site/capacity commitments through power, networking, and hardware bring-up, to production-ready GPU/compute capacity landing in the hands of internal teams or customers. You'll be the person who turns ambitious infrastructure roadmaps into predictable, on-time, delivery. RESPONSIBILITIES Own delivery of on-prem infrastructure builds — colocation expansions, power/cooling readiness, rack-and-stack, network fabric bring-up, and hardware acceptance testing — coordinating across colo providers and partners, network engineering, hardware ops, and vendor teams. Drive neo cloud delivery programs — manage capacity delivery from GPU cloud and neo cloud partners (e.g., colocation/bare-metal/GPU cloud providers), including contract milestones, capacity ramps, SLAs, and go-live readiness. Build and maintain master delivery schedules across concurrent, multi-site, multi-vendor programs, integrating power/shell timelines, hardware lead times, logistics, and software/platform readiness into a single critical path.
About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal and VMs dependable & scalable foundations for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Software Engineer, Host Assurance to build and operate the services, APIs, and host software that establish and maintain trust in our compute infrastructure. You will own production software from design and implementation through testing, rollout, observability, and operation. Your work will support capabilities such as machine identity, certificate issuance and enrollment, secure bootstrap, and host attestation across bare-metal and VM environments. Success in this role requires strong technical judgment, the ability to reason across software and host-system boundaries and learn unfamiliar parts of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—inc
About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About Mixpanel Mixpanel turns data clarity into innovation. Trusted by more than 29,000 companies, including Workday, Pinterest, LG, and Rakuten Viber, Mixpanel’s AI-first digital analytics help teams accelerate adoption, improve retention, and ship with confidence. Powering this is an industry-leading platform that combines product and web analytics, session replay, experimentation, feature flags, and metric trees. Mixpanel delivers insights that customers trust. Visit mixpanel.com to learn more. About The Team Mixpanel Engineering is a small, fast-moving team focused on delivering real value to customers. We build powerful AI-powered product analytics while obsessing over clarity, simplicity, and delight. Engineers here own problems end to end. You can move across the stack to ship impact without being blocked by silos or heavy process. Product innovation drives our business, and product engineering teams own that responsibility. Our OLAP engine queries over 500 trillion events; a typical blob storage system we interact with processes 300 PiB/month at 1.2 Tbps sustained, and we run many of them across the world. The Data Runtime team owns the data execution layer that powers every Mixpanel product. We ensure that every customer query runs fast, cheap, and reliably, at any scale. This is an exciting time to join. Mixpanel's agentic and AI-first products are driving rapid growth in query volume, and Data Runtime is making the big bets that power it. We’re investing in elastic query compute and a distributed file cache that will let us scale query workloads dramatically without scaling cost with them. We
Other cities to consider
More places hiring for this role
Get new compute strategy and transactions jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime