About the Team OpenAI’s mission is to ensure that artificial general intelligence (AGI) benefits all of humanity. A key part of achieving that mission is training models that deeply understand and reflect human preferences — the Human Data team is at the heart of that effort. The Human Data engineering team creates the systems that enable scalable, high-quality human feedback. These systems are essential to how OpenAI trains and improves its most advanced models. Engineers on this team collaborate closely with world-class researchers to bring alignment techniques to life — from experimental ideas to production-ready feedback loops. About the Role We’re looking for software engineers to join the Human Data team and build the platforms, prototypes, tools, and infrastructure that power how our AI models are trained, aligned, and evaluated. You’ll partner with researchers and cross-functional teams to bring alignment ideas to life, influence future model training, and shape how models interact with the real world. We’re looking for people who are excited by technical ownership, enjoy working across the stack, and are eager to solve ambiguous problems in a high-impact, fast-paced environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build and maintain robust full-stack systems for feedback collection, data labeling, and evaluation pipelines, while maintaining high levels of security. Translate experimental alignment research into scalable production infrastructure, including inference and model training stacks. Design and iterate on user-facing tools and backend services to support high-quality data workflows Partner with researchers, engineers, and program leads to shape feedback loops and model interaction paradigms Drive infrastructure improvements that enable faster iteration and scaling across OpenAI’s frontier models, from internal r
Jobs in United States
Inference Engineering And Product Lead in San Francisco
268 active opportunities · Updated October 2026
Showing
15 jobs
Explore current inference engineering and product lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Operations Program Manager (OPM) to serve as the single-threaded operational leader for new hardware introductions (NPI) and production ramps across OpenAI’s AI infrastructure systems. This role combines hands-on execution with strategic ownership. You will be responsible for defining the operating model, aligning cross-functional stakeholders, setting the critical path, making informed tradeoffs, escalating decisively, and ensuring hardware programs deliver on schedule, quality, cost, and scalability. Success in this role requires comfort operating in ambiguity, influencing without authority, and driving alignment across internal teams and external partners—while keeping eyes firmly on long-term system scalability and repeatability. In this role, you will: Strategic & Leadership Ownership Act as the single-threaded owner for operational readiness across NPI and ramp, accountable for outcomes from early bring-up through sustained production Translate OpenAI’s infrastructure strategy and engineering objectives into clear operating plans, execution priorities, and decision frameworks Drive alignment across Engineering, Operations, Strategic Sourcing, Finance, Capacity Planning, and Executive stakeholders by framing tradeoffs, risks, and recommendations Proactively identify inflection points where decisions or investments are required to protect long-term scale, reliability, or cost targets Influence operational strategy with manufacturing par
About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve people's lives. About the Role We are seeking a lead thermal simulation engineer to help accelerate the design and development of next-generation robotic systems through modeling, simulation, and analysis. You will work closely with mechanical, electrical, controls, and robotics engineers to evaluate designs before hardware is built, identify risks early, and guide critical architecture decisions. This role spans structural and thermal analysis and design across robotic subsystems including actuators, mechanisms, structures, electronics, and integrated systems. You will develop simulation workflows that improve engineering velocity, increase confidence in design decisions, and help us build more capable, reliable, and manufacturable robotic platforms. This role is based in San Francisco, CA. This role will be expected to be in office 4 days per week and offer relocation assistance to new employees. In this role, you will: Perform thermal simulations to assess heat generation, cooling strategies, thermal interfaces, and system-level thermal performance Partner with mechanical, electrical, and controls engineers to influence design decisions early in development Build simulation models to evaluate robotic actuators, transmissions, mechanisms, structures, soft goods, and integrated assemblies Correlate simulation results with physical testing and develop methodologies to improve model accuracy Support architecture trade studies by evaluating design concepts before hardware is built Develop simulation workflows, standards, and best practices that scale across the robotics o
About the Team OpenAI's Strategic Sourcing team helps the company scale responsibly, efficiently, and at speed. We partner with leaders across Engineering, Research, IT, Systems, Finance, Legal, and other teams to shape commercial strategies, negotiate critical agreements, and build resilient supplier ecosystems. Our work connects technical strategy, commercial judgment, financial discipline, and execution in support of OpenAI's mission. About the Role OpenAI is seeking a Strategic Technology Negotiations Lead to personally lead some of the company's most complex and consequential technology negotiations. This is not a people manager role. It is a senior strategic IC operator role for an expert negotiator who wants to remain close to the work and personally drive high-impact outcomes. We are especially interested in leaders who have managed teams and are intentionally seeking an individual contributor role where their impact comes through judgment, influence, and direct ownership. Rather than owning a fixed category, you will be deployed against high-priority opportunities where deal complexity, commercial stakes, executive visibility, or time pressure require exceptional negotiation leadership. Your initial focus will include data platforms and infrastructure, including data lake and lakehouse technologies, observability, and enterprise SaaS, with flexibility to work across other strategic technology areas. You will lead negotiations from strategy through execution, aligning decision-makers and driving agreements to closure. Many of these negotiations exist within broader supplier and partner ecosystems. You will look beyond the immediate transaction to account for interconnected cost, equity, revenue, partnership, risk, and long-term strategic implications. Success requires strong economics, sound judgment under pressure, executive credibility, and the ability to bring stakeholders with you through difficult decisions. This role is based in San Francisco, CA. We u
About the Team Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training. Our mission is to make some of the world's largest GPU clusters work reliably for frontier training. We bring new platforms and clusters online, safely maintain installed fleets, and partner with hardware, infrastructure, and research teams to resolve the system-level issues that keep jobs from running. That means building and maintaining the software closest to the machine: Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation. We make these components reproducible, compatible, and safe to operate across heterogeneous fleets. About the Role We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate. You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models
About the Team The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to support our vision as we advance forward in our mission of building AGI that benefits all of humanity. About the Role As a Technical Lead on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. This role is based in San Francisco, CA. We follow a hybrid model with 3 days a week in the office and offer relocation assistance to new employees. In this role, you will: Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models. Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth. Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities. Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads. Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems. Run through the necessary walls to take nascent research capabilities and turn them into capabilities we can build on top of. You might thrive in this role if you: Have experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators. Understand the performance characteristics of transformer models, including attention, KV-cache behavior, and memory bandwidth requirements. Have designed or optimized high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines. Have experience building or leading teams work
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for a seasoned Frontend Engineer to craft performant and delightful user experiences across Baseten’s core platform. You’ll own critical parts of our web application stack and collaborate cross-functionally with product, design, and backend teams to launch impactful features that help users deploy and manage AI systems at scale. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Core Product team: Rolling Deployments Model APIs for frontier models Model training built for production inference RESPONSIBILITIES Design, implement, and maintain responsive, accessible, and user-friendly frontend interfaces using React and TypeScript Collaborate closely with product designers to turn complex ideas into elegant, intuitive UIs Optimize application performance and reliability, with a focus on rendering speed and responsiveness Drive major frontend initiatives, including partnering with backend teams to define APIs and test and refine end-to-end flows Establish best practices, and mentor other engineers on frontend technologies Build reusable component libraries and frontend infrastructure that accelerate product development Partner with backend and platform teams to define and refine APIs and end-to-end flows REQUIREMENTS 5+ years of experience building production-grade web applications Deep expertise in React, TypeScript, and modern web development tooling Track record of bu
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten’s Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We operate at the intersection of distributed systems, model performance, infrastructure, and developer experience. We enable customers to deploy and operate cutting-edge LLM models with industry-leading performance, scalability, reliability, and ease of use. As a Software Engineer on the Inference Stack team, you’ll work across the stack - from the developer experience customers use to deploy models, the libraries used for features like tool calling and reasoning, all the way down to the systems we use to orchestrate deployments in Kubernetes and route traffic efficiently. This is an ideal role for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users. EXAMPLE INITIATIVES Blog Posts https://www.baseten.co/blog/nvidia-dynamo-day-baseten-inference-stack/ https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/ https://www.baseten.co/blog/how-baseten-multi-cloud-capacity-management-mcm-powers-cloud-self-hosted-and-hybr/#comparing-deployment-options-cloud-vs-self-hosted-vs-hybrid RESPONSIBILITIES Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference Work across the stack, from customer-facing features to low-le
About the Team OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, and scalable in production, and we partner closely with Research to bring the next generation of models into the world. We're a small, fast-moving team of engineers focused on delivering a world-class developer experience while pushing the boundaries of what AI can do. We’re expanding into multimodal inference, building the infrastructure needed to serve models that handle image, audio, and other non-text modalities. These workloads are inherently more heterogeneous and experimental, involving diverse model sizes and interactions, more complex input/output formats, and tighter coordination with product and research. About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal models at scale. You’ll be part of a small team responsible for building reliable, high-performance infrastructure for serving real-time audio, image, and other MM workloads in production. This work is inherently cross-functional: you’ll collaborate directly with researchers training these models and with product teams defining new modalities of interaction. You'll build and optimize the systems that let users generate speech, understand images, and interact with models in ways far beyond text. In this role, you will: Design and implement inference infrastructure for large-scale multimodal models. Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs. Enable experimental research workflows to transition into reliable production services. Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities. Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers. You might thrive in t
About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. Our team works across silicon, embedded systems, operating systems, and cloud services to build reliable consumer devices and the novel platforms required to support them. We partner closely with research to bring advanced AI capabilities into the physical world. About the Role As an Operating Systems Engineer focused on on-device inference, you will design, develop, and ship the OS stack that makes advanced AI capabilities reliable, responsive, and energy efficient on consumer devices. Your work will span OS services and frameworks, inference runtime integration, model fitting, scheduling, and performance and power management. You’ll partner with research to adapt models to device constraints, make design decisions across the stack, and carry solutions from early exploration through integration and production. In this role, you will: Build the inference platform: Design and implement maintainable OS services, frameworks, and clear interfaces for inference execution, model loading and lifecycle, and resource management. Fit models to device constraints: Partner with researchers on quantization, runtime integration, and memory optimization to meet memory, compute, and energy budgets while evaluating model quality and product behavior. Coordinate system resources: Develop scheduling and resource policies that balance inference with other device act
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We are hiring at a rapid rate, and every one of our new hires deserves a great new hire experience from the moment they sign their offer letter through their entire onboarding journey. As our People Operations Coordinator, you'll own the tactical side of onboarding: collecting and tracking the completion of required documentation, greeting new hires, assisting in running sessions, and keeping the operations of onboarding on track as we grow. The role is heavily onboarding-focused today and will broaden across the employee lifecycle as our new people systems take over more of the routine work. Things change quickly here, and this role will too. This role is based in our San Francisco office. RESPONSIBILITIES Own the full onboarding experience for every new hire, including: new hire communication from offer signature through Day 1; ensuring completion of onboarding tasks like background checks, Form I-9s, setting up HRIS profiles, etc; coordinating travel logistics; partnering with the Global Mobility Manager on immigration matters that may affect start dates Support day one and San Francisco week one onboarding program, including coordinating with managers and buddies, room booking and coordination of start location Partner with IT and Workplace so new hires have the right access, equipment, and desk waiting on day one Manage the logistics behind our onboarding platform and manage the day-to-day vendor relationsh
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We’re looking for a Corporate Accounting Manager to support the general ledger, the close, and our accounting processes and controls as Baseten scales. This is a hands-on, individual-contributor role for someone who wants full ownership of the core accounting function at a company where the entity structure, transaction volume, and reporting requirements are growing quickly. You’ll join the corporate accounting team and take on the general ledger, vendor contract review, the close calendar, and our internal control environment as the business grows. We are focused on tightening close procedures, strengthening controls, and building reporting rigor to support the scale ahead. You’ll partner closely with FP&A, Data, and cross-functional partners to make sure the books close on time and accurately every cycle. Baseten is building the infrastructure layer for AI-native companies, and we're scaling quickly - in headcount, entity structure, transaction volume, and contract complexity and volume and vendor size. If you want real ownership over a growing function on a lean team, this role offers real scope. RESPONSIBILITIES General Ledger and Financial Reporting Own day-to-day execution of general ledger journal entries, account reconciliations, and monthly financial statement preparation as the entity structure grows Own recurring close areas, including accruals, prepaids, fixed assets, and equity compensation Lead
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We are bringing our people systems in-house on Workday, and we are hiring our first dedicated Workday Analyst to make it excellent. You will join while the implementation is underway, ramp alongside our deployment partner, and own the tenant from go-live onward. This is a hands-on configuration role: you will build business processes, manage security, load data, and test releases yourself, and you will teach others on the People team to do the same as we grow. If you want to shape a Workday environment from its first day in production instead of inheriting years of someone else's decisions, this is that rare opening. RESPONSIBILITIES Own day-to-day Workday configuration: business processes, security groups and roles, custom reports, calculated fields, and tenant settings Build and run EIB loads for data changes, mass updates, and audits Own the twice-yearly Workday release cycle: evaluate new features, regression-test, and roll out changes safely Shadow the implementation build, then take over tenant ownership at go-live Monitor integrations and triage issues with our IT team and vendors Support payroll configuration in partnership with our Accounting team and payroll services provider Field and resolve system requests from employees, managers, and the People team Mentor teammates so Workday administration becomes a team capability, not a single point of failure REQUIREMENTS 4+ years administering a live Workday
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE Our Sales and Solutions teams navigate hard technical conversations spanning inference performance, GPU economics, latency budgets, deployment shape. As Baseten’s platform matures, we need a dedicated owner to translate launch velocity into field readiness. As our first Product Enablement Lead, you'll sit between Product, Marketing, and Sales GTM and own how Baseten's products, features, campaigns, and market moments like the launch of GLM-5.2 or Kimi K3 or the sudden evolution of Tokenomics as a discipline get translated into field execution. You will own how these launches land with the field, how AEs and SAs stay credible on a highly dynamic technical ecosystem, and how what the field hears from customers makes it back to Product. This is a hands-on individual contributor role. You are the bridge between product, marketing, and sales. You'll build the system and run it, which includes cross-functional program leadership, direct training and enablement of in-seat reps, and content and curriculum development for managers, sellers, and new hires. Success here will depend on your ability to build repeatable systems and rhythms and to partner across the business and with your enablement colleagues to ensure alignment and speed of execution. RESPONSIBILITIES Own launch readiness: partner with Product and Marketing on positioning, write internal launch comms, and run readiness sessions so AEs and SAs can sell new pr
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. ABOUT THE TEAM Supply is responsible for knowing everything happening in the compute market: who's building, who's buying, and on what terms. This role owns a specific and fast-moving slice of that map — emerging clouds and international markets — and owns the full relationship lifecycle in that space, from first outreach through to closed terms. RESPONSIBILITIES Build and maintain a real-time picture of the emerging cloud and international compute landscape — who's active, what they're building, and what terms are available Own the full partnership lifecycle in this space — from identifying and sourcing new providers, to negotiating terms, to ongoing relationship management Develop and manage relationships across a broad set of emerging and international providers, from account reps up through leadership Identify, structure, and help close opportunities where Baseten can move quickly to secure favorable capacity terms Define compelling value propositions tailored to different types of providers, rather than a one-size-fits-all pitch Partner closely with others in the team already covering this space to build out a durable, well-organized intelligence and relationship function Collaborate with the broader Supply and Deals functions to bring opportunities to the table and support negotiation when it's time to close WHAT WE’RE LOOKING FOR Equal parts relationship-builder and operator — you can open a door and also drive it
Other cities to consider
More places hiring for this role
Get new inference engineering and product lead jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime