ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. ABOUT BASE LABS Base Labs is a research lab pushing the frontier of open-source LLMs. We think intelligence should be democratized, not controlled by a handful of closed labs and we think very few teams are actually positioned to do something about that. Backed by Baseten's training and inference infrastructure, we have the compute, resources, and talent to take on hard problems at the frontier and open-source what we learn along the way. Our mission is to help build a world where intelligence isn't concentrated, but spread across an ecosystem of models that anyone can build on. That mission shapes what we choose to work on, how we work on it, and who we want in the room. We are accepting applications on a rolling basis for our first cohort of Base Labs Fellows, which is expected to start in late September. Apply using this link. BASE LABS FELLOWSHIP OVERVIEW The Base Labs Fellowship is designed to give researchers exposure to what frontier research looks like in industry. We provide funding, mentorship, and full support to our fellows, with the goal of producing rigorous, published research that shapes both the open-source ecosystem and Baseten's technical roadmap. We run multiple cohorts of Fellows each year and review applications on a rolling basis. This application is for cohorts starting in Sept 2026 and beyond. WHAT TO EXPECT 3 months of full-time research from our San Francisco office A dedicated 1:1 mentorship
Jobs in United States
Inference Technical Lead in San Francisco
268 active opportunities · Updated October 2026
Showing
15 jobs
Explore current inference technical lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. PRODUCT AT BASETEN Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the people who defines it. You'll work directly with our founders and with some of the best systems and infrastructure engineers in the world, and you'll set the standard for what product looks like here. You earn trust by being technical, finding the truth in front of customers, building great cross-functional relationships, and shipping great product experiences. THE ROLE The largest, most demanding Enterprises are starting to run on Baseten and they come with a range of security, compliance, and procurement requirements. Today that readiness is assembled deal-by-deal. You'll own the enterprise-readiness surface end to end and turn it into product: the deployment options customers can choose and buy, compliance posture they can trust, access and security controls their IT teams require, and the billing and spend controls their finance teams expect. What does a complete Baseten Enterprise Product offering look like? RESPONSIBILITIES Drive the Enterprise Readiness customer experience end to end: Partner with GTM and Enterprise Engineering to make "enterprise-ready" a platform-wide capability, not a deal-by-deal scramble. Outcome: readiness becomes a supported, priced product instead of bespoke work asse
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. We are looking for an engineer with strong experience in machine learning and solid foundations in maths and computer science to join our growing Post-Training team at Baseten. Custom models are instrumental to the success of Baseten customers. By inference volume, the overwhelming majority of traffic at Baseten is to and from models that have been post-trained in some way, whether that be through reinforcement learning, supervised finetuning, a recent technique from the literature, or an in-house research technique from Baseten. The Post-Training team is responsible for the success of our customers’ post-trained models, and we employ a wide array of techniques to produce models that are more efficient and higher quality than even the biggest closed source models for the customer’s specific needs. Your role as a research engineer is to build the in-house tooling to support all of this. We care about training a wide spectrum of different model architectures with a variety of techniques efficiently and at scale. At times this involves zooming deep into a particular technical topic, but more often if involves working across the stack as a whole - systems-level concepts like Kubernetes, cgroups, storage systems, and networking topologies, as well as PyTorch distributed tensor computation, and GPU kernels. RECENT RESEARCH Dense, on-policy or both? Repeated kv cache for long-running agents Distillation without the dark – rep
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto
About the Team Our mission at OpenAI is to discover and enact the path to safe, beneficial AGI. To do this, we believe that many technical breakthroughs are needed in generative modeling, reinforcement learning, large-scale optimization, active learning, and other areas. The team builds the performance-critical systems that allow OpenAI's models to run efficiently across a diverse set of AI accelerators. We work across the inference stack, from low-level kernels and compilers through model execution, to unlock the full capabilities of the underlying hardware. About the Role As a Software Engineer, Trainium, you will help bring OpenAI's inference workloads to AWS Trainium and build the software stack required to run cutting-edge frontier models efficiently on the platform. This is a deeply technical, cross-stack role spanning kernels, compilers, and model execution. You will work on the systems needed to support OpenAI's inference stack on Trainium, including developing and optimizing high-performance kernels, improving compiler support, and enabling efficient execution of the model forward pass. You'll work closely with engineers across inference, compilers, kernels, and ML systems to identify performance bottlenecks and build the software needed to take full advantage of Trainium. The work may range from low-level hardware-specific optimization to compiler and runtime improvements to integrating new model architectures into the inference stack. If you enjoy working at the intersection of ML systems, compilers, kernels, and accelerator hardware, this role is for you. We're looking for engineers who are self-directed, comfortable operating across abstraction layers, and excited to solve challenging performance problems for frontier-scale AI systems. In This Role, You Will Build and optimize OpenAI's inference stack for AWS Trainium. Develop high-performance kernels for critical model operations and workloads. Extend and improve compiler support to efficiently target
About the Team Our economics team is continuously working to improve our understanding of an AI-driven economy. About the Role We are seeking a highly technical Economist to join the OpenAI Economic Research team studying the real-world economic impacts of AI. This role is designed for economists with up to 5 years of professional experience post-Ph.D. who are interested in using novel, large-scale datasets to study how AI is reshaping economic systems. We are looking for candidates with deep expertise in at least one core domain relevant to AI’s economic impact, and an interest in contributing to a broader research agenda spanning labor markets, firm behavior, market dynamics, and macroeconomic change. This is an individual contributor role where the candidate will organize and execute on their own data-oriented projects. You will work at the intersection of economic research, data science, and public policy to produce rigorous empirical work that informs decision-makers across the public, industry, and government. Research Areas of Interest We are particularly interested in candidates with demonstrated expertise in one or more of the following areas: Economic Measurement of AI Impact (e.g., adoption trajectories, labor market transitions, productivity growth, and forecasting/scenario modeling for AI-driven economic change) Macroeconomic Implications of AI (e.g., productivity, technology diffusion, economic growth) AI and the Labor Market (e.g., employment, wages, job search, task-level impacts, skill acquisition) Applicants are not expected to have experience across all domains. We aim to build a team with complementary strengths across these areas. In this role, you will: Design and execute empirical research using large-scale observational or experimental data. Apply causal inference and/or structural modeling techniques to study AI-driven economic change. Collaborate with cross-functional teams to translate research questions into testable frameworks and applic
About the Team We’re hiring software engineers to make OpenAI’s Model Performance teams more productive. These teams work on the systems, tooling, and infrastructure that help improve model performance across OpenAI’s training and inference workloads at frontier scale. About the Role We’re looking for an autonomous, high-ownership developer productivity engineer who cares deeply about helping other engineers move faster, safer, and with more confidence. This role will sit within OpenAI’s Model Performance organization, contributing to developer infrastructure, CI systems, testing workflows, tooling, and broader performance infrastructure efforts. There is also a strong opportunity to contribute to the Triton project and help improve the systems that support performance-critical engineering work across OpenAI. In this role you will: Improve development workflows for engineers working on model performance infrastructure Design and improve CI/CD, release, validation, and testing pipelines Build and maintain tools that improve reliability, iteration speed, and engineering confidence Partner closely with engineers to identify friction in testing, debugging, deployment, and development workflows Contribute to infrastructure efforts that support performance-critical training and inference systems Help improve developer experience across Python-heavy codebases and performance-oriented infrastructure Work in a high-context, ambiguous environment where ownership and good judgment matter You might thrive in this role if: You are motivated by enabling the people around you and helping engineers do their best work You have strong experience with CI/CD, developer infrastructure, testing systems, tooling, or build/release workflows You are highly collaborative, empathetic, and comfortable partnering deeply with technical teams You are strong in Python and enjoy building reliable, scalable developer tools and infrastructure You have experience improving large-scale engineering work
About the Team API Multimodal builds the developer-facing products and infrastructure that bring OpenAI’s image, audio, and real-time model capabilities into the world. We are responsible for high-scale APIs for image generation, speech transcription, speech generation, and low-latency voice interactions. We partner closely with Research and Inference to bring frontier model capabilities to developers and use customer feedback to improve our models. About the Role As a software engineer on API Multimodal, you will build and operate the products and distributed systems behind OpenAI’s image, audio, and real-time APIs. You will work across model integration, API design, and production infrastructure to turn new research capabilities into reliable developer experiences. This hands-on role combines backend and systems depth with product judgment: you will own projects end to end, partner with Research, Inference, and Safety, and help make multimodal AI useful at scale. Model training experience is not required. In this role, you will: Design, build, and ship developer-facing APIs and backend services that serve frontier models. Architect low-latency streaming, request, session, and model integration systems that make complex multimodal interactions reliable and intuitive at scale. Work directly with Research to bring new model capabilities into production, shape the systems around them, and incorporate feedback from real-world developers and customers. Own the availability, latency, scalability, and cost efficiency of the services you build. Own projects from technical design and implementation through launch and ongoing iteration, while raising the team’s engineering standards. Your background might look something like: 7+ years of professional experience, excluding internships, in backend, infrastructure, platform, or product engineering roles. A track record of designing, building, and operating production backend services, developer-facing APIs, or distributed syste
About the Team The Applied AI team safely brings OpenAI's technology to the world. We released ChatGPT, Plugins, DALL·E, and the APIs for GPT-4, GPT-3, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate horizon. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. We serve end-users directly through ChatGPT, and serve developers through our APIs, which power product features that were never before possible. About the Role The Engineering Acceleration team designs, builds and maintains the foundational systems that engineers use to build ChatGPT and the API. This is a fast-growing team and you will get a chance to own and define the strategy, vision, and plan for how to increase developer productivity. In this role, you will: Drive the design, development, and implementation of tools, systems, and processes that accelerate engineering velocity, reduce manual effort, and increase the quality of output. Use our latest AI tools to re-think how we can be the most productive team in the industry. Work closely with various teams within OpenAI to understand their workflows, challenges, and needs, and ensure the tools and systems built by the Engineering Acceleration team address these requirements. Bring new features and research capabilities to the world by partnering with product engineers to lay the necessary technical foundations. Guide and advise product engineering teams on best practices for ensuring observable, scalable systems. Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years of experience in engineering, including 3+ years of experience in infrastructure building tooling for developers. Have experi
About the Team The Core Models team helps shape how OpenAI’s frontier models are built, measured, and launched. We work across Research, Engineering, Model Design, Data Science, and Product to turn advances in model capabilities into reliable, useful experiences for people. Our scope includes model planning and launches as well as building data flywheels, evaluations and measurement systems to ensure our models have strong capabilities and behavior. About the Role As a Product Manager for the Core Models team, you'll be at the forefront of defining and guiding the future of how our AI models work in real-world applications. You will connect user needs to model and systems decisions: how prompts are understood; how information is aggregated and made useful for training and evaluation data; and how capabilities move from research prototypes into the mainline model and launch stack. You will operate comfortably across research, infrastructure, and consumer product surfaces, creating clarity where ownership and technical boundaries are still emerging. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate user and product goals into clear model requirements, system architecture choices, and research priorities across query understanding, indexing, retrieval, ranking, tool boundaries, data, training, inference, and evaluation. Build closed learning loops that turn product usage, explicit feedback, and other user signals into datasets, evaluations, experiments, training priorities, and launch decisions. Define success across offline evaluations and online product metrics, balancing model quality, usefulness, latency, safety, reliability, and cost. Partner closely with post-training research, applied product engineering, Model Design, and Data Science to integrate capabilities into the mainline model stack. Create reusable platforms and operatin
About the Team At OpenAI, we’re building safe and beneficial artificial general intelligence. We deploy our models through ChatGPT, our APIs, and other cutting-edge products. Behind the scenes, making these systems fast, reliable, and cost-efficient requires world-class infrastructure. The Caching Infrastructure team is responsible for building a caching layer that powers many critical use cases at OpenAI. We aim to provide a high-availability, multi-tenant cache platform that scales automatically with workload, minimizes tail latency, and supports a diverse range of use cases. We’re looking for an experienced engineer to help design and scale this critical infrastructure. The ideal candidate has deep experience in distributed caching systems (e.g., Redis, Memcached), networking fundamentals, and Kubernetes-based service orchestration. In This Role, You Will: Design, build, and operate OpenAI’s multi-tenant caching platform used across inference, identity, quota, and product experiences. Define the long-term vision and roadmap for caching as a core infra capability, balancing performance, durability, and cost. Collaborate with other infra teams (e.g., networking, observability, databases) and product teams to ensure our caching platform meets their needs. You Might Thrive In This Role If You: Have 5+ years of experience building and scaling distributed systems, with a strong focus on caching, load balancing, or storage systems. Have deep expertise with Redis, Memcached, or similar solutions, including clustering, durability configurations, client-side connection patterns, and performance tuning. Have production experience with Kubernetes, service meshes (e.g., Envoy), and autoscaling systems. Think rigorously about latency, reliability, throughput, and cost in designing platform capabilities. Thrive in a fast-paced environment and enjoy balancing pragmatic engineering with long-term technical excellence. About OpenAI OpenAI is an AI research and deployment company d
About the Team The Compute Strategy team works across research, engineering, product, finance, legal, and go-to-market teams to develop the partnerships, infrastructure capacity, and commercial models needed to advance AI infrastructure. About the Role As a member of the Compute Strategy team, you will develop commercial strategies for AI infrastructure partnerships and offerings. You’ll translate technical infrastructure opportunities into partnerships, transactions, and revenue. We’re looking for a commercially minded strategist who combines knowledge of semiconductors and AI infrastructure with strong financial judgment and the ability to execute complex partnerships. This role is based in San Francisco, CA. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you will: Develop strategies for compute partnerships, vendor access, and infrastructure capacity. Structure and execute transactions with chipmakers, compute providers, and other infrastructure partners. Develop pricing frameworks and business cases for infrastructure-related partnerships. Evaluate partner technologies, strategic fit, commercial terms, and execution risks. Coordinate work across research, engineering, product, finance, legal, and go-to-market teams. Turn partnership learnings into repeatable operating models that can scale. You might thrive in this role if you: Have experience in strategy, corporate development, partnerships, or infrastructure transactions. Understand semiconductors and AI infrastructure. Can evaluate complex technical and commercial opportunities. Bring strong financial, analytical, and strategic judgment. Can influence and align technical and business stakeholders. Have negotiated or executed complex partnerships. Are comfortable operating in a fast-paced environment. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence ben
About the Team The Proactivity Research team, within OpenAI’s broader Personal AGI team, is focused on making our models in ChatGPT and future potential products proactive in ways that are truly useful. We're laying the technical foundations for AI that can anticipate what users need in real time, adapt as their goals and preferences shift, and build a deeper, evolving understanding of the person it's helping. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models’ personalization and agentic capabilities. Our team works on reinforcement learning, dataset creation, evaluations, and other post-training methods. We partner closely with research and product teams across the company to realize the vision of a highly personalized, collaborative, and proactive assistant. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve the proactivity and ability of our models to further user goals. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. Collaborate closely with the other research and product teams to influence the shape of technical solutions in the product You might thrive in this role if you: Have a deep understanding of machine learning and machine learning applications. Have a working knowledge of LLM post-training and evaluation approaches Are passionate about, or have experience thinking about, personalization and enabling users to achieve their goals Are comfortable diving into a large ML codebase to debug. Thrive in a dynamic and technically complex environment. About OpenAI
About the team The Computer Use and New Interfaces team is focused on discovering and building the next generation of AI-native interfaces. We believe that the value of AI is increasingly constrained not by model capabilities, but by the ways people interact with those capabilities. Our mission is to create new interaction paradigms that unlock the full potential of AI and integrate it more deeply into people's lives and work. Our team does both near-term product development and longer-term product incubation that can influence experiences across ChatGPT, Codex, future OpenAI products, and emerging device platforms. We work in a highly collaborative, design-driven environment where engineering, product, and design operate as one team. We value rapid experimentation, prototyping, and iteration, creating the shortest possible path between an idea, a working system, and a product decision. About the role We're looking for exceptional engineers who are excited to invent entirely new ways for people to interact with AI. You'll work at the intersection of engineering, product, and design to explore, prototype, and build novel interface concepts that push beyond traditional software paradigms. This role requires comfort with ambiguity, strong product instincts, and a willingness to move fluidly between experimentation and production systems. You'll help shape both the capabilities and the user experiences that define how people engage with AI. In this role you will: Design, prototype, and build novel AI-native interfaces and interaction models. Develop foundational technologies and frameworks that enable generative UI and computer use experiences. Collaborate closely with designers, product thinkers, and engineers to rapidly explore and validate new concepts. Build end-to-end prototypes and production systems that can influence future OpenAI products. Contribute to platform technologies that can be adopted across multiple product surfaces. Help establish technical directio
We’re looking for a Software Engineer to architect and build backend systems that enforce data privacy and automate compliance at scale. You’ll work closely with product, infrastructure, security, and legal teams to embed privacy-by-design into our data and access layers. This is a hands-on, high-impact role for an experienced engineer who is passionate about protecting user data while enabling innovation. What You’ll Do Design, build, and operate backend services that enforce policy-driven data access, lifecycle controls, and privacy protections. Develop distributed authorization and identity-aware enforcement mechanisms integrated directly into data services and control planes. Implement auditability, policy hooks, and enforcement observability to ensure compliance is continuously verifiable. Partner with Security, Legal, and Compliance to convert privacy requirements into scalable technical designs and developer-friendly APIs. Harden data platforms and backend services through schema-level controls and data handling constraints by default. Collaborate with infrastructure teams to ensure consistent enforcement across systems while minimizing duplicated implementations. Contribute patterns, libraries, and education that elevate trustworthy data access patterns across the organization. You Might Thrive in This Role If You Have 5+ years of industry experience building and operating backend or infrastructure systems in production. Strong software engineering fundamentals , with fluency in at least one major programming language (e.g., Python, Go, Rust, C++, Java). Experience with distributed authorization, RBAC/ACL systems, encryption-based access, or policy engines. Familiarity with global privacy regulations and their architectural implications. Ability to influence and collaborate with teams across legal, compliance, product, and engineering. A bias toward practical, impactful solutions that balance privacy protections with product needs. Nice to Have Experience wi
Other cities to consider
More places hiring for this role
Get new inference technical lead jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime