Jobs in United States

Inference Engineering And Product Lead in San Francisco

268 active opportunities · Updated October 2026

Explore current inference engineering and product lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for a Head of Environment to build the physical expression of Baseten as we grow from roughly 300 employees today to more than 2,000 over the next couple of years. You’ll own the environments, spaces, and experiences that define what it feels like to work at Baseten. This is equal parts operations, hospitality, design, and strategy. You’ll partner closely with our founders and leadership team to create an environment that scales with the company while remaining unmistakably Baseten. You’ll inherit a strong Workplace Experience team and continue to evolve the function into one of the defining strengths of the company. We’re looking for someone who has built and scaled world-class workplace functions before; someone who has seen what’s ahead and can help us get there faster. RESPONSIBILITIES Own Baseten’s global workplace and real estate strategy, developing the long-term roadmap for how our physical footprint evolves as we scale from hundreds to thousands of employees. Lead real estate planning, site selection, expansions, and significant capital investments across our headquarters, growing network of smaller offices, and future international locations. Build and operate exceptional workplaces. Own end-to-end workplace operations, office buildouts, relocations, and launches, ensuring every office runs smoothly with an uncompromising bar for quality, hospitality, and attention to detail. Define the

Machine LearningAIRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an OS / K8s Systems Engineer at Baseten, you’ll build the automation and systems that turn raw GPU hardware into production-ready compute. From provisioning to orchestration, you’ll own the software layer that makes our infrastructure reproducible, scalable, and reliable across data centers. This is a senior, hands-on role focused on building systems not operating them. You’ll work close to the metal designing OS images, building provisioning pipelines, and automating cluster bring-up from scratch. Your work will define how quickly we can turn new capacity into usable compute. EXAMPLE INITIATIVES Zero-to-cluster automation Build workflows that take new hardware from unprovisioned to fully operational cluster. Provisioning systems Design PXE-based or equivalent systems for imaging and lifecycle management. Reproducible infrastructure — Ensure clusters deploy consistently across data centers. RESPONSIBILITIES Own the end-to-end automation of cluster bring-up and lifecycle management. Build and maintain OS images, provisioning systems, and configuration pipelines. Deploy and operate cluster orchestration platforms (Kubernetes, Slurm, or similar). Design systems for reproducibility across sites and hardware generations. Automate upgrades, rollouts, and failure recovery. Optimize system performance, including GPU utilization and networking. Partner with hardware and network teams to validate and improve system b

PythonKubernetesLinuxMachine Learning
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. ABOUT BASE LABS Base Labs is a research lab pushing the frontier of open-source LLMs. We think intelligence should be democratized, not controlled by a handful of closed labs and we think very few teams are actually positioned to do something about that. Backed by Baseten's training and inference infrastructure, we have the compute, resources, and talent to take on hard problems at the frontier and open-source what we learn along the way. Our mission is to help build a world where intelligence isn't concentrated, but spread across an ecosystem of models that anyone can build on. That mission shapes what we choose to work on, how we work on it, and who we want in the room. We are accepting applications on a rolling basis for our first cohort of Base Labs Fellows, which is expected to start in late September. Apply using this link. BASE LABS FELLOWSHIP OVERVIEW The Base Labs Fellowship is designed to give researchers exposure to what frontier research looks like in industry. We provide funding, mentorship, and full support to our fellows, with the goal of producing rigorous, published research that shapes both the open-source ecosystem and Baseten's technical roadmap. We run multiple cohorts of Fellows each year and review applications on a rolling basis. This application is for cohorts starting in Sept 2026 and beyond. WHAT TO EXPECT 3 months of full-time research from our San Francisco office A dedicated 1:1 mentorship

Machine LearningAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an Inbound Sales Development Representative at Baseten, you'll be the first point of contact for prospects who've shown interest in Baseten - quickly engaging, understanding their needs, and qualifying them against our ideal customer profile. You'll turn inbound interest into qualified pipeline. RESPONSIBILITIES Build revenue pipeline by setting introductory meetings with potential businesses and key decision makers. Diligently respond to inbound inquiries and determine potential product fit. Help influence Baseten's product roadmap for customers and prospects. Identify high-potential businesses and verticals and develop and execute outbound strategies to bring them to Baseten. Stay up-to-date on market trends, competition, and industry developments. Manage and document the progression of the sales pipeline. Drive pre- & post-engagement at industry events (will attend multiple events in person). REQUIREMENTS Ability to develop strong, long-lasting relationships both internally and externally. Collaborative and coachable, always looking to improve your skills and impact. Ability to handle rejection and stay persistent in pursuing leads. Excellent verbal and written communication skills. A basic understanding of a standard SaaS seller's technical stack (CRM, Outreach/Salesloft, Prospect Research, etc.). Excellent time management, process, and prioritization skills. Preferred experience with Cloud or Secur

RestMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are looking for an IT Support / Operations Engineer to join Baseten as we continue to scale our IT team. In this role, you will play a critical part in bringing our technical support entirely in-house to provide a seamless, high-touch experience for all Baseten employees. As we continue to scale, you will be the primary point of contact for day-to-day technical issues, allowing you to have a direct impact on our team's productivity and overall office environment. This position is ideal for a hands-on problem solver who enjoys a mix of hardware and software troubleshooting, user lifecycle management, and maintaining the physical IT infrastructure of a modern office. While you will focus heavily on elevating our internal support standards, you will also assist with systems administration and workflow automation as our company evolves. This is a hybrid role based out of our San Francisco or New York office, following our standard policy of three days per week in-person to ensure our physical office and AV systems remain high-performing and reliable. RESPONSIBILITIES Serve as the escalation point for day-to-day technical support, diagnosing and resolving hardware and software issues across our Mac and Windows fleet Manage user lifecycle administration including provisioning, deprovisioning, and access management across all systems and services Own the IT onboarding experience for new employees — from laptop set

Machine LearningAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab

KubernetesMachine LearningAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. We are looking for an engineer with strong experience in machine learning and solid foundations in maths and computer science to join our growing Post-Training team at Baseten. Custom models are instrumental to the success of Baseten customers. By inference volume, the overwhelming majority of traffic at Baseten is to and from models that have been post-trained in some way, whether that be through reinforcement learning, supervised finetuning, a recent technique from the literature, or an in-house research technique from Baseten. The Post-Training team is responsible for the success of our customers’ post-trained models, and we employ a wide array of techniques to produce models that are more efficient and higher quality than even the biggest closed source models for the customer’s specific needs. Your role as a research engineer is to build the in-house tooling to support all of this. We care about training a wide spectrum of different model architectures with a variety of techniques efficiently and at scale. At times this involves zooming deep into a particular technical topic, but more often if involves working across the stack as a whole - systems-level concepts like Kubernetes, cgroups, storage systems, and networking topologies, as well as PyTorch distributed tensor computation, and GPU kernels. RECENT RESEARCH Dense, on-policy or both? Repeated kv cache for long-running agents Distillation without the dark – rep

KubernetesMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. ABOUT THE ROLE We are seeking a strategic and creative Partnerships Product Marketing Manager to drive the success of our partner ecosystem. In this role, you will work closely with our partnerships team and with leading technology companies including NVIDIA, Cloud Hyperscalers, Neocloud providers, and emerging ecosystem players, to bring differentiated joint solutions to market. You will be responsible for shaping partner narratives, creating impactful content, and enabling joint go-to-market programs that amplify our reach. The ideal candidate combines storytelling skills, technical acumen, and the ability to collaborate across partner and internal teams. RESPONSIBILITIES Partner Storytelling & Content Develop compelling blog posts, white papers, solution briefs, and customer success stories highlighting the value of joint solutions. Translate technical integrations into clear, differentiated messaging for both technical and business audiences. Enablement & Training Build and deliver enablement programs to arm partner sales teams with the tools, messaging, and collateral they need to position and sell our solutions. Create sales playbooks and battlecards tailored to each partner’s ecosystem. Joint Go-to-Market Campaigns Plan and execute co-marketing campaigns with partners, including webinars, events, solution launches, and demand-generation activities. Manage partner marketing calendars and ensure alignment

Machine LearningAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for a customer-obsessed software engineer to come ship with us. You’ll own features like multi-node training and products like serverless reinforcement learning (RL) from conception to MVP (and from MVP to GA!). You’ll work through the stack, architecting solutions from API and UI down to our infrastructure layer. You’ll fine tune models yourself to develop an understanding of user workflows. You’ll work closely with research engineers leveraging state-of-the-art training techniques to build experiences that accelerate model development and solve for real pain points. If you’re excited to dive deep into the training, let’s talk! THE PRODUCT Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done EXAMPLE INITIATIVES Checkpointing Pipeline: Our checkpointing pipeline starts with automated checkpointing, a feature that ensures that versions of models created during training are automatically backed up to the cloud. Users are able to then deploy checkpoints seamlessly into inference servers, providing point-and-click integrations into inference frameworks like vLLM and Baseten’s Inference Stack. This enables customers to quickly evaluate the performance of their checkpoints with real traffic. Multinode training: Multinode training enables customers to easily run training jobs across multiple compute nodes, enablin

KubernetesRestMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Baseten’s Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost‑efficient. As part of this team, you’ll focus on Model APIs — the infrastructure powering our hosted API endpoints for the latest open‑source models. This work spans distributed systems, model serving, and developer experience. You’ll join a small, high‑impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale. RESPONSIBILITIES: Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations Productionize performa

KubernetesMachine LearningAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re hiring a People Business Partner to support our team during a critical phase of growth. This is a highly strategic, high-impact role for someone who has partnered closely with leadership teams and helped organizations scale with intention. You will be deeply embedded with managers, bringing clarity and rigor to how teams are structured, how leaders operate, and how talent is developed across the organization. You’ll help shape team effectiveness, identify critical talent gaps, drive talent and performance strategies that enable high-performing teams, and build the people practices and change management approaches that allow us to scale with both speed and discipline. This role requires strong business judgment and the ability to operate with deep context. You’ll partner closely with leaders to navigate complex organizational decisions, anticipate challenges before they surface, and bring a clear point of view on what great looks like at every level of the organization. RESPONSIBILITIES Strategic partnership to leadership Serve as the trusted people partner to leadership, maintaining deep business context and translating it into people priorities by proactively surfacing systemic issues, risks, and opportunities before they become urgent. Bring data-driven insights to advise management on org design, succession planning, performance, retention, and engagement. Build management capacity across the org, equ

Machine LearningAIGoRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building for talent density. We believe attracting and retaining exceptional people, and ensuring they feel recognized and valued for their impact, is core to becoming the best place to work. Compensation is a critical lever in that mission. As our Compensation Manager, you will own compensation programs company-wide. You’ll be a trusted advisor to senior leaders, shaping our compensation philosophy, leveling framework, and equity programs to ensure we remain competitive, principled, and performance-oriented as we scale. This role blends strategy and execution: designing clear, fair systems while moving quickly in a high-growth environment. RESPONSIBILITIES Own and evolve Baseten’s company-wide compensation strategy, philosophy, and programs. Collaborate with leadership and HRBP to create and evolve job architecture and leveling frameworks. Build and maintain compensation bands. Conduct regular market benchmarking to ensure comp bands and strategy remain competitive in a fast moving industry. Partner closely with Talent to design and approve competitive new hire offers, advising on negotiation strategy within our compensation principles. Lead bi-annual leveling and compensation review cycles to ensure market competitiveness and reward high performance across teams. Manage new hire equity grants in partnership with Finance and Legal. Design and administer a thoughtful equity refresher program for ten

Machine LearningAIGoRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building a world-class team, anchored in our San Francisco and New York offices and increasingly growing across the globe. We believe the best talent can come from anywhere, and our ability to hire and support that talent is mission-critical. As our Immigration and Mobility Manager, you will make global hiring operationally seamless. You’ll be Baseten’s in-house expert on immigration, relocation, and international employment strategy, ensuring that exceptional candidates from all over the world can confidently build their careers with us. This is a high-stakes, high-impact role. Speed, clarity, and correctness matter deeply in immigration and mobility. You will own the end-to-end experience, navigate a rapidly evolving U.S. immigration landscape, and design the systems and policies that enable Baseten to hire globally while delivering an exceptional employee experience. Over time, you’ll also help shape where and how we expand internationally, advising on global hiring models, new hubs, and employment structures that support our long-term growth. RESPONSIBILITIES Own the end-to-end immigration lifecycle, managing all U.S. visa processes (e.g., H-1B, O-1, J-1, TN, E-3, L-1, EB-2/3 PERM) from offer stage through renewals and permanent residency. Oversee and project manage visa sponsorships executed through Employer of Record (EOR) partners Serve as Baseten’s internal immigration expert, partnering clo

Machine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg

PythonAWSGCPKubernetes
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI’s AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production. In this role, you will: Design and implement the LLM inference runtime for frontier models running on custom silicon. Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization. Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads. Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack. Enable new

PythonAWSRestAI
🔔

Get new inference engineering and product lead jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime