Jobs in United States

Inference Technical Lead in United States

672 active opportunities · Updated October 2026

Explore current inference technical lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We’re looking for a Staff Forward Deployed Engineer who’s excited to build with the engineers using the AI computers Tenstorrent makes. You will create continuity between customers, engineering, and AI inference service products. This is an engineering role first: you contribute production code, operate deployments, and you can explain a trade-off to customer leadership as clearly as to core engineering teams. This is a high-autonomy role with direct customer impact. This role is remote, based out of North America, with preference near one of our main hubs: Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You understand how accelerator compute, memory, and networking topology constrain AI workloads, and don't treat hardware as a black box. You're an early adopter of AI for your work from coding to building agentic workflows that multiply your impact. You work directly with customers to understand their challenges and provide effective solutions. You are comfortable debugging across the full inference stack: from failing requests, through the serving layer, down to OOMs or kernel dispatch if n

AWSKubernetesMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. Our mission is to push the frontier of code generation and agentic reasoning, and deploy these capabilities in real-world products such as ChatGPT and the API, as well as in next-generation tools specifically designed for agentic coding. We operate across research, engineering, product, and infrastructure—owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. About the Role As a Performance & Systems Engineer on the Codex team, you will be responsible for whole-system optimization across a complex, evolving stack. Codex spans LLM inference, cloud orchestration, agentic work management, and multiple product surfaces. Your job will be to identify and land high-leverage changes—across infrastructure, modeling, and product layers—that make Codex agents significantly faster and cheaper to serve. We’re looking for generalists who thrive in ambiguity and love chasing performance bottlenecks to ground. This is a high-ownership role where your work will directly improve the experience of millions of users. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Hunt down and address inefficiencies across the Codex system stack, from agent behavior to LLM inference to container orchestration, and beyond. Build tooling to measure, profile, and optimize system performance at scale. Collaborate with researchers and engineers to land high-ROI changes that improve latency and cost. You might thrive in this role if you: Have experience operating across both ML systems and cloud infrastructure. Enjoy diving into messy, ambiguous problems and emerging with clear wins. Think holistically about performance, balancing spee

AWSRestAIRust
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We are hiring at a rapid rate, and every one of our new hires deserves a great new hire experience from the moment they sign their offer letter through their entire onboarding journey. As our People Operations Coordinator, you'll own the tactical side of onboarding: collecting and tracking the completion of required documentation, greeting new hires, assisting in running sessions, and keeping the operations of onboarding on track as we grow. The role is heavily onboarding-focused today and will broaden across the employee lifecycle as our new people systems take over more of the routine work. Things change quickly here, and this role will too. This role is based in our San Francisco office. RESPONSIBILITIES Own the full onboarding experience for every new hire, including: new hire communication from offer signature through Day 1; ensuring completion of onboarding tasks like background checks, Form I-9s, setting up HRIS profiles, etc; coordinating travel logistics; partnering with the Global Mobility Manager on immigration matters that may affect start dates Support day one and San Francisco week one onboarding program, including coordinating with managers and buddies, room booking and coordination of start location Partner with IT and Workplace so new hires have the right access, equipment, and desk waiting on day one Manage the logistics behind our onboarding platform and manage the day-to-day vendor relationsh

Machine LearningAILogistics
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We’re looking for a Corporate Accounting Manager to support the general ledger, the close, and our accounting processes and controls as Baseten scales. This is a hands-on, individual-contributor role for someone who wants full ownership of the core accounting function at a company where the entity structure, transaction volume, and reporting requirements are growing quickly. You’ll join the corporate accounting team and take on the general ledger, vendor contract review, the close calendar, and our internal control environment as the business grows. We are focused on tightening close procedures, strengthening controls, and building reporting rigor to support the scale ahead. You’ll partner closely with FP&A, Data, and cross-functional partners to make sure the books close on time and accurately every cycle. Baseten is building the infrastructure layer for AI-native companies, and we're scaling quickly - in headcount, entity structure, transaction volume, and contract complexity and volume and vendor size. If you want real ownership over a growing function on a lean team, this role offers real scope. RESPONSIBILITIES General Ledger and Financial Reporting Own day-to-day execution of general ledger journal entries, account reconciliations, and monthly financial statement preparation as the entity structure grows Own recurring close areas, including accruals, prepaids, fixed assets, and equity compensation Lead

Machine LearningAIAccountingFinance
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We are bringing our people systems in-house on Workday, and we are hiring our first dedicated Workday Analyst to make it excellent. You will join while the implementation is underway, ramp alongside our deployment partner, and own the tenant from go-live onward. This is a hands-on configuration role: you will build business processes, manage security, load data, and test releases yourself, and you will teach others on the People team to do the same as we grow. If you want to shape a Workday environment from its first day in production instead of inheriting years of someone else's decisions, this is that rare opening. RESPONSIBILITIES Own day-to-day Workday configuration: business processes, security groups and roles, custom reports, calculated fields, and tenant settings Build and run EIB loads for data changes, mass updates, and audits Own the twice-yearly Workday release cycle: evaluate new features, regression-test, and roll out changes safely Shadow the implementation build, then take over tenant ownership at go-live Monitor integrations and triage issues with our IT team and vendors Support payroll configuration in partnership with our Accounting team and payroll services provider Field and resolve system requests from employees, managers, and the People team Mentor teammates so Workday administration becomes a team capability, not a single point of failure REQUIREMENTS 4+ years administering a live Workday

Machine LearningAIAccountingPayroll
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. ABOUT THE ROLE This role owns Baseten's relationships and market intelligence across the hardware and chip layer of the compute stack: NVIDIA directly and key OEM partners such as Dell, Lenovo, Pegatron, and Supermicro. As Baseten's compute strategy increasingly depends on hardware access and terms, this role is central to keeping Baseten ahead of the market. WHAT YOU'LL DO Build and maintain relationships across NVIDIA and key OEM partners (e.g. Dell, Supermicro) Track market intelligence on hardware availability, roadmaps, and terms to keep Baseten informed and strategically well-positioned Support deal structuring and negotiation in partnership with Baseten's deal-making function Work closely with Infrastructure and Hardware Platform engineering teams to ensure consistent, high-quality provider relationships and engineering partnerships Represent Baseten credibly across senior relationships in the hardware ecosystem, escalating to company leadership when strategically valuable WHAT WE'RE LOOKING FOR Existing relationships and credibility within the NVIDIA, OEM, and HPC ecosystem Strong relationship-management instincts, with the judgment to know when to bring in senior leadership for maximum impact Comfort operating in a fast-moving, high-stakes market where hardware access can be a major competitive differentiator Collaborative style — this role depends on close coordination with engineering counterparts, not just ex

Machine LearningAIGoHR
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. ABOUT THE TEAM Supply is responsible for knowing everything happening in the compute market: who's building, who's buying, and on what terms. This role owns a specific and fast-moving slice of that map — emerging clouds and international markets — and owns the full relationship lifecycle in that space, from first outreach through to closed terms. RESPONSIBILITIES Build and maintain a real-time picture of the emerging cloud and international compute landscape — who's active, what they're building, and what terms are available Own the full partnership lifecycle in this space — from identifying and sourcing new providers, to negotiating terms, to ongoing relationship management Develop and manage relationships across a broad set of emerging and international providers, from account reps up through leadership Identify, structure, and help close opportunities where Baseten can move quickly to secure favorable capacity terms Define compelling value propositions tailored to different types of providers, rather than a one-size-fits-all pitch Partner closely with others in the team already covering this space to build out a durable, well-organized intelligence and relationship function Collaborate with the broader Supply and Deals functions to bring opportunities to the table and support negotiation when it's time to close WHAT WE’RE LOOKING FOR Equal parts relationship-builder and operator — you can open a door and also drive it

Machine LearningAIGoHR
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE This is a sourcing-first role, not a deal-closing role. Baseten needs someone who can build and maintain deep relationships across the long tail of data center and powered land providers, well beyond the handful of large, well-known players that everyone in the market is already competing for. This coverage area is a key differentiator for Baseten's broader compute strategy, so we're looking for the best possible person in this specific lane rather than a generalist. You'll own the full lifecycle of a sourcing relationship — from first outreach to ongoing management — not just the introduction. WHAT YOU'LL DO Build and maintain a comprehensive map of data center and powered land opportunities, with a particular focus on the long tail rather than the handful of major, oversubscribed players Own the full sourcing lifecycle for each relationship — from identifying and reaching out to new providers, through negotiation support, to ongoing relationship management — not just the initial introduction Develop and manage sourcing relationships across neoclouds, hyperscalers, brokers, and independent operators Quickly and independently evaluate new sites and spaces to determine fit and priority Prepare business cases and cost analysis to support new data center and powered land opportunities, partnering with Finance where needed Maintain accurate records of suppliers, contracts, and commercial terms so the team has a reli

Machine LearningAIGoProject Management
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE Baseten's Compute org is in hyper growth. As it scales, the systems and workflows that keep supply and demand balanced across our GPU fleet need to get more sophisticated, and this role exists to make sure they do. Compute sits at the center of how Baseten allocates, forecasts, and manages the capacity that powers every customer inference request. The team that supports this work, C3, runs on a mix of internal tooling, manual processes, and systems that haven't fully kept pace with the scale of the problem. This role exists to close that gap. You'll design, build, and ship AI-powered workflows that give the Compute and C3 teams real leverage, automating the manual, repetitive, and error-prone parts of the capacity lifecycle so the team can focus on judgment calls that actually need a human. We want someone who can walk in, audit what exists today, identify what's missing or broken, and start shipping fast. You know when to reach for an existing internal tool and when to build something custom in Claude Code. You think two to three steps ahead about how the thing you build today fits into the broader capacity systems architecture tomorrow. And you bring a point of view on our stack, on what we should be building, and on where AI can do something existing tooling simply can't. RESPONSIBILITIES Ship AI-powered workflows for Compute and C3 : build the agents and automations that give capacity analysts, ops leads, an

Machine LearningAIGoRust
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Manager focused on TPUs at Baseten, you will lead the "engine room" for our non-NVIDIA accelerator fleet, architecting, securing, and optimizing the Google Cloud TPU (and broader emerging accelerator) capacity that powers our customers' AI workloads. You'll own the end-to-end journey of capacity management for this fleet, from securing large-scale TPU pod allocations to building the automation that ensures reliable uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering, with a specific focus on the TPU ecosystem. You will act as the fleet orchestrator for Google's TPU architecture, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics as we diversify beyond NVIDIA. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the latest generation of TPU hardware, like Google's Trillium (v6e) architecture, and partnering closely with the Model Performance (MP) team to ensure workloads are tuned for TPU-specific execution. EXAMPLE INITIATIVES The TPU Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's TPU clusters, including pod slicing and topology planning Global Workload Orchestration: Bui

PythonAWSAzureGCP
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a member of the Capacity Strategy & Operations team, you will sit at the intersection of supply intelligence, demand forecasting, and cross-functional execution, turning a complex, fast-moving hardware market into a predictable, reliable foundation for our customers and internal engineering teams. This is not a purely analytical role. You will own the end-to-end capacity planning process: from translating customer commitments and growth forecasts into concrete supply requirements, to coordinating fulfillment across vendors, finance, and the infrastructure team, to building the systems that make all of this repeatable and scalable. When supply is constrained and tradeoffs are unavoidable, you are the person in the room who can model the options, make a clear recommendation, and drive alignment fast. You are a strong fit if you have operated at the intersection of strategy and execution before — someone who is equally comfortable building a capacity model in a spreadsheet and running a cross-functional war room when a customer deployment is at risk. EXAMPLE INITIATIVES Demand-Supply Alignment Framework: Build and own the process that translates customer pipeline, signed commitments, and growth projections into a forward-looking GPU demand signal — so the team is never caught flat-footed when a customer scales faster than expected. Constrained Allocation Playbook: Define the decision framework for how Basete

Machine LearningAIGoExcel
B
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We’re looking for a Recruiting Coordinator to help create a seamless, welcoming, and well-organized interview experience for every candidate who engages with our team. You’ll work closely with our recruiters to coordinate both virtual and in-person interviews, support executive involvement when needed, and ensure candidates have everything they need during throughout their interview process. This role is ideal for someone who thrives on operational excellence, loves solving logistics problems on the fly, and brings both warmth and precision to every interaction. RESPONSIBILITIES Work closely with recruiters and hiring managers to coordinate interview loops and debriefs for candidates and the internal team members conducting interviews Ensure every candidate has a smooth, well-communicated, and positive experience Manage logistics for onsite interviews, including candidate arrival and workspace setup Proactively identify and solve day-of issues, including last-minute changes or scheduling conflicts Communicate clearly and promptly with candidates and internal teams about interview logistics and updates REQUIREMENTS 1+ year of recruiting or HR experience Detail-oriented and operationally strong—you know how to keep things moving Clear and professional written and verbal communication skills Personable and warm—you're great at making candidates feel welcome and supported Ability to think on your feet and respond to

Machine LearningAIExcelLogistics
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We’re looking for a high-performing strategic finance professional to join our growing GTM Finance team. Our business grows with our customers' usage, which makes the finance function highly strategic at Baseten: growth, pricing, margin, and capacity decisions are business model decisions. You'll sit at the center of them, partnering directly with GTM leadership and reporting into a finance team with a seat at the table for the calls that shape the company's trajectory. This role is ideal for someone with 3 to 7 years of experience across strategic finance, investing, and/or investment banking who wants broad exposure to company-building inside a fast-scaling AI infrastructure company. Experience at a usage-based software company is a plus. RESPONSIBILITIES Own financial planning, forecasting, and budgeting processes for the GTM org Build and maintain financial models across revenue, S&M spend, headcount, and strategic bets Analyze the metrics that define a usage-based business – ARR, gross margin, consumption trends, retention, and GTM efficiency Partner with GTM leaders to set targets, evaluate growth initiatives, shape pricing, and design sales compensation Help prepare board materials, investor updates, and fundraising analyses Improve financial reporting, dashboards, and operational rigor so our infrastructure scales as fast as our revenue Work cross-functionally to turn ambiguous business questions int

SQLRestMachine LearningAI
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Lead at Baseten, you will lead the "engine room" of the company, architecting, securing, and optimizing the global GPU fleet that powers our customers' AI workloads. You’ll own the end-to-end journey of capacity management, from securing multi-million dollar GPU clusters to building the automation that ensures 99.9% uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering. You will act as the fleet orchestrator for the world's most advanced chips, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the next generation of hardware, like NVIDIA’s Blackwell (B200) architecture. EXAMPLE INITIATIVES The B200 Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's first Blackwell GPU clusters. Global Workload Orchestration: Building "Multi-cloud Capacity Management" systems to move customer workloads seamlessly across regions to optimize cost and latency. Precision GPU Triage: Developing automated Go-based operators to identify, cordon, and repair unhealthy H100 nodes in under an hour. The Supply Chain of Intelligence: Partnering with lead

PythonAWSAzureGCP
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments. We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment. This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment.. RESPONSIBILITIES Core Responsibilities: Drive suppliers to keep the maximum amount of the GPU fleet online and healthy. Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs. Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier. SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short. Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability. Scope and

Machine LearningAIGoFinance
🔔

Get new inference technical lead jobs in United States by email

Daily job updates · Unsubscribe anytime