ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Container runtimes were designed for general-purpose software workloads. AI inference is not a general-purpose workload. Running large models at production scale exposes cracks in every layer of the container stack: runtimes unaware of GPU memory constraints, images that take minutes to pull when a model needs to scale to thousands of replicas, and isolation mechanisms that weren't designed for the multi-tenant serving environments that production AI requires. The tools the industry has relied on for a decade weren't built for this, and patching around those limitations at higher layers only goes so far. Baseten owns the entire pipeline, from the moment a developer pushes a model to the moment a request gets a response. That vertical ownership means we can fix these problems at the root. The Runtime Fabrics team is doing exactly that: purpose-building the container runtime and storage layers for AI inference workloads, led by some of the world's top containerd maintainers. As Engineering Manager of the Runtime Fabrics team, you will lead this work, setting technical direction, growing a world-class team of systems engineers, and ensuring the team's output shapes not just Baseten's infrastructure but the open-source container ecosystem at large. If you've contributed to containerd, runc, or related OCI projects and are ready to lead a team solving some of the hardest problems in infrastructure today, we'd love
Jobs in United States
Inference Technical Lead in United States
672 active opportunities · Updated October 2026
Showing
15 jobs
Explore current inference technical lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an Inbound Sales Development Representative at Baseten, you'll be the first point of contact for prospects who've shown interest in Baseten - quickly engaging, understanding their needs, and qualifying them against our ideal customer profile. You'll turn inbound interest into qualified pipeline. RESPONSIBILITIES Build revenue pipeline by setting introductory meetings with potential businesses and key decision makers. Diligently respond to inbound inquiries and determine potential product fit. Help influence Baseten's product roadmap for customers and prospects. Identify high-potential businesses and verticals and develop and execute outbound strategies to bring them to Baseten. Stay up-to-date on market trends, competition, and industry developments. Manage and document the progression of the sales pipeline. Drive pre- & post-engagement at industry events (will attend multiple events in person). REQUIREMENTS Ability to develop strong, long-lasting relationships both internally and externally. Collaborative and coachable, always looking to improve your skills and impact. Ability to handle rejection and stay persistent in pursuing leads. Excellent verbal and written communication skills. A basic understanding of a standard SaaS seller's technical stack (CRM, Outreach/Salesloft, Prospect Research, etc.). Excellent time management, process, and prioritization skills. Preferred experience with Cloud or Secur
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As the Engineering Manager for Baseten's Cloud Platform team, you will directly manage a team of cloud platform engineers responsible for building the systems and processes that keep our infrastructure scalable, reliable, and efficient — from automated deployments and monitoring to performance optimization and incident response. You are a people-first leader with a strong cloud infrastructure background. You set a high bar for reliability and operational excellence, engage credibly in technical discussions and code reviews, and know how to build a culture of ownership and accountability. You'll spend most of your time close to the work: unblocking your team, shaping technical direction on day-to-day decisions, and developing your engineers. At Baseten, we work closely with our users to understand their struggles operationalizing ML — you'll keep your team connected to that mission and translate user learnings into better infrastructure. RESPONSIBILITIES Recruit, hire, and grow a high-performing team of cloud platform engineers; provide ongoing coaching, feedback, and career development through regular 1:1s. Set clear performance expectations, hold a high bar, and create an environment where engineers do their best work. Foster a culture of ownership, accountability, and continuous improvement. Drive day-to-day technical decisions through design reviews, code reviews, and architectural discussions; translate th
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for an Enterprise Account Executive to help build and scale Baseten’s Enterprise go-to-market motion. You’ll own new business in verticals where AI adoption is accelerating including enterprise software, financial services, and big tech. This is a high-autonomy role where you’ll contribute directly to our Enterprise GTM strategy—identifying new use cases within your verticals, winning lighthouse customers that become references for their industries, and feeding signal back to product and engineering to shape our roadmap. You’ll work alongside Baseten’s founders, forward-deployed engineering team, and GTM leadership. WHAT YOU’LL DO Own a revenue target and all aspects of the sales cycle from prospecting to close, including outbounding and engaging Tier 1 accounts in your assigned verticals Drive new logo acquisition and strategic expansion within key accounts, prioritizing organizations that can serve as lighthouse customers within their industries Become a trusted advisor to customers—understand their unique infrastructure needs, co-innovate on solutions, and lead with conviction by providing clear recommendations grounded in deep industry expertise Partner with Baseten’s forward-deployed engineers to run technical evaluations, POCs, and architecture reviews that earn customer confidence Collaborate cross-functionally with Product, Engineering, Legal, and Marketing to bring new solutions to marke
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: As a Software Engineer at Baseten, you will own one of the most critical surfaces of our business: pricing, billing, and revenue infrastructure. As we launch more and more products— billing is no longer just operational plumbing. It is a strategic lever for growth. This role will establish clear ownership of billing as a function and create leverage for Finance, Sales, and GTM teams while maintaining a seamless customer experience. RESPONSIBILITIES: Own Baseten’s end-to-end billing and revenue infrastructure, including pricing, invoicing, metering, and reporting foundations. Build and evolve our billing platform and integrations (including Orb), ensuring correctness, auditability, and a high-trust experience for customers and internal teams. Partner closely with Finance, Sales, GTM, and Forward Deployed Engineering to turn real-world workflows into reliable internal tooling and automation (quoting, approvals, renewals, usage reconciliation, revenue reporting). Design systems that scale with new products, packaging, and go-to-market motions, making billing a strategic lever for growth. Drive reliability and operational excellence for revenue-critical workflows: monitoring, alerting, incident response, backfills, and clear runbooks. Lead from the front on high-impact projects: clarify requirements, propose crisp technical approaches, ship iteratively, and raise the bar on quality and velocity. Debug and resolve
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for Forward Deployed ML Engineers who want to work at the intersection of deep technical work and direct customer impact. As an ML FDE, you'll partner with leading AI companies and foundation model labs to help them achieve state-of-the-art performance on their most demanding workloads — LLM serving, model training (SFT, RLHF), audio pipelines, scientific computing, and more. You're helping teams reach outcomes most engineers can't on their own. The FDE team today includes world-class software engineers, computational scientists, ML engineers, and former founders. We're looking for people with strong engineering fundamentals, deep curiosity across the AI stack, and energy for working directly with customers on hard problems. You will: Work hands-on with companies like Suno, Lovable, Cognition, and Meta to architect and optimize production AI workloads
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: Modal builds AI infrastructure products that developers love. That's how we grew so quickly and why word-of-mouth remains one of our most important channels today. From powering one of the largest vibe-coding platforms at Lovable to enabling teams like Ramp to build their own internal coding agents , Modal Sandboxes are used by developers to safely execute AI-generated code at scale. We're now hiring our first developer relations engineer focused on Modal Sandboxes. Whether it’s banger tweets , in-depth technical resources or long-form talks , we want to meet developers by any medium necessary and empower them to build and ship novel AI products. In this role, you will primarily be creating and distributing technical content that is unique, educational, and practical. This content will be the first Modal touchpoint for many of our users. We want to not only showcas
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an
Ready to help developers and researchers get more from open AI models? At NVIDIA, our team improves support for leading community models such as Nemotron, Llama, Gemma, DeepSeek, and Qwen. As a Product Manager for Open Models, you will help coordinate model enablement, developer experiences, technical content, and product launches across NVIDIA’s accelerated computing platform. You will work alongside experienced product managers, engineers, model builders, and product marketing teams. This role is ideal for someone early in their product-management career who has a strong technical foundation, enjoys working across teams, and is excited about the open-model ecosystem. What you’ll be doing: Support collaboration with community model builders around model access, technical enablement, launch readiness, and go-to-market activities. Track emerging open-model releases and summarize their capabilities, technical differentiators, hardware requirements, and ecosystem impact. Maintain product requirements, launch plans, readiness checklists, and supporting documentation for assigned models. Partner with engineering, infrastructure, and developer-experience teams to support open models across inference, fine-tuning, evaluation, and deployment. Gather feedback from model builders and developers, identify recurring issues, and translate findings into actionable product requirements. Review and test developer workflows, sample applications, notebooks, and Python code used in demonstrations and technical content. Collaborate with product marketing on blogs, documentation, presentations, case studies, and social-media content. Support product announcements, demonstrations, keynote materials, and other high-visibility launches while working onsite at least three days per week. What we need to see: 2+ years of relevant experience in product ma
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our
Job Details: Job Description: This is a high-visibility, commissioned sales leadership role within Intel's US Sales organization, specifically focused on our most disruptive AI-Native and Strategic CSP accounts. You will be the primary architect of Intel's relationship with industry titans who are redefining the boundaries of AI model training, AIaaS solutions and deployment at scale. This is not a traditional sales role. You will operate at the intersection of deep technical engineering and executive business strategy, ensuring Intel's silicon and software roadmap aligns with the world's most demanding AI-as-a-Service and SaaS platforms across on-prem, Tier1 CSP and NeoCloud environments. Key Responsibilities Executive Orchestration: Act as the One Intel lead, building deep-rooted partnerships with C-suite executives and Principal Engineers at world-class AI and SaaS firms. Technical Value Synthesis: Translate complex hardware architectures (CPU, GPU, Accelerator, Networking, and Packaging) into business outcomes for customers running massive-scale distributed training and inference workloads. Strategic Growth: Drive Intel's data-centric growth strategy by identifying and securing design wins within the core infrastructure of the world's leading AI models and solution providers. Cross-Functional Leadership: Partner closely with Cloud Solution Architects (CSAs), Intel Business Units, Cloud and OEM partners to influence future product roadmaps based on the unique needs of AI-native disruptors. Market Evangelism: Serve as a technical and business evangelist, articulating Intel's vision for the future of AI and compute infrastructure in a highly competitive landscape. <p style="text-align:inhe
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
$293K – $325K/yr
About the Team The Statsig team at OpenAI builds and operates the experimentation platform that powers product development, measurement, and decision-making across the company. We partner closely with product, engineering, and infrastructure teams to ensure experiments are trustworthy, statistically rigorous, and scalable to the needs of frontier AI products. Our mission is to help teams make better decisions through reliable experimentation. We care deeply about statistical correctness, pragmatic solutions, and building systems that researchers and engineers can trust at massive scale. The team operates at the intersection of experimentation methodology, data infrastructure, causal inference, and product analytics. We are looking for experienced experimentation experts who want to shape the future of experimentation in the AI era. About the Role We are hiring a Staff-level Data Scientist to help lead the evolution of OpenAI’s core experimentation platform. This role is focused on improving the statistical rigor, reliability, and practical usability of experimentation across the company. You’ll work on some of the hardest problems in online experimentation: sample ratio mismatch detection, variance reduction, bias mitigation, metric design, triggered analysis, heterogeneous treatment effects, sequential testing, and experimentation in complex ML systems. You’ll also help translate advanced statistical concepts into pragmatic systems and product experiences that teams can actually use. This is a highly technical individual contributor role with significant influence across methodology, platform architecture, and experimentation best practices. The ideal candidate combines deep statistical expertise with strong systems intuition and hands-on experience building or operating experimentation platforms at scale. In this role, you will: Drive the statistical direction and technical strategy for OpenAI’s experimentation platform Design and improve experimentation methodolo
About the Team OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical. The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly. About the Role We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity. In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity. This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions. This role is based in San Francisco, CA, with travel as needed. Key Responsibilities Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint Drive readiness to convert contracted compute capacity into schedulable production clusters Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives Build integrated schedules spanning procurement, logistics, installation, st
Other cities to consider
More places hiring for this role
Get new inference technical lead jobs in United States by email
Daily job updates · Unsubscribe anytime