Jobiba hiring network

Inference Technical Lead Jobs

1,491 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current inference technical lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
1mo ago

About the Team OpenAI’s Stargate and 3P Engineering teams are responsible for building and scaling the external infrastructure ecosystem that powers advanced AI systems. We work across hyperscalers, colocation providers, cloud partners, and strategic third-party operators to turn contracted capacity into production-ready compute. Our scope spans the full lifecycle of external deployments: commercial alignment, technical readiness, network integration, hardware enablement, operational readiness, and long-range scaling strategy. As OpenAI’s infrastructure footprint expands globally, we need leaders who can convert complex partner environments into reliable, high-velocity capacity for training and inference workloads. About the Role We are seeking a Technical Program Manager, Token-as-a-Service (TaaS) to lead delivery of external compute capacity that directly serves OpenAI model workloads. In this role, you will own complex cross-functional programs that transform third-party infrastructure into usable tokens at scale. You will partner across engineering, capacity planning, networking, hardware, finance, product, and external providers to ensure that deployed capacity translates into real production throughput. This role sits at the intersection of infrastructure execution, systems readiness, and business impact. Success requires strong technical fluency, elite program management, and the ability to drive accountability across internal teams and external partners. This is a high-visibility role with direct impact on OpenAI’s ability to scale model training and inference globally. This role is based in San Francisco, CA, with a hybrid work model of 3 days in office per week. Relocation assistance is available. Key Responsibilities Lead end-to-end delivery programs that convert external infrastructure capacity into production-ready token supply. Own readiness across compute, storage, networking, security, and operational dependencies for third-party environments. Build

awsrestai
View job →
D
Datalab
📍 New York• Full-time• $300K – $350K/yr
1mo ago

Salary range: $300k - $350k | Equity: 0.4% - 0.6% | In-Person: NYC About Datalab Datalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right. We’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, Chandra, Surya, Marker, and Lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face. Role Overview We're looking for an engineering lead to guide our team while staying hands-on in the code. You'll set the technical direction and standards for how we build the interfaces, tools, and infrastructure behind our OCR, extraction, and document-understanding systems. This includes everything from optimizing agent loops and interfaces to helping to speed up inference. This is a player-coach role. You'll manage and grow a team of three engineers, own engineering delivery and quality, and spend a large share of your time writing code - focused on architecture, infrastructure, and the hard problems rather than routine feature work. You’ll partner closely with the research team to define the handoff between experimentation and production. As a small and fast-moving team, roles are fluid and ownership is high. You'll work directly with the founder to set priorities, ship features, and make our technology accessible to a global community of builders. Day to day, you will: Manage and grow a team of three engineers - 1:1s, prioritization, feedback, and hiring as we scale. Own engineering delivery, quality, and technical standards across code, testing, infrastruct

O
1mo ago

About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation

awsrestai
View job →
R
Roblox
📍 San Mateo• Full-time• From $280.5K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. About the Role Inference Platform is Roblox's multi-tier orchestration system for microservices, powering AI Inference across Roblox. Through a simple developer API, we deliver a "deploy and forget" runtime, hiding the complexity of scheduling, scaling, and reliably running services across our on-prem and multi-cloud footprint at global scale. As Principal Product Manager, Jobs Platform, you'll take the helm at a defining moment - leading the charge as we scale the platform to become the default runtime for critical Roblox services worldwide. You Will Own Inference Platform end-to-end - set the multi-year vision for how Roblox engineers deploy, run, and scale AI and other services across our Core and Edge Datacenters, and cloud. Power Roblox's AI future - build the platform that brings frontier models and next-gen AI workloads to life, with the primitives, scheduling guarantees, and resource classes AI teams need to move fast. Evolve the platform's techni

reactawsazure
View job →
L
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Growth Products team drives rider and driver acquisition to scale the business and balance the marketplace. We specialize in incentive and messaging targeting, budget optimization, and paid media measurement, and move rapidly to test new ideas and products. As a Data Scientist expert in causal inference and marketing mix models (MMM), you will lead our efforts to measure and optimize investments across marketing channels. Responsibilities: Deliver results across the entire lifecycle of data science solutions for Growth: from defining the problem with cross-functional stakeholders to deploying production models that address key business problems. Own complex domains and develop long-term roadmaps to maximize business impact. Build statistical pipelines, write production code, and design/analyze experiments. Participate in the science on-call rotation to ensure automated campaigns operate successfully. Experience: Advanced degree in statistics, economics, mathematics, or equivalent industry experience. 4+ years of industry experience in causal inference or data science. Proven ability to apply statistics to unstructured problems and deliver measurable results. Deep technical expertise in causal inference and tackling challenging measurement problems. Expertise in marketing mix modeling is highly preferred. Expertise in SQL and experience with large-scale data platforms. Proficiency in Python and working within production coding environments. Benefits: Great medical, dental, and vision insurance options with additional programs available when enrolled Mental health benefits Family building benefits Child care and pet benefits 401(k) plan with company match to help save for your future In addition to 12 observed holidays, salaried team members have discretionary paid time off, hourly team membe

pythonsqlai
View job →
L
Lyft
📍 San Francisco• Full-time
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Growth Products team drives rider and driver acquisition to scale the business and balance the marketplace. We specialize in incentive and messaging targeting, budget optimization, and paid media measurement, and move rapidly to test new ideas and products. As a Data Scientist expert in causal inference and marketing mix models (MMM), you will lead our efforts to measure and optimize investments across marketing channels. Responsibilities: Deliver results across the entire lifecycle of data science solutions for Growth: from defining the problem with cross-functional stakeholders to deploying production models that address key business problems. Own complex domains and develop long-term roadmaps to maximize business impact. Build statistical pipelines, write production code, and design/analyze experiments. Participate in the science on-call rotation to ensure automated campaigns operate successfully. Experience: Advanced degree in statistics, economics, mathematics, or equivalent industry experience. 4+ years of industry experience in causal inference or data science. Proven ability to apply statistics to unstructured problems and deliver measurable results. Deep technical expertise in causal inference and tackling challenging measurement problems. Expertise in marketing mix modeling is highly preferred. Expertise in SQL and experience with large-scale data platforms. Proficiency in Python and working within production coding environments. Benefits: Great medical, dental, and vision insurance options with additional programs available when enrolled Mental health benefits Family building benefits Child care and pet benefits 401(k) plan with company match to help save for your future In addition to 12 observed holidays, salaried team members have discretionary paid time off, hourly team membe

pythonsqlai
View job →
L
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Growth Products team drives rider and driver acquisition to scale the business and balance the marketplace. We specialize in incentive and messaging targeting, budget optimization, and paid media measurement, and move rapidly to test new ideas and products. As a Data Scientist expert in causal inference and marketing mix models (MMM), you will lead our efforts to measure and optimize investments across marketing channels. Responsibilities: Deliver results across the entire lifecycle of data science solutions for Growth: from defining the problem with cross-functional stakeholders to deploying production models that address key business problems. Own complex domains and develop long-term roadmaps to maximize business impact. Build statistical pipelines, write production code, and design/analyze experiments. Participate in the science on-call rotation to ensure automated campaigns operate successfully. Experience: Advanced degree in statistics, economics, mathematics, or equivalent industry experience. 4+ years of industry experience in causal inference or data science. Proven ability to apply statistics to unstructured problems and deliver measurable results. Deep technical expertise in causal inference and tackling challenging measurement problems. Expertise in marketing mix modeling is highly preferred. Expertise in SQL and experience with large-scale data platforms. Proficiency in Python and working within production coding environments. Benefits: Great medical, dental, and vision insurance options with additional programs available when enrolled Mental health benefits Family building benefits Child care and pet benefits 401(k) plan with company match to help save for your future In addition to 12 observed holidays, salaried team members have discretionary paid time off, hourly team membe

pythonsqlai
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
16 days ago

About the Role OpenAI’s Industrial Compute organization is responsible for ensuring our compute infrastructure scales efficiently to support millions of users and increasingly sophisticated AI models. We’re looking for a Data Scientist to partner closely with Capacity Systems Engineering, Infrastructure, Product, and Research to optimize inference capacity across our global GPU fleet. This role combines statistical modeling, large-scale data analysis, forecasting, and systems thinking to drive critical decisions around infrastructure investments, performance-efficiency trade-offs, and customer experience. You’ll transform complex operational data into actionable insights that directly influence how OpenAI allocates and scales one of the world’s largest AI compute environments. Key Responsibilities Build statistical and machine learning models to profile and improve GPU utilization, latency, throughput, and overall fleet efficiency. Develop forecasting models for inference demand across products, regions, and model families. Analyze production workloads to identify latency bottlenecks and capacity constraints, highlighting optimization opportunities. Partner with Capacity Systems Engineering to inform infrastructure planning and long-term GPU investment strategies. Design experiments and simulations to evaluate scheduling policies, serving strategies, and infrastructure tradeoffs. Build dashboards and operational metrics that enable leadership to make data-driven capacity decisions. Collaborate with Product, Research, Finance, and Infrastructure teams to align compute planning with business growth and model roadmaps. Communicate technical findings clearly to both engineering teams and executive leadership. Qualifications MS or PhD in Statistics, Computer Science, Operations Research, Applied Mathematics, Economics, or related quantitative discipline (or equivalent industry experience). 5+ years of experience working in the infrastructure data science space. Strong ex

REMOTEpythonsqlaws
View job →
A
Airbnb
📍 - USA• Full-time• Remote• From $179K/yr
1mo ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: “ Our real innovation is not allowing people to book a home; it’s designing a framework to allow millions of people to trust one another. Trust is the real energy source that drives Airbnb… ” - Brian Chesky, Airbnb Co-Founder & CEO (2019) Data science is the engine behind Airbnb's most impactful decisions. The Platform Data Science team accelerates product evolution and business outcomes by combining scientific rigor with deep domain expertise, spanning experimentation, machine learning, causal inference and scalable intelligence. We partner closely with product, engineering, policy, and operations teams across Trust to detect and defend against the adversarial behavior that threatens guest and host trust: fraudulent listings and fake inventory, review and content manipulation, account takeover, and other bad-actor activity on the platform. Whether measuring the impact of a new listing integrity defense, modeling risk at the listing or account level, or evaluating the effectiveness of an enforcement policy, our work helps guests and hosts experience an Airbnb that is safer, smarter and more personalized. The Difference You Will Make: This role sits at the heart of some of Airbnb's most consequential data science challenges, where rigorous statistical thinking and applied ML directly shape platform outcomes. You will own high-visibility initiatives that require both technical depth and strong business judgment - work that is visible to leadership and has measurable impact on Airbnb's users and bottom line. A Typical Day: The ideal candidate is a technically exce

REMOTEpythonsqlmachine learning
View job →

About Paytm Paytm is a pioneer of digital payments in India, serving over 450 million consumers and 45 million merchants across payments, financial services, and commerce. Over the years, Paytm has built deep in-house capabilities across technology, data, and operations to operate at scale with high reliability. Paytm is building a full stack AI platform focussed on Inference and Agents, enabling large enterprises to deploy AI driven automation across sales, service, operations, and analytics. The Inference and Agentic AI team operates as a cross functional unit spanning engineering, product, data science, business management, and sales, and owns the full lifecycle of AI solutions from opportunity discovery to deployment and scale. Role Overview Paytm is looking to hire a Client Onboarding Director to own implementation delivery and client onboarding governance for Paytm’s AI Inference and Agentic AI products across enterprise clients. This is a managerial role that will lead Client Onboarding Managers and ensure that enterprise deployments move smoothly from sales closure to go live and early adoption. The role sits at the intersection of client teams, product, engineering, business, and sales, and is responsible for converting signed enterprise deals into successful, timely, and scalable deployments. The candidate will own delivery planning, integration governance, risk management, stakeholder communication, and post go live stabilization across enterprise AI agent deployments. The role requires strong program management, technical understanding, client handling, and ability to drive execution across multiple internal and external teams. Key Responsibilities Own end to end delivery governance for enterprise AI agent deployments from sales handoff to go live and stabilization. Lead the implementation planning process across scope, timelines, milestones, dependencies, risks, and success metrics. Ensure every enterprise deployment has a clear project plan, own

gitrestai
View job →

The AI platform is responsible for all AI infrastructure across Datadog. Our mission is to provide tools and platforms that enable data scientists and engineers to conduct large-scale training and inference with ease. We support products such as Bits AI , LLMObs and all our AI research . As an engineering manager for the Evaluation & Annotation team, you’ll join a new and fast growing team and organization. You will support building and scaling the team, define our technical vision and help shape the roadmap. Your team will lead the charge on multiple critical technical challenges: AI model evaluation both offline and online, designing tooling and processes around human annotation, and establishing the standard around synthetics and AI generated datasets. You’ll work closely with sister teams in the AI platform organization ensuring a seamless AI development cycle. You’ll also partner with the Applied AI org and with Datadog infrastructure & tooling teams to build out systems from the ground up. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Manage and grow the Evaluation & Annotation team, directly managing 4-6 engineers Define our technical roadmap in alignment with AI platform goals and the Applied AI team roadmap. Work with our core platform teams to tailor Datadog's storage and data pipelines to our needs Create a strong team culture aligned with our engineering standards and our customer focus Participate in hands-on work: Code reviews, design reviews and some coding Who You Are: A Software Engineer at heart with a previous experience leading software engineering teams, as a tech lead or people manager Excellent leader with strong interpersonal skills, and the

restaigo
View job →

The AI platform is responsible for all AI infrastructure across Datadog. Our mission is to provide tools and platforms that enable data scientists and engineers to conduct large-scale training and inference with ease. We support products such as Bits AI , LLMObs and all our AI research . As an engineering manager for the Training & Serving team, you’ll join a new and fast growing team and organization. You will support building and scaling the team, define our technical vision and help shape the roadmap. Your team will lead the charge on multiple critical technical challenges: distributed training of foundation models, serving at scale, designing the user experience. You’ll work closely with sister teams in the AI platform organization ensuring a seamless AI development cycle. You’ll also partner with the Applied AI org and with Datadog infrastructure & tooling teams to build out systems from the ground up. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Manage and grow the Training & Serving team, directly managing 10+ engineers Define our technical roadmap in alignment with AI platform goals and the Applied AI team roadmap. Work with our core platform teams to tailor Datadog's storage, infrastructure and data pipelines to our needs Create a strong team culture aligned with our engineering standards and our customer focus Participate in hands-on work: Code reviews, design reviews and some coding Who You Are: Previous experience (1+ years) leading software engineering teams, as a tech lead or people manager Strong technician with a mix of backend, data engineer and infrastructure experience who is interested in remaining a hands-on leader Excellent leader with strong

restaigo
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
14 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As our TT-Distributed Software Engineer, you will develop and optimize distributed software systems that power the most efficient and highest-performing AI and HPC clusters. In this role, you'll work on distributed programming across multiple nodes, utilizing systems programming, inter-node communication, and Tenstorrent’s scalable architectures to advance the state-of-the-art distributed inference and training infrastructure. This role is hybrid, based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Strong C or C++ engineer with solid foundations in systems programming, operating systems, and distributed systems principles. Enthusiastic about distributed computing, including IPC, socket programming, and cluster resource coordination. Comfortable reasoning about scalability, fault tolerance, and performance across multi-node environments. Curious and first-principles thinker who challenges conventional approaches to distributed system design. Motivated to grow into a deep technical expert in large-scale distributed AI infrastructure. What We Need Architect, implement, and optim

awsaic++
View job →
T
14 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a skilled Software Engineer with a passion for building high-performance, low-level systems software. In this role, you’ll contribute to the development and optimization of the infrastructure that powers our cutting-edge processors, with a primary focus on C/C++ development and low-level programming. You'll work closely with large inference and training model development to further drive Scale Out software and hardware performance. This role is hybrid, based out of Toronto, ON. Who You Are Strong C or C++ systems engineer with a deep understanding of memory, threading, I/O, and low-level execution models. Experienced building low-level software, drivers, embedded systems, or performance-critical infrastructure. Comfortable working close to hardware and curious about how systems behave under the hood. Proficient with Linux systems programming and debugging tools such as gdb, strace, and perf. Structured problem solver who thrives in fast-paced, highly technical environments. What We Need Design, develop, and maintain core infrastructure software that interfaces directly with Tenstorrent hardware. Build low-level libraries and APIs for communication and synchronization across compute nodes. Optimize system-level software for performance, scalability, and reliability in distributed environments. Support hardware

awslinuxai
View job →

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an Engineering Manager (Player & Coach), you will lead and mentor a team of Forward Deployed Engineers focused on building, scaling, and optimizing LLM inference workloads for Baseten customers. Applying both hands-on technical ownership and managerial leadership, you will guide your team through the processes of designing, deploying, and managing high performance, low latency AI applications on Baseten’s platform. FDE at Baseten is not a sales function – we are a mix of engineering, product, and customer architects who contribute to the core Baseten codebase, drive large portions of our feature roadmap, and execute on complicated customer engagements. You will also partner with product, infrastructure, and other customer engineering teams to ensure that large language models (LLMs) and other generative AI systems deliver best-in-class performance, reliability, and cost efficiency in production environments. EXAMPLE INITIATIVES Take a look at these blog posts written by members of our Forward Deployed Engineering team: Forward Deployed Engineering on the frontier of AI The fastest, most accurate Whisper transcription Deploy production-ready model servers from Docker images Deploy custom ComfyUI workflows as APIs RESPONSIBILITIES Leadership & Team Management Lead, mentor, and grow a team of Forward Deployed Engineers, providing guidance on technical direction, project execution, and professional deve

pythondockermachine learning
View job →
🔔

Get new inference technical lead jobs by email

Daily job updates · Unsubscribe anytime