Jobiba hiring network

Inference Technical Lead Jobs

1,448 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current inference technical lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for a seasoned Frontend Engineer to craft performant and delightful user experiences across Baseten’s core platform. You’ll own critical parts of our web application stack and collaborate cross-functionally with product, design, and backend teams to launch impactful features that help users deploy and manage AI systems at scale. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Core Product team: Rolling Deployments Model APIs for frontier models Model training built for production inference RESPONSIBILITIES Design, implement, and maintain responsive, accessible, and user-friendly frontend interfaces using React and TypeScript Collaborate closely with product designers to turn complex ideas into elegant, intuitive UIs Optimize application performance and reliability, with a focus on rendering speed and responsiveness Drive major frontend initiatives, including partnering with backend teams to define APIs and test and refine end-to-end flows Establish best practices, and mentor other engineers on frontend technologies Build reusable component libraries and frontend infrastructure that accelerate product development Partner with backend and platform teams to define and refine APIs and end-to-end flows REQUIREMENTS 5+ years of experience building production-grade web applications Deep expertise in React, TypeScript, and modern web development tooling Track record of bu

typescriptreactmachine learning
View job →
B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Voice is becoming the internet’s next interface, but a production-grade Voice AI system is "hard to build" . You’ll join a small founding team of Baseten Voice AI, focused on bringing state-of-the-art open source models into production for Voice AI customers across productivity, customer service, clinical conversation, creator tools, education, and more. You’ll make a meaningful impact on people’s daily lives and help reshape these industries. This is a high-impact, high-ownership role. You will be the primary owner of Baseten Voice AI - our in-house inference stack to power Voice AI models - from product roadmap through engineering implementation. You’ll partner closely with Forward Deployed Engineers, Model Performance Engineers, and sister engineering teams to push the boundaries of Voice AI. EXAMPLE INITIATIVES: Develop world-class model serving stack for state-of-the-art open-source voice models - reduce end-to-end and tail latency (p95/p99), increase throughput, and improve GPU efficiency via profiling, runtime tuning, and server-level optimizations. Build large-scale, real-time infrastructure for multi-model voice agents - orchestrate STT, TTS, and agent components with streaming I/O to meet customer SLOs. Design tight training and inference iteration loops for voice model customization - enable fast evaluation, safe rollout, and rapid experimentation for custom voice model development. Past projects:

pythondockerkubernetes
View job →
B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten’s Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We operate at the intersection of distributed systems, model performance, infrastructure, and developer experience. We enable customers to deploy and operate cutting-edge LLM models with industry-leading performance, scalability, reliability, and ease of use. As a Software Engineer on the Inference Stack team, you’ll work across the stack - from the developer experience customers use to deploy models, the libraries used for features like tool calling and reasoning, all the way down to the systems we use to orchestrate deployments in Kubernetes and route traffic efficiently. This is an ideal role for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users. EXAMPLE INITIATIVES Blog Posts https://www.baseten.co/blog/nvidia-dynamo-day-baseten-inference-stack/ https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/ https://www.baseten.co/blog/how-baseten-multi-cloud-capacity-management-mcm-powers-cloud-self-hosted-and-hybr/#comparing-deployment-options-cloud-vs-self-hosted-vs-hybrid RESPONSIBILITIES Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference Work across the stack, from customer-facing features to low-le

kubernetesci/cdrest
View job →
A
1mo ago

About Anyscale At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure. As part of this role, you will Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices We'd love to hear from you if you have Familiarity with running ML inference at large scale with high throughput and low latency Familiarity with deep learning and deep learning frameworks (e.g. PyTorch) Solid understanding of distributed systems, ML inference challenges Bonus points

machine learningai
View job →
L
Lyft
📍 Toronto• Full-time• From C$1.3M/yr
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Safety and Customer Care (SCC) team at Lyft manages over 1.7 million monthly human and AI interactions and serves as Lyft's primary direct touchpoint with riders and drivers. We handle critical infrastructure that powers both human associates and AI agents to make riders and drivers feel safe and comfortable while riding or driving with Lyft, transforming every support interaction into a moment of genuine connection. As a Data Scientist working on Causal Inference in SCC, you'll partner with a strong team of engineers, product managers, designers, and operations leaders to deliver a personalized and exceptional experience for Lyft customers, using rigorous causal inference to guide the highest-stakes decisions we make. We're looking for a motivated and talented Data Scientist with deep causal inference expertise to join the SCC Data Science team. You'll partner closely with the area's tech lead on high-impact work spanning AI-powered support products, differentiated service, and operations optimization. The ideal candidate brings sharp applied inference intuition, a bias toward impact, and the ability to cut through ambiguity in complex problem spaces. You'll work on projects like: Design rigorous experiments and quasi-experiments to measure the causal impact of SCC product and AI-agent launches, and drive data-informed launch decisions. Build causal ML models to optimize concession budget allocation, targeting the right support credit, to the right rider or driver, at the right moment to maximize trust and business impact. Quantify the long-term effects of support-experience changes on rider and driver retention, and uncover heterogeneous treatment effects across our community. Deliver strategic insights on quality–cost tradeoffs, empowering leadership to balance service quality, coverage,

pythonsqlai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role We are looking for an engineer who wants to take the world's largest and most capable AI models and optimize them for use in a high-volume, low-latency, and high-availability production and research environment. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Work alongside researchers to enable advanced research through awesome engineering. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 5 years of professional software engineering experience. Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc. Have experience architecting, building, observing, and debugging production distributed systems. Bonus point if worked on performance-critical distributed systems. Have need

awsazurerest
View job →
O
1mo ago

About the Team OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, and scalable in production, and we partner closely with Research to bring the next generation of models into the world. We're a small, fast-moving team of engineers focused on delivering a world-class developer experience while pushing the boundaries of what AI can do. We’re expanding into multimodal inference, building the infrastructure needed to serve models that handle image, audio, and other non-text modalities. These workloads are inherently more heterogeneous and experimental, involving diverse model sizes and interactions, more complex input/output formats, and tighter coordination with product and research. About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal models at scale. You’ll be part of a small team responsible for building reliable, high-performance infrastructure for serving real-time audio, image, and other MM workloads in production. This work is inherently cross-functional: you’ll collaborate directly with researchers training these models and with product teams defining new modalities of interaction. You'll build and optimize the systems that let users generate speech, understand images, and interact with models in ways far beyond text. In this role, you will: Design and implement inference infrastructure for large-scale multimodal models. Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs. Enable experimental research workflows to transition into reliable production services. Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities. Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers. You might thrive in t

awsrestai
View job →
PE
1mo ago

Client Onboarding Manager – Inference & Agentic AI | Paytm (Noida) About the Role Paytm is a pioneer of digital payments in India, serving over 450 million consumers and 45 million merchants across payments, financial services, and commerce. Over the years, Paytm has built deep in-house capabilities across technology, data, and operations to operate at scale with high reliability. Paytm is building a full stack AI platform focussed on Inference and Agents, enabling large enterprises to deploy AI driven automation across sales, service, operations, and analytics. The Inference and Agentic AI team operates as a cross functional unit spanning engineering, product, data science, business management, and sales, and owns the full lifecycle of AI solutions from opportunity discovery to deployment and scale. Key Responsibilities Own client onboarding from sales handover to go-live. Understand client workflows, systems, and integration requirements. Coordinate with Product, Engineering, and Client teams for seamless deployment. Manage onboarding timelines, milestones, and stakeholder communication. Conduct client training sessions and drive product adoption. Track onboarding KPIs, client satisfaction, and implementation success. Gather client feedback and support continuous product improvements. Ideal Candidate 2–5 years of experience in Client Onboarding, Implementation, Customer Success, or Solutions Engineering. Experience in SaaS, Fintech, Enterprise Technology, or AI products preferred. Good understanding of APIs, integrations, CRM systems, and enterprise workflows. Strong project management, problem-solving, and stakeholder management skills. Excellent communication and client-facing abilities. Bachelor’s degree in Engineering, Business, or related field. Location: Noida Why Join? Be part of Paytm’s fast-growing AI business and work closely with enterprise clients to deliver cutting-edge AI-driven automation solutions at scale.

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for talented systems developers and researchers to join the Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization . Our mission is to build the next generation of high-performance and intelligent inference systems . We optimize not only how fast and efficiently models run, but also how quickly inference systems can adapt to new models, architectures, hardware, and workloads. Our work spans the full inference stack—from distributed serving and runtime systems to GPU kernels and model-system co-design. We explore techniques such as adaptive parallelism, speculative and parallel decoding, disaggregated inference, scheduling and batching, KV-cache optimization, model swapping, quantization, and GPU kernel optimization to push the frontier of latency, throughput, scalability, and cost. Beyond optimizing individual models, we are building intelligent and adaptive inference systems that can automate performance optimization—rapidly profiling new models and workloads, identifying bottlenecks, selecting effective execution strategies, and adapting system configurations with minimal manual tuning. We embrace AI-native engineering , using AI not only as the workload we optimize, but also as a tool to accelerate system deve

machine learningaiswift
View job →

About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, then turn that understanding into performance optimizations and models that project performance and capacity needs for future launches. About the Role In this role, you will model inference performance across application, model, and fleet layers with higher fidelity. You will build cost-to-serve estimates from microbenchmarks and create tools that help cross-functional teams reason about latency, capacity, utilization, and cost tradeoffs. In this role, you will Build and refine performance models that translate microbenchmark results into cost-to-serve estimates. Analyze inference workloads end to end across applications, models, and fleet infrastructure. Enhance tooling to identify bottlenecks across layers for latency and throughput. Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference. You might thrive in this role if you: Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency. Are comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling. Have deep expertise with performance profiling, benchmarking, analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve o

awsrestai
View job →
N
Nuro
📍 Mountain View• Full-time• From $160.4K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role The ML Infrastructure team is responsible for building & improving the core infrastructure for autonomy teams at Nuro. In this role, you will work closely with teams across Nuro, to design, build and deploy core infrastructure components in machine learning model life cycle, to push the autonomous future forward. You will have an opportunity to work across the full stack of machine learning solutions - from designing robust and scalable model & data pipelines to building to deploying the optimized models on Nuro’s fleet of self-driving robots! About the Work Design and develop ML workflow pipelines to train, optimize, validate, and deploy Nuro autonomy models. Develop and maintain continuous testing and monitoring systems for core ML infrastructure components. Develop observability to track ML model lifecycles from data generation to on-road validation. Maintain an in-house ML inference platform to serv

pythonmachine learningai
View job →
C
Cloudflare
📍 Hybrid• Full-time• Hybrid
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations New York, US San Francisco, US About the Role Cloudflare is building one of the most differentiated AI inference platforms in the market. We aren’t building another commoditized GPU-by-the-hour provider, but a globally distributed, serverless inference platform integrated into the world's largest developer ecosystem. We need someone to own the go-to-market strategy for this business. As Head of GTM, AI Inference, you will be th

🔔

Get new inference technical lead jobs by email

Daily job updates · Unsubscribe anytime