At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Data Science is at the heart of Lyft’s products and decision-making. Data Scientists at Lyft operate in dynamic environments, moving quickly to build the world’s best transportation solutions.We tackle a wide range of challenges - from shaping long-term business strategy with data, to making critical short-term decisions, to developing algorithms and models that power both internal systems and customer-facing products. Lyft Business builds products that help organizations move the people who matter most - employees, customers, patients, and guests - easily and efficiently. Our offerings include Business Travel, Lyft Pass, and Concierge (for healthcare and non-healthcare rides), enabling companies to manage transportation at scale through APIs, integrations (e.g., Concur, Expensify), and dedicated tools. These platforms power high-impact B2B use cases across corporate travel, healthcare access, customer experience, and community programs. We are seeking a Senior Data Scientist to lead technical initiatives across the entire Lyft Business product suite. In this role, you will shape the technical vision, define algorithmic roadmaps, and drive execution for data science projects that accelerate growth, improve operational efficiency, and deliver measurable value to our enterprise partners. You’ll collaborate closely with Product, Engineering, Design, and Go-to-Market teams to build production ML models, experimentation frameworks, and advanced analytics that inform strategy and power product innovation. This is a high-visibility, high-impact role with direct influence on Lyft’s enterprise offerings. The ideal candidate will bring deep expertise in algorithm development, machine learning, causal inference, and experimentation, alongside strong business acumen in B2B contexts and a proven track record of t
Jobiba hiring network
Inference Technical Lead Jobs
1,491 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current inference technical lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Data Science is at the heart of Lyft’s products and decision-making. Data Scientists at Lyft operate in dynamic environments, moving quickly to build the world’s best transportation solutions.We tackle a wide range of challenges - from shaping long-term business strategy with data, to making critical short-term decisions, to developing algorithms and models that power both internal systems and customer-facing products. Lyft Business builds products that help organizations move the people who matter most - employees, customers, patients, and guests - easily and efficiently. Our offerings include Business Travel, Lyft Pass, and Concierge (for healthcare and non-healthcare rides), enabling companies to manage transportation at scale through APIs, integrations (e.g., Concur, Expensify), and dedicated tools. These platforms power high-impact B2B use cases across corporate travel, healthcare access, customer experience, and community programs. We are seeking a Senior Data Scientist to lead technical initiatives across the entire Lyft Business product suite. In this role, you will shape the technical vision, define algorithmic roadmaps, and drive execution for data science projects that accelerate growth, improve operational efficiency, and deliver measurable value to our enterprise partners. You’ll collaborate closely with Product, Engineering, Design, and Go-to-Market teams to build production ML models, experimentation frameworks, and advanced analytics that inform strategy and power product innovation. This is a high-visibility, high-impact role with direct influence on Lyft’s enterprise offerings. The ideal candidate will bring deep expertise in algorithm development, machine learning, causal inference, and experimentation, alongside strong business acumen in B2B contexts and a proven track record of t
We are seeking a mission-driven Developer Relations Manager focused on Foundational AI Research to engage leading academic labs advancing the next generation of AI models, systems, and methods. In this role, you will work directly with top researchers building frontier AI systems, including large language models, multimodal models, reasoning systems, training methods, inference systems, model serving, and scalable AI infrastructure. You will help researchers adopt NVIDIA’s AI and accelerated computing platforms to push the boundaries of model performance, efficiency, and scale. The ideal candidate brings deep technical credibility in foundational AI, strong research engagement experience, and hands-on expertise in either AI inference research or AI training research. What you'll be doing: Serve as a trusted technical advisor to leading academic AI labs working on foundation models, LLMs, multimodal AI, reasoning, training, inference, and AI systems. Identify high-impact research workloads where NVIDIA software, systems, and accelerated computing platforms can advance model performance, scale, and efficiency. Engage principal investigators, postdocs, graduate researchers, and lab leadership to understand research goals, technical blockers, infrastructure needs, and collaboration opportunities. Track frontier AI research across papers, benchmarks, open-source projects, and academic labs to identify emerging trends and future platform opportunities. Partner with Research Account Managers, Solution Architects, Product, Engineering, and Business Development teams to support researcher adoption and long-term engagement. Represent researcher needs internally by translating academic feedback into actionable insights for product roadmaps, developer programs, education, and platform strategy. Support NVIDIA participation in major AI, ML, and systems research venues through technical content,
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Staff Software Engineer on the Data Platform team within Platform , you'll define and lead the technical strategy for Coinbase's data infrastructure, spanning ingestion, transformation, warehousing, streaming, and serving systems. This is a foundational role at the intersection of distributed systems, data engineering, and AI-readiness, reporting to the Senior Director of Engineering. You'll set architectural direction, drive multi-quarter roadmaps, and transition the organization from managed-service dependency toward engineering-built, platform-grade infrastructure that powers everything from fraud detection to modern multi-agent AI architectures. What you'll do: Own the technical strategy and architecture for Data Platform, setting direction across data ingestion, transformation, warehousing, streaming, and serving systems while driving engineering-led cost reduction at the infrastructure layer. Architect data infrastructure to natively support AI and ML workloads, ensuring pipelines, data lake systems, and compute can power ML training, feature stores, real-time inference, and multi-agent AI architectures at scale. Drive the evolution to near-real-time data availability, enabling downstream teams across Coinbase to act on fresher data for fraud detection, financial reporting, and analytics. Build alignment and secure commitment from senior leadership
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role Modal's LLM inference platform delivers frontier performance for open-source models with best-in-class elasticity and developer experience, made in part possible by our custom runtime with GPU memory snapshots and multi-cloud substrate . We're looking for a leader to own the direction and execution of this platform to continue to establish us as the clear market leader, working closely with customers like Cognition, Doordash, Ramp, and many more. You'll be leading a group of highly talented engineers working on our market-leading LLM inference offering, spanning the serving stack, routing infrastructure, internal agentic optimization platform, and the user-facing product surface area. This is a hands-on leadership role — expect to split your time between technical contribution, product shaping and people management depending on what the team needs. You'll set direct
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: Most of the value of owning a model shows up at serving time. We're building a platform that covers the whole life of an LLM -- train it, deploy it, observe it -- and inference is where teams feel the difference every day. We already run elastic inference, sandboxes, distributed volumes, and multi-node training, and we control the infrastructure underneath, so the serving stack is ours to shape rather than something we resell. You will do hands-on inference research at Modal, working with the research lead to pick high-impact bets and owning them end to end. The bets that matter most are the ones that move cost per token and tail latency on the workloads our customers actually run. What you'll do: Own end-to-end inference research bets: speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, autoscaling for spik
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE Our Sales and Solutions teams navigate hard technical conversations spanning inference performance, GPU economics, latency budgets, deployment shape. As Baseten’s platform matures, we need a dedicated owner to translate launch velocity into field readiness. As our first Product Enablement Lead, you'll sit between Product, Marketing, and Sales GTM and own how Baseten's products, features, campaigns, and market moments like the launch of GLM-5.2 or Kimi K3 or the sudden evolution of Tokenomics as a discipline get translated into field execution. You will own how these launches land with the field, how AEs and SAs stay credible on a highly dynamic technical ecosystem, and how what the field hears from customers makes it back to Product. This is a hands-on individual contributor role. You are the bridge between product, marketing, and sales. You'll build the system and run it, which includes cross-functional program leadership, direct training and enablement of in-seat reps, and content and curriculum development for managers, sellers, and new hires. Success here will depend on your ability to build repeatable systems and rhythms and to partner across the business and with your enablement colleagues to ensure alignment and speed of execution. RESPONSIBILITIES Own launch readiness: partner with Product and Marketing on positioning, write internal launch comms, and run readiness sessions so AEs and SAs can sell new pr
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Our team is a fast-growing group of researchers and engineers focused on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop techniques that improve how models execute in production, driving lower latency, higher throughput, and consistent quality across diverse workloads. As an engineer on this team, you’ll work across the inference stack to improve core performance metrics by diving deep into model execution, identifying bottlenecks, and developing innovative optimizations. You’ll collaborate closely with modeling and systems teams to experiment, measure, and ship improvements that meaningfully accelerate inference. As the team evolves, you’ll have opportunities to build expertise in advanced performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model execution strategies for MoE and large-scale architectures. Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, e
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE This role owns Baseten's relationships and market intelligence across hyperscalers and strategic neoclouds, including NVIDIA cloud partners. This is a technical and commercial role in equal measure: you'll evaluate capacity from the GPU to the data center, negotiate cost and terms with suppliers, and stay close enough to the market to develop and defend a real point of view on where it's heading. Given current market conditions, Baseten needs a much stronger pulse on this part of the market so we can track pricing, stay close to the right relationships, and move fast the moment more capacity is needed. This is a senior, experienced hire who will also help pair with and develop 1-2 junior to mid-level teammates covering the same space. WHAT YOU'LL DO Build and maintain deep relationships across hyperscalers and strategic neoclouds (including NVIDIA cloud partners), working each organization from top to bottom rather than a single point of contact Maintain a consistent, "top of mind" presence with key accounts so Baseten is positioned to move quickly when capacity needs arise Evaluate capacity from the GPU to the data center — hardware generation, rack and node configuration, interconnect, power density, and cooling — so you know what a configuration will actually deliver, not just what the spec sheet claims Live in compute pricing daily: track rates by GPU generation, region, and contract term to keep Baseten inf
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for a Finance Systems Lead to own and evolve the systems that support Baseten's finance operations, including ERP, planning tools, and the broader finance and adjacent systems stack, spanning HRIS/payroll, CPQ, and commissions, as we implement, integrate, and automate across them. As Baseten scales, so does the number of systems finance depends on and touches, from the ERP to planning to compensation and quote to cash tooling. This role exists to bring a dedicated owner to that stack, someone who can lead implementations (partnering with contractors or vendor CSMs for hands-on technical configuration where needed), keep systems integrated and clean as we grow, and find automation opportunities that save the team time. This is a highly cross-functional role. You will partner with Accounting, FP&A, People, Sales and Revenue Operations, Legal, and IT to understand how each team uses its systems, and design integrations and workflows that serve the business without creating more manual work. RESPONSIBILITIES Core Finance Systems Own finance systems strategy, configuration, and roadmap, partnering with contractors or vendor CSMs on hands-on technical configuration as needed Lead the implementation and ongoing administration of our finance systems and tools Own finance system integrations end to end, ensuring clean data flow between ERP, planning, billing, and other systems, partnering with contrac
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE The Field Productivity & Enablement Lead is responsible for making our sales motion clear, practical, and repeatable. This person will help define how we sell at Baseten: how managers run the business, how reps qualify and advance deals, how sales works with FDE, product, marketing, and support, and how those expectations get turned into training and day-to-day habits. This is a senior individual contributor role for someone who excels at both strategy and execution. You'll shape the system, but you won't stop at slides or frameworks. You'll build the playbooks, run the training, coach to the behaviors, and help managers ensure the process is followed consistently. We're not looking for someone who wants to force-fit a single methodology onto the business. We're looking for someone who is fluent in MEDDIC/MEDDPICC, Command of the Message, Challenger, and similar approaches, and can take the best ideas from each, adapt them to a technical sales motion, and build an approach that fits how Baseten actually sells. You'll report to the Head of Enablement and Productivity. Frontline managers and reps are your primary customers. RESPONSIBILITIES Manage operating model. Define the core rhythms for frontline managers, including 1:1s, forecast calls, pipeline reviews, deal reviews, and coaching cadences. Create clear expectations for how managers inspect deals, coach reps, and drive consistency across the team. Sale
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE At Baseten, we’re looking for a Technical Program Manager to drive our most complex, cross-cutting infrastructure programs. This role will operate across all domains of AI infrastructure, from the GPUs up to the multi-cluster orchestration layer. This is an execution-first role. The work is less about owning a single system and more about imposing order on ambiguity: standing up the right structures, driving decisions to closure, and making sure nothing falls through the cracks across dozens of stakeholders. If you take satisfaction in turning a chaotic, half-defined initiative into a predictable, well-governed program, this role is for you. RESPONSIBILITIES Own complex migrations end to end. Lead large-scale infrastructure migrations across teams and domains. This will involve scoping the work, sequencing dependencies, managing risk, and driving them to completion without surprises. Drive process across infrastructure. Establish and run the operating rhythms that keep programs healthy: planning cadences, status reporting, decision logs, risk reviews, and escalation paths. Make the process light enough that teams adopt it and rigorous enough that it actually works. Help managers build the right structures. Partner with engineering managers and leads to design the team structures, ownership boundaries, and working models a program needs to succeed. Spot gaps in accountability before they become problems. Own fo
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc
Who we are Graviton Research Capital is a privately funded quantitative trading firm. We trade across a multitude of asset classes and trading venues using a diverse range of concepts, from time series analysis and stochastic models to machine learning and statistical inference. We analyse terabytes of data to identify pricing anomalies and drive innovation in financial markets. Role Overview We are looking for a Program Manager who thrives at the intersection of rigorous engineering and predictable delivery. You will not just "manage tasks" — you will orchestrate the development lifecycle for mission-critical systems. Your goal is to ensure that our elite engineering teams can focus on high-performance code while you own the execution strategy, dependency mapping, and release discipline. Key Responsibilities Lead Agile ceremonies (Sprint Planning, Stand-ups, Retrospectives) tailored for deep-tech engineering teams. Transform high-level trading requirements into granular, executable backlogs. Own capacity planning and burn-down metrics to provide high-visibility delivery timelines. Navigate the complex interplay between engineering teams (e.g., Connectivity, Core Infrastructure, Simulation) to prevent bottlenecks. Build and maintain advanced Jira dashboards, automated roadmaps, and Confluence documentation that serve as the "single source of truth" for stakeholders. Proactively identify technical debt, architectural blockers, or resource gaps that threaten release stability. Continuously refine Agile methodologies to suit low-latency, performance-sensitive development cycles (where "Definition of Done" includes rigorous performance benchmarking). Eligibility & Required Skills 5+ years of experience as a TPM, Program Manager, or Scrum Lead in a product-engineering environment (HFT, FinTech, Networking, or Kernels/Systems). A strong grasp of the software development lifecycle for high-performance systems. While you won't write code, you must understand concepts li
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity As a Member of Technical Staff on AI Infrastructure, you will build and maintain the foundational systems and distributed infrastructure that power AI model post training, inference, and data pipelines. You will collaborate with engineering and research teams to ensure performance, scalability, and reliability of critical AI systems. What You’ll Do Design and implement large-scale, distributed AI infrastructure and services Optimize performance for GPU/xPU accelerators and cloud environments Build tools for observability, reliability, and scaling of AI workloads Partner with cross-functional teams to define AI infrastructure requirements and roadmap Contribute to architectural design and system longevity About You Have experience with GenAI infrastructure systems, distributed systems, cloud computing, and high-performance infrastructure Are proficient in programming languages like Python, Go, or similar Understand scaling challenges specific to AI workloads and accelerators Thrive in fast-paced, collaborative engineering environments The reasonably estimated base salary for this role ranges from $256,000.00 to $276,
Get new inference technical lead jobs by email
Daily job updates · Unsubscribe anytime