Jobiba hiring network

Inference Technical Lead Jobs

1,448 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current inference technical lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

N
Notion
📍 San Francisco• Full-time
1mo ago

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role: You will apply your expertise in statistical inference and experimentation to understand the impact of product changes and drive decision-making. You will explore ambiguous areas and conduct analysis to identify opportunities for the Growth org, helping shape the strategy and roadmap. You'll partner closely with product, engineering, and marketing leaders at every step of the development process, driving forward data science and analytics work. You'll work with your teams to determine north star and operational metrics, and build the foundational datasets and dashboards to monitor progress and make key decisions. You'll influence the broader company by communicating your findings and driving change in our product and business. (Insights are useful. Impact is even better!) Skills You'll Need to Bring: You are comfortable with ambiguity and motivated by business results. You have expertise in SQL and at least one scripting language (ideally Python or R). You know how to use statistical inference and experimentation to drive actionable recommendations. You are comfortable transforming raw data to build your own data sets

pythonsqlai
View job →
L
Lyft
📍 Toronto• Full-time• From C$108K/yr
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Data Science is at the heart of Lyft’s products and decision-making. Data Scientists at Lyft operate in dynamic environments, moving quickly to build the world’s best transportation solutions. We tackle a wide range of challenges - from shaping long-term business strategy with data, to making critical short-term decisions, to developing algorithms and models that power both internal systems and customer-facing products. Driver Incentives Science owns the algorithms and systems behind incentive design, influencing driver engagement and marketplace efficiency — from real-time supply positioning to longer-horizon earnings and engagement programs. The team is responsible for designing pay and incentive mechanisms that are efficient and good for driver experience over the long run. As a Data Scientist specializing in Algorithms, you'll partner closely with product, engineering, and operations leaders to build and scale incentive systems, shape long-term mechanism design strategy, and deliver on critical business goals tied to marketplace efficiency and driver earnings. Candidates with strong optimization backgrounds — think mathematical programming, control theory, or operations research — are a great fit, though we welcome strong candidates from machine learning or causal inference as well. The ideal candidate thrives in a fast-paced environment and brings a hands-on, entrepreneurial mindset to drive results. Responsibilities: Collaborate with engineering and product teams to design, implement, and iterate on new features and algorithmic improvements for driver incentives and pay mechanisms. Design, develop, and deploy optimization models, algorithms, and systems for problems such as budget allocation, multidimensional cost-curve development, and incentive targeting. Write production model code; collabor

pythonmachine learningai
View job →
S
Synthesia
📍 United Kingdom• Full-time• Remote
1mo ago

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. About the role As an Applied Research Engineer in our Video team, you will help build the next generation of production-grade foundation models for human-centric video generation. You will join a highly focused team working at the intersection of large-scale generative modeling, distributed systems, and production engineering. Our mission is to develop and optimize video base models that power realistic, controllable, and emotionally expressive synthetic humans at scale. This is not pure research. This is applied research with direct product impact. You will work on advancing training recipes, scaling distributed systems, improving evaluation frameworks, and optimizing inference to ensure our models are high quality, stable, and efficient enough for real-world deployment. Your work will directly influence models used by tens of thousands of businesses worldwide. What you’ll do You will own and execute end-to-end research and engineering projects, from hypothesis to production impact. This includes: Developing and scaling latent video diffusion models tailored for human-centric video generation Designing conditioning mechanisms to improve control (pose, emotion, script, camera) without sacrificing fidelity Advanc

REMOTEpythonawsdocker
View job →

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect -- for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author

pythonartificial intelligenceai
View job →

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each brings together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We are searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will be instrumental in ensuring our data center solutions deliver industry-leading performance for accelerated computing workloads. What you will be doing: Design and execute comprehensive performance benchmarking strategies for our data center platforms and products Characterize real-world AI training, inference, and HPC workloads at scale Define, track, and report key performance indicators (throughput, latency, efficiency, scaling) Build automation tools and frameworks for performance monitoring and analysis Identify and analyze performance bottlenecks across compute, memory, network and storage subsystems Work closely with architecture, hardware,

pythondockerkubernetes
View job →
A
Anyscale
📍 Remote• Full-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About Ray Data Team: Ray Data is Python-native data processing engine that is a one stop shop for all AI data processing needs. Ray Data provides performant, first-class integration with cutting edge AI frameworks using both multi-modal and structured data. The Ray Data team currently develops and maintains Ray Data . We are a team of engineers passionate about building a Data processing engine which is a one-stop shop for all of your ML/AI needs. We are looking for exceptional engineers to build, optimize, and scale Ray for modern and increasingly complex AI workloads. As part of this role, you will: Improve the performance of Ray Data and multi-modal batch inference use cases. Ensure efficient scaling across different stages of the Data pipeline in a heterogeneous environment. Building data loading solutions for production training workloads. Focus on stability and fault tolerance at high scale Working with customers and new age AI native companies in scaling their AI workloads. We'd love to hear from you if have: At least 3-4 years of relevant work experience Solid background in building scalable and fault-tolerant distributed systems Experience with data processing, database internals. Passionate about large

pythonmachine learningai
View job →
D
Datalab
📍 New York• Full-time• $250K – $350K/yr
1mo ago

Salary range - $250k - $350k | Equity - up to 0.5% | In-person NYC About Datalab Datalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right. We’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, Chandra, Surya, Marker, and Lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face. Role Overview We're looking for a Research Engineer to own problems end to end across our models, inference service, and product. You won't just train a model and hand it off. You'll take it from training through benchmarking, into our inference stack, and work with the team to integrate it into our products. We're a small team that has shipped the current state of the art OCR model, Chandra. Our models collectively have 70k+ Github stars. Our tools are used internally at frontier AI labs like Anthropic, and Fortune 500 enterprises like Siemens. Our team focuses on training small, efficient models that outperform much larger LLMs on domain-specific tasks (like OCR, structured extraction, tables). We move fast, prioritize practical results, and build tools that are open, reproducible, and built to last. You'll test hypotheses quickly, iterate on results, and balance experimental rigor with shipping to customers. Day to day: A typical project might look like: identify a gap in extraction quality on long documents, train and benchmark a new model, optimize it for inference, and work with the team to ship it to users. Concretely: Train and evaluate models: Train task-

pythongitai
View job →
D
Datalab
📍 New York• Full-time• $225K – $300K/yr
1mo ago

Salary range: $225k - $300k | Equity: 0.15% - 0.35% | In-Person: NYC About Datalab Datalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right. We’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, Chandra, Surya, Marker, and Lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face. Role Overview We’re looking for a fullstack engineer who wants to build the interfaces, tools, and infrastructure that help developers and enterprises use our models. You’ll work across the stack to shape how people interact with OCR, extraction, and document-understanding systems. That includes building core inference workflows, creating intuitive UI for complex parsing tasks, and improving the developer experience across our open-source repos and API. This is a high-ownership role that blends engineering, product thinking, and community engagement. You will work closely with the founders and the rest of the team to ship features, improve performance, and make our technology accessible to a global community of builders. As a small and fast-moving team, roles are fluid. You should enjoy working across backend, frontend, performance, and user-facing surfaces. Your work will directly influence how teams evaluate and deploy our models. Day to day, you will: Ship features to our open source repos, API, and internal tooling. Design and build frontend features that make document parsing more interactive and understandable. Optimize inference performance and improve th

pythongitrest
View job →
P
Plaid
📍 San Francisco• Full-time
1mo ago

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid Protect is a real-time fraud intelligence product built on a unique advantage: Plaid’s network-level visibility across bank accounts, devices, identities, sessions, institutions, applications, and financial behavior. Protect helps customers detect first-party fraud, synthetic identities, account takeovers, and coordinated attacks that are difficult to see from a single application, account, or transaction. Trust Index turns that fraud intelligence into real-time fraud scores and actionable attributes. This team builds the systems that make this intelligence possible: low-latency inference, new data and model integrations, customer-facing APIs and attributes, safe rollouts, and feedback loops. Ti3 expanded Plaid’s fraud graph nearly 10x and, in early testing, detected up to 41% more fraud at the same false-positive rate. Learn more about Ti2 and Ti3 . We are a small, high-agency team working closely with Product, Data Science, and Machine Learning. We value demos over docs, conviction over consensus/alignment, builder schedule over meeting-heavy calendars. We’re scrappy and a talent-dense team that has high agency and high ownership. As a Staff Software Engineer on the Protect Core team, you wi

awsrestmachine learning
View job →
A
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About Ray Data Team: Ray Data is Python-native data processing engine that is a one stop shop for all AI data processing needs. Ray Data provides performant, first-class integration with cutting edge AI frameworks using both multi-modal and structured data. The Ray Data team currently develops and maintains Ray Data . We are a team of engineers passionate about building a Data processing engine which is a one-stop shop for all of your ML/AI needs. We are looking for exceptional engineers to build, optimize, and scale Ray for modern and increasingly complex AI workloads. As part of this role, you will: Improve the performance of Ray Data and multi-modal batch inference use cases. Ensure efficient scaling across different stages of the Data pipeline in a heterogeneous environment. Building data loading solutions for production training workloads. Focus on stability and fault tolerance at high scale Working with customers and new age AI native companies in scaling their AI workloads. We'd love to hear from you if have: At least 3-4 years of relevant work experience Solid background in building scalable and fault-tolerant distributed systems Experience with data processing, database internals. Passionate about large

pythonmachine learningai
View job →

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. About the role As a Research Engineer in our Video team, you will help build the next generation of production-grade foundation models for human-centric video generation. You will join a highly focused team working at the intersection of large-scale generative modeling, distributed systems, and production engineering. Our mission is to develop and optimize video base models that power realistic, controllable, and emotionally expressive synthetic humans at scale. This is not pure research. This is applied research with direct product impact. You will work on advancing training recipes, scaling distributed systems, improving evaluation frameworks, and optimizing inference to ensure our models are high quality, stable, and efficient enough for real-world deployment. Your work will directly influence models used by tens of thousands of businesses worldwide. What you’ll do You will own and execute end-to-end research and engineering projects, from hypothesis to production impact. This includes: Developing and scaling latent video diffusion models tailored for human-centric video generation Designing conditioning mechanisms to improve control (pose, emotion, script, camera) without sacrificing fidelity Advancing distr

pythonawsdocker
View job →
S
Supabase
📍 Remote• Full-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. Edge Functions are server-side TypeScript functions, distributed globally at the edge - close to your users. They power use cases such as webhook receivers, AI inferences, OG image generation and real-time bots (Slack, Discord, etc). Built on top of Supabase Edge Runtime : an open-source, Deno-based runtime written in Rust that runs JavaScript, TypeScript, and WASM services. We want developers to be able to build truly global applications by distributing both compute and data globally. Infrastructure concerns like regions, cold starts, CPU and memory provisioning should fade into the background so developers can focus on iterating on business logic. We are looking for experienced and passionate engineers to help us go further in this vision. What You’ll Own Evolving Supabase Edge Runtime - an Open-sourced Rust-based host that runs the Deno isolate, manages the main/user runtime split, and enforces per-request memory and CPU limits. Implementing monitoring, alerting, and OpenTelemetry tracing across the runtime, then using that visibility to drive optimizations that improve latency and reliability of the service. Working closely with the Deno and other open-source teams, contributing to upstream and relaying our users' requirements. Participating in an on-call rotation to keep Edge Functions healthy in production. Help manage and improve features like scheduled functions, background tasks, WebSockets streaming, ephemeral file storage, and custom routing. Integrating functions more tightly with the rest of the Supabase stack - Auth, Postgres, Storage, and Realtime. Expanding functions to support more use cases (AI inference, MCP servers, hosting simple websites, URL shorteners). Improving the DX

javascripttypescriptjava
View job →
S
Supabase
📍 Remote• Full-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We're looking for a SDK Engineer - JavaScript to join the SDK team and own a core part of how developers talk to Supabase from JavaScript and TypeScript. Our JS/TS SDKs are the front door to the platform: database queries, auth, storage, realtime subscriptions and edge functions, used by a very large number of developers and by most of the AI coding tools building on Supabase today. This is a library-authoring role, and the work happens in the open: public repos, public issue trackers, and API decisions whose consequences are permanent. The type system here is a design surface rather than a formality. If you get satisfaction from an inference that Just Works, from a migration guide that saves people an afternoon, and from an issue tracker that isn't a graveyard, you'll like it here. This role is ideal for someone who thrives in async, fast-paced environments and is excited about building developer tools that scale to millions. What You'll Be Responsible For Build and evolve our JavaScript/TypeScript SDKs. Start new things. Alongside the maintenance work, you're expected to find what's missing, make the case for it, and build it. Some of what we ship next doesn't exist yet, and nobody will hand you the list. Work in the open. These are public repos: you'll triage inbound issues, review and shepherd outside contributions, keep CI and release automation healthy, and hold a clear, kind line on scope. Own the type experience. Keep generated database types flowing correctly through the client APIs so autocomplete and inference stay correct in real codebases. Keep us honest across runtimes. Node, Deno, Bun, browsers, Cloudflare Workers, React Native/Expo, including dual ESM/CJS publishi

javascripttypescriptjava
View job →
C
Cohere
📍 San Francisco• Full-time
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We're building the foundational infrastructure that will define how the world thinks about and deploys AI, and we want the sharpest, most curious people to help us do it. As a Lead Data Scientist on our Analytics and Data Insights team, you'll tackle problems that don't have textbook answers yet; shaping go-to-market strategy for technology that's still being invented, designing the experiments that prove or kill our biggest bets, and helping enterprises understand what foundational AI actually means for their bottom line. You'll own the full analytical lifecycle, from framing the right questions and building the models, to leading a team that delivers answers leadership can act on. As a Lead Data Scientist, you will: Drive the mission forward. Own the science: design and lead experimentation programs including A/B tests, multi-armed bandits, causal inference studies, that directly map to product and go-to-market decisions. Build predictive models that matter: develop and deploy models for forecasting, segmentation, propensity scoring, and opportunity sizing across Cohere's core business lines. Lead and grow a tea

pythonsqlgit
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Security Clearance: Active Secret+ clearance strongly preferred; candidates eligible and willing to obtain clearance will also be considered. More information about Canadian Security Clearance is available here . As an Infrastructure Security Engineer, your key responsibilities include: Deploy, and manage infrastructure for Protected B classified environments, ensuring compliance with ITSG-33 and Canadian government standards Design and implement security controls for cloud (AWS, GCP, Azure) and hybrid/multi-cloud deployments Evaluate, implement, and manage security tools and technologies for training cluster and inference infrastructure hardening Implement security best practices including IAM, encryption, logging, and monitoring Participate in security incident response activities, including detection, analysis, containment, and remediation Conduct regular vulnerability assessments and penetration testing of infrastructure components Maintain comprehensive security documentation, procedures, and configurations for classified environments Maintain active Secret+ security clearance and adhere to all Canadian government security

awsazuregcp
View job →
🔔

Get new inference technical lead jobs by email

Daily job updates · Unsubscribe anytime