A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We are a software engineering team with expertise in enabling ML models in production. We deploy AI models to run in variety of environments: air-gapped government networks, forward-deployed defense environments, edge nodes, and enterprises with strict data sovereignty requirements. Our customers rely on us for frontier AI capabilities running on hardware they control, often with constrained GPU resources and limited direct access. Rising to that challenge and meeting those expectations is what Palantir's excels at. We treat models like any other software: continuously tested, continually delivered, packaged for reproducible deployment, and built for long-term maintainability. You will own services end-to-end, and work across the full stack, from inference engines, GPU scheduling to deployment pipelines, observability, and integration with Palantir's platform. The goal is to deliver new models and capabilities quickly and continuously. Join us if you want to solve problems at the intersection of infrastructure and machine learning that directly enable critical customers.
Jobiba hiring network
Inference Technical Lead Jobs
1,448 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current inference technical lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We are a software engineering team with expertise in enabling ML models in production. We deploy AI models to run in variety of environments: air-gapped government networks, forward-deployed defense environments, edge nodes, and enterprises with strict data sovereignty requirements. Our customers rely on us for frontier AI capabilities running on hardware they control, often with constrained GPU resources and limited direct access. Rising to that challenge and meeting those expectations is what Palantir's excels at. We treat models like any other software: continuously tested, continually delivered, packaged for reproducible deployment, and built for long-term maintainability. You will own services end-to-end, and work across the full stack, from inference engines, GPU scheduling to deployment pipelines, observability, and integration with Palantir's platform. The goal is to deliver new models and capabilities quickly and continuously. Join us if you want to solve problems at the intersection of infrastructure and machine learning that directly enable critical customers.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We are a software engineering team with expertise in enabling ML models in production. We deploy AI models to run in variety of environments: air-gapped government networks, forward-deployed defense environments, edge nodes, and enterprises with strict data sovereignty requirements. Our customers rely on us for frontier AI capabilities running on hardware they control, often with constrained GPU resources and limited direct access. Rising to that challenge and meeting those expectations is what Palantir's excels at. We treat models like any other software: continuously tested, continually delivered, packaged for reproducible deployment, and built for long-term maintainability. You will own services end-to-end, and work across the full stack, from inference engines, GPU scheduling to deployment pipelines, observability, and integration with Palantir's platform. The goal is to deliver new models and capabilities quickly and continuously. Join us if you want to solve problems at the intersection of infrastructure and machine learning that directly enable critical customers.
About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About the Team The Proactive Insights team is a newly formed team at the center of Mixpanel's AI-first analytics vision. With a greenfield charter, we're building the intelligent layer that transforms Mixpanel from a tool you query into a partner that works for you. We answer the question every data-driven team asks: "What changed, why, and what should I do about it?" We proactively keep users informed about what matters in their data, delivering the right insights and recommendations at the right time, to the right places, both inside and outside of Mixpanel. Some examples of what we are building: Signals : Statistical analysis that automatically identifies which user behaviors cause downstream business outcomes — such as which actions genuinely improve 30-day retention — using causal inference to move beyond correlation Forecasting : Time-series modeling that projects whether a KPI (e.g. Signups) will hit its goal by end of quarter — including trend decomposition, seasonality adjustment, and confidence bands against a target. Simulation : Causal impact modeling that estimates how moving one metric (e.g. weekly sharing rate) by a given amount will ripple through to downstream KPIs like retention or revenue — giving teams a quantified basis for prioritization Cohort Detection : Automated identification of at-risk user cohorts by finding active users who resemble known churned segments across both behavioral patterns and descriptive characteristics, before they churn. Offline batch survival analysis models that estimate each user's probability of a future outcome (e.g. likelihood to churn or convert withi
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role In this role, you will be a member of the Perception & Behavior team, leveraging the cutting edge of machine learning research to solve challenging real-world robotics problems. This role is focused on bringing advancements in the field of ML and large-scale learning to the AV domain, moving towards a more end-to-end autonomous driving system. This role requires working with and developing large models for perception and behavior, keeping up-to-date and experimenting with state-of-the-art architectures quickly and efficiently, collaborating with other teams to determine data and infrastructure support needs, and working to improve model optimization and inference speeds. You will use your applied research skills to think through the creation and deployment of these ML models on our autonomous vehicles while working alongside talented researchers in the field. If you love solving fundamental AI challenges and deploying your s
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role The Autonomy ML Infrastructure team is responsible for building & improving the core infrastructure for autonomy teams at Nuro. In this role, you will work closely with teams across Nuro, to design, build and deploy core infrastructure components in machine learning model life cycle, to push the autonomous future forward. You will have an opportunity to work across the full stack of machine learning solutions - from designing robust and scalable model pipelines to building to deploying the optimized models on Nuro’s fleet of self-driving robots! About the Work Optimize Nuro’s autonomy stack with cutting-edge optimization techniques like quantization, low precision inference, and model pruning. Work with autonomy engineers to optimize, validate, and deploy large language models. Develop and maintain a world-class model compiler framework, FTL . Write robust, high-quality software to increase our confidence in our vehicl
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Job Description/ Responsibilities: Designing, developing and maintaining stable and reliable AI/ML Ops platforms / pipelines Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production. Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow. Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data Technology Eva
Sendbird is building AI agents for customer experience. Our platform already powers billions of conversations every month across chat, voice, video, and messaging APIs. We are now using that foundation to build agents that understand customer context, reason over business data, and take reliable action in production. We are looking for a Machine Learning Engineer to research, build, and productionize new capabilities for those agents. This role sits at the intersection of agent product development, applied AI research, and production engineering. You will work on systems that enterprise customers depend on every day, not demos or isolated prototypes. About Sendbird and delight.ai Sendbird has spent more than a decade building communication infrastructure for in-app chat, voice, video, and messaging APIs. More than 4,000 brands use our platform, including DoorDash, Match Group, Noom, Yahoo Sports, and Rakuten. Our systems support more than 7 billion messages every month. In 2024, we made a strategic shift toward AI-first customer experience. In 2025, we launched our enterprise AI agent product, delight.ai. Delight.ai helps businesses deliver customer support and engagement that is faster, more contextual, and more personal. Unlike simple FAQ bots, our agents are built to remember customer context, use tools, retrieve relevant knowledge, connect across channels, and handle real customer workflows with accuracy and control. The Role As a Machine Learning Engineer, you will design, build, evaluate, and ship new capabilities for our AI agents. You will work across agent architecture, retrieval, memory, planning, tool use, workflow automation, voice, evaluation, data pipelines, model adaptation, inference, and production integration. This is a hands-on engineering role for someone who can turn AI research and product ideas into reliable customer-facing features. Some problems will require training, fine-tuning, or adapting models. Others will require better retrieval, bet
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet is hiring a Senior Machine Learning Operations Engineer to architect our machine learning production lifecycle. Your mission is to maintain and deploy ML models to a scalable, reliable, and secure production environment. You will design and maintain the infrastructure, automation, and monitoring systems that ensure our AI products are high-performing and cost-effective. You will report to our Director, Analytics Engineering & Data Governance and work from our Bangalore, India office. You Will: Model and Pipeline Automation Automate the deployment and retraining of ML models, from training through to production inference, by building and managing complete CI/CD/CT (Continuous Training) pipelines, adhering to MLOps best practices. Build, fine-tune, or use pre-trained LLMs, deep learning models or traditional machine learning models. Evaluate and recommend AI or ML solutions for the product using any combination of vendor solutions and/or custom-built models. Governance & Compliance Implement model versioning, lineage tracking, and auditing to ensure compliance with security and ethical standards. Performance Monitoring Continuously monitor the health and performance of production machine learning models, proactively identifying and correcting model drift, staleness, and performance degradation. Incorporate user feedback for iterative improvements and manage necessary model retraining cycles. Cross-Functional Collaboration Act as the "glue" between Data Scientists (who build models
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX (Hybrid) About the role You'll design and build the core infrastructure that powers AI inference across Cloudflare's global network — real-time voice, frontier open LLMs, and customer-deployed models running on a heterogeneous fleet of GPUs and next-generation accelerators in hundreds of cities worldwide. Working alongside AI/ML engineers, hardware partners, and Cloudflare product teams, you'll solve hard prob
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Machine Learning Engineer. About the job This position is to design and engineer AI powered features that makes every customer conversation smarter. As a Machine Learning Engineer on the Conversation Intelligence team, you'll develop and deploy solutions that extract meaning from voice and messaging data at Twilio scale. You'll work alongside experienced ML practitioners to ship real features - from model pipelines to production inference - that directly shape how businesses understand their customers. Responsibilities In this role, you’ll: Design and development of machine learning solutions, ensuring accuracy, performance, security, and scalability. Implement and maintain end-to-end AI/ML pipelines - from data ingestion and feature engineering through to model development, validation, and deployment with guidance from senior engineers on complex architectural decisions Instrument AI/ML services with appropriate metrics
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Machine Learning Engineer. About the job This position is to design and engineer AI powered features that makes every customer conversation smarter. As a Machine Learning Engineer on the Conversation Intelligence team, you'll develop and deploy solutions that extract meaning from voice and messaging data at Twilio scale. You'll work alongside experienced ML practitioners to ship real features - from model pipelines to production inference - that directly shape how businesses understand their customers. Responsibilities In this role, you’ll: Design and development of machine learning solutions, ensuring accuracy, performance, security, and scalability. Implement and maintain end-to-end AI/ML pipelines - from data ingestion and feature engineering through to model development, validation, and deployment with guidance from senior engineers on complex architectural decisions Instrument AI/ML services with appropriate metrics
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Millions of people across the world come to Pinterest to find new ideas every day. It’s where they get inspiration, dream about new possibilities and plan for what matters most. Our mission is to help those people find their inspiration and create a life they love. As a Pinterest employee, you’ll be challenged to take on work that upholds this mission and pushes Pinterest forward. As a Principal Engineer on the AI Platform team, you'll help architect the infrastructure that powers both Generative AI and Recommender Systems across Pinterest's entire product suite. Our team builds the end-to-end engines for petabyte-scale data orchestration, model training and fine-tuning, and high-performance inference, ensuring our models scale seamlessly to hundreds of millions of inferences per second in service of over 600 million monthly active users.
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! Figma is seeking a versatile and experienced Machine Learning / AI Engineer to join our growing AI team, working at the intersection of applied machine learning, infrastructure, and product innovation. Whether you’re building intelligent search systems, crafting scalable data pipelines, or enhancing AI-powered creativity tools, your work will drive user productivity, shape new product experiences, and advance the state of AI at Figma. You’ll collaborate closely with engineers, researchers, designers, and product managers across multiple teams to deliver high-quality ML-driven features and infrastructure. This is a high-impact, cross-functional role where you’ll shape both foundational systems and user-facing capabilities. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you’ll do at Figma: Design, build, and productionize ML models for Search, Discovery, Ranking, Retrieval-Augmented Generation (RAG), and generative AI features. Build and maintain scalable data pipelines to collect high-quality training and evaluation datasets, including annotation systems and human-in-the-loop workflows. Collaborate with AI researchers to iterate on datasets, evaluation metrics, and model architectures to improve quality and relevance. Work with product engineers to define and deliver impactful AI features across Figma’s platform. Partner with infrastructure engineers to develop and optimize systems for training, inference, monitoring, and deployment. Explo
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! We're looking for a research-minded Data Scientist to join the Core Data team. This team is a group of analytics professionals and Engineers building the foundational platforms for data science at Figma. We build the experimentation, analytics, and AI tooling that every product team relies on to make confident, data-driven decisions, partnering closely with Data Infra, ML, and Applied Science to evolve our platforms and embed AI into the daily workflows of data scientists across the company. This role is for someone who thrives at the intersection of rigorous research and real-world impact. You'll bring PhD-level depth to problems that matter. This includes advancing our experimentation platform and developing machine learning-based analytical systems. You will also help craft how we measure AI-powered features through causal inference and statistical modeling. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: Partner across teams to define and track important metrics, develop experiments, and uncover insights that inform strategic decisions Accelerate Figma's experimentation platform and methodology, including A/B testing frameworks and causal inference techniques Construct models and analytical frameworks based on machine learning to support product, platform, and business initiatives Create tools, datasets, and systems that enable others to work with data more efficiently
Get new inference technical lead jobs by email
Daily job updates · Unsubscribe anytime