Jobiba hiring network

Inference Technical Lead Jobs

1,448 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current inference technical lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse using Apache Iceberg — at dramatically lower cost. OpenTelemetry is the foundation of how customers send data to Observe, and we are making significant investment in upstream OTel contributions to ensure Observe is the best destination for OTel-instrumented environments. As a Senior Software Engineer on the OpenTelemetry team, you'll play a key role in driving Observe's open-source strategy within the OpenTelemetry ecosystem. You'll work on some of the most exciting challenges in telemetry collection and instrumentation — expanding coverage into emerging domains like LLM/AI observability and Browser RUM — while collaborating closely with the OTel community to shape the standards that the industry builds on. AS A SENIOR SOFTWARE ENGINEER - OBSERVE BY SNOWFLAKE, OPENTELEMETRY, YOU WILL: Pioneer new instrumentation capabilities in the OpenTelemetry community, defining and building next-generation telemetry collection for emerging domains including LLM/AI inference, Browser RUM, and mobile. Design and implement OTel Collector components (receivers, processors, exporte

pythonjavaaws
View job →
PE
Private Employer
📍 Uttar Pradesh, India• Full-time
1mo ago

About Us Paytm is India's payment Super App offering consumers and merchants the most comprehensive payment services. As the pioneer of the mobile QR payments revolution in India, today, Paytm stands as India’s largest payment company by Users, Merchants, Payment Transactions, and Revenue. Paytm’s mission is to drive financial inclusion in India and bring half a billion Indians into the mainstream economy through technology-led financial Services. Paytm enables commerce for small merchants and distributes various financial services offerings to its consumers and merchants in partnership with financial institutions. Paytm has been a pioneer in the merchant space by introducing innovative solutions like QR codes to accept payments and Sound-box to reconcile payments via voice alerts. We are also distributing loans to these partners via our ‘Paytm for Business’ App. About The Team Paytm Intelligence is building next-generation AI platforms across two key areas: AI inference infrastructure and agentic AI solutions. Our focus is on enabling enterprises to move from AI experimentation to production-grade deployment with measurable business outcomes. In the agentic AI space, we are building solutions such as Outreach Manager, CLM agents, and domain- specific AI agents that automate complex customer journeys across sales, service, and operations. These systems go beyond traditional chatbots, orchestrating workflows, decisions, and actions across enterprise systems. About the Role We are hiring an AI Agentic Solutions Architect to drive enterprise adoption of AI-powered workflow automation. This role sits at the intersection of business process transformation, AI systems, and enterprise sales. You will work directly with enterprise customers to understand their workflows, identify automation opportunities, and design agentic solutions that deliver measurable outcomes such as cost reduction, conversion uplift, and improved customer experience. Key Responsibilities Customer

gitagileai
View job →
PE
Private Employer
📍 Toronto• Full-time• Hybrid
1mo ago

About the role There’s a wide gap between an agent that works in a demo and one that works across millions of live transactions. Closing it is the job. You’ll embed with the teams and customers who depend on AI — risk, fraud, collections, payments, support, developer experience — and design, build, and ship agentic systems into their production environments. You’ll also help build the platform underneath: Paytm’s AI inference platform (Pi) and the agentic runtime, orchestration, and tooling that lets agents reason, plan, use tools, and run multi-step workflows safely.

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Job Description/ Responsibilities: Designing, developing and maintaining stable and reliable AI/ML Ops platforms / pipelines Minimum experience of 4-6 Years required in AI ML Ops Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production. Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow. Security & Compliance: Implementing data security measures, ensuring compliance with data governance

pythonsqlaws
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. About our Team: Micron’s Industrial and Physical AI team is driving the transformation of semiconductor manufacturing through Autonomous Operations, AI, robotics, and digital twin technologies! We develop and deploy innovative solutions across Micron’s global fabrication and assembly/test facilities, enabling smarter, safer, and more efficient operations at scale. Position Overview: We are seeking a hands-on Full-Stack AI Engineer to design, build, and deploy production-grade AI applications that support Micron's Autonomous Operations initiatives. This role owns the end-to-end development lifecycle, from data pipelines and AI models to APIs, web applications, digital twin integrations, and cloud/edge deployments, delivering impactful solutions for engineers, operators, and business leaders worldwide. Responsibilities: Design, architect, and deliver end-to-end AI products, including data ingestion pipelines, feature engineering, model training/inference, APIs, user interfaces, and application monitoring. Build and maintain modern front-end applications using React, Angular, or Streamlit, supported by backend services in Python and FastAPI. Develop scalable integrations between manufacturing systems, robotics platforms, AMRs, sensor networks, and enterprise applications to enable intelligent factory operations. Design and implement digital twin environments using platforms such as NVIDIA Omniverse, Gazebo, or Unity Robotics Hub to support simulation, validation, and o

javascripttypescriptpython
View job →
N
Notion
📍 San Francisco• Full-time• Remote
12 days ago

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: Notion has been at the cutting edge of AI since before ChatGPT launched. The job of the Model Capabilities team is to keep us there. We own the model layer of Notion AI: integrating frontier models as they ship, keeping inference reliable and economical at scale, and building new capabilities that other teams take advantage of. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: Bring new frontier models into production quickly, making them available for our users and our engineers. Make inference reliable: better error categorization, self-healing retries, and cross-provider failover. Own observability for the model layer, driving down both time to detection and time to fix. Build new model-level capabilities and help product teams adopt them. Act as connective tissue across Notion's AI teams: find the gaps, unblock people, and make sure fixes land with the r

REMOTEai
View job →

Reolink , a leader in intelligent visual technology for homes and businesses, was founded in 2009 by a group of engineers with a strong commitment to and passion for smarter security solutions. Our products are now trusted by millions of users across more than 110 countries and regions worldwide. Building on this trust, we continue expanding our presence and bringing our innovations to more markets around the globe. Reolink remains committed to delivering advanced, reliable, and user‑centric solutions that empower people to protect what matters most. AI Algorithms Specialist (PHD Holder Only) 5 Work Days Per Week Office Near Tai Seng MRT, Singapore Medical Benefits Provided Entitled to Yearly Bonus & Performance Bonus Job Requirements PHD Holder in Computer Science, Applied Mathematics, Electrical Engineering, Pattern Recognition, Artificial Intelligence, Automatic Control, Operations Research, Biology, Physics / Quantum Computing, Neuroscience, Statistics or a related field. At least 2-5 years of workplace working experiences is preferable for this post. Familiar with common machine learning and deep learning algorithms and keeping track with the latest SOTA implementations. Strong programming skill in Python, C / C++, proficient in mathematical / statistical concepts and exceptional coding skills Hands-on experience with AI / ML frameworks be familiar such as Caffe, PyTorch, TensorFlow, MxNet etc. Have rich project experience in machine learning and deep learning, be familiar with common algorithm models, such as CNN, RNN, LSTM, Transformer, ViT, etc., and be able to improve and innovate models according to actual problems. Experience in familiar the design, parameter tuning and optimization methods of neural network models is a plus Experience in model compression and in the transplantation and optimization of deep learning forward inference on various platforms, including NPU / GPU / DSP / ARM on mobile platforms and CPU / GPU on server platforms is also a p

pythonmachine learningartificial intelligence
View job →
P
Pinterest
📍 San Francisco• Remote
13 days ago

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Millions of people across the world come to Pinterest to find new ideas every day. It’s where they get inspiration, dream about new possibilities and plan for what matters most. Our mission is to help those people find their inspiration and create a life they love. As a Pinterest employee, you’ll be challenged to take on work that upholds this mission and pushes Pinterest forward. As a Principal Engineer on the AI Platform team, you'll help architect the infrastructure that powers both Generative AI and Recommender Systems across Pinterest's entire product suite. Our team builds the end-to-end engines for petabyte-scale data orchestration, model training and fine-tuning, and high-performance inference, ensuring our models scale seamlessly to hundreds of millions of inferences per second in service of over 600 million monthly active users.

REMOTEjavaaic++
View job →
N
Nvidia
📍 Santa Clara, United States
14 days ago

NVIDIA Research is seeking extraordinary networking innovators to join our NVResearch team. As a research intern on this team, you will contribute to the development of future high-performance networking and computing systems. We are seeking a balanced background of research excellence in building systems and a deep understanding and broad perspective across the fields of computer architecture and communication systems for distributed computation. NVIDIA has pioneered programmable GPUs and the CUDA language, and this visionary Research team will take those technologies to the next level with its creative ideas and new inventions. This position offers you the opportunity to have a real impact while working with some of the most creative and forward-thinking people in the world who are here at this dynamic, technology-focused company. What you'll be doing: Develop algorithms and design hardware and software, extending the state of the art in computing, networking, and other technology areas surrounding NVIDIA's business. Invent new techniques, technologies, methodologies, processes, and devices, to enable new products or types of products. Deliverable results include prototypes, patents, and publications. Contribute to research that informs NVIDIA's technology direction 5-10 years out. Work focuses on long-horizon problems rather than products currently shipping or in development, except as to how they can be extended and improved. Projects can include but are not limited to: optimizing communication stacks for AI training and inference, designing network protocols and congestion control, co-designing AI systems across software and hardware, developing circuits and microarchitecture for network controllers and switches, and architecting networks built on optical switching and silicon photonics. What we need to see: Pursuing a

pythonartificial intelligenceai
View job →

NVIDIA is seeking outstanding Research Interns to join the Data-Driven AI for Robotics (DAIR) group. The focus is on learning embodied skills from large-scale human data. Our objective is to develop AI systems that capture, understand, and reproduce complex human motion and interaction skills across physical and digital embodiments, including humanoid robots and animated characters. Our research spans the full stack: reconstructing human motion and human-object interactions from video; generating diverse, controllable character behaviors; transferring motion across embodiments; and training physically grounded controllers for humanoid robots and interactive virtual characters. You will collaborate with a passionate and supportive research team that consistently produces influential work published at leading computer vision, machine learning, graphics, and robotics conferences. You will also have the opportunity to collaborate with world-class research and product teams across NVIDIA, following our strong “one-team” culture. What you'll be doing: Innovate and implement novel AI algorithms that transform large-scale human data into controllable motion and interaction skills across physical and digital embodiments. Develop robust, scalable training and inference pipelines for motion reconstruction, generation, retargeting, and character and robot control. Build methods that transfer human skills to humanoid robots, including whole-body loco-manipulation and dexterous manipulation. Maintain a close, collaborative relationship with your mentor(s). Publish your research findings at leading computer vision, machine learning, graphics, and robotics conferences. Partner with product teams to enable effective technology transfer of your work. Research Topics Include: Human motion and human-object interaction reconstruction, synthesis, and generatio

machine learningai
View job →
S
17 days ago

We're looking for an ML Data & Platform Engineer to own the infrastructure that powers our speech AI models: the pipelines that source and prepare training data, and the platform that trains, evaluates, and serves them in production. Speech AI has a data problem most ML teams don't, and you'll be at the centre of solving it, working as part of our ML team to remove friction across the entire lifecycle and get better models into production faster. This is a broad, cross-functional role suited to someone who enjoys working across the full stack: data infrastructure, distributed systems, and production ML, and who takes ownership of problems end to end rather than waiting to be told what to fix. What you'll do Designing, building, and maintaining scalable data pipelines for ingesting, transforming, validating, and storing large datasets used to train our models Developing and maintaining web scraping and data acquisition solutions to keep training datasets fresh, high-quality, and available at scale Building and operating the infrastructure that lets the ML team deploy and evaluate new models quickly, and that serves models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning data infrastructure, training, and inference Continuously improving our data and MLOps practices, and helping shape the roadmap for how our platform evolves as we scale What you'll need Strong proficiency in Python and SQL, with a solid backend or data engineering foundation Hands-on experience with containerisation and orchestration (Docker, Kubernetes), and working with a major cloud provider Experience building data pipelines and ETL/ELT processe

pythonsqldocker
View job →
R
17 days ago

Reolink , a leader in intelligent visual technology for homes and businesses, was founded in 2009 by a group of engineers with a strong commitment to and passion for smarter security solutions. Our products are now trusted by millions of users across more than 110 countries and regions worldwide. Building on this trust, we continue expanding our presence and bringing our innovations to more markets around the globe. Reolink remains committed to delivering advanced, reliable, and user‑centric solutions that empower people to protect what matters most. AI Algorithms Engineer (PHD Holder Only) 5 Work Days Per Week Office Near Tai Seng MRT, Singapore Medical Benefits Provided Entitled to Yearly Bonus & Performance Bonus Job Requirements: PHD Holder in Computer Science, Applied Mathematics, Electrical Engineering, Pattern Recognition, Artificial Intelligence, Automatic Control, Operations Research, Biology, Physics / Quantum Computing, Neuroscience, Statistics or a related field. Familiar with common machine learning and deep learning algorithms and keeping track with the latest SOTA implementations. Strong programming skill in Python, C / C++, proficient in mathematical / statistical concepts and exceptional coding skills Hands-on experience with AI / ML frameworks be familiar such as Caffe, PyTorch, TensorFlow, MxNet etc. Have rich project experience in machine learning and deep learning, be familiar with common algorithm models, such as CNN, RNN, LSTM, Transformer, ViT, etc., and be able to improve and innovate models according to actual problems. Experience in familiar the design, parameter tuning and optimization methods of neural network models is a plus Experience in model compression and in the transplantation and optimization of deep learning forward inference on various platforms, including NPU / GPU / DSP / ARM on mobile platforms and CPU / GPU on server platforms is also a plus. Strong logical thinking and problem-solving ability, able to independen

pythonmachine learningai
View job →
R
17 days ago

Reolink , a leader in intelligent visual technology for homes and businesses, was founded in 2009 by a group of engineers with a strong commitment to and passion for smarter security solutions. Our products are now trusted by millions of users across more than 110 countries and regions worldwide. Building on this trust, we continue expanding our presence and bringing our innovations to more markets around the globe. Reolink remains committed to delivering advanced, reliable, and user‑centric solutions that empower people to protect what matters most. AI Algorithms Engineer (PHD Only) 5 Work Days Per Week Office Near to Kaki Bukit MRT, Singapore Relocate Near Tai Seng MRT in Mid-August 2026 Medical & Dental Benefits Provided Entitled to Yearly Bonus & Performance Bonus Job Requirements: PHD Holder in Computer Science, Applied Mathematics, Electrical Engineering, Pattern Recognition, Artificial Intelligence, Automatic Control, Operations Research, Biology, Physics / Quantum Computing, Neuroscience, Statistics or a related field. Familiar with common machine learning and deep learning algorithms and keeping track with the latest SOTA implementations . Strong programming skill in Python, C / C++ , proficient in mathematical / statistical concepts and exceptional coding skills Hands-on experience with AI / ML frameworks be familiar such as Caffe, PyTorch, TensorFlow, MxNet etc. Have rich project experience in machine learning and deep learning, be familiar with common algorithm models, such as CNN, RNN, LSTM, Transformer, ViT, etc., and be able to improve and innovate models according to actual problems. Experience in familiar the design, parameter tuning and optimization methods of neural network models is a plus Experience in model compression and in the transplantation and optimization of deep learning forward inference on various platforms, including NPU / GPU / DSP / ARM &nbs

pythonmachine learningai
View job →
S
SpaceXAI
📍 London• Full-time• £107K – £262K/yr
17 days ago

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: As an ideal candidate you have a good understanding of how highly scalable and reliable production infrastructure is built. Most of our backend infrastructure is written in Rust. So familiarity with a compiled language such as C++, Rust, or Go is highly beneficial. RESPONSIBILITIES: Build the SpaceXAI API that serves our models to developers worldwide Own the end-to-end system responsible for high-throughput inference, handling billions of tokens per minute with low latency and high availability, including model serving infrastructure, request routing, SDK development, rate limiting, observability, and efficient scaling BASIC QUALIFICATIONS: Expert knowledge of either Rust or C++ Experience in designing, implementing, and maintaining reliable and horizontally scalable distributed systems Knowledge of service observability and reliability best practices Experience in operating commonly used databases such as PostgreSQL, Clickhouse, and MongoDB PREFERRED SKILLS AND EXPERIENCE: Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM) Experience designing or building with agent SDKs and agent orchestration frameworks Experience with Docker, Kubernetes, and containerized applicatio

sqlpostgresqlmongodb
View job →
HI
HP IQ
📍 San Francisco• Full-time• C$45 – C$51/hr
17 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role HP IQ's AI Machine Learning (AML) team is building the foundational platform powering a new generation of agentic devices. This platform orchestrates the complete lifecycle of AI models: from creation and fine-tuning through optimized inference and intelligent orchestration. The team works to make complex AI capabilities run efficiently on-device, enabling locally-deployed agentic experiences that reduce token costs and improve privacy. In this internship role, you will contribute to one or more core pillars of the AML platform: model inference optimization, orchestration and agent workflows, or model creation and fine-tuning. You will partner closely with experienced engineers to ship features that directly impact HP's next-generation devices and gain visibility into how each component of an end-to-end AI system integrates and scales. What You Might Do Work on model hosting and inference optimization, learning how to profile, benchmark, and accelerate model execution on resource-constrained devices; experiment with quantization, distillation, or other optimization techniques to reduc

pythonredismachine learning
View job →
🔔

Get new inference technical lead jobs by email

Daily job updates · Unsubscribe anytime