Jobiba hiring network

Inference Technical Lead Jobs

1,448 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current inference technical lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

HI
HP IQ
📍 San Francisco• Full-time• $190K – $270K/yr
17 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role The AI team is building cutting-edge solutions that bring the power of AI directly to edge devices while seamlessly integrating with cloud infrastructure. We are looking for a Lead Software Engineer to design and develop high-performance, scalable services to support AI workloads across edge and cloud environments. What You Might Do Design, build, and maintain services that power AI-driven applications, ensuring scalability and performance. Develop APIs and microservices that facilitate seamless integration between cloud-based AI models and edge devices. Optimize data pipelines and storage solutions for real-time AI inference and processing. Implement security and privacy best practices for distributed AI systems. Work closely with AI researchers, infrastructure engineers, and frontend developers to deliver end-to-end AI-driven solutions. Build and optimize an agent orchestration runtime that enables tool use, memory management, and multi-step reasoning across LLMs, APIs, and edge-connected systems. Develop robust logging, monitoring, and alerting systems to ensure system reliability.

pythonjavasql
View job →
HI
HP IQ
📍 San Francisco• Full-time• $140K – $225K/yr
17 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role HP IQ's Connectivity team is seeking an Embedded Firmware Engineer with strong hands-on experience across RTOS and embedded Linux platforms. You'll bring up new hardware, develop and support firmware across core device subsystems, and help scale products from prototype to fleet deployment. A key part of this role is enabling on-device intelligence — bringing AI models, sensing algorithms, and local processing to lightweight, power-constrained devices at the edge. What You Might Do Design, develop, and debug firmware across RTOS and embedded Linux platforms. Lead hardware bring-up — board bring-up, driver integration, and firmware support through to production. Develop and maintain firmware for subsystems such as connectivity, power and other sensors. Integrate lightweight model inference and sensing/proximity algorithms within tight compute, memory, and power budgets. Support fleet-scale deployment: OTA updates, field diagnostics, and post-launch sustainment. Collaborate with hardware, systems, and QA teams to troubleshoot system-level issues and drive root-cause fixes to closure. Ess

redislinuxai
View job →
HI
HP IQ
📍 San Francisco• Full-time• $149K – $240K/yr
17 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role The AI team is building cutting-edge solutions that bring the power of AI directly to edge devices while seamlessly integrating with cloud infrastructure. We are looking for a Senior Software Engineer to design and develop high-performance, scalable services to support AI workloads across edge and cloud environments. What You Might Do Design, build, and maintain services that power AI-driven applications, ensuring scalability and performance. Develop APIs and microservices that facilitate seamless integration between cloud-based AI models and edge devices. Optimize data pipelines and storage solutions for real-time AI inference and processing. Implement security and privacy best practices for distributed AI systems. Work closely with AI researchers, infrastructure engineers, and frontend developers to deliver end-to-end AI-driven solutions. Build and optimize an agent orchestration runtime that enables tool use, memory management, and multi-step reasoning across LLMs, APIs, and edge-connected systems. Develop robust logging, monitoring, and alerting systems to ensure system reliabilit

pythonjavasql
View job →
E
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Our Senior Applied AI Engineer builds and operate production-grade AI systems that extract meaning from large-scale unstructured document collections, enabling enterprise data discovery classification, and governance. This role owns the full lifecycle of graph intelligence solutions — from problem definition and data modelling, to building and enriching knowledge graphs, and deploying ML- and LLM-assisted analytics in production. The focus is on semantic and contextual analysis of unstructured data to uncover relationships, patterns, and insights that support AI safety, security, and compliance requirements. WHAT YOU'LL DO Design, build, and deploy graph-based AI solutions, combining knowledge graphs , LLMs, and ML models applied to large-scale unstructured data Define and own data pipelines that extract, transform, and enrich entity relationships into production-grade knowledge graphs Integrate LLMs and ML models into text processing pipelines for classification, embedding generation, document similarity, and semantic analysis Design, deploy, and operate graph and vector databases to support retrieval, reasoning, and analytics Optimize models and inference pipelines for production constraints including latency, throughput, cost, and infrastructure Deploy, monitor, and iterate on ML systems in production environments ensuring reliability and continuous integration Drive architectural decisions and tech

pythonawsdocker
View job →
E
Everpure
📍 Bengaluru• Full-time
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... We’re in an unbelievably exciting area of tech, fundamentally reshaping the cloud-native, modern virtualization, and AI infrastructure landscape. Here, you’ll lead with innovative thinking, grow alongside us, and work with some of the smartest minds in the industry. We’re looking for engineers passionate about system testing, distributed systems, Kubernetes, storage, modern virtualization, AI workloads, and automation. You’ll work on complex, real-world scenarios involving HA, resiliency, disaster recovery, scalability, and failure testing across large-scale Kubernetes environments. As enterprises modernize their infrastructure, containers, virtual machines, and AI/ML workloads are increasingly converging on Kubernetes. From traditional enterprise applications and VMs to GPU-accelerated AI training, inference, and data-intensive workloads, Kubernetes is rapidly becoming the common platform for running the next generation of applications. This role gives you the opportunity to test and influence how Portworx delivers enterprise-grade storage, data protection, availability, and resiliency across these workloads—including containerized applications, KubeVirt/OpenShift Virtualization VMs, and demanding AI/ML workloads running on Kubernetes. WHAT YOU WILL DO: Own system-level quality for Portworx Enterprise across Kubernetes, storage, and modern virtualization environments. Develop comprehens

pythonawsazure
View job →
E
Everpure
📍 Bengaluru• Full-time
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... We’re in an unbelievably exciting area of tech, fundamentally reshaping the cloud-native, modern virtualization, and AI infrastructure landscape. Here, you’ll lead with innovative thinking, grow alongside us, and work with some of the smartest minds in the industry. We’re looking for engineers passionate about system testing, distributed systems, Kubernetes, storage, modern virtualization, AI workloads, and automation. You’ll work on complex, real-world scenarios involving HA, resiliency, disaster recovery, scalability, and failure testing across large-scale Kubernetes environments. As enterprises modernize their infrastructure, containers, virtual machines, and AI/ML workloads are increasingly converging on Kubernetes. From traditional enterprise applications and VMs to GPU-accelerated AI training, inference, and data-intensive workloads, Kubernetes is rapidly becoming the common platform for running the next generation of applications. This role gives you the opportunity to test and influence how Portworx delivers enterprise-grade storage, data protection, availability, and resiliency across these workloads—including containerized applications, KubeVirt/OpenShift Virtualization VMs, and demanding AI/ML workloads running on Kubernetes. WHAT YOU WILL DO: Own system-level quality for Portworx Enterprise across Kubernetes, storage, and modern virtualization environments. Develop comprehens

pythonawsazure
View job →
E
Everpure
📍 Bengaluru• Full-time
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Pure Solutions team as a Senior MLOps Solutions Engineer to architect and build high-scale, enterprise-grade AI/ML solutions. You will be instrumental in integrating Pure Storage platforms with the evolving open-source MLOps ecosystem (Kubeflow, MLflow, Ray) to operationalize the complete machine learning lifecycle. This role requires a creative technologist with deep Python expertise to drive innovation and enable our customers and partners to achieve production AI success. WHAT YOU'LL DO Design and Automate MLOps Pipelines: Lead the development of end-to-end MLOps workflows using CI/CD tools (Git/Jenkins) and orchestration platforms (MLflow/Kubeflow), specifically integrating Pure Storage's FlashBlade, FlashArray, and Portworx as the high-performance data plane for data ingestion, training, and inference. Build High-Performance AI/ML Reference Architectures: Create validated, repeatable deployment models using Infrastructure as Code (e.g., Ansible, Terraform) for AI/ML environments spanning bare metal, virtual machines, and GPU-accelerated Kubernetes clusters, ensuring optimal performance for distributed training. Optimize and Operationalize GPU Inference: Architect and implement solutions for high-throughput, low-latency model serving, utilizing technologies like NVIDIA Triton Inference Server and advanced optimization techniques (quantization, model sharding like DeepSpeed/Megatron-LM, and dynamic bat

pythonawskubernetes
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
17 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a senior High Speed Interconnect / Signal Integrity Engineer to design and validate high-bandwidth links for large-scale AI systems. You will define, model, and qualify interconnect solutions across copper and optical technologies for next-generation AI inference and training clusters. This role is on-site in Santa Clara, CA, Austin, TX, or Toronto, Canada. We welcome candidates at various experience levels. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting Who You Are An experienced electrical engineer with a Bachelor’s or Master’s in Electrical Engineering. 5+ years working on high-speed communications (100G–1.6T), including signal integrity, channels, and links. Comfortable building and owning link and channel budgets and making clear tradeoffs between reach, loss, BER, and margin. Hands-on with SI tools and lab equipment such as Keysight ADS, VNAs, TDRs, BERTs, and protocol analyzers. Familiar with cable specification and testing, as well as accelerated life testing, mating life, and failure analysis. Able to collaborate across hardware, systems, and manufacturing teams; manufacturing/DFM/DTM experience is a plus. What We Need Define and specify high-speed interconnect architectures

awsaisem
View job →
G
Graphcore
📍 Cambridge• Full-time
17 days ago

About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary As a research engineer at Graphcore, you will contribute to the advancement of AI research, investigating new ideas that push the limits on important AI/ML problems. Specialised hardware has been the key driver of the progress of AI over the last decade, and we believe that hardware-aware AI algorithms and AI-aware hardware developments will continue to be critical to advancing this exciting field. We are therefore looking for individuals who combine strong machine learning experience with practical engineering skills to deliver impactful AI research. We are seeking AI researchers with strong software engineering experience, particularly in lower-level programming and performance optimisation for hardware efficiency. Our research spans a broad range of topics, including efficient training and inference, world models, life sciences, reinforcement learning, and beyond. You will work closely with researchers to generate ideas and translate them into scalable implementations, contributing to publications and projects that help to steer the future of AI hardware. The Team Graphcore Research participates in both fundamental and applied research, to characterise the computational requirements of machine intelligence a

pythonrestmachine learning
View job →
G
17 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe

pythonaigo
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

REMOTEkubernetesaigo
View job →

About the Role Adobe is seeking a Machine Learning Engineer to join the Adobe Genuine Engineering team. This group protects Adobe's ecosystem from fraud, abuse, and misuse using intelligent systems worldwide. In this position, you will build and develop machine learning models from scratch, including custom transformer-based frameworks, to identify fraudulent actions, stop account sharing, and protect the experience of hundreds of millions of users. You will manage the entire model lifecycle: raw behavioral data and feature engineering, architecture development, large-scale GPU training, deployment, and monitoring. The team is actively building in-house behavioral foundation models that learn identity-preserving representations from long sequences of user activity. This is a role for an engineer who wants to own deep learning systems end-to-end — not consume pre-built ones. Key Responsibilities Build and train deep learning models from scratch, including custom transformer and attention-based architectures for long behavioral event sequences. Own the full training stack: event tokenization, temporal and positional embeddings, self-supervised pretraining (e.g., masked modeling, contrastive learning), and downstream fine-tuning. Train large models efficiently on GPU infrastructure using mixed-precision training, gradient accumulation/checkpointing, efficient attention, and distributed strategies (DDP, FSDP, or equivalent). Build and optimize feature pipelines on Databricks and Spark, transforming raw behavioral events into high-quality model inputs. Translate prototypes into production ML systems — scalable, reliable, and observable — and drive inference performance through architectural and serving-side optimization. Contribute to MLOps practices: experiment tracking, model versioning, CI/CD, automated retraining, and

pythonmachine learningai
View job →
A
1mo ago

About the Role Adobe is seeking a Machine Learning Engineer to join the Adobe Genuine Engineering team. This group protects Adobe's ecosystem from fraud, abuse, and misuse using intelligent systems worldwide. In this position, you will build and develop machine learning models from scratch, including custom transformer-based frameworks, to identify fraudulent actions, stop account sharing, and protect the experience of hundreds of millions of users. You will manage the entire model lifecycle: raw behavioral data and feature engineering, architecture development, large-scale GPU training, deployment, and monitoring. The team is actively building in-house behavioral foundation models that learn identity-preserving representations from long sequences of user activity. This is a role for an engineer who wants to own deep learning systems end-to-end — not consume pre-built ones. Key Responsibilities Build and train deep learning models from scratch, including custom transformer and attention-based architectures for long behavioral event sequences. Own the full training stack: event tokenization, temporal and positional embeddings, self-supervised pretraining (e.g., masked modeling, contrastive learning), and downstream fine-tuning. Train large models efficiently on GPU infrastructure using mixed-precision training, gradient accumulation/checkpointing, efficient attention, and distributed strategies (DDP, FSDP, or equivalent). Build and optimize feature pipelines on Databricks and Spark, transforming raw behavioral events into high-quality model inputs. Translate prototypes into production ML systems — scalable, reliable, and observable — and drive inference performance through architectural and serving-side optimization. Contribute to MLOps practices: experiment tracking, model versioning, CI/CD, automated retraining, and

pythonmachine learningai
View job →

About the Role Adobe is seeking a Machine Learning Engineer to join the Adobe Genuine Engineering team. This group protects Adobe's ecosystem from fraud, abuse, and misuse using intelligent systems worldwide. In this position, you will build and develop machine learning models from scratch, including custom transformer-based frameworks, to identify fraudulent actions, stop account sharing, and protect the experience of hundreds of millions of users. You will manage the entire model lifecycle: raw behavioral data and feature engineering, architecture development, large-scale GPU training, deployment, and monitoring. The team is actively building in-house behavioral foundation models that learn identity-preserving representations from long sequences of user activity. This is a role for an engineer who wants to own deep learning systems end-to-end — not consume pre-built ones. Key Responsibilities Build and train deep learning models from scratch, including custom transformer and attention-based architectures for long behavioral event sequences. Own the full training stack: event tokenization, temporal and positional embeddings, self-supervised pretraining (e.g., masked modeling, contrastive learning), and downstream fine-tuning. Train large models efficiently on GPU infrastructure using mixed-precision training, gradient accumulation/checkpointing, efficient attention, and distributed strategies (DDP, FSDP, or equivalent). Build and optimize feature pipelines on Databricks and Spark, transforming raw behavioral events into high-quality model inputs. Translate prototypes into production ML systems — scalable, reliable, and observable — and drive inference performance through architectural and serving-side optimization. Contribute to MLOps practices: experiment tracking, model versioning, CI/CD, automated retraining, and

pythonmachine learningai
View job →

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Data Science is at the heart of Lyft’s products and decision-making. Data Scientists at Lyft operate in dynamic environments, moving quickly to build the world’s best transportation solutions. We tackle a wide range of challenges - from shaping long-term business strategy with data, to making critical short-term decisions, to developing algorithms and models that power both internal systems and customer-facing products. Driver Incentives Science owns the algorithms and systems behind incentive design, influencing driver engagement and marketplace efficiency — from real-time supply positioning to longer-horizon earnings and engagement programs. The team is responsible for designing pay and incentive mechanisms that are efficient and good for driver experience over the long run. As a Data Scientist specializing in Algorithms, you'll partner closely with product, engineering, and operations leaders to build and scale incentive systems, shape long-term mechanism design strategy, and deliver on critical business goals tied to marketplace efficiency and driver earnings. Candidates with strong optimization backgrounds — think mathematical programming, control theory, or operations research — are a great fit, though we welcome strong candidates from machine learning or causal inference as well. The ideal candidate thrives in a fast-paced environment and brings a hands-on, entrepreneurial mindset to drive results. Responsibilities: Collaborate with engineering and product teams to design, implement, and iterate on new features and algorithmic improvements for driver incentives and pay mechanisms. Design, develop, and deploy optimization models, algorithms, and systems for problems such as budget allocation, multidimensional cost-curve development, and incentive targeting. Write production model code; collabor

pythonmachine learningartificial intelligence
View job →
🔔

Get new inference technical lead jobs by email

Daily job updates · Unsubscribe anytime