Jobiba hiring network

Inference Technical Lead Jobs

1,448 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current inference technical lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
1mo ago

About the Team We’re hiring software engineers to make the Workload team more productive. The Workload team maintains the core components of OpenAI’s training and inference frameworks and helps execute frontier experiments. About the Role We’re looking for someone who cares about the developer experience of working in and around OpenAI’s core training and inference frameworks. In this role you will: Be responsible for optimizing the development workflows of the engineers around you Work within various Workload teams to address their specific needs, but collaborate with the centralized teams that own various aspects of development experience Optimize iteration speed, both broadly, and in particular by optimizing specific teams’ CI Improve reliability, for instance, by driving testing strategy for particular components Work through the long tail of things that it takes to build libraries and systems that will delight researchers You might thrive in this role if: You are motivated by helping people. You believe a thing that separates great teams from good teams are the players willing to do whatever work it takes, without ego. You believe in the power of developer experience. Something magical happens when people can quickly and confidently iterate on a simple codebase, but this magic is fragile and must be fought for. When you see someone trip over something, no matter how small, your first instinct is asking yourself what it would take for that to not happen again. Your second instinct is clicking merge on the PR you’ve already written to make it so. You are pragmatic. You have the ability to see the world through a perfectionist’s eyes, but are not yourself a perfectionist. You know which problems to pick and when to switch to making progress on a different problem. You like going end-to-end on things. You love co-design — that feeling when you were only able to find the right solution because you both deeply understand the users that interact with a system and the

pythonawsrest
View job →
O
1mo ago

About the Team We’re hiring software engineers to make OpenAI’s networking teams more productive. These teams build and operate the high-performance networking systems that support OpenAI’s training and inference infrastructure at frontier scale. About the Role We’re looking for someone who cares deeply about the developer experience of engineers working on complex infrastructure systems — especially around build systems, test architecture, release pipelines, and reliable development workflows. This role will be embedded with OpenAI’s networking team: making it faster, safer, and easier for engineers to build, test, validate, and ship changes across multi-server, networked, and hardware-adjacent environments. In this role you will: Improve development workflows for engineers building and operating OpenAI’s networking systems Design and improve continuous deployment, release, and validation pipelines Build and maintain test harnesses for multi-server, networked, and hardware-backed environments Improve iteration speed across C++, Python, and build-system-heavy codebases Partner with engineers to identify friction in CI, testing, debugging, and deployment workflows Drive testing and reliability strategy for infrastructure components that support large-scale training and inference workloads Work closely with centralized developer experience teams while staying deeply embedded with the networking engineers closest to the systems You might thrive in this role if: You are motivated by helping other engineers move faster and with more confidence You have experience with CI/CD, release pipelines, testing infrastructure, or build systems You are comfortable moving between C++, Python, and build systems such as CMake, Bazel, or Blaze You enjoy building test harnesses, automation, and workflow improvements for complex systems You do not need to be a networking expert, but you are excited to learn enough about the domain to make the team meaningfully more effective When you see

pythonawsci/cd
View job →

About the Team DoorDash is building the world’s most reliable on-demand logistics engine for delivery! We’re looking for machine learning engineer interns to join our fast-growing engineering team to help us develop a 24x7 global infrastructure system that powers DoorDash’s three-sided marketplace of consumers, merchants, and dashers. About the Role As a Machine Learning Engineer intern at Doordash, you’ll work on tackling new challenges in machine learning and artificial intelligence. You’ll conduct research that can be applied across Doordash engineering teams and engage in external collaborations and mentoring, while also performing research in any of the following areas: Auction, Game theory, Recommender systems, Ranking, AdTech, Computer Vision, Causal Inference, and Big data analytics. We offer a 12-week summer internship program in our San Francisco, Sunnyvale, New York, or Seattle offices. You’re excited about this opportunity because you will… Use cutting-edge research in ML/AI, NLP, RecSys, Ranking, Computer Vision, Causal Inference, Ad Tech, Graph analysis to solve real-world problems across discovery, ads,forecasting, fulfillment and search experiences at Doordash. Contribute and execute on research ideas that can be applied and used to improve product experience at Doordash. Collect, analyze, and synthesize findings from data and use these insights to build relevant ML models. Write clean, efficient, and sustainable code We’re excited about you because you… Are working towards a Masters degree in Computer Science, ML, NLP, Statistics, Information Sciences or related field and are graduating between Fall 2027 & Summer 2028 Have a mastery of at least one systems languages (Java, C++, Python, Kotlin, GoLang) or one ML framework (Tensorflow, Pytorch, MLFlow) Have experience in research and in solving analytical problems Are a strong communicator and team player. Have a passion for applied ML and the Doordash product Ideally have

pythonjavamachine learning
View job →

About the Team DoorDash is building the world’s most reliable on-demand logistics engine for delivery! We’re looking for machine learning engineer interns to join our fast-growing engineering team to help us develop a 24x7 global infrastructure system that powers DoorDash’s three-sided marketplace of consumers, merchants, and dashers. About the Role As a Machine Learning Engineer intern at Doordash, you’ll work on tackling new challenges in machine learning and artificial intelligence. You’ll conduct research that can be applied across Doordash engineering teams and engage in external collaborations and mentoring, while also performing research in any of the following areas: Auction, Game theory, Recommender systems, Ranking, AdTech, Computer Vision, Causal Inference, and Big data analytics. We offer a 12-week summer internship program in our San Francisco, Sunnyvale, New York, or Seattle offices. You’re excited about this opportunity because you will… Use cutting-edge research in ML/AI, NLP, RecSys, Ranking, Computer Vision, Causal Inference, Ad Tech, Graph analysis to solve real-world problems across discovery, ads,forecasting, fulfillment and search experiences at Doordash. Contribute and execute on research ideas that can be applied and used to improve product experience at Doordash. Collect, analyze, and synthesize findings from data and use these insights to build relevant ML models. Write clean, efficient, and sustainable code We’re excited about you because you… Are working towards a PhD degree in Computer Science, ML, NLP, Statistics, Information Sciences or related field and are graduating between Fall 2027 & Summer 2028 Have a mastery of at least one systems languages (Java, C++, Python, Kotlin, GoLang) or one ML framework (Tensorflow, Pytorch, MLFlow) Have experience in research and in solving analytical problems Are a strong communicator and team player. Have a passion for applied ML and the Doordash product Ideally have work

pythonjavamachine learning
View job →
MT
11 days ago

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Job Summary We are seeking a motivated engineer to join the DRAM Systems Engineering team, focusing on the development, evaluation, and optimization of next-generation memory systems for AI accelerators. This role emphasizes research and development across hardware architecture, operating systems, and performance analysis to support Agentic AI inference workloads. If you are ambitious and eager to make an impact in the exciting world of AI and memory systems, this is the perfect opportunity for you! Responsibilities Characterize AI inference workloads and examine memory behavior Build and evaluate tiered memory hierarchies for AI accelerators Study KV cache lifecycles, MoE models, and data placement strategies Compare and optimize explicit versus hardware-assisted data movement Develop, test, debug, and detail system-level and OS components Prototype and evaluate agentic AI systems by building agents and multi-agent workflows using modern frameworks and orchestration patterns (planning, tool use, memory, and context management). Apply these technologies both as workloads under study and as accelerators for internal engineering workflows <h2 style="color:!importan

pythonlinuxai
View job →
M
Modal
📍 New York• Full-time
12 days ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. About Modal Data: We’re growing our Data team and are looking for our first few key hires to build self-serve data tools and drive business strategy in the right direction. The mission of the Modal Data team is to make it easy to track company goals, make evidence-backed decisions, and prioritize the right work. We do this via: Self-serve AI analytics tools (Hex, Snowflake) Embedding with teams as a “data adviser”, providing strategic analysis and consulting What You'll Do: Contribute to building the most modern analytics stack in Data today to support AI-driven self-serve analysis, key metrics tracking, and external customer reporting Influence work on new products like LLM Inference Endpoints through product analytics tracking Identify millions of dollars of cost savings and optimization across our tools and financial operations Write data pipelines that power the operatio

pythonsqlai
View job →
A
Abbott
📍 United States
12 days ago

Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: Position Overview The AI Platform Engineer builds and operates the machine learning and generative AI platform used by teams across Abbott Cancer Diagnostics. You'll own the full model lifecycle in production — data and feature pipelines, training and experimentation, evaluation and promotion, serving, and monitoring — along with the platform services, compute and tooling underneath it. This is hands-on infrastructure work backed by solid platform engineering practice: making inference fast and cheap, making the path from experiment to production repeatable and auditable, and shipping interfaces other engineers can build on — in support of software that ultimately reaches patients. Essential Duties Include, but are not limited to, the following: Build and maintain data, feature, and training pipelines for ML and LLM workloads — ingestion, transformation, fine-tuning, distributed training, and reproducible experiment execution with lineage tracked from dataset and code to resulting model. Implement automated evaluation and promotion gates — performance benchmarks, regression checks, and validation criteria that determine whether a model advances toward production. Automate the model lifecycle end to end through CI/CD and GitOps: packaging, promotion across environments, progressive rollout, and rollback. Build and operate production model-serving infrastructure for LLMs and predictive models, including inference optimization, autoscaling,

pythonjavaaws
View job →
T
Tenstorrent
📍 Toronto• $100K – $500K/yr
13 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As a Datacenter Liquid Cooling Architect, you will define, design, and architect next-generation liquid cooling infrastructure for Tenstorrent’s large-scale AI training and inference clusters. You will partner with systems engineering, mechanical engineering, software, and cross-functional design teams to develop chassis-, rack-, and cluster-scale cooling solutions, including CDU integration, telemetry and control, leak detection, and resilient operating strategies. This role will help shape reliable AI datacenter architectures and deployments for both internal and external customers. This role is on-site, based out of Toronto, Canada, Austin, Texas or Santa Clara, California. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A datacenter and system thermal design professional with 10+ years of experience architecting cooling infrastructure for complex computing environments. An experienced liquid cooling architect who can design chassis- and rack-scale solutions for large AI training and inference clusters. A systems thinker who understands how mechanical, electrical, software, facility, and systems engineering decisions come toge

N
13 days ago

We are looking for an enthusiastic software engineer to join our AI networking acceleration team, to work on a groundbreaking open-source library, using hardware offloads, GPU Kernels and RDMA network cards. Our product is a performance-oriented low-level infrastructure, crafted to change the way inference works. We thrive as a team in a deeply strong environment, and we're passionate about innovation. The rewards are sweet and include working with some of the brightest people in the industry, an aggressive compensation plan that rewards top performers, and the opportunity to collaborate on products that transform daily the way people work and play. What you'll be doing: Developing a highly optimized inference framework Running on the world’s largest supercomputers and data centers. The work environment is dynamic and challenging as our employees work on innovative, next-generation products at the forefront of technology in terms of performance, scalability, and features. What we need to see: B.Sc. or equivalent experience in Computer Science or Software Engineering 6&#43; years of experience in modern C&#43;&#43; / C / Rust development 3 years of experience in Linux environment and familiarity with development tools Deep knowledge of the TCP/IP network stack Understanding of computer architecture and operating systems concepts Ways to stand out from the crowd: Background in Linux internals and low-level software optimizations (benchmarking, bottleneck research, performance tuning) Experience in programming CUDA kernels is an advantage Familiarity with ML frameworks and LLMs Background in parallel programming / high-performance computing / RDMA t

REMOTElinuxai
View job →
N
13 days ago

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. NVIDIA is seeking best-in-class ASIC Verification Engineers to verify the design and implementation of the world’s leading inference accelerator. This position offers the opportunity to have real impact in a dynamic, technology-focused company impacting product lines ranging from consumer graphics to self-driving cars and the growing field of artificial intelligence. We have crafted a team of outstanding people around the globe. Their mission is to push the frontiers of what is possible today and define the platform for the future of computing. In this position, you will help to build the high-performance processor elements that implement programmable compute and graphics functionality. What you'll be doing: As a key member of our ASIC Verification team, you will verify the design and implementation of inference accelerator You will be responsible for verification of the ASIC design, architecture, reference models and micro-architecture using advanced verification methodologies Understand the design and implementation of your unit, define the verification scope, develop the verification infrastructure and verify the correctness of the design Collaborate with architects, designers, and pre and post silicon verifi

pythonmachine learningartificial intelligence
View job →
NS
NK Securities Research
📍 Gurugram• Full-time
18 days ago

NK Securities Research is a leading financial firm that leverages cutting-edge technology and sophisticated algorithms to trade the financial markets. Founded in 2011, we have gained invaluable experience in the field of High-Frequency Trading (HFT) across different asset classes. Role Overview We’re looking for engineers who can take AI work beyond experiments and make it hold up in production. You’ll work closely with quant researchers and infra engineers to build AI systems that actually get used improving research speed and internal tooling without slowing down the core stack. We value engineers who think about trade-offs, test what they build, and care about how things run in production. What You’ll Build Production AI Ship models that meet defined latency and reliability expectation Add monitoring, rollback, and guardrails before anything goes live Optimise inference across CPU/GPU environments when it matters Integration into Real Systems Plug AI into data-heavy workflows without hurting performance Work within existing low-latency architecture instead of fighting it Profile and remove bottlenecks rather than guessing AI for Engineers & Researchers Build tools that genuinely speed up research and development Improve code understanding, review workflows, and internal knowledge retrieval Keep systems auditable and predictable LLM & Retrieval Systems Implement structured RAG and embedding pipelines with validation in place Create safe integration layers between models and internal systems Performance & Standards Track latency, drift, and stability — not just accuracy Build observability into everything you ship Help raise the bar for how AI is engineered here What We’re Looking For Strong Python fundamentals Clear thinking around system design and performance trade-offs Experience deploying AI systems in production (1–5 years is typical) Familiarity with transformers, embeddings, or LLM deployment Nice to have: Exposure to C++ / Rust / Go E

pythonaic++
View job →
J
Jumio
📍 India• Full-time• Remote
18 days ago

Role Purpose: At Jumio, you will work for one of the market leaders in the global identity verification space that is helping to make the digital world a safer place for everyone. As a Software Development Engineer in the MLOpsTeam, you will develop the blueprint for highly scalable and performant ML model serving. Role Value: As a Software Engineer (SDE III), you will drive the continuous improvement of the infrastructure and applications to manage the lifecycle of ML assets (data, models) to better developer experience and strengthen governance capabilities. Secondly, you will design and implement robust ML infrastructure for model deployment, serving, and optimization. You will work on efficient CI/CD pipelines for ML models and leverage advanced compilers or hardware optimization to maximize inference performance while optimizing costs. We welcome you to challenge us to impact our software development processes and tools. Example Responsibilities: Upgrade ML assets (models, data) management systems for better developer experience and robust governance capabilities Build and optimize model serving infrastructure with a focus on inference latency and cost optimization Architect efficient inference pipelines that balance latency, throughput, and cost across various acceleration options Implement cost-efficient, enterprise-scale solutions Collaborate in a cross-functional, distributed team for continuous system improvement Work with MLEs, QA Engineers, and DevOps Engineers Evaluate and implement new technologies and tools Contribute to architectural decisions for distributed ML systems Experience and Qualifications : 5+ years of experience in software engineering with Python Experience with model lifecycle management (MLFlow, Weights & Biases or equivalent) Experience with data management ecosystem (quality, transformation, catalog) Experience with ML frameworks, particularly PyTorch Experience optimizing ML models with hardwar

REMOTEpythonawsdocker
View job →
GR
18 days ago

Role: VPN Engineer Location: Gurgaon Graviton is a privately funded quantitative trading firm striving for excellence in financial markets' research. We are seeking a Network Engineer for our team in Gurgaon. Graviton trades across a multitude of asset classes and trading venues using a gamut of concepts and techniques ranging from time series analysis, filtering, classification, stochastic models, pattern recognition to statistical inference analysing terabytes of data to come up with ideas to identify pricing anomalies in financial markets. Responsibilities Manage and support corporate network infrastructure across multiple locations, including routers, firewalls, switches, wireless access points, VPN gateways, Internet links, and LAN/WAN connectivity. Configure and troubleshoot VPN technologies such as IPsec, SSL VPN, site-to-site VPN, remote-access VPN, WireGuard, OpenVPN, and FortiClient/FortiGate VPN, including issues related to authentication, tunnels, routing, DNS, packet loss, performance, split tunnelling, and firewall policies. Manage secure connectivity between offices, data centres, cloud environments, and remote users. Configure and maintain office LAN infrastructure, including VLANs, trunk/access ports, inter-VLAN routing, DHCP, DNS, NAT, ACLs, static routing, BGP, and OSPF where required. Manage multiple ISP connections, including primary and backup Internet links, automatic failover, and monitoring of utilization, latency, jitter, packet loss, and link availability. Coordinate with ISPs and telecom providers for new circuits, link failures, bandwidth upgrades, routing issues, packet-loss investigations, and service escalations. Manage firewall policies, NAT rules, VPN policies, network objects, and routing, while regularly reviewing and removing unnecessary access. Implement network segmentation across user, server, management, guest, and other business networks, while maintaining secure administrative access to network equip

pythonlinuxrest
View job →
GR
18 days ago

Description: Graviton is a privately funded quantitative trading firm striving for excellence in financial markets' research. We are seeking a Quant Analyst - Risk for our team in Gurugram. Graviton trades across a multitude of asset classes and trading venues using a gamut of concepts and techniques ranging from time series analysis, filtering, classification, stochastic models, pattern recognition to statistical inference analysing terabytes of data to come up with ideas to identify pricing anomalies in financial markets. As a Quant Analyst - Risk you will be responsible Work as a team with senior traders to operate and implement/improve our automated trading strategies. Analysing production trades and developing ideas to improve our trading strategies. Implement monitoring tools which highlight potential issues in the production strategies. Write comprehensive and scalable scripts in both C++ and python analysing production strategies for risk attribution, performance break-ups along various buckets and so on. Build ‘cool’ scalable post-trade systems analysing multitude of statistics across all production strategies. Implementing tools for analysing Market Data centrally across various exchanges. Managing deployments and release cycle, with working along with a senior trader. Requirements : Possess a degree in a highly analytical field, such as Engineering, Mathematics, or Computer Science from top-ranked universities 3+ years of experience in Python, Shell/Bash scripting. Basic knowledge of Linux and shell command-line tools Basic programming and scripting (Python/Shell) skills Strong problem-solving, and analytical skills Excellent communication skills Have a strong work ethics Mentorship experience in guiding junior developers. Benefits: Our open and collaborative work culture gives you the freedom to innovate and experiment. Our cubicle free offices, non-hierarchical work culture and insistence to hire the very best creates a melting pot for great idea

pythonlinuxc++
View job →
GR
18 days ago

Who are we: Graviton is a privately funded quantitative trading firm striving for excellence in financial markets research. We trade across a multitude of asset classes and trading venues using a gamut of concepts and techniques ranging from time series analysis, filtering, classification, stochastic models, pattern recognition, to statistical inference analyzing terabytes of data to come up with ideas to identify pricing anomalies in financial markets. As part of this team you will be tasked to apply machine learning and specifically deep learning techniques to trading problems while staying connected to broader research community. The researcher will put theory into practice and can immediately impact the global trading landscape with the expanding presence of Graviton in various markets. Description Lead research in applying machine learning to a wide variety of datasets and trading problems Follow latest developments in academic research and incorporating research techniques from different fields of applications to our problems Improve tick-by-tick order book based time series feature sets using latest preprocessing techniques Work on current and develop new deep learning models to exploit large pool of in-house features and computing infrastructure Develop scalable pipeline for building predictive models across global markets Discover and implement new sources of predictive alpha, verify that they improve existing models, and integrate them into the firm's strategy development pipeline Partner with quant researchers and software developers in implementation of conducted research to production using Python / C++ Advise infrastructure support team on latest developments on hardware and software to improve computing infrastructure for ML based research Qualifications Masters or PhD in Computer Science, Mathematics, Statistics, or a related field At least two years of demonstrated experience of ML/AI research in a professional setting or at a repu

pythonmachine learningai
View job →
🔔

Get new inference technical lead jobs by email

Daily job updates · Unsubscribe anytime