Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an
Jobiba hiring network
Inference Technical Lead Jobs
1,491 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current inference technical lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Client Onboarding Director, Inference and Agentic AI Location: Noida Company: Paytm About Paytm Paytm is a pioneer of digital payments in India, serving over 450 million consumers and 45 million merchants across payments, financial services, and commerce. Over the years, Paytm has built deep in-house capabilities across technology, data, and operations to operate at scale with high reliability. Paytm is building a full stack AI platform focussed on Inference and Agents, enabling large enterprises to deploy AI driven automation across sales, service, operations, and analytics. The Inference and Agentic AI team operates as a cross functional unit spanning engineering, product, data science, business management, and sales, and owns the full lifecycle of AI solutions from opportunity discovery to deployment and scale. Role Overview Paytm is looking to hire a Client Onboarding Director to own implementation delivery and client onboarding governance for Paytm’s AI Inference and Agentic AI products across enterprise clients. This is a managerial role that will lead Client Onboarding Managers and ensure that enterprise deployments move smoothly from sales closure to go live and early adoption. The role sits at the intersection of client teams, product, engineering, business, and sales, and is responsible for converting signed enterprise deals into successful, timely, and scalable deployments. The candidate will own delivery planning, integration governance, risk management, stakeholder communication, and post go live stabilization across enterprise AI agent deployments. The role requires strong program management, technical understanding, client handling, and ability to drive execution across multiple internal and external teams. Key Responsibilities Delivery Ownership and Implementation Governance Own end to end delivery governance for enterprise AI agent deployments from sales handoff to go live and stabilization. Lead the implementation planning process across scope
Ready to help developers and researchers get more from open AI models? At NVIDIA, our team improves support for leading community models such as Nemotron, Llama, Gemma, DeepSeek, and Qwen. As a Product Manager for Open Models, you will help coordinate model enablement, developer experiences, technical content, and product launches across NVIDIA’s accelerated computing platform. You will work alongside experienced product managers, engineers, model builders, and product marketing teams. This role is ideal for someone early in their product-management career who has a strong technical foundation, enjoys working across teams, and is excited about the open-model ecosystem. What you’ll be doing: Support collaboration with community model builders around model access, technical enablement, launch readiness, and go-to-market activities. Track emerging open-model releases and summarize their capabilities, technical differentiators, hardware requirements, and ecosystem impact. Maintain product requirements, launch plans, readiness checklists, and supporting documentation for assigned models. Partner with engineering, infrastructure, and developer-experience teams to support open models across inference, fine-tuning, evaluation, and deployment. Gather feedback from model builders and developers, identify recurring issues, and translate findings into actionable product requirements. Review and test developer workflows, sample applications, notebooks, and Python code used in demonstrations and technical content. Collaborate with product marketing on blogs, documentation, presentations, case studies, and social-media content. Support product announcements, demonstrations, keynote materials, and other high-visibility launches while working onsite at least three days per week. What we need to see: 2+ years of relevant experience in product ma
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our
Role Purpose We’re looking for a Staff/Senior Machine Learning Engineer with deep expertise in computer vision and biometrics to lead the design and scaling of face recognition systems in production. You’ll build and train models, and own ML systems end-to-end on AWS. The final job level for this role will be determined following the interview process. What You’ll Do Lead the design and development of computer vision systems for biometrics (face attributes, detection, quality, and recognition) Rigorous fairness analysis and benchmarking of biometric models across various datasets and operating conditions. Architect, train, and optimize models using PyTorch, Tensorflow, and/or JAX Own and evolve end-to-end ML pipelines, from data ingestion to deployment. Design automated pipelines (Airflow) for data ingestion and cleaning. You will be responsible for curating balanced training sets and generating synthetic data to address both quality and diversity gaps. Production Engineering: Own the path to production. Optimize models for low-latency inference (quantization, distillation, TensorRT/ONNX) and manage deployment on AWS. Mentor ML engineers, conduct code/design reviews, and drive technical best practices across the Computer Vision team. What We’re Looking For Experience: 5+ years of industry experience in Machine Learning, with at least 3 years dedicated to Biometrics or Face Analysis. Deep expertise in computer vision and biometrics, especially face recognition. Fairness & Ethics: You understand the sources of algorithmic bias in Computer Vision and have practical experience measuring and mitigating disparate impact. Strong Engineering: Expert proficiency in Python (both machine learning and vision libraries such as Pillow, OpenCV, PyTorch, etc). You write clean, modular, production-ready code. Systems Architecture: Experience designing end-to-end ML pipelines (Data to Train to Deploy) and working with workflow orchestrators like Airflow. Cloud Native: Hands-on ex
Job Details: Job Description: This is a high-visibility, commissioned sales leadership role within Intel's US Sales organization, specifically focused on our most disruptive AI-Native and Strategic CSP accounts. You will be the primary architect of Intel's relationship with industry titans who are redefining the boundaries of AI model training, AIaaS solutions and deployment at scale. This is not a traditional sales role. You will operate at the intersection of deep technical engineering and executive business strategy, ensuring Intel's silicon and software roadmap aligns with the world's most demanding AI-as-a-Service and SaaS platforms across on-prem, Tier1 CSP and NeoCloud environments. Key Responsibilities Executive Orchestration: Act as the One Intel lead, building deep-rooted partnerships with C-suite executives and Principal Engineers at world-class AI and SaaS firms. Technical Value Synthesis: Translate complex hardware architectures (CPU, GPU, Accelerator, Networking, and Packaging) into business outcomes for customers running massive-scale distributed training and inference workloads. Strategic Growth: Drive Intel's data-centric growth strategy by identifying and securing design wins within the core infrastructure of the world's leading AI models and solution providers. Cross-Functional Leadership: Partner closely with Cloud Solution Architects (CSAs), Intel Business Units, Cloud and OEM partners to influence future product roadmaps based on the unique needs of AI-native disruptors. Market Evangelism: Serve as a technical and business evangelist, articulating Intel's vision for the future of AI and compute infrastructure in a highly competitive landscape. <p style="text-align:inhe
About the Role & Team We’re looking for an Engineering Manager to lead the Data Infrastructure team within Statsig Experiment at Amplitude. You will lead a multidisciplinary team of software engineers, data engineers, and data scientists responsible for the systems that power experimentation at scale. The team owns three critical areas: Data ingestion: Collecting and importing experiment exposures, custom events, OpenTelemetry data, and real user monitoring data across SDKs, streaming systems, cloud storage, and customer data warehouses. Data computation: Building distributed computation systems that transform raw data into accurate, timely experiment results. Stats engine: Developing and productionizing the statistical methods that help customers make trustworthy decisions from their experiments. This is not a traditional data engineering management role. We are looking for a leader with a solid data science and statistical foundation who can connect advances in experimentation methodology with scalable production systems. You will help set our technical and scientific direction, translating new statistical methods and machine learning research into capabilities that customers can use reliably at scale. You’ll partner closely with data scientists, engineers, product managers, and customers to advance the state of experimentation. The ideal candidate is equally comfortable discussing causal inference and statistical power with data scientists, distributed computation architectures with engineers, and experimentation strategy with customers. What You’ll Do Lead and grow the team responsible for Statsig’s data ingestion, experiment computation, and stats engine. Define the technical and scientific strategy for advancing experimentation across both Statsig Cloud and warehouse-native deployments. Partner with data scientists and engineers to turn new statistical and causal inference methods into scalable, reliable product capabilities. Evolve our data and computatio
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
Machine Learning Engineer We’re looking for a Machine Learning Engineer with deep expertise in computer vision and biometrics to lead the design and scaling of face recognition systems in production. You’ll build and train models, and own ML systems end-to-end on AWS. The final job level for this role will be determined following the interview process. What You’ll Do Lead the design and development of computer vision systems for biometrics (face attributes, detection, quality, and recognition) Rigorous fairness analysis and benchmarking of biometric models across various datasets and operating conditions. Architect, train, and optimize models using PyTorch, Tensorflow, and/or JAX Own and evolve end-to-end ML pipelines, from data ingestion to deployment. Design automated pipelines (Airflow) for data ingestion and cleaning. You will be responsible for curating balanced training sets and generating synthetic data to address both quality and diversity gaps. Production Engineering: Own the path to production. Optimize models for low-latency inference (quantization, distillation, TensorRT/ONNX) and manage deployment on AWS. Mentor ML engineers, conduct code/design reviews, and drive technical best practices across the Computer Vision team. What We’re Looking For Experience: Multiple years of industry experience in Machine Learning, with at least 3 years dedicated to Biometrics or Face Analysis. Deep expertise in computer vision and biometrics, especially face recognition. Fairness & Ethics: You understand the sources of algorithmic bias in Computer Vision and have practical experience measuring and mitigating disparate impact. Strong Engineering: Expert proficiency in Python (both machine learning and vision libraries such as Pillow, OpenCV, PyTorch, etc). You write clean, modular, production-ready code. Systems Architecture: Experience designing end-to-end ML pipelines (Data to Train to Deploy) and working with workflow orchestrators like Airflow. Cloud Native:
Machine Learning Engineer IV – (Computer Vision) We’re looking for a Staff/Senior Machine Learning Engineer with deep expertise in computer vision and biometrics to lead the design and scaling of face recognition systems in production. You’ll build and train models, and own ML systems end-to-end on AWS. The final job level for this role will be determined following the interview process. What You’ll Do Lead the design and development of computer vision systems for biometrics (face attributes, detection, quality, and recognition) Rigorous fairness analysis and benchmarking of biometric models across various datasets and operating conditions. Architect, train, and optimize models using PyTorch, Tensorflow, and/or JAX Own and evolve end-to-end ML pipelines, from data ingestion to deployment. Design automated pipelines (Airflow) for data ingestion and cleaning. You will be responsible for curating balanced training sets and generating synthetic data to address both quality and diversity gaps. Production Engineering: Own the path to production. Optimize models for low-latency inference (quantization, distillation, TensorRT/ONNX) and manage deployment on AWS. Mentor ML engineers, conduct code/design reviews, and drive technical best practices across the Computer Vision team. What We’re Looking For Strong industry experience in Machine Learning, dedicated to Biometrics or Face Analysis. Deep expertise in computer vision and biometrics, especially face recognition. Fairness & Ethics: You understand the sources of algorithmic bias in Computer Vision and have practical experience measuring and mitigating disparate impact. Strong Engineering: Expert proficiency in Python (both machine learning and vision libraries such as Pillow, OpenCV, PyTorch, etc). You write clean, modular, production-ready code. Systems Architecture: Experience designing end-to-end ML pipelines (Data to Train to Deploy) and working with workflow orchestrators like Airflow. Cloud Native: Hands-on exper
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Host Pricing & Settings team builds the platform and tools that help hosts run their business — with pricing strategies informed by market intelligence, comparable listings, and demand signals. We partner with Search, Listings, Tax, and Payments to ensure our guidance is accurate, timely, and trusted. Behind every pricing recommendation is a sophisticated ML system undergoing a fundamental rearchitecture. Our north star: a serving infrastructure where training, inference, and evaluation are consistent by design — features from a centralized store, model composition in one place, and backfills available on demand so data scientists and MLEs can evaluate candidates in days, not weeks. The Difference You Will Make: As a senior technical individual contributor, you will own the technical strategy for the full Modeling → ML Serving → API interface across the Host Pricing org. Although you will be at one of our highest levels of seniority, all individual contributors at Airbnb are Software Engineers — you are expected to be hands-on and contribute code. Define the architecture and contracts governing how models move from development to production — feature store design, model schema management, online/offline inference consistency, and multi-version support. Lead the buildout of a unified serving stack that eliminates per-model one-off implementations and gives data scientists a turnkey path from training to production. Architect backfill and evaluation infrastructure so the modeling team can simulate production inference over historical data in days, not weeks. Establish do
About the Team The Statsig team at OpenAI builds and operates the experimentation platform that powers product development, measurement, and decision-making across the company. We partner closely with product, engineering, and infrastructure teams to ensure experiments are trustworthy, statistically rigorous, and scalable to the needs of frontier AI products. Our mission is to help teams make better decisions through reliable experimentation. We care deeply about statistical correctness, pragmatic solutions, and building systems that researchers and engineers can trust at massive scale. The team operates at the intersection of experimentation methodology, data infrastructure, causal inference, and product analytics. We are looking for experienced experimentation experts who want to shape the future of experimentation in the AI era. About the Role We are hiring a Staff-level Data Scientist to help lead the evolution of OpenAI’s core experimentation platform. This role is focused on improving the statistical rigor, reliability, and practical usability of experimentation across the company. You’ll work on some of the hardest problems in online experimentation: sample ratio mismatch detection, variance reduction, bias mitigation, metric design, triggered analysis, heterogeneous treatment effects, sequential testing, and experimentation in complex ML systems. You’ll also help translate advanced statistical concepts into pragmatic systems and product experiences that teams can actually use. This is a highly technical individual contributor role with significant influence across methodology, platform architecture, and experimentation best practices. The ideal candidate combines deep statistical expertise with strong systems intuition and hands-on experience building or operating experimentation platforms at scale. In this role, you will: Drive the statistical direction and technical strategy for OpenAI’s experimentation platform Design and improve experimentation methodolo
About the Team OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical. The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly. About the Role We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity. In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity. This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions. This role is based in San Francisco, CA, with travel as needed. Key Responsibilities Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint Drive readiness to convert contracted compute capacity into schedulable production clusters Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives Build integrated schedules spanning procurement, logistics, installation, st
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking an experienced SoC Architect to lead the definition and development of next-generation custom AI silicon for edge deployments. This role will be responsible for shaping the architecture of highly efficient, high-performance SoCs optimized for machine learning inference and on-device intelligence. You will work cross-functionally with internal engineering teams and external ecosystem partners to translate product requirements into scalable silicon solutions, driving execution from concept through delivery. In this role you will: Define the architecture and technical roadmap for custom SoCs targeted for edge applications. Drive system-level tradeoff analysis across compute, memory, interconnect, power, thermal, and cost constraints. Architect energy-efficient ML compute subsystems optimized for inference workloads and real-world deployment environments. Collaborate with internal hardware, software, systems, and product teams to align architecture with platform needs. Partner with external silicon vendors, IP providers, and manufacturing partners to execute development plans. Lead hardware/software co-design efforts to maximize performance per watt and end-to-end system efficiency. Guide implementation teams through microarchitecture, RTL development, validation, and bring-up phases. Operate effectively in agile development environments and help teams deliver against aggressive schedules and milestones. You might thrive in this role if: Proven exper
Get new inference technical lead jobs by email
Daily job updates · Unsubscribe anytime