Jobs in United States

Inference Technical Lead in United States

672 active opportunities · Updated October 2026

Explore current inference technical lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%
Quick readStrong listing-quality and freshness signals

Datadog's integrations are the connective tissue between our platform and the technologies our customers run in the real world. As a Sr. PM on the Agent Integrations team, you will own the vision, prioritization, and execution for 100+ integrations that run directly inside the Datadog Agent from foundational infrastructure (MySQL, Kafka, Kubernetes) to the rapidly growing landscape of self-hosted AI and on-premise enterprise technologies. This is a high-impact, breadth-first role at the intersection of infrastructure observability and the frontier of AI-native workloads. At Datadog, we place value in our office culture; the relationships it builds, the creativity it brings, and the collaboration of being together. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Own the Agent Integrations roadmap. Determine which new integrations to build and which existing ones to improve, balancing customer demand, business impact, and engineering capacity across a catalog of 100+ technologies. Drive the expanding AI integration surface. Lead product strategy for self-hosted AI workloads, including LLM inference frameworks (e.g., Hugging Face TGI, BentoML), AI agents, MCP servers, and model orchestration tools, so Datadog customers can monitor every layer of their AI stack. Expand on-prem and hybrid coverage. Prioritize and execute new integrations for on-prem technologies including storage systems, HPC schedulers, network devices, and legacy enterprise platforms where customers run critical workloads. Build observability for ERP systems. Define and drive Datadog's strategy for monitoring enterprise ERP platforms (SAP, Oracle EBS/Fusion, Microsoft Dynamics) covering performance, job execution health, and integration layer telemetry so enterprise customers can observe their ERP stack alongside the rest of their infrastructure. Analyze adoption and customer feedback at scale. Use data from multiple sources to

SQLPostgreSQLMySQLMongoDB
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI's Training team is responsible for producing the large language models that power our research, our products, and ultimately bring us closer to AGI. Achieving this goal requires combining deep research into improving our current architecture, datasets and optimization techniques, alongside long-term bets aimed at improving the efficiency and capability of future generations of models. We are responsible for integrating these techniques and producing model artifacts used by the rest of the company, and ensuring that these models are world-class in every respect. Recent examples of artifacts with major contributions from our team include GPT4-Turbo, GPT-4o and o1-mini. About the Role As a member of the architecture team, you will push the frontier of architecture development for OpenAI's flagship models, enhancing intelligence, efficiency, and adding new capabilities. Ideal candidates have a deep understanding of LLM architectures, a sophisticated understanding of model inference, and a hands-on empirical approach. A good fit for this role will be equally happy coming up with a creative breakthrough, investing in strengthening a baseline, designing an eval, debugging a thorny regression, or tracking down a bottleneck. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, prototype and scale up new architectures to improve model intelligence Execute and analyze experiments autonomously and collaboratively Study, debug, and optimize both model performance and computational performance Contribute to training and inference infrastructure You might thrive in this role if you: Have experience landing contributions to major LLM training runs Can thoroughly evaluate and improve deep learning architectures in a self-directed fashion Are motivated by safely deploying LLMs in the real world Are well-versed in the state of the art tran

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Workload team is responsible for designing and running OpenAI’s LLM training and inference infrastructure that powers frontier models at massive scale. Our systems unify how researchers train and serve models, abstracting away the complexity of performance, parallelism, and execution across vast GPU/accelerator fleets. By providing this foundation, the Workload team ensures that researchers can focus on advancing model capabilities while we handle the scale, efficiency, and reliability required to bring those models to life. About the Role We are looking for an engineer to design and implement the dataset infrastructure that powers OpenAI’s next-generation training stack. You will be responsible for building standardized dataset interfaces, scaling pipelines across thousands of GPUs, and proactively testing performance bottlenecks. In this role, you will collaborate closely with the multimodal researchers, and other infra groups to ensure datasets are unified, efficient, and easy to consume. In this role, you will: Design and maintain standardized dataset APIs, including for multimodal (MM) data that cannot fit in memory. Build proactive testing and scale validation pipelines for dataset loading at GPU scale. Collaborate with teammates to integrate datasets seamlessly into training and inference pipelines, ensuring smooth adoption and a great user experience. Document and maintain dataset interfaces so they are discoverable, consistent, and easy for other teams to adopt. Establish safeguards and validation systems to ensure datasets remain reproducible and unchanged once standardized. Debug and resolve performance bottlenecks in distributed dataset loading (e.g., straggler systems slowing global training). Provide visualization and inspection tools to surface errors, bugs, or bottlenecks in datasets. You might thrive in this role if you: Have strong engineering fundamentals with experience in distributed systems, data pipelines, or infrastructure.

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,

AWSKubernetesLinuxRest
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model. About the Role We’re looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You’ll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production — directly improving how quickly and reliably research moves forward. In this role, you will: Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance. Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention. You might thrive in this role if you: Are excited about the potential of personal AGI and want to build the infrastructure that enables it. Have strong software engineering fundamentals and experience building or operating large-scal

AWSRestAIRust
A
📍 San Francisco, CA, United States
✓ Quality checkedCompany trend -100%

The Opportunity Typography is central to how ideas are communicated. If you're passionate about beautifully created design, have deep curiosity about what AI can do, and take personal responsibility for creating products that people love; then this may be the role for you. Adobe Fonts supports millions of creatives in choosing and using typefaces across fonts.adobe.com, Express, Photoshop, Illustrator, Acrobat, and more Creative Cloud platforms. Our Internal Services team provides the platform engineering and deployment backbone for all of these. We manage CI/CD, deployment approaches, the services and data layers our engineers depend on, our observability and security stance, and increasingly the agentic tools that transform how our entire organization delivers software. We're seeking a Senior Software Development Engineer to lead this exciting journey in our San Francisco location. What you'll do Own and evolve our deployment platform. Lead strategy for CI/CD, PR environments, and release safety across a mixed fleet that includes containerized services, serverless services, and static front ends. Build the foundation for AI-accelerated development. Help build our agent factory and grow our internal agentic toolkit and skill library. Ship inference applications at scale. Take greenfield services from spec to production and standardize our ML/inference footprint. Modernize our services for the AI era. Identify where an existing service is held back by its current build and lead the fix. Rethink our security posture for agentic threats. Lead how we secure autonomous agents and their tool use. Expose Adobe Fonts to the agentic ecosystem. Extend our Model Context Protocol (MCP) surface and conversational, intent-based font discovery. Work higher up the stack, too. Contribute directly to search, browse, discovery, and the customer-facing experie

JavaScriptTypeScriptPythonReact
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

PythonAWSLinuxRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

AI Systems Engineer - Codex Core Agents About The Team The Codex Core Agents team builds the agent harness that turns model capability into real-world action. We own the systems around the model: prompting and interpreting model outputs, executing actions safely in real environments, and feeding production experience back into better models and better agent behavior. This team sits close to research and works across the stack: harness, model interaction, inference, sandboxed execution, orchestration, evals, production reliability, and the performance envelope around tokens, latency, cost, capacity, and quality. The harness is open source and increasingly part of how models are trained and evaluated, making this one of the highest-leverage layers in Codex. About The Role We’re looking for engineers to build the AI systems that make Codex agents dependable in production. The ideal candidate is an agent-systems builder: hands-on across low-level systems and ML workflows, able to debug Codex behavior end to end across the harness, model behavior, inference/runtime stack, GPU fleet, and product surface. You’ll work with research, infrastructure, and product to design agent harness capabilities, run experiments and ablations across the model + system prompt + harness stack, build frameworks for assessing production agent performance, and turn messy failures into durable improvements. What You’ll Do Design and build the core agent harness and execution loop that lets Codex agents interpret model outputs, use tools, execute code, and complete long-horizon tasks safely. Build sandboxing, isolation, orchestration, state, and workflow infrastructure for agents operating in real development environments. Develop evaluation, experimentation, and debugging systems that distinguish harness issues, model behavior, inference/runtime issues, and product failures. Run ablations across prompts, model-facing interfaces, context construction, tool-use strategies, and harness behavior to

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team We’re hiring software engineers to make the Workload team more productive. The Workload team maintains the core components of OpenAI’s training and inference frameworks and helps execute frontier experiments. About the Role We’re looking for someone who cares about the developer experience of working in and around OpenAI’s core training and inference frameworks. In this role you will: Be responsible for optimizing the development workflows of the engineers around you Work within various Workload teams to address their specific needs, but collaborate with the centralized teams that own various aspects of development experience Optimize iteration speed, both broadly, and in particular by optimizing specific teams’ CI Improve reliability, for instance, by driving testing strategy for particular components Work through the long tail of things that it takes to build libraries and systems that will delight researchers You might thrive in this role if: You are motivated by helping people. You believe a thing that separates great teams from good teams are the players willing to do whatever work it takes, without ego. You believe in the power of developer experience. Something magical happens when people can quickly and confidently iterate on a simple codebase, but this magic is fragile and must be fought for. When you see someone trip over something, no matter how small, your first instinct is asking yourself what it would take for that to not happen again. Your second instinct is clicking merge on the PR you’ve already written to make it so. You are pragmatic. You have the ability to see the world through a perfectionist’s eyes, but are not yourself a perfectionist. You know which problems to pick and when to switch to making progress on a different problem. You like going end-to-end on things. You love co-design — that feeling when you were only able to find the right solution because you both deeply understand the users that interact with a system and the

PythonAWSRestAgile
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team We’re hiring software engineers to make OpenAI’s networking teams more productive. These teams build and operate the high-performance networking systems that support OpenAI’s training and inference infrastructure at frontier scale. About the Role We’re looking for someone who cares deeply about the developer experience of engineers working on complex infrastructure systems — especially around build systems, test architecture, release pipelines, and reliable development workflows. This role will be embedded with OpenAI’s networking team: making it faster, safer, and easier for engineers to build, test, validate, and ship changes across multi-server, networked, and hardware-adjacent environments. In this role you will: Improve development workflows for engineers building and operating OpenAI’s networking systems Design and improve continuous deployment, release, and validation pipelines Build and maintain test harnesses for multi-server, networked, and hardware-backed environments Improve iteration speed across C++, Python, and build-system-heavy codebases Partner with engineers to identify friction in CI, testing, debugging, and deployment workflows Drive testing and reliability strategy for infrastructure components that support large-scale training and inference workloads Work closely with centralized developer experience teams while staying deeply embedded with the networking engineers closest to the systems You might thrive in this role if: You are motivated by helping other engineers move faster and with more confidence You have experience with CI/CD, release pipelines, testing infrastructure, or build systems You are comfortable moving between C++, Python, and build systems such as CMake, Bazel, or Blaze You enjoy building test harnesses, automation, and workflow improvements for complex systems You do not need to be a networking expert, but you are excited to learn enough about the domain to make the team meaningfully more effective When you see

PythonAWSCI/CDRest
MT
📍 San Jose, California, Canada
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Job Summary We are seeking a motivated engineer to join the DRAM Systems Engineering team, focusing on the development, evaluation, and optimization of next-generation memory systems for AI accelerators. This role emphasizes research and development across hardware architecture, operating systems, and performance analysis to support Agentic AI inference workloads. If you are ambitious and eager to make an impact in the exciting world of AI and memory systems, this is the perfect opportunity for you! Responsibilities Characterize AI inference workloads and examine memory behavior Build and evaluate tiered memory hierarchies for AI accelerators Study KV cache lifecycles, MoE models, and data placement strategies Compare and optimize explicit versus hardware-assisted data movement Develop, test, debug, and detail system-level and OS components Prototype and evaluate agentic AI systems by building agents and multi-agent workflows using modern frameworks and orchestration patterns (planning, tool use, memory, and context management). Apply these technologies both as workloads under study and as accelerators for internal engineering workflows <h2 style="color:!importan

PythonLinuxAIRecruitment
M
📍 New York, new york, United States· Full-time
✓ High-confidence listingCompany trend -67.9%
Quick readStrong listing-quality and freshness signals

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. About Modal Data: We’re growing our Data team and are looking for our first few key hires to build self-serve data tools and drive business strategy in the right direction. The mission of the Modal Data team is to make it easy to track company goals, make evidence-backed decisions, and prioritize the right work. We do this via: Self-serve AI analytics tools (Hex, Snowflake) Embedding with teams as a “data adviser”, providing strategic analysis and consulting What You'll Do: Contribute to building the most modern analytics stack in Data today to support AI-driven self-serve analysis, key metrics tracking, and external customer reporting Influence work on new products like LLM Inference Endpoints through product analytics tracking Identify millions of dollars of cost savings and optimization across our tools and financial operations Write data pipelines that power the operatio

PythonSQLAIProject Management
A
📍 United States
✓ High-confidence listingCompany trend +365.2%
Quick readStrong listing-quality and freshness signals

Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: Position Overview The AI Platform Engineer builds and operates the machine learning and generative AI platform used by teams across Abbott Cancer Diagnostics. You'll own the full model lifecycle in production — data and feature pipelines, training and experimentation, evaluation and promotion, serving, and monitoring — along with the platform services, compute and tooling underneath it. This is hands-on infrastructure work backed by solid platform engineering practice: making inference fast and cheap, making the path from experiment to production repeatable and auditable, and shipping interfaces other engineers can build on — in support of software that ultimately reaches patients. Essential Duties Include, but are not limited to, the following: Build and maintain data, feature, and training pipelines for ML and LLM workloads — ingestion, transformation, fine-tuning, distributed training, and reproducible experiment execution with lineage tracked from dataset and code to resulting model. Implement automated evaluation and promotion gates — performance benchmarks, regression checks, and validation criteria that determine whether a model advances toward production. Automate the model lifecycle end to end through CI/CD and GitOps: packaging, promotion across environments, progressive rollout, and rollback. Build and operate production model-serving infrastructure for LLMs and predictive models, including inference optimization, autoscaling,

PythonJavaAWSKubernetes
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. NVIDIA is seeking best-in-class ASIC Verification Engineers to verify the design and implementation of the world’s leading inference accelerator. This position offers the opportunity to have real impact in a dynamic, technology-focused company impacting product lines ranging from consumer graphics to self-driving cars and the growing field of artificial intelligence. We have crafted a team of outstanding people around the globe. Their mission is to push the frontiers of what is possible today and define the platform for the future of computing. In this position, you will help to build the high-performance processor elements that implement programmable compute and graphics functionality. What you'll be doing: As a key member of our ASIC Verification team, you will verify the design and implementation of inference accelerator You will be responsible for verification of the ASIC design, architecture, reference models and micro-architecture using advanced verification methodologies Understand the design and implementation of your unit, define the verification scope, develop the verification infrastructure and verify the correctness of the design Collaborate with architects, designers, and pre and post silicon verifi

PythonMachine LearningArtificial IntelligenceAI
M
📍 New York, new york, United States· Full-time
✓ High-confidence listingCompany trend -67.9%
Quick readStrong listing-quality and freshness signals

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. About Talent at Modal Modal is growing fast, and the programs that bring people in and set them up to succeed are still being built. You'll join the Talent team as one of its first hires focused purely on programs; working closely with recruiting and leadership to build the events, internship, and campus presence that shape how the best people discover and experience Modal for the first time. The Role As Talent Programs Manager, you will own Modal's talent events, our intern program, and our presence at career fairs, end-to-end. This is a build-from-the-ground-up role for someone who wants full ownership rather than an existing playbook to execute. You'll work directly with recruiters, hiring managers, and marketing to make sure every program ladders up to real hiring outcomes, and you'll be the person who makes candidates' and interns' first experience of Modal a great one.

AIGoExcelMarketing
🔔

Get new inference technical lead jobs in United States by email

Daily job updates · Unsubscribe anytime