Jobs in United States

Inference Technical Lead in United States

672 active opportunities · Updated October 2026

Explore current inference technical lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

B
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for a Marketing Operations Manager who can own and harden the systems layer of Baseten's Marketing engine. Marketing at Baseten is scaling fast — more spend, more campaigns, more model launches, more inbound. The systems underneath (Our ESP, CRMs, forms, tracking, routing, alerting) need an owner who treats them like production infrastructure that cannot go down. When something breaks, it costs us time, pipeline, and trust in the data. This role exists so it doesn't break. You’ll simultaneously build for the future and re-think assumptions about our tech stack in the age of agents. This is an offensive play that gives the rest of the team leverage and superpowers to hit our ambitious goals. This is NOT an IT or service role. This is a core member of the marketing team who implements technology to achieve outcomes. RESPONSIBILTIES Own the marketing tech stack end-to-end: ad platforms, email systems, tracking, pixels, forms, connectors. Build defense-in-depth on inbound: spam/bot protection, rate limiting, email/domain validation, sync gating — and the alerting to catch anomalies before they hit sales or leadership dashboards. Enforce data integrity: UTM governance, campaign membership, lifecycle stages, lead scoring and routing logic, field-level hygiene, canonical metric definitions. Operationalize the web request pipeline with our dev agency: structured briefs, tickets, SLAs, and launch-day runb

PythonRestMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking an experienced and detail-oriented GRC (Governance, Risk, and Compliance) Manager to build, support, and continuously enhance Baseten’s security governance, compliance, and privacy programs. As one of the early members of our security organization, you will play a key role in ensuring our platform meets and exceeds the highest standards for privacy, trust, and regulatory compliance. In this role, you’ll work cross-functionally with engineering, operations, legal, and leadership teams to develop policies, manage audits, and implement controls aligned with frameworks such as SOC 2, ISO 27001, ISO 27701, and FedRAMP. You’ll be instrumental in building scalable processes to manage risk, support customer assurance, and uphold Baseten’s commitment to security and compliance as we grow. RESPONSIBILITIES Governance & Policy Development: Design, implement, and maintain security governance frameworks, policies, and procedures that align with Baseten’s risk posture and industry best practices. Risk Management: Build and manage the company-wide risk assessment program, identifying, tracking, and mitigating key security and compliance risks. Compliance Operations: Lead efforts to achieve and maintain compliance with SOC 2, ISO 27001/27701, HIPAA, FedRAMP and other applicable standards and regulations. Audit & Certification Management: Coordinate external audits and certification processes, ensuring e

AWSGCPMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for a Procurement Lead to build Baseten's procurement function from the ground up. This is a foundational, high-ownership role for someone who wants to define how a company runs procurement, not inherit an existing process. As the company scales across engineering, GTM, and corporate functions, we need to build purchasing workflows and vendor relationships. This is a rare, critical-moment hire: you'll set the leading-practice processes, systems, and policies that the company runs on as it continues to grow, with the mandate and dedicated focus to build something built to scale. This role exists to make procurement a source of leverage for the business, not friction. Your goal is to build scalable process, take on vendor management and RFP work that today sits informally with individual teams, and free up business owners' time and capacity, all while building the guardrails the company needs as it grows. You'll partner closely with Finance, Legal, Security, Compliance, IT, and business owners across the company, and you'll be judged on whether teams feel supported and unblocked, not slowed down. RESPONSIBILITIES Procurement Strategy and Process Design and implement Baseten's first company-wide procurement process, approval routing, and other workflows Establish procurement policies (spend thresholds, approval authority, competitive bidding requirements) tailored to a high-growth business Design pr

Machine LearningAIGoRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for a customer-obsessed software engineer to come ship with us. You’ll own features like multi-node training and products like serverless reinforcement learning (RL) from conception to MVP (and from MVP to GA!). You’ll work through the stack, architecting solutions from API and UI down to our infrastructure layer. You’ll fine tune models yourself to develop an understanding of user workflows. You’ll work closely with research engineers leveraging state-of-the-art training techniques to build experiences that accelerate model development and solve for real pain points. If you’re excited to dive deep into the training, let’s talk! THE PRODUCT Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done EXAMPLE INITIATIVES Checkpointing Pipeline: Our checkpointing pipeline starts with automated checkpointing, a feature that ensures that versions of models created during training are automatically backed up to the cloud. Users are able to then deploy checkpoints seamlessly into inference servers, providing point-and-click integrations into inference frameworks like vLLM and Baseten’s Inference Stack. This enables customers to quickly evaluate the performance of their checkpoints with real traffic. Multinode training: Multinode training enables customers to easily run training jobs across multiple compute nodes, enablin

KubernetesRestMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Baseten’s Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost‑efficient. As part of this team, you’ll focus on Model APIs — the infrastructure powering our hosted API endpoints for the latest open‑source models. This work spans distributed systems, model serving, and developer experience. You’ll join a small, high‑impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale. RESPONSIBILITIES: Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations Productionize performa

KubernetesMachine LearningAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re hiring a People Business Partner to support our team during a critical phase of growth. This is a highly strategic, high-impact role for someone who has partnered closely with leadership teams and helped organizations scale with intention. You will be deeply embedded with managers, bringing clarity and rigor to how teams are structured, how leaders operate, and how talent is developed across the organization. You’ll help shape team effectiveness, identify critical talent gaps, drive talent and performance strategies that enable high-performing teams, and build the people practices and change management approaches that allow us to scale with both speed and discipline. This role requires strong business judgment and the ability to operate with deep context. You’ll partner closely with leaders to navigate complex organizational decisions, anticipate challenges before they surface, and bring a clear point of view on what great looks like at every level of the organization. RESPONSIBILITIES Strategic partnership to leadership Serve as the trusted people partner to leadership, maintaining deep business context and translating it into people priorities by proactively surfacing systemic issues, risks, and opportunities before they become urgent. Bring data-driven insights to advise management on org design, succession planning, performance, retention, and engagement. Build management capacity across the org, equ

Machine LearningAIGoRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building for talent density. We believe attracting and retaining exceptional people, and ensuring they feel recognized and valued for their impact, is core to becoming the best place to work. Compensation is a critical lever in that mission. As our Compensation Manager, you will own compensation programs company-wide. You’ll be a trusted advisor to senior leaders, shaping our compensation philosophy, leveling framework, and equity programs to ensure we remain competitive, principled, and performance-oriented as we scale. This role blends strategy and execution: designing clear, fair systems while moving quickly in a high-growth environment. RESPONSIBILITIES Own and evolve Baseten’s company-wide compensation strategy, philosophy, and programs. Collaborate with leadership and HRBP to create and evolve job architecture and leveling frameworks. Build and maintain compensation bands. Conduct regular market benchmarking to ensure comp bands and strategy remain competitive in a fast moving industry. Partner closely with Talent to design and approve competitive new hire offers, advising on negotiation strategy within our compensation principles. Lead bi-annual leveling and compensation review cycles to ensure market competitiveness and reward high performance across teams. Manage new hire equity grants in partnership with Finance and Legal. Design and administer a thoughtful equity refresher program for ten

Machine LearningAIGoRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building a world-class team, anchored in our San Francisco and New York offices and increasingly growing across the globe. We believe the best talent can come from anywhere, and our ability to hire and support that talent is mission-critical. As our Immigration and Mobility Manager, you will make global hiring operationally seamless. You’ll be Baseten’s in-house expert on immigration, relocation, and international employment strategy, ensuring that exceptional candidates from all over the world can confidently build their careers with us. This is a high-stakes, high-impact role. Speed, clarity, and correctness matter deeply in immigration and mobility. You will own the end-to-end experience, navigate a rapidly evolving U.S. immigration landscape, and design the systems and policies that enable Baseten to hire globally while delivering an exceptional employee experience. Over time, you’ll also help shape where and how we expand internationally, advising on global hiring models, new hubs, and employment structures that support our long-term growth. RESPONSIBILITIES Own the end-to-end immigration lifecycle, managing all U.S. visa processes (e.g., H-1B, O-1, J-1, TN, E-3, L-1, EB-2/3 PERM) from offer stage through renewals and permanent residency. Oversee and project manage visa sponsorships executed through Employer of Record (EOR) partners Serve as Baseten’s internal immigration expert, partnering clo

Machine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re hiring a Data Engineer to build and scale Baseten’s internal data platform. This role sits at the intersection of data engineering, analytics, and data science, transforming raw product and business data into reliable datasets that power decision-making. You’ll design the data models, pipelines, and analytics infrastructure that enable teams across Product, Engineering, Finance, Marketing, and Sales to understand usage and performance. This includes working with AI inference, infrastructure, and observability data to generate insights about the product, business operations and platform economics. You’ll partner closely with stakeholders to build robust, scalable pipelines, define company-wide metrics that inform strategy and planning. RESPONSIBILITIES Design and maintain core data models and semantic layers Develop and orchestrate batch and streaming data pipelines using technologies such as Apache Beam, Kafka, Airflow, or similar frameworks Analyze inference and infrastructure telemetry , including data from OpenTelemetry, Grafana, and other observability tools Define and maintain company-wide metrics across product usage, performance, and customer lifecycle Enable self-service analytics through agents and tools, with well-structured semantic layers and context Ensure data reliability and quality through testing, documentation, and governance PREFERRED QUALIFICATIONS Understanding of inference metrics s

Machine LearningAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. At Baseten, we are building the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale, the network is the computer. We are looking for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building block in our infrastructure and unlocking the next generation of distributed inference optimizations. THE OPPORTUNITY Networking and compute are no longer separate disciplines; they are converging. The massive throughput of H100, B200, and NVL72 architectures enables and demands a new approach where communication is co-optimized alongside computation. We are entering an era where the network is an active accelerator, leveraging smart hardware offloads and direct interconnects to ensure that data movement operates at wire-speed. In this role, you will go beyond network configuration to architect the software fabric that unifies thousands of GPUs into a cohesive operating system. While you will leverage the best of the open-source ecosystem, you won't be limited by it. Where off-the-shelf solutions stop, you will build from scratch, engineering the primitives required to co-optimize communication and compute for Disaggregated Serving, Wide Expert Parallelism (WideEP), and lightening cold starts. WHAT YOU'LL DO Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack,

PythonKubernetesMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE This role sits at the frontier of our research agenda. You will pursue open problems at the intersection of post-training methodology and performant inference, and then collaborate with research engineering to translate findings into production systems. A meaningful portion of your time will be dedicated to research that deepens our understanding of how models learn, alignment, and architectural efficiency — questions that may not have immediate product application. The remainder will be directed toward research that solves concrete problems for Baseten's platform and customers, who are the fastest growing AI companies in the world like Cursor, Lovable, and Notion. We are looking for someone with sharp research taste and genuine creative instinct for problem selection. Someone who can identify questions that matter, design clean experiments to answer them, and push the state of the art. The environment here is not theoretical, but rather research that can be validated with eager customers who are serving billions of tokens a second. RECENT RESEARCH Towards infinite context windows: neural KV cache compaction Dense, on-policy or both? Repeated kv cache for long-running agents Distillation without the dark – replicating black-box on-policy distillation on Baseten RESPONSIBILITIES Define and pursue a research agenda spanning both foundational and applied work, with the applied component connected to Baseten's pla

Machine LearningAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE The largest, most demanding enterprises are starting to run on Baseten, and they arrive with a range of security, compliance, and procurement requirements. As a Senior Engineer on Baseten's enterprise engineering team, you'll build the capabilities that enable large organizations like Writer, HubSpot, and Notion to succeed on Baseten. Enterprise engineering authors the core building blocks, APIs, and user experiences powering the Baseten platform: identity and access management, billing, regional isolation, and self-hosted and single-tenant deployment options. This is deep product and systems work across the full stack, from designing authentication and authorization systems using standards like OAuth and OIDC to shipping the admin experiences enterprise IT teams use to manage their organization. EXAMPLE INITIATIVES Recent and upcoming work on the team: Fine-grained authorization for users, service accounts, and agentic workloads SSO and SCIM support, allowing customers to centralize and automate access to Baseten Expanding the billing platform to support evolving pricing models, advanced data exports, and controls to manage spend In-product management and enforcement of customer compliance requirements like data residency and HIPAA Securing network paths in and out of a customer's models with private connectivity and ingress and egress restrictions Allowing customers to run Baseten inside their own VPC, on-pr

KubernetesRestMachine LearningAI
C
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Large Language Models (LLMs) continue to push the boundaries of what AI systems can do — but inference is still the bottleneck. The Model Efficiency team is responsible for pushing the limits of LLM inference efficiency across our foundation models. We explore and ship breakthroughs across the model execution stack, including: model architecture and MoE routing optimization decoding and inference-time algorithm improvements software/hardware co-design for GPU acceleration performance optimization without compromising model quality Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. As a Staff Research Engineer, you will develop, prototype, and deploy techniques that materially improve how fast and efficiently our models run in production. You may be a good fit

GitRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The GPT Infrastructure team builds systems that turn advances in model inference and optimization into reliable production capabilities. We enable OpenAI workloads to be qualified and optimized across new accelerator platforms without requiring a one-off port and tuning effort for every hardware target. Our work spans distributed systems, model execution, compilers and runtimes, performance engineering, secure partner integrations, evaluation systems, and developer tooling. We build the infrastructure that makes optimization workflows automated, reproducible, and trustworthy. About the Role We are seeking a software engineer to help build the platform that qualifies and optimizes inference workloads across heterogeneous compute environments. You will develop both OpenAI-hosted services and secure partner-side software for running long-lived optimization workflows. These workflows generate candidate kernels, runtime configurations, and serving-stack changes; compile and execute them on target hardware; verify their correctness; measure their performance; and use the results to guide further optimization. You will work across model architecture, distributed execution, compilers, runtimes, networking, and accelerator systems. A central part of the role is turning research prototypes and one-off hardware bring-up efforts into reliable, reusable infrastructure with clear contracts, reproducible results, strong observability, and well-defined security boundaries. Key Responsibilities Design, build, and operate APIs and control-plane services for long-running workload qualification and optimization campaigns, including scheduling, retries, checkpointing, resource budgets, and observability. Build secure partner-side execution and evaluation software that can compile, run, verify, profile, and benchmark candidate artifacts on accelerator hardware. Integrate model workloads, hardware profiles, compiler toolchains, runtimes, serving engines, and distributed-exe

PythonAWSLinuxRest
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI’s AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production. In this role, you will: Design and implement the LLM inference runtime for frontier models running on custom silicon. Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization. Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads. Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack. Enable new

PythonAWSRestAI
🔔

Get new inference technical lead jobs in United States by email

Daily job updates · Unsubscribe anytime