About the Team The Integrity team at OpenAI is dedicated to ensuring that our cutting-edge technology is not only revolutionary, but also secure from a myriad of adversarial threats. We strive to maintain the integrity of our platforms as they scale. The Integrity team is at the front lines of defending against misuse in all its forms: content abuse, scaled attacks, and other actions that could undermine the user experience or harm our operational stability. About the Role As a Machine Learning Engineer in OpenAI's Integrity team, you will have the opportunity to work with some of the brightest minds in AI. You’ll work on state-of-the-art models and classifiers, experiment with new architecture and approaches, and push forward our abilities in content and user understanding. You’ll help turn research breakthroughs into tangible solutions that improve the trust and safety of our platform. If you're excited about training LLMs and building ML models, this role is your chance to make a significant mark. In this role, you will: Innovate and Deploy: Design and deploy advanced machine learning models that solve real-world problems. Bring OpenAI's research from concept to implementation, creating AI-driven applications with a direct impact. Collaborate with the Best: Work closely with researchers, software engineers, and product managers to understand complex business challenges and deliver AI-powered solutions. Be part of a dynamic team where ideas flow freely and creativity thrives. Optimize and Scale: Implement scalable data pipelines, optimize models for performance and accuracy, and ensure they are production-ready. Contribute to projects that require cutting-edge technology and innovative approaches. Learn and Lead: Stay ahead of the curve by engaging with the latest developments in machine learning and AI. Take part in code reviews, share knowledge, and lead by example to maintain high-quality engineering practices. Make a Difference: Monitor and maintain deployed m
Jobs in United States
Machine Learning Engineer Ii Core Engineering in San Francisco
238 active opportunities · Updated October 2026
Showing
15 jobs
Explore current machine learning engineer ii core engineering jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. Our team works across silicon, embedded systems, operating systems, and cloud services to build reliable consumer devices and the novel platforms required to support them. We partner closely with research to bring advanced AI capabilities into the physical world. About the Role As an Operating Systems Engineer focused on on-device inference, you will design, develop, and ship the OS stack that makes advanced AI capabilities reliable, responsive, and energy efficient on consumer devices. Your work will span OS services and frameworks, inference runtime integration, model fitting, scheduling, and performance and power management. You’ll partner with research to adapt models to device constraints, make design decisions across the stack, and carry solutions from early exploration through integration and production. In this role, you will: Build the inference platform: Design and implement maintainable OS services, frameworks, and clear interfaces for inference execution, model loading and lifecycle, and resource management. Fit models to device constraints: Partner with researchers on quantization, runtime integration, and memory optimization to meet memory, compute, and energy budgets while evaluating model quality and product behavior. Coordinate system resources: Develop scheduling and resource policies that balance inference with other device act
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE Baseten's Compute org is in hyper growth. As it scales, the systems and workflows that keep supply and demand balanced across our GPU fleet need to get more sophisticated, and this role exists to make sure they do. Compute sits at the center of how Baseten allocates, forecasts, and manages the capacity that powers every customer inference request. The team that supports this work, C3, runs on a mix of internal tooling, manual processes, and systems that haven't fully kept pace with the scale of the problem. This role exists to close that gap. You'll design, build, and ship AI-powered workflows that give the Compute and C3 teams real leverage, automating the manual, repetitive, and error-prone parts of the capacity lifecycle so the team can focus on judgment calls that actually need a human. We want someone who can walk in, audit what exists today, identify what's missing or broken, and start shipping fast. You know when to reach for an existing internal tool and when to build something custom in Claude Code. You think two to three steps ahead about how the thing you build today fits into the broader capacity systems architecture tomorrow. And you bring a point of view on our stack, on what we should be building, and on where AI can do something existing tooling simply can't. RESPONSIBILITIES Ship AI-powered workflows for Compute and C3 : build the agents and automations that give capacity analysts, ops leads, an
About the Team The Health team, within OpenAI’s broader Personal AGI organization, has a mission to ensure AGI improves health for all humanity. Improving human health will be one of the defining impacts of AGI. Hundreds of millions of people already turn to ChatGPT for questions about their health and millions of clinicians use it weekly to support care delivery. Increasingly capable models create an opportunity to make high-quality medical intelligence more accessible across patients and clinicians—raising the floor of human health—and accelerate the new capabilities and scientific advances that raise the ceiling of human health. Our job is to make those benefits real. We work across the full model stack—pretraining, midtraining, reinforcement learning, post-training, evaluations, harnessing, and deployment—and connect that research to the patients, clinicians, and real-world outcomes we aim to improve. About the Role We’re looking for an exceptional, hands-on researcher who wants to build frontier health capabilities and turn them into impact at scale. This is a role for someone who can take an important, underdefined problem from 0→1: identify the right bet, build what’s needed to test it, and drive it all the way to a measurable improvement in the models and products we actually ship. We’re especially excited about two kinds of people: researchers with the technical depth to move the frontier in pretraining, reinforcement learning (RL) / post-training, or evals; and researchers with real depth in developing frontier biomedical AI capabilities. Prior experience in healthcare is helpful but not required. Research excellence, velocity, ownership, and alignment with the mission are most important to us. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own a high-leverage research direction end to end—from deciding which problem matters and h
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are looking for an IT Support / Operations Engineer to join Baseten as we continue to scale our IT team. In this role, you will play a critical part in bringing our technical support entirely in-house to provide a seamless, high-touch experience for all Baseten employees. As we continue to scale, you will be the primary point of contact for day-to-day technical issues, allowing you to have a direct impact on our team's productivity and overall office environment. This position is ideal for a hands-on problem solver who enjoys a mix of hardware and software troubleshooting, user lifecycle management, and maintaining the physical IT infrastructure of a modern office. While you will focus heavily on elevating our internal support standards, you will also assist with systems administration and workflow automation as our company evolves. This is a hybrid role based out of our San Francisco or New York office, following our standard policy of three days per week in-person to ensure our physical office and AV systems remain high-performing and reliable. RESPONSIBILITIES Serve as the escalation point for day-to-day technical support, diagnosing and resolving hardware and software issues across our Mac and Windows fleet Manage user lifecycle administration including provisioning, deprovisioning, and access management across all systems and services Own the IT onboarding experience for new employees — from laptop set
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re hiring a Data Engineer to build and scale Baseten’s internal data platform. This role sits at the intersection of data engineering, analytics, and data science, transforming raw product and business data into reliable datasets that power decision-making. You’ll design the data models, pipelines, and analytics infrastructure that enable teams across Product, Engineering, Finance, Marketing, and Sales to understand usage and performance. This includes working with AI inference, infrastructure, and observability data to generate insights about the product, business operations and platform economics. You’ll partner closely with stakeholders to build robust, scalable pipelines, define company-wide metrics that inform strategy and planning. RESPONSIBILITIES Design and maintain core data models and semantic layers Develop and orchestrate batch and streaming data pipelines using technologies such as Apache Beam, Kafka, Airflow, or similar frameworks Analyze inference and infrastructure telemetry , including data from OpenTelemetry, Grafana, and other observability tools Define and maintain company-wide metrics across product usage, performance, and customer lifecycle Enable self-service analytics through agents and tools, with well-structured semantic layers and context Ensure data reliability and quality through testing, documentation, and governance PREFERRED QUALIFICATIONS Understanding of inference metrics s
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: As a Software Engineer at Baseten, you will own one of the most critical surfaces of our business: pricing, billing, and revenue infrastructure. As we launch more and more products— billing is no longer just operational plumbing. It is a strategic lever for growth. This role will establish clear ownership of billing as a function and create leverage for Finance, Sales, and GTM teams while maintaining a seamless customer experience. RESPONSIBILITIES: Own Baseten’s end-to-end billing and revenue infrastructure, including pricing, invoicing, metering, and reporting foundations. Build and evolve our billing platform and integrations (including Orb), ensuring correctness, auditability, and a high-trust experience for customers and internal teams. Partner closely with Finance, Sales, GTM, and Forward Deployed Engineering to turn real-world workflows into reliable internal tooling and automation (quoting, approvals, renewals, usage reconciliation, revenue reporting). Design systems that scale with new products, packaging, and go-to-market motions, making billing a strategic lever for growth. Drive reliability and operational excellence for revenue-critical workflows: monitoring, alerting, incident response, backfills, and clear runbooks. Lead from the front on high-impact projects: clarify requirements, propose crisp technical approaches, ship iteratively, and raise the bar on quality and velocity. Debug and resolve
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role This isn’t a typical engineering role. You won’t be embedded in a single product team or siloed in one product area. Instead, you’ll sit within Platform Engineering, own the AI-assisted coding domain, and work across all of engineering at Sentry, focused specifically on how AI coding agents participate in our software development lifecycle. For AI coding agents to work well in our repo, the internal systems they depend on need to be accessible via API, not locked behind UIs that require human interaction. Right now, many of those systems aren’t agent-ready. You’ll audit and prioritize that gap, expose those systems programmatically, and build the connections that let tools like Claude Code operate on them end-to-end. From there, the scope expands to improving the quality of AI-generated pull requests and automating the engineering work that’s important but consistently deprioritized. You will look from context engineering standpoint to see what to send to our model; you will look from harness engineering standpoint to see the tools it can use, the permissions it has, the state it carries forward, the tests it has to pass, the logs you capture, the retries, checkpoints, guardrails, and evals. You’ll work closely with the dev infrastructure team as your home base, then collaborate across every product team coding in our repo once the tooling foundation is in place. It’s a broad role with real impact, and the work you do will directly change how Sentry engineers ship software. What You’ll Do Audit Sentry’s internal developer systems and make them API-ready for AI agents. You’ll prioritize and drive the work of ex
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Are you the person on your team who builds the agent everyone else ends up using? We're looking for an AI Engineer to join our Training Product team and do that at Baseten. You'll build AI-driven product features for the customers training and post-training frontier models on our platform, and you'll raise the ceiling on how Baseten itself uses AI internally, turning manual workflows into agentic ones that make every other team faster. You'll work directly with our research engineers to scope and build products, taking ideas from a research loop that already works internally to something customers can run themselves. This is a hands-on role with real autonomy. You'll pick the problems worth solving, build the harnesses, execution flows, and guardrails that make AI systems reliable, and own the results. If you've been shipping agents and want that to be the job, let's talk. EXAMPLE INITIATIVES: Take a look at these blog posts written by members of our team: Baseten Training: an autoresearch substrate Introducing Baseten Loops Harnesses are everything. Here's how to optimize yours. Building with NVIDIA Nemotron 3 Ultra and LangChain Deep Agents Code on Baseten RESPONSIBILITIES: Build and ship agentic product experiences, including chat-style and assistant-like interfaces, from prototype to GA. Design the harnesses, execution flows, and guardrails that make AI systems reliable in production. Build internal autom
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to join our team at the cutting edge of AI acceleration, where your code directly impacts the performance of state-of-the-art machine learning models. As a GPU Kernel Engineer, you'll craft the foundation that powers modern AI workloads, optimizing every microsecond of computation to enable breakthrough applications. You'll work in a fast-paced, intellectually stimulating environment where technical excellence is paramount and your contributions directly influence production systems serving millions of users across numerous products. This role offers exceptional growth potential for engineers passionate about low-level optimization and high-impact systems work. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Model Performance team: Baseten Embeddings Inference: The fastest embeddings solution available The Baseten Inference Stack Driving model performance optimization RESPONSIBILITIES Core Engineering Responsibilities Design and implement high-performance GPU kernels for key ML operations, including matrix multiplications, attention mechanisms, and mixture-of-experts routing Write and optimize code using CUDA, PTX assembly, and architecture-specific techniques Apply advanced performance optimization methods such as memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap Performance & Innovation Impl
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. We are looking for an engineer with strong experience in machine learning and solid foundations in maths and computer science to join our growing Post-Training team at Baseten. Custom models are instrumental to the success of Baseten customers. By inference volume, the overwhelming majority of traffic at Baseten is to and from models that have been post-trained in some way, whether that be through reinforcement learning, supervised finetuning, a recent technique from the literature, or an in-house research technique from Baseten. The Post-Training team is responsible for the success of our customers’ post-trained models, and we employ a wide array of techniques to produce models that are more efficient and higher quality than even the biggest closed source models for the customer’s specific needs. Your role as a research engineer is to build the in-house tooling to support all of this. We care about training a wide spectrum of different model architectures with a variety of techniques efficiently and at scale. At times this involves zooming deep into a particular technical topic, but more often if involves working across the stack as a whole - systems-level concepts like Kubernetes, cgroups, storage systems, and networking topologies, as well as PyTorch distributed tensor computation, and GPU kernels. RECENT RESEARCH Dense, on-policy or both? Repeated kv cache for long-running agents Distillation without the dark – rep
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Baseten’s Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost‑efficient. As part of this team, you’ll focus on Model APIs — the infrastructure powering our hosted API endpoints for the latest open‑source models. This work spans distributed systems, model serving, and developer experience. You’ll join a small, high‑impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale. RESPONSIBILITIES: Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations Productionize performa
About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training: ML Framework Engineer, you will work on improving the training throughput for our internal training framework, while enabling researchers to experiment with new ideas. This requires good engineering (for example designing, implementing, and optimizing state-of-the-art AI models), writing bug-free machine learning code (surprisingly difficult!), and acquiring deep knowledge of the performance of supercomputers. In all the projects this role pursues, the ultimate goal is to push the field forward. We’re looking for people who love optimizing performance, understanding distributed systems, and who cannot stand having bugs in their code. Since our training framework is used for large runs with massive numbers of GPUs, performance improvements here will have a large impact. This role is based in San Francisco, CA. We use a
About the Team The Codex Core Agent team builds the kernel of Codex. We own making the agent better, accelerating research, and making those improvements real in production for our users. That means working across the systems that make Codex actually function as an agent in the real world: the production performance envelope around tokens, latency, reliability, cost, and capacity; the core execution loop and interfaces that turn models into useful behavior; the shared infrastructure that enables other teams to build on Codex; and the feedback loops that turn real-world usage into better models and better agent behavior over time. About the Role We’re looking for applied AI engineers to help bring Codex agents from impressive demos to dependable tools. This role is about improving agent performance on real software engineering tasks and closing the gap between research capability and real-world usefulness. You’ll work closely with research, infrastructure, and product to ensure agents are not just powerful, but useful, steerable, and reliable in practice. The job is not only to improve model behavior in isolation, but to turn those improvements into measurable gains in solve rate, usefulness, and economic value for users. What You’ll Do Design and iterate on agent behaviors across real-world coding tasks and long-horizon workflows. Work closely with research to develop and run evals to measure agent performance, regressions, failure modes, and edge cases. Improve performance through prompting, tool-use strategies, context construction, and model-facing experimentation. Analyze failures in production and systematically improve robustness and reliability. Build feedback loops and data systems that get better real-task data into evaluation and research. Work with product teams to shape user-facing agent experiences and the interfaces the agent depends on. Help define what “good” looks like for agents completing complex tasks end-to-end. You Might Be a Good Fit If You Ha
Related career options
Similar roles with stronger pay
Demand 46/100 · 8 jobs
$840K – $840K/yr
Salary →Demand 43/100 · 5 jobs
$840K – $840K/yr
Salary →Demand 43/100 · 6 jobs
$382.5K – $382.5K/yr
Salary →Demand 43/100 · 8 jobs
$300K – $300K/yr
Salary →Demand 42/100 · 7 jobs
$300K – $300K/yr
Salary →Demand 43/100 · 22 jobs
$278.9K – $278.9K/yr
Salary →Other cities to consider
More places hiring for this role
Get new machine learning engineer ii core engineering jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime