Jobs in United States

Ml Platform Engineer in San Francisco

91 active opportunities · Updated October 2026

Explore current ml platform engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight: Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful autonomy (for example future versions of auto-review ). About the Role We’re looking for strong executors with excellent judgment, comfort with ambiguity, and an understanding of frontier model research. You don’t need prior safety or alignment experience, we also welcome people that recently realized that alignment and safety is a critical area to contribute to. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Train and evaluate frontier models to reduce harmful or misaligned agent actions, forming clear hypotheses and executing independently through ambiguity. Mine incidents and build scalable measurement, data-processing, and evaluation systems that turn real failures into repeatable safety signals. Collaborate closely with post-training, capabilities, oversight, and pre-training partners to ship research-backed mitigations into large-scale training and agent systems. You might thrive in this role if you: Have demonstrated strength in research engineering, ML en

AWSRestAIRust
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an

KubernetesGitAIGo
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Our Fraud team's mission is to help companies detect and prevent fraud using Plaid's financial network data. We believe that transaction patterns, device signals, identity linkages, and behavioral data are dramatically underleveraged tools in fraud prevention. Our products — including Protect and Signal — operate at network scale and depend on real-world investigation and research to stay ahead of adaptive adversaries. As a Senior Fraud Researcher, you will sit at the intersection of live fraud investigation, applied data science, and product innovation. You will lead complex investigations, translate findings into detection improvements, and collaborate tightly with Data Science, ML, and Product teams to shape the next generation of Plaid's fraud capabilities. This is not a purely operational role — your research directly drives features, model inputs, and product design. Responsibilities: Live Fraud Investigation & Reconstruction Lead investigations into complex fraud cases across identities, accounts, devices, and transaction surfaces Provide support to day-to-day fraud operations including SEVs and alert triage Reconstruct attacker sequences and hypothesize actor intent and tooling Distill p

PythonSQLAWSGit
D
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $200K/yr

Quick readStrong listing-quality and freshness signals

The Datadog for Startups (DDFS) program helps the next generation of fast-scaling companies adopt best-in-class observability and security from day one. We're looking for the technical engine of this program - someone who can sit across from a startup CTO, earn credibility in the first five minutes, and help them see how Datadog fits into their stack before they've even finished describing it. You'll be the first technical member on a lean, five-person team, owning the technical motion end-to-end: discovery calls, demos, startup enablement, forward-deployed engineering projects, and representing Datadog at founder events across San Francisco. This isn't a traditional SE seat - it's part solutions architect, part technical consultant, part startup evangelist, and it requires someone adaptable, proactive, and ready to take initiative without being told what to do next. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Run discovery calls with startup CTOs and engineering leads to identify quick wins, validate technical needs, and position Datadog against alternatives like Grafana, New Relic, Sentry, and Clickhouse Deliver tailored Datadog demos and help startups get instrumented quickly - removing friction, showcasing value, and ensuring smooth technical onboarding, particularly around AI/ML observability, infrastructure scaling, and security Build automation to improve internal team workflows - EX: outreach, reporting, the application process, and website updates Represent Datadog for Startups at accelerator demo days, hackathons, conferences, founder dinners, and workshops across San Francisco Build relationships across SF's startup ecosystem - founders, VCs, accelerator partners, and technical communities - and develop thought leadership content for tech

KubernetesCI/CDRestAI
A
📍 San Francisco, United States· Full-time
✓ High-confidence listingCompany trend -98.8%

From $168K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Total Rewards Compensation team serves as a strategic advisor to business leaders, managers, and employees across Airbnb. We are compensation experts with a deep understanding of our stakeholders, problem solvers who use data and insights to drive value, and partners who build the tools, models, and frameworks that support sound compensation decision-making across the company. This role sits within the Compensation function and will work closely with the Technology organization and cross-functional partners including Recruiting, People Analytics, Finance, Legal and Talent. The Difference You Will Make: We are looking for a Technical Compensation Partner to serve as the dedicated compensation partner for the Technology organization, covering Engineering, Infrastructure, Machine Learning, and Data Science. This role will support VP and Director level tech leaders and their Talent Partners as the primary day-to-day compensation resource. The ideal candidate will be a strong business partner and a builder of compensation programs, tools, and data infrastructure, and must be comfortable operating with autonomy in a fast-moving environment. A Typical Day: Serve as the dedicated compensation partner for the Technology organization, supporting VP and Director level leaders, Talent Directors, and People Partners across Engineering, Infrastructure, ML/AI, and Data Science. Act as the primary point of contact for Talent Directors and senior tech leaders on new hire offers, internal equity reviews, leveling decisions, and out-of-cycle requests. Own compensation cycle execution f

PythonSQLMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The CoT Monitorability team at OpenAI studies whether and when the chain-of-thought of frontier reasoning models is monitorable enough to support scalable oversight. We study how to measure monitorability , which training mechanisms affect monitorability, and speculative methods to improve monitorability. While we mostly focus on CoT monitorability at the moment, we care more generally about any form of monitorability, auditing methods, and improving alignment. We were the first to show that chain-of-thought monitoring can be a practical additional safety mechanism, and today our monitoring systems are actively used on OpenAI’s largest RL training runs to detect misbehavior. The issues we surface are then used to help improve our reward functions, environments, etc (without directly training against a CoT monitor). Our work sits in Alignment and intersects with model training, alignment evaluations, monitoring, and frontier-risk research.We care most about monitorability where the stakes are high, and about preserving useful oversight signals as models become more capable. About the Role We’re looking for a researcher with strong empirical ML expertise and a deep interest in model behavior, alignment, or interpretability. Direct chain-of-thought interpretability experience is welcome but not required; strong candidates may come from broader interpretability, alignment, model training, or investigative model-behavior work. As a researcher on the Alignment team, you will design and run experiments that improve our understanding of model monitorability. You will investigate how training interventions across the model-development pipeline influence whether reasoning remains legible, build evaluations that make those questions measurable, and help translate findings into practical oversight and training recommendations. You may also help develop new monitoring models or methods and apply them to OpenAI’s largest training runs. This role is especially well

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Personal AGI team is responsible for training and improving pre-trained models to be deployed into ChatGPT, the API, and potential future products. In the Model Experience team, we shape the default character and behavior of ChatGPT: how the model communicates, responds to users, uses its capabilities, and behaves across different contexts and languages. Our goal is to make every interaction with ChatGPT thoughtful, helpful, and trustworthy. We take an opinionated view of what good human–AI interaction should look like, then turn that vision into real model behavior through human data, evaluations, reward models, and post-training. Our work sits at the intersection of research, product, and model design. We partner closely with teams across OpenAI to conduct research and ensure our models are thoughtful, safe, reliable to serve millions of users. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models. Our team works in research areas combining reinforcement learning and products. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research and the quality of human-AI interaction. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve model capability and performance. Collaborate closely with the other research and product teams, allowing customers to optimize their own models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. You might thrive in this role if you: Have a deep understanding of machine learning and machine learning applications. Have good judgment about model behavior and can communicate this judgment effec

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Monetization team is a cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products, including next-generation ads experiences that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to build measurement systems that connect ad interactions to meaningful advertiser outcomes while protecting user privacy. In this foundational role, you’ll design infrastructure for conversion signals, attribution, reporting, and feedback loops across OpenAI’s ads products. This role is ideal for engineers who have built large-scale ads measurement, data, experimentation, marketplace, or distributed systems and want to apply that experience in a highly ambiguous 0→1 environment. You’ll work across event collection and normalization, deduplication and matching, attribution and modeled measurement, privacy-safe aggregation, reporting, and high-quality labels for ads optimization. We are hiring engineers who can independently own complex systems, make sound technical tradeoffs, and help define what should be built. You’ll work closely with Ads Delivery, Ads ML, Product, Research, Privacy, Data Sc

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Personal AGI team seeks to empower all of humanity to benefit from frontier intelligence in whatever way they choose. We are responsible for training models to deploy to millions of users globally via ChatGPT, the API, and future products. We aim to evolve ChatGPT from a chatbot to an infinitely capable and personalized superassistant supporting human flourishing. We work on defining, measuring, and improving capabilities across the training stack. Our focus areas include but are not limited to model behavior, personalization, safety, factuality, instruction following, personality, interactivity, multilingual fluency, world interaction, and bringing agents to everyone. We chart the course for what to strive towards. We partner closely with research and product teams across the company ensuring that our models are safe, efficient, and reliable. About the Role You’ll work as a Research Engineer / Scientist on the North Stars team within the broader Personal AGI research org. You will work on bringing the next generation of AI-enabled experiences to all of humanity by closing the capability overhang between power users and the average consumer, including areas like tool-use, feature discovery, connectors, and instruction following. You will think deeply about the current bottlenecks in model behavior, translate these insights into robust evals, training data, reward signals, and model and harness improvements. We're looking for individuals with strong ML engineering skills and research experience passionate about creative, product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve model capability and performance. Collaborate closely with the other research and product teams, allowing customers to optimize their own models. Build robust evaluations for tracking modelin

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The GTM Data Science team partners with Go-to-Market, Technical Success, Product, Engineering, RevOps, and Strategic Finance to build the shared intelligence layer for OpenAI's B2B business. The team turns product usage, customer behavior, revenue, field activity, and customer feedback into rigorous insight products that help leaders and field teams understand where customers are succeeding, where adoption is blocked, and what actions will accelerate durable growth. We are building systems that make customer intelligence proactive: surfacing risk, expansion potential, product gaps, and repeatable playbooks before they show up as escalations or missed opportunities. About the Role As the Applied Data Science & Insights Lead for GTM Intelligence Solutions and Technical Success, you will be a hands-on technical leader responsible for shaping how OpenAI measures, understands, and improves customer adoption across our B2B products. You will build AI/ML-powered intelligence products that connect account health, product usage, customer lifecycle, support tier, qualitative sentiment, commercial context, and field actions into a practical operating system for GTM and Technical Success. This role will build the data science foundation for Technical Success: defining the metrics, models, operating insights, and decision systems that help the team scale customer adoption and expansion with rigor. You will also be expected to build and lead a small mighty team over time: setting direction, hiring and developing talent, creating operating cadences, and holding a high bar for technical rigor and business impact. You will lead the development of models, metrics, and decision systems that recommend what GTM and Technical Success teams should do next, explain why, and measure whether those interventions worked. Your work will help customers move from pilots to production, deepen usage across products, identify high-value use cases, reduce churn risk, and create a f

PythonSQLAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role We are looking for an engineer who wants to take the world's largest and most capable AI models and optimize them for use in a high-volume, low-latency, and high-availability production and research environment. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Work alongside researchers to enable advanced research through awesome engineering. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 5 years of professional software engineering experience. Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc. Have experience architecting, building, observing, and debugging production distributed systems. Bonus point if worked on performance-critical distributed systems. Have need

AWSAzureRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

AI Systems Engineer - Codex Core Agents About The Team The Codex Core Agents team builds the agent harness that turns model capability into real-world action. We own the systems around the model: prompting and interpreting model outputs, executing actions safely in real environments, and feeding production experience back into better models and better agent behavior. This team sits close to research and works across the stack: harness, model interaction, inference, sandboxed execution, orchestration, evals, production reliability, and the performance envelope around tokens, latency, cost, capacity, and quality. The harness is open source and increasingly part of how models are trained and evaluated, making this one of the highest-leverage layers in Codex. About The Role We’re looking for engineers to build the AI systems that make Codex agents dependable in production. The ideal candidate is an agent-systems builder: hands-on across low-level systems and ML workflows, able to debug Codex behavior end to end across the harness, model behavior, inference/runtime stack, GPU fleet, and product surface. You’ll work with research, infrastructure, and product to design agent harness capabilities, run experiments and ablations across the model + system prompt + harness stack, build frameworks for assessing production agent performance, and turn messy failures into durable improvements. What You’ll Do Design and build the core agent harness and execution loop that lets Codex agents interpret model outputs, use tools, execute code, and complete long-horizon tasks safely. Build sandboxing, isolation, orchestration, state, and workflow infrastructure for agents operating in real development environments. Develop evaluation, experimentation, and debugging systems that distinguish harness issues, model behavior, inference/runtime issues, and product failures. Run ablations across prompts, model-facing interfaces, context construction, tool-use strategies, and harness behavior to

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The RL and Reasoning team drives the core reasoning paradigm and has created groundbreaking innovations such as o1 and o3. They focus on pushing the boundaries of reinforcement learning research, building next-generation generative models, and deploying them at scale. About the Role As a Research Engineer/Research Scientist at OpenAI, you will advance the frontier of AI alignment and capabilities through cutting-edge RL methods. Your work will sit at the heart of training intelligent, aligned, and general-purpose agents, including the systems that power various models. We’re looking for people who have a background in reinforcement learning research, are able to iterate quickly, and are proficient at coding. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. You might thrive in this role if: You love being on the cutting edge of RL and language model research. You’re a self-starter who takes initiative and ownership of ideas, driving them to completion. You value principled approaches, simple experiments in tightly-controlled settings, and reaching trustworthy conclusions which stand the test of time. You thrive in a fast-paced, dynamic, and technically complex environment where rapid iteration is key. You’re comfortable diving into a large ML codebase to debug and improve it. You have a deep understanding of machine learning and machine learning applications. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the ful

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are hiring a Sim Infrastructure Engineer to turn simulation systems into reliable, automated, production-quality pipelines that power model training, evaluation, and hardware-in-the-loop validation. This role owns the automation, orchestration, and tool integration that apply simulation to concrete robotics tasks: building CI/CD for SIL/HIL, presubmit checks, automatic model evaluation, metric computation and reporting, and the runtime infrastructure to run simulations at scale. You will collaborate closely with Sim Realism, Sim Environments, Research, and Ops to make simulation an integrated, reproducible, and measurable part of our ML and robotics workflows. This role is based in San Francisco, CA, and requires in-person 4 days a week. In this role, you will: Build and maintain presubmit checks, continuous integration and deployment pipelines for simulation code, environments, and tasks so simulation artifacts are testable, versioned, and reproducible. Implement end-to-end automation to run model evaluation in sim (SIL) and orchestrate HIL runs; compute realism and task metrics, generate dashboards and alerts, and ensure evaluation is repeatable and auditable. Create robust APIs and connectors so research, training, and data-collection systems can schedule, seed, and evaluate batches of simulations; support RL rollouts, imitation-data collection, and presubmit model checks. Build scheduling, batching and orchestration for running very large numbers of concurrent rollouts (target tens of thousands of rollouts / large RL workloads), sol

PythonAWSKubernetesCI/CD
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. Our mission is to push the frontier of code generation and agentic reasoning, and deploy these capabilities in real-world products such as ChatGPT and the API, as well as in next-generation tools specifically designed for agentic coding. We operate across research, engineering, product, and infrastructure—owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. About the Role As a Performance & Systems Engineer on the Codex team, you will be responsible for whole-system optimization across a complex, evolving stack. Codex spans LLM inference, cloud orchestration, agentic work management, and multiple product surfaces. Your job will be to identify and land high-leverage changes—across infrastructure, modeling, and product layers—that make Codex agents significantly faster and cheaper to serve. We’re looking for generalists who thrive in ambiguity and love chasing performance bottlenecks to ground. This is a high-ownership role where your work will directly improve the experience of millions of users. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Hunt down and address inefficiencies across the Codex system stack, from agent behavior to LLM inference to container orchestration, and beyond. Build tooling to measure, profile, and optimize system performance at scale. Collaborate with researchers and engineers to land high-ROI changes that improve latency and cost. You might thrive in this role if you: Have experience operating across both ML systems and cloud infrastructure. Enjoy diving into messy, ambiguous problems and emerging with clear wins. Think holistically about performance, balancing spee

AWSRestAIRust
🔔

Get new ml platform engineer jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime