Jobs in United States

Machine Learning Research Scientist in San Francisco

238 active opportunities · Updated October 2026

Explore current machine learning research scientist jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Safety Systems team is dedicated to ensuring the safety, robustness, and reliability of AI models and their deployment in the real world. Learn more about OpenAI’s approach to safety. Building on the many years of our practical alignment work and applied safety efforts, Safety Systems addresses emerging safety issues and develops new fundamental solutions to enable the safe deployment of our most advanced models and future AGI, to make AI that is beneficial and trustworthy. About the Role At OpenAI, we're dedicated to advancing artificial intelligence, and we know that creating a secure and reliable platform is vital to our mission. That's why we're seeking a software engineer to help us build out our trust and safety capabilities. In this role, you'll work with our entire engineering team to design and implement systems that detect and prevent abuse, promote user safety, and reduce risk across our platform. You'll be at the forefront of our efforts to ensure that the immense potential of AI is harnessed in a responsible and sustainable manner. Your Responsibilities: Architect, build, and maintain anti-abuse and content moderation infrastructure designed to protect us and end users from unwanted behavior. Work closely with our other engineers and researchers to utilize both industry standard and novel AI techniques to measure, monitor and improve AI models’ alignment to human values. . Diagnose and remediate active incidents on the platform and build new tooling and infrastructure that address the root causes of system failure. You might thrive in this role if: You have built and run production services in a high growth, rapidly scaling environment. You can debug live issues and restore systems quickly. You have worked on content safety, fraud, or abuse, or are motivated and excited to work on present-day (“now-term”) AI safety. You have experience with Python or with modern languages such as C++, Rust, or Go, and are able to quickly ramp up on Py

PythonAWSAzureKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking an experienced SoC Architect to lead the definition and development of next-generation custom AI silicon for edge deployments. This role will be responsible for shaping the architecture of highly efficient, high-performance SoCs optimized for machine learning inference and on-device intelligence. You will work cross-functionally with internal engineering teams and external ecosystem partners to translate product requirements into scalable silicon solutions, driving execution from concept through delivery. In this role you will: Define the architecture and technical roadmap for custom SoCs targeted for edge applications. Drive system-level tradeoff analysis across compute, memory, interconnect, power, thermal, and cost constraints. Architect energy-efficient ML compute subsystems optimized for inference workloads and real-world deployment environments. Collaborate with internal hardware, software, systems, and product teams to align architecture with platform needs. Partner with external silicon vendors, IP providers, and manufacturing partners to execute development plans. Lead hardware/software co-design efforts to maximize performance per watt and end-to-end system efficiency. Guide implementation teams through microarchitecture, RTL development, validation, and bring-up phases. Operate effectively in agile development environments and help teams deliver against aggressive schedules and milestones. You might thrive in this role if: Proven exper

AWSRestAgileMachine Learning
W
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend +8.1%

🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Are you passionate about ensuring the highest quality for cutting-edge generative AI applications? As a software quality engineer at WRITER, you'll play a critical role in shaping the reliability, performance, and trustworthiness of our AI-powered work orchestration platform. You’ll be at the forefront of defining and implementing rigorous quality strategies for our enterprise-grade LLMs and AI agents, directly impacting how hundreds of global companies unlock transformational value through AI. This is a unique chance to dive deep into the unique challenges of AI quality assurance and make a tangible difference in a rapidly evolving field. This is a hybrid role based out of our London, San Francisco, Seattle, and New York City hubs. You will report directly to the director of engineering. 🦸🏻‍♀️ What you'll do Define and implement comprehensive quality assurance strategies and test plans for our AI agents and LLM-powered applications, ensuring exceptional prod

TypeScriptPythonAWSAzure
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -80.2%

About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal and VMs dependable & scalable foundations for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Software Engineer, Host Assurance to build and operate the services, APIs, and host software that establish and maintain trust in our compute infrastructure. You will own production software from design and implementation through testing, rollout, observability, and operation. Your work will support capabilities such as machine identity, certificate issuance and enrollment, secure bootstrap, and host attestation across bare-metal and VM environments. Success in this role requires strong technical judgment, the ability to reason across software and host-system boundaries and learn unfamiliar parts of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—inc

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training. Our mission is to make some of the world's largest GPU clusters work reliably for frontier training. We bring new platforms and clusters online, safely maintain installed fleets, and partner with hardware, infrastructure, and research teams to resolve the system-level issues that keep jobs from running. That means building and maintaining the software closest to the machine: Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation. We make these components reproducible, compatible, and safe to operate across heterogeneous fleets. About the Role We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate. You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models

AWSLinuxRestAI
M
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -67.9%
Quick readStrong listing-quality and freshness signals

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll automate the integration of new capacity from a growing set of hardware providers; from auditing and benchmarking hosts and clusters, to maintaining our machine images, configuring GPUs, RDMA, networking, and storage, and getting machines into production. You'll build the automation that keeps the fleet healthy without human intervention: detecting bad GPUs, thermals, and disks. You'll dig into whatever is between the hardware and the software that runs on

PythonLinuxAIAuditing
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Safety Systems team is dedicated to ensuring the safety, robustness, and reliability of AI models and their deployment in the real world. Building on the many years of our practical alignment work and applied safety efforts, Safety Systems addresses emerging safety issues and develops new fundamental solutions to enable the safe deployment of our most advanced models and future AGI, to make AI that is beneficial and trustworthy. Learn more about OpenAI’s approach to safety About the Role As an Analytics Engineer in Safety Systems, you will play a pivotal role in building a data-centric culture, enhancing decision-making processes, and driving strategic initiatives through analytics. You will partner closely with Engineering, Research, and Data Science to develop and maintain canonical data sources and source-of-truth dashboards that enable both people and AI agents across the organization to derive trustworthy, actionable insights. You will own the consumption layer for safety metrics: defining intuitive, reliable ways for stakeholders across Safety Systems, partner teams, and leadership to understand the safety of our products, answer safety-related questions independently, and inform product decisions and company strategy. Most importantly, you will be a core member of the Safety Systems team, collaborating with researchers and engineers to advance our goals of safe, robust, and reliable AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and maintain canonical datasets that serve as sources of truth for safety metrics. Develop and refine data products such as dashboards, reports, agent-enabled workflows, and machine-readable interfaces that empower stakeholders to extract and analyze data independently. Work closely with stakeholders in Engineering, Research, and Data Science to understand their decision-making n

SQLAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world. Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure. About the Role We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology. You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more. There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future. In this role, you will: Build, lead, and grow high-performing infrastructure engineering teams. Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search. Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach. Champion pragmatic use of agent technology to amplify execution velocity. Reduce operational toil and incident frequency through better abstractions, gua

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,

AWSKubernetesLinuxRest
FC
📍 San Francisco, CA, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Fanatics Commerce is the global leader in licensed sports merchandise, operating a vertically integrated platform that designs, manufactures, and delivers officially licensed apparel, jerseys, headwear, and collectibles for major leagues, teams, and events worldwide. With more than 900 e-commerce sites and a global omnichannel presence across digital, in-venue, and retail, Fanatics Commerce reaches fans in over 180 countries and powers official fan experiences for many of the world's most iconic sports properties. At Fanatics, we bring our BOLD Leadership Principles to life every day - building championship teams, obsessing over fans, acting with entrepreneurial speed, and delivering with a determined and relentless mindset. ROLE OVERVIEW The Retail Lead delivers business and fan impact through BOLD leadership and execution excellence, leveraging data, automation, and AI-enabled insights. Home game shifts are considered a core responsibility of the position and may include pre-game, in-game, and post-game operational support. HOW WILL YOU DRIVE IMPACT Success is measured by the ability to deliver results through BOLD capabilities and measurable outcomes. Fan & Customer Impact (Obsessed with Fans) Drive sales results by consistent execution of daily operations Support back of house operations; maintain stockroom organization Knowledge of retail operation systems including but not limited to OpSuite5.0 (online inventory manager), POS (point of sales) and CAYAN (credit card machines) Partner with Leadership team when making decisions including but not limited to; revenue targets, per cap and UPT (unit per time) Work with Retail Associates to ensure an exemplary fan experience Ownership & Execution (Determined & Relentless Mindset) Communicate expectations for assignments and projects to Retail Associates Provide training and assistance to Retail Assoc

GitAIGoExcel
FC
📍 San Francisco, CA, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Fanatics Commerce is the global leader in licensed sports merchandise, operating a vertically integrated platform that designs, manufactures, and delivers officially licensed apparel, jerseys, headwear, and collectibles for major leagues, teams, and events worldwide. With more than 900 e-commerce sites and a global omnichannel presence across digital, in-venue, and retail, Fanatics Commerce reaches fans in over 180 countries and powers official fan experiences for many of the world's most iconic sports properties. At Fanatics, we bring our BOLD Leadership Principles to life every day - building championship teams, obsessing over fans, acting with entrepreneurial speed, and delivering with a determined and relentless mindset. ROLE OVERVIEW The Retail Lead delivers business and fan impact through BOLD leadership and execution excellence, leveraging data, automation, and AI-enabled insights. Home game shifts are considered a core responsibility of the position and may include pre-game, in-game, and post-game operational support. HOW WILL YOU DRIVE IMPACT Success is measured by the ability to deliver results through BOLD capabilities and measurable outcomes. Fan & Customer Impact (Obsessed with Fans) Drive sales results by consistent execution of daily operations Support back of house operations; maintain stockroom organization Knowledge of retail operation systems including but not limited to OpSuite5.0 (online inventory manager), POS (point of sales) and CAYAN (credit card machines) Partner with Leadership team when making decisions including but not limited to; revenue targets, per cap and UPT (unit per time) Work with Retail Associates to ensure an exemplary fan experience Ownership & Execution (Determined & Relentless Mindset) Communicate expectations for assignments and projects to Retail Associates Provide training and assistance to Retail Assoc

GitAIGoExcel
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a Mechanical Engineer to design, build, and own the mechanical side of our robotic actuator dynamometer and test infrastructure. You will create the test stands, couplings, fixtures, load paths, guarding, and serviceable lab hardware that enable rigorous characterization of robotic actuators. This role combines precision mechanical design with hands-on lab work. You will take robotic actuator test infrastructure from requirements and analysis through CAD, fabrication, assembly, commissioning, and iteration, partnering closely with electrical and software engineers to deliver safe, flexible, high-uptime test cells. In this role, you will Own the mechanical architecture of dynamometer and actuator test cells, including frames, bases, load paths, alignment, guarding, and serviceability. Design dynamometer structures, robotic actuator fixtures, load-motor mounts, couplings, shafts, bearings, adapters, and torque-reaction hardware. Translate robotic actuator test requirements into robust mechanical systems for torque, speed, thermal, durability, backdrive, efficiency, and failure testing. Perform first-principles analysis and simulation for stiffness, strength, fatigue, vibration, thermal growth, critical speed, and safety factors. Create precise, repeatable alignment strategies that protect test articles, load machines, sensors, and couplings. Design modular fixturing that supports rapid changeover across actuator and motor variants without compromising measurement quality. Work closely with electrical engineers on cable routing

ReactAWSRestAI
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -80.2%

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking an Electrical Engineer to build and own the electrical backbone of our robotic actuator dynamometer and test infrastructure. You will design, integrate, and operate the load motor drives, power distribution, instrumentation wiring, DAQ interfaces, and safety systems that make high-performance robotic actuator testing repeatable, safe, and scalable. This role spans hands-on lab execution and system architecture: selecting and commissioning power electronics, designing robust test-cell electrical systems, bringing up sensors and DAQ, and partnering with mechanical and software engineers to turn robotic actuator hardware into trustworthy data. In this role, you will Own the electrical architecture of dynamometer and actuator test cells, from mains distribution and protection through load motor drives, braking, and auxiliary power. Specify, integrate, commission, and tune motor drives and load machines for robotic actuator torque, speed, efficiency, thermal, and durability testing. Design power distribution, grounding, shielding, cable routing, and connectorization for high-current, high-voltage, and low-level measurement systems. Integrate torque, position, speed, temperature, voltage, current, vibration, and other instrumentation from robotic actuators into DAQ and control systems. Develop electrical schematics, wiring diagrams, panel layouts, harness documentation, and test-cell interface definitions. Build, debug, and maintain test-cell electrical hardware, rapidly diagnosing noise, EMI, grounding, drive, sensor, and power-q

Artificial IntelligenceAI
🔔

Get new machine learning research scientist jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime