About the Team The Frontier Assurance team brings independent scrutiny into OpenAI’s safety decisions and helps the public understand and assess our safety work. We lead third-party assessments and safeguard testing for OpenAI’s flagship launches, pilot new assurance mechanisms such as embedded auditing, run our misalignment disclosure process, and incorporate independent expert input as evidence for critical safety decisions. About the Role As a Research Program Manager on the Frontier Assurance team, you will build programs that bring independent expertise into frontier AI safety decisions and make the evidence behind those decisions understandable to the public. You will lead external research partnerships and third-party assessments, coordinate public safety documentation, and develop new approaches to independent scrutiny and transparency. Working across research, engineering, product, policy, and communications, you will help ensure external findings inform concrete decisions and that our public explanations accurately reflect the evidence, limitations, and remaining uncertainty. We’re looking for people with deep experience in research partnerships and program management with technical and research teams. This role combines partnership management, cross-functional coordination, an understanding of AI safety research, alignment, and evaluations, and strong communication skills. You will work with researchers and engineers within OpenAI and across the external community to initiate projects, set ambitious goals and milestones, and drive execution across multiple teams. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and run third-party assessment programs for frontier models and safeguards, including independent evaluations, adversarial testing, and new approaches such as embedded auditing. Work with researchers and external part
Jobs in United States
Ai Systems Engineer in San Francisco
1,456 active opportunities · Updated October 2026
Showing
15 jobs
Explore current ai systems engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll automate the integration of new capacity from a growing set of hardware providers; from auditing and benchmarking hosts and clusters, to maintaining our machine images, configuring GPUs, RDMA, networking, and storage, and getting machines into production. You'll build the automation that keeps the fleet healthy without human intervention: detecting bad GPUs, thermals, and disks. You'll dig into whatever is between the hardware and the software that runs on
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on the distributed object storage system that underpins every container image, volume, and checkpoint on Modal: hundreds of petabytes of data, replicated across multiple cloud object stores and a CDN, cached on local NVMe across a large fleet of workers in many datacenters, and shared peer-to-peer within each datacenter. You'll make cold starts feel local when the data is hundreds of milliseconds away, designing the caching, preloading, and peer-to-peer layers that hide object-store latency and keep public ingress off saturated uplinks. You'll own durability and cost at petabyte scale, from streaming and batch replication between origins, to garbage collecti
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for a strong technical lead to guide the engineers designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. You'll lead the team responsible for the distributed object storage system that underpins every container image, volume, and checkpoint on Modal: hundreds of petabytes of data, replicated across multiple cloud object stores and a CDN, cached on local NVMe across a large fleet of workers in many datacenters, and shared peer-to-peer within each datacenter. You'll set technical direction for the primitives that other teams (filesystems, training, sandboxes) build on, balancing durability, latency, throughput, and cost. You'll own the roadmap from today's hardest problems (garbage collection at petabyte scale, active-active replication, rate limiting that protects the upstream without wasting ut
About the Team Safety Systems manages the complete lifecycle of safety efforts for OpenAI’s frontier models, ensuring our models are deployed responsibly and have a positive impact on society. Our work spans diverse research and engineering initiatives—from system-level safeguards and model training to evaluation and red-teaming—all aimed at mitigating misuse, misalignment, and maintaining our high bar for safety. We lead OpenAI's commitment to developing and deploying safe Artificial General Intelligence (AGI), fostering a culture of trust, responsibility, and transparency. Our goal is to continuously learn from deployments, distribute AI’s benefits widely, and ensure that powerful tools remain aligned with human values and safety considerations. Within Safety Systems, the Model Policy team works to ensure that frontier models behave safely and reliably in real-world environments by designing policies that define safe model behavior. Some of our publications include: Safety at every step OpenAI GPT6 System Card OpenAI Model Spec About the Role We’re hiring a Model Policy Manager to shape model behavior for U.S. government use, with a focus on national security applications. You’ll define nuanced policies and translate them into training and evaluation criteria, helping models navigate high-stakes scenarios while preserving their usefulness and capabilities. In this role, you will: Develop model policies that guide safe and useful behavior. Build evaluations, identify policy gaps and model failures, and use findings to improve policies and training. Work with research, engineering, and domain experts to support safe, reliable deployment. You might thrive in this role if you: Bring relevant experience in AI safety, policy, or risk assessment. Have strong judgment and can turn complex safety questions into clear, practical policies. Have the technical fluency to work hands-on with model data and evaluations. Are motivated by OpenAI’s mission and the responsible use of
About the Team Safety Systems manages the complete lifecycle of safety efforts for OpenAI’s frontier models, ensuring our models are deployed responsibly and have a positive impact on society. Our work spans diverse research and engineering initiatives—from system-level safeguards and model training to evaluation and red-teaming—all aimed at mitigating misuse, misalignment, and maintaining our high bar for safety. We lead OpenAI's commitment to developing and deploying safe Artificial General Intelligence (AGI), fostering a culture of trust, responsibility, and transparency. Our goal is to continuously learn from deployments, distribute AI’s benefits widely, and ensure that powerful tools remain aligned with human values and safety considerations. Within Safety Systems, the Model Policy team works to ensure that frontier models behave safely and reliably in real-world environments by designing policies that define safe model behavior. Our relevant publications include: Safety at every step OpenAI GPT6 System Card OpenAI Model Spec GPT-Live ChatGPT Images 2.5 About the Role We are hiring a Model Policy Manager to focus on the safety of multimodal models. In this role, you will shape how OpenAI identifies, evaluates, and addresses risks in multimodal AI models - such as GPT-Live and ChatGPT Images - as well as multimodal capabilities in frontier AI models. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and maintain model policies for audio, image, video, and omni-modal behavior. Translate theories of harm and threat models into behavioral safety policies, evaluation criteria, grading guidance, and safeguards. Identify and analyze safety regressions and failure patterns to identify gaps in existing policies and inform policy iteration. Develop policy artifacts that support model training, evaluation, and deployment, including behavior i
About the Team The Safety Training research team aims to fundamentally advance our capabilities for precisely implementing safe behavior in AI models, and to leverage these advances to make OpenAI’s deployed models safe and beneficial. This requires a breadth of new ML research to address the growing set of safety challenges as AI becomes more powerful and used in more settings. Key focus areas include how to train nuanced safety behaviors, how to make the model robust to bad actors, how to address privacy and security risks, and how to make the model trustworthy in safety-critical situations. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. About the Role We’re seeking a researcher to train and evaluate models for U.S. government use, with a focus on national security applications. You’ll advance safety post-training and robustness, helping models follow nuanced policies while preserving their usefulness and capabilities. In this role, you will: Research and implement methods for safety training, reinforcement learning, and adversarial robustness. Develop evaluations, identify model failure modes, and use findings to improve training. Work with research, engineering, security, and policy partners to support safe, reliable deployment. You might thrive in this role if you: Bring 4+ years of relevant AI safety research experience, including RLHF, adversarial training, or robustness. Have a degree in computer science, machine learning, or a related field, and strong deep learning research or engineering skills. Have experience improving model safety for deployment and enjoy collaborative research. Are motivated by OpenAI’s mission and the responsible use of AI in safety-critical settings. Security Requirements Active TS/SCI clearance or equivalent. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefi
About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight : Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful agent autonomy (for example future versions of auto-review ). About the Role This role focuses on oversight and system-level mitigations that enable increasingly capable agents to operate safely and autonomously in real environments. We prioritize building oversight systems that are used in practice today, both internally and externally (see our recent work on action monitoring for codex and former code review ). We also study longer-term questions about how increasingly capable agentis systems can be supervised, constrained, and corrected. We’re looking for a safety&security minded researcher or engineer who can reason rigorously about security boundaries and agent behavior, then build and test practical mitigations. A background in AI control or security is welcome but not required. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and evaluate system-level controls for agent actions like agent-based review. Plan how they fit in a broader syste
About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight: Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful autonomy (for example future versions of auto-review ). About the Role We’re looking for strong executors with excellent judgment, comfort with ambiguity, and an understanding of frontier model research. You don’t need prior safety or alignment experience, we also welcome people that recently realized that alignment and safety is a critical area to contribute to. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Train and evaluate frontier models to reduce harmful or misaligned agent actions, forming clear hypotheses and executing independently through ambiguity. Mine incidents and build scalable measurement, data-processing, and evaluation systems that turn real failures into repeatable safety signals. Collaborate closely with post-training, capabilities, oversight, and pre-training partners to ship research-backed mitigations into large-scale training and agent systems. You might thrive in this role if you: Have demonstrated strength in research engineering, ML en
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is an AI-native developer companion that integrates deeply into every part of the development workflows. We are looking for a Product Builder to join our product team and help shape how developers experience Postman in an AI-first world. This role blends product management, engineering, and customer empathy. You should be as comfortable writing a spec and making a PR as you are building a prototype. You will work closely with engineering and design to imagine, prototype, and ship new experiences that make Postman indispensable for developers. The ideal candidate moves easily between strategy and execution and uses technical skills to accelerate product learning and delivery. You are an engineer first who enjoys solving business problems with a strong bias to action to meet customer needs. What You’ll Do Own and drive the developer journey, focusing on how AI-native experiences. Build functional prototypes and proof-of-concepts to validate ideas quickly and inform roadmap direction. Collaborate closely with engineers to shape APIs, workflows, and systems. Translate business and
About the Team The Growth Platforms team builds the systems and operating foundations that help OpenAI grow responsibly. We partner across the product portfolio to connect customer signals, identity and consent, campaign workflows, measurement, and product experiences into an AI-enabled growth engine. Our work helps teams launch, learn, and scale with a high bar for data quality, privacy, reliability, and customer trust. About the Role We’re looking for an experienced marketing technology and operations leader to drive cross-functional work at the intersection of growth, measurement, data, and automation. Your mission will be to turn fragmented tools, signals, and workflows into reliable, measurable, AI-enabled capabilities that teams can use safely at scale. You’ll work across Growth, Marketing Operations, Product, Engineering, Data Engineering, Data Science, Security, Privacy, Legal, and Revenue Operations, as well as external advertising platforms, measurement providers, and implementation partners. You’ll translate business requirements and privacy constraints into data contracts, integration designs, rollout plans, and reliable first-party data systems. This is a hands-on, high-impact role for someone who brings structure to ambiguity and moves from event schemas, APIs, and data quality assurance to operating cadences, partner enablement, and executive updates. This role is based in San Francisco or New York City with a hybrid office expectation. In this role, you will: Own the operating model for Growth’s marketing technology stack across identity, consent, audiences, activation, measurement, and experimentation. Own and operate the complete paid-media tracking and measurement system, including website pixels, server-to-server conversion events, mobile measurement integrations, identity and consent controls, attribution methods, and timely signal delivery to advertising platforms. Design and implement event schemas, data mappings, APIs, and integrations; valid
About the Team The Statsig team is responsible for the experimentation, feature rollout, dynamic configuration, and analytics systems that help OpenAI ship products with speed, safety, and evidence. Teams across ChatGPT, Codex, model measurement, monetization, business subscriptions, developer products, and shared infrastructure rely on Statsig to introduce capabilities safely, measure their impact, and make high-confidence product decisions. About the Role As a Product Lead on the Statsig team, you will define how experimentation, rollout, configuration, and analytics become a simple, reliable, and trusted part of how every OpenAI product team ships. You will set strategy across multiple product and platform workstreams, translate company-wide needs into durable capabilities, and help Statsig become a core part of OpenAI’s product development system. We’re looking for a product leader who combines strong product judgment, technical fluency, and deep analytical thinking. You should be comfortable navigating ambiguous customer needs, influencing teams across the company, and balancing rapid adoption with reliability, usability, and measurement quality. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Travel requirements should be confirmed with the recruiter before publishing. In this role, you will: Define the product vision, strategy, and roadmap for experimentation, feature management, dynamic configuration, rollout safety, and analytics. Partner with product, engineering, research, data, design, and infrastructure leaders to turn recurring launch and measurement needs into reusable platform capabilities. Develop a deep understanding of workflows across ChatGPT, Codex, model measurement, monetization, subscriptions, and developer products, then establish clear priorities across competing needs. Drive adoption by making sophisticated experimentation and analytics c
About the Team The Growth team drives user and revenue growth across ChatGPT’s consumer and business segments as well as other OpenAI products worldwide. We operate across the full funnel - from awareness and acquisition through activation, retention, and expansion - using a combination of global performance marketing, AI-powered workflows, in-product optimization, insights, experimentation, and creative ops engineering. About the Role We are hiring a Lifecycle Lead to build the company-wide owned-channel capability that helps teams reach users with relevant, timely, and trustworthy experiences. This senior, hands-on leader will set the lifecycle strategy, partner with Engineering to build the orchestration and deployment platform, and establish the operating model that allows teams across the company to launch and improve evergreen programs safely at scale. You will sit at the intersection of platform, product, and campaign strategy. You will partner with Engineering, Product, Data Science, and Analytics on the underlying systems, and with Product Marketing Managers and other client teams to design journeys that help new, active, and returning users reach value and build durable habits. In this role, you will: Partner with product to set the company-wide vision, roadmap, and operating model for lifecycle and owned-channel engagement. Partner with Engineering, Product, Data Science, and Analytics to shape the tooling and infrastructure for identity, audiences, eligibility, consent, triggers, orchestration, decisioning, frequency, experimentation, localization, quality assurance, and observability. Define scalable deployment workflows—including self-service and centrally supported paths, intake, templates, approvals, governance, service levels, and incident response—so teams across the company can launch safely and efficiently. Partner with Product Marketing Managers and other client teams to translate audience, product, and business goals into evergreen journey stra
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE This role sits at the frontier of our research agenda. You will pursue open problems at the intersection of post-training methodology and performant inference, and then collaborate with research engineering to translate findings into production systems. A meaningful portion of your time will be dedicated to research that deepens our understanding of how models learn, alignment, and architectural efficiency — questions that may not have immediate product application. The remainder will be directed toward research that solves concrete problems for Baseten's platform and customers, who are the fastest growing AI companies in the world like Cursor, Lovable, and Notion. We are looking for someone with sharp research taste and genuine creative instinct for problem selection. Someone who can identify questions that matter, design clean experiments to answer them, and push the state of the art. The environment here is not theoretical, but rather research that can be validated with eager customers who are serving billions of tokens a second. RECENT RESEARCH Towards infinite context windows: neural KV cache compaction Dense, on-policy or both? Repeated kv cache for long-running agents Distillation without the dark – replicating black-box on-policy distillation on Baseten RESPONSIBILITIES Define and pursue a research agenda spanning both foundational and applied work, with the applied component connected to Baseten's pla
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: The People Team at Notion exists to help build a generational company that enables people to do their best and most meaningful work. This role operates at all altitudes - one minute, you’ll be building a people strategy with a member of our leadership team, and the next, you’ll be coaching an employee. With Notion's scale this year, you'll play a major role in how we shape, mold, and build the future of Notion for years to come. This role is based in San Francisco, CA. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: Build deep partnerships with Notion’s leadership team (especially across Engineering, Product, and Design), and serve as a conduit between stakeholders up to C-suite and the People Team. Advise on, create, and execute people strategies that drive the business forward. Consult with the leadership team on Talent Planning, Org Design, Retention Strategies, and Change Management. Apply systems thinking to evolve
Other cities to consider
More places hiring for this role
Get new ai systems engineer jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime