Jobs in United States

Agent Ai Engineer in San Francisco

287 active opportunities · Updated October 2026

Explore current agent ai engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team API Agents builds the shared agent harness, tools, and infrastructure that turn OpenAI’s frontier models into systems that can reliably complete real work. We carry the capabilities behind Codex into a much broader set of products and workflows across software engineering, research, finance, healthcare, enterprise operations, and more. Our work spans search and connected context, computer use, memory, delegation and multi-agent coordination, and safe execution. Sitting at the intersection of Research, Codex, infrastructure, and applied product teams, we build reusable agent capabilities that compound across the ecosystem. About the Role We are looking for an experienced backend software engineer to build the core systems behind the next generation of agents. You will design reliable services and abstractions that help agents find the right context, use tools and computers, retain knowledge, coordinate over long-running workflows, and take action safely. The role combines deep backend and infrastructure work with strong product judgment, with opportunities to work across agent runtimes, orchestration, search, execution environments, identity and permissions, observability, and evaluations. This is software and systems engineering rather than model training: success comes from strong backend fundamentals, high agency, and the ability to turn fast-moving research capabilities into dependable production primitives. In this role, you will: Design, build, and operate the shared agent harness and backend infrastructure that power long-running, high-value workflows across OpenAI and third-party products. Build reusable capabilities across search and connected context, computer use, memory, tool execution, delegation, subagents, and multi-agent orchestration. Establish the foundations agents need to operate safely in production, including secure execution environments, identity and permissions, observability, evaluations, reliability, and cost and latency effi

TypeScriptPythonAWSRest
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -86%

$175K – $240K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role As a Forward Deployed Engineer on Notion’s Services team, you will lead the hands on technical delivery of our most complex customer engagements. You will design, build, and deploy production ready solutions that help enterprise customers integrate Notion as the operating layer for their business. You’ll embed with enterprise customers to design and build solutions that integrate Notion deeply into their technical and operational environments. This includes writing and maintaining custom code, designing and deploying production-grade custom agents and AI workflows with MCP, Agent APIs, and Notion’s automation and execution infrastructure, building data pipelines and resolving complex challenges around scale, permissions, and governance. You’ll embed with enterprise customers to design and build solutions that integrate Notion deeply into their technical and operational environments. This is a customer-facing engineering role for someone who is comfortable writing code, debugging technical issues, explaining tradeoffs to stakeholders, and turning ambiguous customer problems into scalable technical solutions. Your work w

JavaScriptPythonJavaNode.js
N
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -86%

From $299K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: This role will be based in San Francisco. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. You’ll build the systems that let any knowledge worker leverage fast, scalable databases without having to become a DBA. You’ll be a hands-on technical leader, helping set the architectural direction and roadmap as we take a new product from early alpha to general availability. You’ll work directly with early customers to shape foundational technical and product decisions. Then, together, we’ll work on scaling as adoption and workloads grow. This is a backend-leaning role with work spanning the stack. You might build the systems that provision and manage a fleet of customer Postgres instances, design how user-defined schemas evolve safely, or make complex queries execute efficiently. You’ll also follow those systems into the product: improving how a table loads, how users understand a slow operation, or how an agent safely works with their data. You’ll

TypeScriptReactNode.jsSQL
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. Our mission is to push the frontier of code generation and agentic reasoning, and deploy these capabilities in real-world products such as ChatGPT and the API, as well as in next-generation tools specifically designed for agentic coding. We operate across research, engineering, product, and infrastructure—owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. About the Role As a Performance & Systems Engineer on the Codex team, you will be responsible for whole-system optimization across a complex, evolving stack. Codex spans LLM inference, cloud orchestration, agentic work management, and multiple product surfaces. Your job will be to identify and land high-leverage changes—across infrastructure, modeling, and product layers—that make Codex agents significantly faster and cheaper to serve. We’re looking for generalists who thrive in ambiguity and love chasing performance bottlenecks to ground. This is a high-ownership role where your work will directly improve the experience of millions of users. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Hunt down and address inefficiencies across the Codex system stack, from agent behavior to LLM inference to container orchestration, and beyond. Build tooling to measure, profile, and optimize system performance at scale. Collaborate with researchers and engineers to land high-ROI changes that improve latency and cost. You might thrive in this role if you: Have experience operating across both ML systems and cloud infrastructure. Enjoy diving into messy, ambiguous problems and emerging with clear wins. Think holistically about performance, balancing spee

AWSRestAIRust
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -86%

$280K – $330K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the role Notion's Search & Context Platform is the substrate that powers how 100M+ users and, increasingly, Notion's agents find and reason over the right information. This team owns the search infrastructure and indexing systems powering lexical and semantic retrieval, the platform primitives for managing agent context and memories , and the scalability, performance, security, and enterprise capabilities that make all of it production-grade. Data growth is faster than ever, and the systems that power this need to rapidly evolve to support an order of magnitude growth—both in the volume of content we index and in the load that agents now place on the retrieval layer. As the Engineering Manager for this team, you'll lead a technically deep group of engineers building platform systems used by multiple product teams. Your most important customers are the Search & Context product team and the AI team (building agents on top of these primitives). You'll set the technical direction, product manage the platform scope on behalf of those customers, and make hard tradeoffs to move fast in service of product velocity — while keepi

B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.1%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE The largest, most demanding enterprises run on Baseten, and they bring exacting requirements for how people, services, and agents access the platform. This is the founding role for our identity and authorization team within enterprise engineering. You'll own the identity and access layer of the Baseten platform: the authorization model, credential systems, and admin experiences that enterprise IT teams use to govern access for organizations like Harvey, HubSpot, and Notion. You'll design and build Baseten's fine-grained authorization system from the ground up to support the workflows customers depend on today while giving them cleaner, more precise ways to manage access as the platform grows. Authorization at Baseten requires low-latency permission checks at high request volume, consistent contracts and behaviors across the product suite, and strong security guarantees for mission-critical, highly regulated workloads. EXAMPLE INITIATIVES Recent and upcoming work in this area: Fine-grained authorization for users, service accounts, and agentic workloads: per-resource permissions at the organization, team, and workload scope to support both common workflows and complex enterprise access policies Programmatic authentication allowing high-compliance customers to connect service principles securely via short-lived, workload-based credentials Agent credentials that grant an agent exactly the access it needs for the gi

PythonKubernetesMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Post-Training Frontiers team is responsible for training the frontier agents OpenAI ships to the world (GPT-Next). We train the flagship agentic models behind Codex, ChatGPT, and the API through large-scale reinforcement learning. The team’s work spans four areas. First, execution and science: working with teams across OpenAI to decide what can go into the final model and how, using scientific experiments and evals that are representative of the final pipeline so issues can be recognized early. Second, RL scaling: executing the final large-scale reinforcement learning run, making sure GPUs are used efficiently and training stays healthy. Third, research: improving horizontal capabilities like instruction following, factuality, memory, and multi-agent behavior, where the team’s broad visibility helps identify cross-cutting improvements across teams and domains. Fourth, engineering: maintaining the infrastructure stack and internal tools to ensure that both the final run and all integrations go as smoothly as possible and that the systems are easy to work with. About the Role This role focuses on keeping our frontier RL training runs fast, reliable, and unblocked. You will work across engineering and infrastructure problems as they emerge, from scaling and orchestration issues to inference bottlenecks, numerical problems, and hardware failures, as well as supporting large horizontal integrations in the big run, like multi-agent capabilities or memory. This is a role for a strong generalist who quickly learns anything needed for the task, has high attention to detail, debugs deeply, and is motivated by fixing the highest-impact problem in front of the team. In this role, you will: Keep large-scale async RL training runs moving by jumping into the most urgent engineering and infrastructure problems. Debug issues across training systems, inference, orchestration, scaling, and distributed infrastructure. Improve the reliability and efficiency of RL trai

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of this API & power-users team, you will improve the capabilities, reliability, and product fit of OpenAI’s agentic models for power users and API developers. You might design evals from real developer workflows, build training environments around production-like tool use, turn qualitative model failures into training data, evals, or post-training interventions, or drive a behavior improvement from discovery through post-training, integration, and launch. This role is intentionally broad. The strongest candidates are comfortable turning ambiguous model behavior problems into concrete progress, whether that means improving tool use, planning, instruction following, recovery from mistakes, or how models behave in API-based workflows. You should be excited to work across research, engineering, data, evals, and product to make models better at acting in real workflows. You will work closely with researchers, engineers, API/product teams, Codex, infrastructure, and safety/align

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Platform Analytics team builds the systems OpenAI researchers use to understand the quality and behavior of the models we train including what models are doing, why they behave in a particular way, and how that behavior changes across experiments. Neptune is a core part of this work. It ingests, stores, queries, and visualizes large volumes of metrics from pretraining, post-training, and reinforcement learning. Hundreds of researchers depend on these systems in their daily work to compare experiments, debug unexpected behavior, and decide what to try next. Our scope is broader than metrics. We also build platforms that help researchers analyze samples, traces, evaluation results, and other structured or unstructured data through dashboards, APIs, and increasingly agent-driven workflows. These systems need to remain fast, reliable, and understandable as the scale and complexity of research change quickly. We are not trying to become a consulting team that builds a separate solution for every research project. We work directly with researchers to understand recurring problems, then turn them into reusable infrastructure and platform capabilities that many teams can build on. About the Role We’re looking for a hands-on experienced software engineer who can take ownership of a critical system and drive it from problem definition through production adoption. This person should be able to own a platform such as CacheHouse end to end: define its technical direction, design its data model and storage architecture, integrate it with several research dashboards and workflows, guide one or two engineers, and ensure the system works reliably for its users. The right candidate should already bring the technical judgment, ownership, and execution expected at this level. The primary learning curve should be OpenAI’s stack and research problem space, not learning how to lead a complex engineering effort or deliver a production system. You will work directly with

AWSRestAIC++
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t

AWSKubernetesRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight : Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful agent autonomy (for example future versions of auto-review ). About the Role This role focuses on oversight and system-level mitigations that enable increasingly capable agents to operate safely and autonomously in real environments. We prioritize building oversight systems that are used in practice today, both internally and externally (see our recent work on action monitoring for codex and former code review ). We also study longer-term questions about how increasingly capable agentis systems can be supervised, constrained, and corrected. We’re looking for a safety&security minded researcher or engineer who can reason rigorously about security boundaries and agent behavior, then build and test practical mitigations. A background in AI control or security is welcome but not required. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and evaluate system-level controls for agent actions like agent-based review. Plan how they fit in a broader syste

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight: Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful autonomy (for example future versions of auto-review ). About the Role We’re looking for strong executors with excellent judgment, comfort with ambiguity, and an understanding of frontier model research. You don’t need prior safety or alignment experience, we also welcome people that recently realized that alignment and safety is a critical area to contribute to. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Train and evaluate frontier models to reduce harmful or misaligned agent actions, forming clear hypotheses and executing independently through ambiguity. Mine incidents and build scalable measurement, data-processing, and evaluation systems that turn real failures into repeatable safety signals. Collaborate closely with post-training, capabilities, oversight, and pre-training partners to ship research-backed mitigations into large-scale training and agent systems. You might thrive in this role if you: Have demonstrated strength in research engineering, ML en

AWSRestAIRust
S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -72.4%

$220K – $450K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role AI and machine learning are reshaping how developers debug, monitor, and ship software, and Sentry is uniquely positioned to lead that shift. We sit on a novel and massive dataset of real production errors, spans, and logs from tens of thousands of engineering organizations — the kind of signal that makes ML genuinely useful, whether it's a clustering model that groups related issues, a ranking system that surfaces the right alert at the right time, or an agent that proposes a fix. We're looking for an Engineering Manager to lead and grow our Machine Learning Engineering team. This team owns the full spectrum of ML at Sentry: classical techniques like clustering, ranking, anomaly detection, and embeddings that quietly power core product surfaces today, alongside the LLM-based and agentic systems shaping where the product is headed. You'll partner closely with product, design, and engineering leaders to decide where ML belongs in our products, what kind of ML actually fits the problem, and how we translate that work into experiences millions of developers rely on every day. In this role you will Set technical direction across the team's full ML surface area — from classical models for clustering, ranking, and anomaly detection to LLM-based and agentic systems — and make sharp calls about which approach fits each problem Define how the team evaluates and monitors ML systems in production, from offline metrics to online experimentation to model and agent observability Stay hands-on enough to review code and model designs, contribute to architecture discussions, and unblock engineers on complex ML problems Define

RestMachine LearningAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Connectors, you will teach models how to interface with the top professional software using code. You will help train agents to use code, APIs, tools, and structured integrations to operate across applications like Slack, Google Workspace, GitHub, Notion, Linear, Salesforce, and other core systems of work. You will help enable models to take useful actions across a user’s digital context: finding information, updating systems, coordinating work, generating artifacts, and completing multi-step workflows through the tools teams already use. You will train models to be supercharged by the world’s most important productivity and enterprise software, turning connected tools into a powerful action surface for our agents. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people.

AWSGitRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, you will improve the capabilities, reliability, and product fit of OpenAI's agentic models. You might own a research direction, build the infrastructure that makes large training runs faster and more trustworthy, create evals that reveal where models fail, or drive a capability from an idea through experimentation, integration, and launch. This role is intentionally broad. The strongest candidates are not defined by one method or subfield; they are people who can take an ambiguous capability problem and make progress across research, engineering, data, evals, and product. You should be excited to work on models that act in the world: writing and debugging code, using tools, calling functions, operating computers, collaborating with other agents, and completing valuable work on behalf of users. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, meas

AWSRestMachine LearningAI
🔔

Get new agent ai engineer jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime