Jobs in United States

Agent Ai Engineer in San Francisco

287 active opportunities · Updated October 2026

Explore current agent ai engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a researcher working on Frontier Evals & Environments, you will help build north star model environments to drive progress towards safe AGI/ASI. Your work will directly guide the research programs of the most ambitious training runs happening at OpenAI. Some prior open-sourced evaluations built by researchers in this role include GDPval , SWE-bench Verified , MLE-bench , PaperBench , and SWE-Lancer . If you are interested in feeling firsthand the fast progress of our models, and steering them towards good outcomes, this is the role for you. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you might Create ambitious RL environments to push our models to their limits, and measure frontie

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role We believe that the final enabler for AGI is spending compute on context. As a Context Researcher on Agent Post-Training, you will scale compute spent on context. You will get to work in our frontier training stack on enabling the next paradigm of model training with a clear product interface for iterative deployment (Codex Chronicle). You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you might Design and run experiments that improve scaling of compute on context. Own end-to-end improvements to the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis. Build evals and environments that expose the next set of model failures,

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Computer Use, you will teach models to operate computers. You will help train models that can navigate browsers and desktops, use tools and applications, reason through complex workflows, collaborate with users and other agents, and complete long-horizon tasks with reliability and judgment. This work sits at the intersection of frontier model training, product behavior, evaluation, and systems engineering, and will directly shape the computer-use capabilities shipped in OpenAI’s next generation of agents. Currently, our models are the best in the world at this behavior! You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you might Design and run experiments th

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Artifacts, you will train frontier models to create polished, useful work products: documents, spreadsheets, slide decks, dashboards, reports, analyses, and other interactive or editable artifacts. You will help teach our models to move from a vague user goal to a finished artifact with strong structure, visual taste, domain judgment, correctness, and low latency. This work will require owning improvements across our post-training stack, including RL, data pipelines, graders, reward signals, evals, and behavioral analysis. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you will: Design and run experiments that improve agentic model behavior for complex so

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Codex Research team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of the Codex Research team, you will improve the capabilities, reliability, and product fit of OpenAI's agentic models. You might own a research direction, build the infrastructure that makes large training runs faster and more trustworthy, create evals that reveal where models fail, or drive a capability from an idea through experimentation, integration, and launch. This role is intentionally broad. The strongest candidates are not defined by one method or subfield; they are people who can take an ambiguous capability problem and make progress across research, engineering, data, evals, and product. You should be excited to work on models that act in the world: writing and debugging code, using tools, calling functions, operating computers, collaborating with other agents, and completing valuable work on behalf of users. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measu

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Platform and Infrastructure Engineering organization advances the mission of deploying artificial general intelligence (AGI) for the benefit of all by delivering secure, scalable, and resilient technology solutions. Our team builds and maintains robust infrastructure that safeguards OpenAI’s data and systems while ensuring employees are well-equipped and seamlessly connected. By prioritizing security, reliability, and user-centric solutions, we empower OpenAI employees to drive impactful AI research, corporate operations, and product innovation. About the Role As a Software Engineer: Internal Applications, Enterprise, you will build internal products that make technology support and administration safer, faster, and less dependent on manual intervention. You will help reduce reliance on broadly privileged human actions, turn recurring technology problems into paved paths, and build agentic systems that can help resolve tickets end to end. A core part of the role is building the interfaces that bring employees, AI agents, and human responders together in a shared ITSM experience, with the right context, controls, and handoffs at each step. We are seeking engineers who enjoy working across frontend and backend layers on ambiguous, high-leverage enterprise problems. You should bring strong product judgment, solid backend engineering fundamentals, and an interest in building software that changes how technology support, system administration, and agent-assisted operations are delivered. The best fit will care as much about the quality of the operator and employee experience as the correctness of the backend systems behind it. In this role, you will: Build frontend experiences that let employees request help, let agents gather context and take safe actions, and let human responders review, approve, or take over without losing the thread. Reduce reliance on broadly privileged manual actions by replacing them with narrow, auditable, policy-aware aut

AWSRestAIRust
P
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity As agentic AI becomes the core fabric of how engineering is done worldwide, we want to be the leaders in how agents interact with APIs. This transformation requires solving agentic identity, extending or building new standards and protocols, building agent-to-agent interactions including authorization, licensing, trust delegation, security and monetization. We are looking for an exceptional Staff Engineer to help us build the next generation of our Business Platform to support the agentic era. You will be responsible for setting the vision for the long term architecture of various components in the Business Platform. You’ll be accountable for strategy, technical roadmap and architecture of the platform. You’ll work with multiple teams, overseeing and contributing to delivery and owning critical operational KPIs such as availability and uptime. Your role involves designing, implementing, and running critical tier 0 services for the company. We’re looking for a seasoned individual contributor leader who can work effectively with product management and engineering leaders within the company and build a high quality pla

JavaScriptPythonJavaNode.js
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Future of Computing Research team is an applied research team within the Consumer Devices group focused on developing new methods, models, and evaluation frameworks that support our vision for the future of computing. We work at the frontier of multimodal AI, helping turn emerging model capabilities into product experiences that are useful, delightful, and worthy of long-term trust. Our work explores a new class of AI systems that can learn over time, adapt to individuals, and support people in the flow of daily life. This includes long-term memory, user modeling, and personalization systems that are aligned not just with immediate satisfaction, but with a person’s broader goals, values, and well-being. We work closely across research, engineering, design, product, and safety to define what it means to build AI systems that know you over time, act at the right moment, and help in ways that are context-aware, respectful, and demonstrably beneficial. About the Role We are looking for a Research Engineer / Scientist to join the Future of Computing Research team to work on RLHF and post-training for personalized, multimodal AI systems. This role will focus on building the learning and evaluation foundations that help models become more context-aware, adaptive, and useful over time. You will work on problems such as reward modeling, preference learning, long-horizon evaluation, and policy improvement for systems that must make high-quality behavioral decisions in realistic user settings. The work is deeply product-grounded: success is not just higher benchmark performance, but better model behavior in real-world use. The ideal candidate is excited about pushing beyond one-turn assistant behavior toward systems that improve through feedback, learn from richer signals, and are trained against meaningful notions of user value. Internally, that maps closely to the need for careful reward design, feedback loops, and evaluation frameworks that test whether i

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI's mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect model weights, customer data, and critical systems across multiple cloud environments. The team partners across OpenAI, including Applied Engineering, Research, IT, Security, Infrastructure, and Engineering, to provide secure and scalable platforms for identity, access management, permissioning, orchestration, and safe AI research. About the Role We’re looking for an engineering leader to lead Identity Infrastructure Engineering, the team building the systems that govern and scale access across OpenAI’s research, engineering, and internal platforms. This role sits at the center of cloud infrastructure, identity, software engineering, and security-critical operations. You’ll lead engineers building control planes, policy systems, workload and agent authorization patterns, infrastructure-as-code, and operational foundations that help OpenAI move quickly while keeping access reliable, auditable, least-privileged, and safe under failure. The ideal candidate has led teams responsible for large-scale, mission-critical infrastructure. They can go deep into code and architecture when needed, while giving engineers and technical leads the clarity and ownership to do their best work. They set technical direction, grow strong teams, make durable architecture decisions, and turn ambiguous 0-to-1 problems into platforms OpenAI can trust and build on for years. In this role, you will: Build and lead a high-performing Identity Infrastructure team, going deep enough technically to set direction while empowering the team to own delivery. Define the strategy for identity platform as the policy plane for access across people, agents, workloads, services, clouds, and internal systems. Scale Acc

AWSGitRestAI
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The ChatGPT product team is a rapidly evolving, high-impact group within OpenAI that builds intuitive, safe, and useful AI-powered experiences for millions of people worldwide. Our team brings together engineering, design, research, and product to explore how conversational AI can help people learn, create, and solve problems. AI has the potential to transform how millions of people learn, teach, and achieve their goals. Learning is one of the top use cases on ChatGPT— not only for students, but also for adults building new skills in a rapidly changing world. The Education team is focused on advancing how humans learn with AI and working to make high-quality learning accessible. We aim to deliver measurable gains in cognition and achievement, partnering closely with students, educators, and country leaders to ensure new tools are safe, effective, and trusted. About the Role As a Product Manager for Education & Learning, you will shape the strategy and build products to advance learning experiences in ChatGPT— defining a category that could reshape classrooms, careers, and lifelong learning. You’ll collaborate with cross-functional partners—from AI research to design to go-to-market—to build ChatGPT into a true learning agent and ensure those products you build meet the needs of stakeholders throughout the ecosystem. You will also partner closely with product teams across growth, youth well-being, and enterprise to ensure alignment and maximize impact. This role is based in San Francisco, CA. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you will: Define and drive the product vision for education and learning in ChatGPT. Partner with research teams to explore how AI can better advance learning outcomes Collaborate with cross-functional teams—including design, engineering, product, and go-to-market—to bring education features to life. Partner with xfn growth, GTM, and

Artificial IntelligenceAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's Enterprise team builds AI-powered enterprise products and shared platform capabilities that help organizations put advanced AI to work securely and at scale. Our work spans enterprise workflows, agent experiences, integrations, identity, administration, security, governance, and deployment. About the Role As a Technical Program Manager on Enterprise, you will lead the technical strategy and execution behind the products and shared capabilities that make ChatGPT, Codex, and future OpenAI products useful, secure, and scalable for organizations. You will translate customer needs, competitive dynamics, and product priorities into actionable plans, influence architectural direction, and deliver durable capabilities across application, platform, and infrastructure layers. The role requires deep technical fluency, strong product judgment, and the ability to move between hands-on execution and broader enterprise strategy. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Drive technical strategy and execution for enterprise product and AI workflow initiatives, from design through implementation, launch, customer rollout, and iteration. Partner with engineering teams to influence architectural direction, interface definitions, and implementation tradeoffs across full-stack products, APIs, integrations, and shared platform systems. Translate enterprise customer requirements into actionable product priorities across AI-powered workflows, agent experiences, integrations, permissions, data access, evaluations, identity, security, governance, and deployment readiness. Represent the needs of enterprise buyers, IT administrators, security teams, business leaders, developers, and end users in product and technical decisions. Identify adoption barriers, competitive gaps, and opportunities to make OpenAI products easier for organizations

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s mission is to build safe artificial general intelligence (AGI) that benefits all of humanity. Achieving this requires bringing together world-class scientists, engineers, and business leaders to translate frontier research into real-world impact. Within OpenAI, the Go-to-Market organization helps customers understand, adopt, and scale our products across their businesses. The team includes Sales, Solutions, Support, Marketing, Partnerships, and Strategic Pursuits professionals who work together to bring the benefits of AI to organizations globally. About the Role At OpenAI, we're building toward an agent-first world where AI systems reason, act, transact, and create alongside people and businesses. As agents become a primary interface for work, the Strategic Pursuits Team leads our most complex enterprise opportunities and helps define how customers adopt OpenAI at scale. We're seeking a Strategic Pursuits Lead to shape and close high-impact enterprise opportunities across Frontier, AI Solutions, ChatGPT, Codex, API, and emerging offerings. This person partners with customers, account teams, Product, FDE, Legal, Finance, GTM, and executives to turn ambiguous demand into a clear vision, business case, commercial structure, delivery path, and repeatable GTM motion. This is a builder role for someone who thrives in ambiguity, can operate without a playbook, and is energized by creating the playbook while running the pursuit. The right person brings executive presence, commercial judgment, product intuition, and the ability to make messy work executable. This role can be based in our San Francisco or Seattle offices. We are also considering applications to work remotely from within the U.S. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead strategic pursuits from early qualification through customer alignment, proposal, commercial structure, and transition in

AWSRestAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. Product at Baseten Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the first people who will help define it. You'll work directly with our founders and with some of the best systems and infrastructure engineers in the world, and you'll set the standard for building great AI Infrastructure. PMs at Baseten don't sit above engineers - you earn ownership by being technical, finding the truth in front of customers, building great cross-functional relationships, and shipping great product experiences. The role Getting a model into production still takes real expertise — choosing a serving engine, sizing hardware, tuning it, wiring it into an app. We want a developer to go from "it runs on my laptop" to "it's serving production traffic" in minutes, on their own. You'll own the entire experience a developer touches to deploy and iterate: the CLI and SDKs, the console, onboarding, model discovery, deployment configuration, truss, and the increasingly agent-driven ways developers build. Your job is to make Baseten synonymous with Great DevEx and make it effortless to drive and self-serve deploy models on Baseten for far more developers than it is today. Impact and outcomes you'll drive You will collapse time-to-production — take a developer from first sign-up to a running, maint

Machine LearningAIGoRust
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why This Role Is Different This is not a typical “Applied Scientist” or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work directly with enterprise customers on problems that push LLMs to their limits. You’ll rapidly understand customer domains, design custom LLM solutions, and deliver production-ready models that solve high-value, real-world problems. Train and customize frontier models — not just use APIs. You’ll leverage Cohere’s full stack: CPT, post-training, retrieval + agent integrations, model evaluations, and SOTA modeling techniques. Influence the capabilities of Cohere’s foundation models. Techniques, datasets, evaluations, and insights you develop for customers will directly shape the next generation of Cohere’s frontier models. Operate with an early-startup level of ownership inside a frontier-model company. This role combines the breadth of an early-stage CTO with the infrastructure and scale of a deep-learning lab. Wear multiple hats, set a high technical bar, and define what Applied ML at Cohere becomes. Few roles in the industry combine application, research, customer-facing engineeri

PythonGitAIGo
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why This Role Is Different This is not a typical “Applied Scientist” or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work directly with enterprise customers on problems that push LLMs to their limits. You’ll rapidly understand customer domains, design custom LLM solutions, and deliver production-ready models that solve high-value, real-world problems. Train and customize frontier models — not just use APIs. You’ll leverage Cohere’s full stack: CPT, post-training, retrieval + agent integrations, model evaluations, and SOTA modeling techniques. Influence the capabilities of Cohere’s foundation models. Techniques, datasets, evaluations, and insights you develop for customers will directly shape the next generation of Cohere’s frontier models. Operate with an early-startup level of ownership inside a frontier-model company. This role combines the breadth of an early-stage CTO with the infrastructure and scale of a deep-learning lab. Wear multiple hats, set a high technical bar, and define what Applied ML at Cohere becomes. Few roles in the industry combine application, research, customer-facing engineeri

PythonGitAIGo
🔔

Get new agent ai engineer jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime