About the Team The Agent Infrastructure team at OpenAI is responsible for building systems that enable training and deployment of highly useful AI agents, both internally and for the world. We work hand-in-hand with researchers to design and scale the environment in which agentic models are trained – providing a workspace for AI models to execute code, debug issues, and develop software just as human SWEs do. Our training environment for agentic models operates at an extremely high scale and has the flexibility to emulate any environment in which an agent might work. At the same time, our team builds and maintains OpenAI’s core platform for the deployment and execution of agents in production. Our systems power products such as Codex, Operator, tool use in ChatGPT, and future agentic products. Some of the most challenging technical problems in scaling the capabilities and utility of agents and agentic models lie in the infrastructure layer – and our team is focused on building the research and production systems that enable OpenAI to train the most capable models in the world, and maximize the utility of our agentic products for users around the world. About the Role As a Software Engineer on the Agent Infrastructure team, you will have the opportunity to work closely with both research and product at OpenAI - building and scaling systems to train highly capable agentic models, and building the platform and integrations to launch new agents to hundreds of millions of users worldwide. Your work will consist of both building new capabilities - standing up the infrastructure and integrations needed to train more complex agentic models - and rapidly scaling these new capabilities to some of the largest compute clusters in the world. At the same time, you’ll be instrumental to the launch of agentic products at OpenAI - building, maintaining, and scaling the production platform on which all agents run. We’re looking for people with deep experience building AI infrastructu
Jobs in United States
Agentic Risk Analyst in San Francisco
137 active opportunities · Updated October 2026
Showing
15 jobs
Explore current agentic risk analyst jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Platform team powers how millions of developers and enterprises build with our models. We provide APIs and agentic solutions used by global startups and fortune 500s. We work closely with product, engineering, design, and go-to-market to build a world-class platform that pushes the frontier of AI capabilities. About the Role As a Data Scientist on the Platform team, you will drive a data-driven culture for OpenAI’s API and B2B solutions. You’ll define the metrics that matter for developer success and enterprise value, measure the impact of new models and features, and partner with PMs and engineers to improve model quality, reliability, latency, and cost. Your work will shape how thousands of products adopt agentic AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will Embed with the Platform product team as a trusted partner, uncovering ways to improve developer experience, reliability, and usage growth Define north-star metrics across the developer funnel (activation, retention, growth), as well as latency/cost guardrails for new features and models Design and interpret A/B tests and controlled rollouts (e.g., new model versions, pricing/limits, new API features, new B2B products) Build source-of-truth dashboards and self-serve data tools for product, engineering, and go-to-market teams Translate product learnings into actionable feedback for Research (e.g., failure modes, eval gaps, model response quality) You might thrive in this role if you have 5+ years in a quantitative role in ambiguous, high-growth environments (platforms, APIs, or B2B products a plus) Depth in SQL and Python, with a track record proposing, designing, and running rigorous experiments Experience defining and operationalizing metrics from scratch (including reliability/latency/cost and safety) Strong cross-functional communication with PMs, enginee
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role Modal's LLM inference platform delivers frontier performance for open-source models with best-in-class elasticity and developer experience, made in part possible by our custom runtime with GPU memory snapshots and multi-cloud substrate . We're looking for a leader to own the direction and execution of this platform to continue to establish us as the clear market leader, working closely with customers like Cognition, Doordash, Ramp, and many more. You'll be leading a group of highly talented engineers working on our market-leading LLM inference offering, spanning the serving stack, routing infrastructure, internal agentic optimization platform, and the user-facing product surface area. This is a hands-on leadership role — expect to split your time between technical contribution, product shaping and people management depending on what the team needs. You'll set direct
From $177.2K/yr
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . We're looking for a Staff Software Engineer to lead the technical direction of the backend systems powering Pinterest's AI-driven product experiences — Pinterest Assistant, visual editing and content creation tools and future LLM-based products. You'll design and ship backend systems while architecting the broader platform strategy that enables these experiences to scale across surfaces and teams. This is a hands-on leadership role where you'll move between deep technical execution, system-level architecture and cross-team technical leadership. What you'll do: Define the backend and platform architecture for AI-driven product experiences — visual-chat, AI image generation and editing, and agentic or LLM-based products — partnering with Engineering, Product, ML and UX leaders to shape the technical vision and roadmap. Architect end-to-end systems
About the Team The Intelligence and Investigations team is dedicated to ensuring the safe, responsible deployment of AI by rapidly detecting and mitigating abuse. Our team leverages the latest testing methodologies to uncover vulnerabilities and emerging threats, helping safeguard OpenAI’s products and users. We work closely with cross-functional partners across product, policy, and engineering to drive a comprehensive defense strategy against evolving adversarial challenges. About the Role As a Red Team Specialist focused on cyber, you will help answer two practical questions: What cyber capabilities can our models provide to real-world attackers, and do our safeguards remain effective when those attackers use increasingly sophisticated techniques? The role combines scaled evaluation with expert-driven testing. You may bring deeper experience in cybersecurity and use that expertise to judge whether a model’s behavior meaningfully changes attacker capability. Alternatively, you may bring deeper experience in model evaluations, automation, or agentic harnesses and apply those skills to building rigorous cyber testing. We do not expect every candidate to be equally deep in both areas, but successful candidates will have a strong foundation in one and enough fluency in the other to work effectively across the boundary. Most of your work will focus on model cyber capabilities and safeguards; you will also spend a portion of your time testing novel abuse risks in agentic systems. This role is located in San Francisco, CA or Seattle, WA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and run rigorous evaluations of model cyber capabilities and safeguards, including policy adherence, correct refusal, over refusal, and resilience to jailbreaking and other adversarial techniques. Conduct hands-on testing to understand what models can enable when used by experienced security practiti
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role As a Senior Software Engineer on Sentry’s AI team, you’ll be directly responsible for developing the platform used by our debugging agents. This role is crucial; you will be at the forefront of integrating AI and machine learning into our core products, from issue triage and resolution to predictive analytics for application performance monitoring. Your work will help companies around the globe gain actionable insights into their software, enabling them to build better products, faster. In this role you will Build state-of-the-art agentic AI platforms to triage, debug, and solve real production issues Leverage Sentry’s novel (and massive) dataset of errors, spans, and profiles Own the development of major initiatives in the AI/ML space You'll love this job if you Are driven by impact and enjoy working on high-stakes, high-visibility projects Enjoy building things. You will have the opportunity to join the AI/ML team as one of its foundational members Thrive in cross-functional teams and enjoy building features alongside developers and product teams Qualifications Minimum 5+ years of professional experience with Bachelor’s degree in computer science, machine learning, or a related field Demonstrated expertise building production-grade agentic systems and tools You are comfortable writing production quality code (we use Python and Typescript) Familiarity with deep learning frameworks (we use PyTorch) Familiarity in deploying machine learning models at scale in production environments The base salary range (or hourly wage range, if applicable) that Sentry reasonably expects to pay for this position is $155,000 to
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role As a Staff Machine Learning Engineer on Sentry’s AI/ML team, you’ll be directly responsible for developing the models and agents used to make our product smarter and more capable. This role is crucial; you will be at the forefront of integrating AI and machine learning into our core products, from issue triage and resolution to predictive analytics for application performance monitoring. Your work will help companies around the globe gain actionable insights into their software, enabling them to build better products, faster. In this role you will Build state-of-the-art agentic AI systems to triage, debug, and solve real production issues Leverage Sentry’s novel (and massive) dataset of errors, spans, and profiles Own the development of major initiatives in the AI/ML space You'll love this job if you Are driven by impact and enjoy working on high-stakes, high-visibility projects Enjoy building things. You will have the opportunity to join the AI/ML team as one of its foundational members Thrive in cross-functional teams and enjoy building features alongside developers and product teams Qualifications Minimum 4+ years of professional experience with a MS/PhD degree in computer science, machine learning, or a related field Minimum 6+ years of professional experience with Bachelor’s degree in computer science, machine learning, or a related field Demonstrated expertise building production-grade agentic systems and tools You are comfortable writing production quality code (we use Python) Expertise with deep learning frameworks (we use PyTorch) Familiarity in deploying machine learning models at scale in production
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role At Sentry, Support is an engineering discipline. Our customers are the greatest technical minds in the world—developers at elite enterprises building the future of software—and they deserve answers that go deeper than a knowledge base link. We're looking for an APAC Technical Support Engineer based in San Francisco to join our global Support Engineering team. This role is designed to provide APAC coverage to our users; with the shift being Sunday through Thursday 4PM-12AM PST. We are architecting the Technical Support engine . We’re looking for an experienced engineer to help us redefine the standard of technical support by combining deep human expertise with autonomous agentic systems. You are a debugger of both code and systems. You will treat support volume as a data signal to build automated resolution paths, ensuring our human engineers only touch the most complex, high-impact architectural puzzles. Sentry Support Engineers aren't just clearing queues; they are Orchestrators . You will engage with our users across GitHub, Discord, and our internal systems, while acting as the Technical Lead for our Agentic Ops. You ensure that when a developer asks a complex question, our systems have the right context and a seamless "Human-in-the-Loop" path to you when deep, nuanced expertise is required. In this role you will Master the Sentry Ecosystem & Support Elite Developers Deep-Dive Debugging: Perform root-cause analysis on complex issues and distributed tracing gaps across polyglot environments. Support the Great Minds: Act as a strategic consultant for senior engineers at our largest enterprise customers, s
From $183K/yr
About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. At Flexport, we are building the unified Supply Chain Operating System for global trade. While traditional SaaS companies sell static dashboards and walk away, Flexport's Forward Deployed Engineers (FDEs) embed directly on the frontlines with enterprise clients (such as Fortune 500 retailers, automotive manufacturers, and tech giants) to solve high-stakes supply chain challenges. You are an operational strike-team leader who sits at the intersection of full-stack software engineering, applied AI, operations research, and executive client strategy. Deployed to client command centers and logistics hubs, you will connect fragmented enterprise data, architect solutions, and build agentic AI systems that optimize real-world cargo moving across ocean, air, and land. You won't just write code in a vacuum, you will embed with client engineering leaders and VPs of Supply Chain, map messy operational reality into production software, and deploy autonomous workflows that directly eliminate millions in logistics waste. You Will Embed & Map: Travel to client sites to map end-to-end supply chain processes, audit legacy systems, uncover hidden financial leaks, and translate c
About the Team OpenAI’s mission is to ensure that general-purpose artificial intelligence benefits all of humanity. The Payments team works across product, engineering, design, and finance to build the financial infrastructure that makes OpenAI’s products accessible to consumers and enterprises around the world. As AI introduces new ways for people and organizations to work, the team is defining how to support and monetize emerging forms of product usage, from usage-based pricing to agentic work. We’re building the foundational systems that help OpenAI products deliver clear, reliable, and scalable payment experiences while ensuring that this powerful technology is deployed responsibly. About the Role In this role, you’ll lead design for one of OpenAI’s most foundational product areas: the payments and monetization infrastructure that supports our consumer and enterprise products. You’ll partner closely with product, engineering, and cross-functional teams to shape how customers understand, manage, and pay for entirely new kinds of AI usage. Your work will extend beyond traditional checkout and billing. You’ll help define the systems, frameworks, and experiences behind durable pay-as-you-go models, Codex usage, and agentic workflows, translating complex business and technical requirements into intuitive experiences. As a product designer in a highly ambiguous and rapidly evolving space, you’ll influence both product strategy and the underlying infrastructure that OpenAI products depend on. This role is based in our San Francisco HQ. We offer relocation assistance to new employees. In this role, you will: Lead the design direction for foundational payments, billing, and monetization experiences across OpenAI’s consumer and enterprise products. Design and ship high-quality, end-to-end product experiences, from early systems and interaction concepts to high-fidelity prototypes and production-ready designs. Shape the infrastructure and product frameworks that support em
About the Team OpenAI’s GTM Data Science team helps shape how our products are adopted, monetized, and scaled across organizations. We work at the intersection of Product, Go-to-Market, Finance, Research, and Data, turning product usage, customer evidence, and market signals into decisions that grow durable enterprise value. We’re looking for a senior Data Scientist to own the analytical strategy for enterprise knowledge-worker adoption. As ChatGPT Work and Codex become capable of research, analysis, document creation, spreadsheets, presentations, internal knowledge synthesis, and other agentic workflows, you will help determine how these products become embedded in everyday work—not merely tried once. About the Role You will define how we measure activation, retained usage, workflow depth, and value across ChatGPT Work, Codex, and connected enterprise systems. You will explain why adoption succeeds or stalls and identify the product, enablement, and commercial interventions most likely to create durable usage. This is a hands-on, zero-to-one role. You will work through imperfect telemetry, overlapping product surfaces, evolving definitions, and ambiguous business questions. You will partner closely with GTM, Product, Finance, Research, Customer Deployment, Analytics Engineering, and Data Science. Your work is successful when it changes a product, GTM, or investment decision. In This Role, You Will Define a trusted measurement framework for knowledge-worker adoption, including identity, eligible populations, activation, retained usage, penetration, workflow depth, feature adoption, and monetization. Map the knowledge-worker journey from initial exposure through first successful task, repeated workflows, multi-surface usage, and durable adoption. Identify which personas, functions, use cases, product capabilities, and account conditions are associated with deep and retained usage. Design and evaluate experiments and quasi-experiments across onboarding, enablement, wo
About the Team OpenAI’s API Multicloud team is responsible for extending OpenAI’s API platform into strategic cloud environments, starting with AWS . The team’s mission is to distribute OpenAI’s API broadly and safely by enabling key API technologies in cloud-native environments, in close partnership with Amazon and internal teams across Codex, Research, Safety Systems, and Applied. The team is focused on bringing core developer and enterprise capabilities into cloud-native environments, including cloud-hosted Codex, model customization / post-training as a service, and new stateful runtime environments for agentic workloads. This work sits at the intersection of production ML systems, developer platforms, model behavior, and large-scale infrastructure. About the Role We’re looking for a backend engineer who can quickly understand OpenAI’s models, products, and systems, then adapt first-party deployments for other cloud platforms. You’ll build backend services, APIs, SDK integrations, authentication flows, and cloud service infrastructure that let developers use OpenAI capabilities in the cloud environments where they already build. This role involves working across teams, sometimes embedded with partner product groups, to ship products quickly and across multiple platforms at the same time. It’s a strong fit for engineers who have built developer tools, especially AI-powered tools, communicate clearly across technical boundaries, and can shape architectures that support different deployment models; experience building cloud services is a strong plus. In this role, you will: Build backend and infrastructure systems that extend OpenAI’s API platform into cloud-native environments, like AWS. Design and ship cloud-contained products that allow customers to use OpenAI capabilities while keeping workloads and data within cloud environments. Help stand up cloud-hosted Codex experiences powered by the OpenAI Responses API. Build the infrastructure and runtime abstractions
About the Team The Cloud Agents team builds product infrastructure for long-running agents in the cloud: orchestration, sandboxing and isolation, secure environment connectivity, secrets and identity, observability, reliability, and cost controls. These agents securely connect to diverse developer and customer environments and use tools to accomplish goals. We partner closely with product, research, and infrastructure teams to turn agentic capabilities into dependable platforms for OpenAI products and developers building on OpenAI. About the Role We are looking for an experienced software engineer to help build and scale our cloud agent platform. You will design and operate systems for orchestrating agents at scale. You will work closely with product engineers on ChatGPT, API, and Codex to define the right abstractions and enable them to ship products quickly. Strong backend or infrastructure experience is important; experience with Python, Rust, distributed systems, cloud infrastructure, or product platforms is especially helpful. In this role, you will: Design and scale the orchestration, sandboxing and storage systems that run agentic workloads for Codex, ChatGPT, and the OpenAI API. Partner with product engineers to build a platform that enables them to ship quickly and turn feedback into robust abstractions. Improve reliability, security, performance, and cost efficiency for long-running agents. Deploy services that can operate across different environments and clouds. Your background might look something like: 9+ years of professional engineering experience, excluding internships, in relevant roles at technology and product-driven companies. Experience leading large-scale backend, platform, or infrastructure projects from ambiguous problem statements to production systems. Proficiency in one or more backend languages such as Python, Go, Rust, TypeScript, or similar, and the ability to move across service, platform, and product boundaries. Strong understanding
About the Team The Personalization-Memory team, within OpenAI's broader Personal AGI organization, is focused on developing agents that can learn from prior interactions in order to become more helpful and efficient over time. We build general-purpose memory and personalization capabilities that transfer across ChatGPT and other agentic products, and we collaborate with applied engineering on the product surfaces that allow users to interact with memory. About the Role As a Research Engineer / Research Scientist on the Personalization-Memory team, you will research and develop improvements to memory usage and personalization in OpenAI's frontier models. Our team works on reinforcement learning, dataset creation, evaluations, and other post-training methods. We partner closely with research and product teams across the company to realize the vision of a truly personalized ChatGPT. We're looking for individuals who have a background in frontier model post-training, are able to iterate quickly, and who are passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda for improving memory use and personalization in frontier models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. Collaborate closely with the research and product teams to influence the shape of technical solutions in the product. You might thrive in this role if you: Are passionate about personalization and building personalized assistants. Have experience working with user signals and human data to turn feedback into reliable signals for training and evaluation. Have a deep understanding of frontier model post-training and machine learning applications. Value principled approaches and research craftsmanship. Are comfortable diving into a lar
About the Team The Codex Web Layer team provides the web-based systems and user experiences for Codex across the entire stack, from the Electron-like application framework that powers the application, to the user-facing in-app browser. About the Role In this role, you will be responsible for designing and implementing infrastructure and features end-to-end for the Codex desktop client application. You will help define what it means to be a hybrid agentic/interactive web browser. The team embodies “full stack” development from the lowest-level OS integration to the highest-level interaction design. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role you will: Partner closely with product and design to conceive, design, and build features for Codex web browsing features on macOS and Windows. This role will focus mostly on the backend C++ layer and Chromium, but many features cross the full stack including some TypeScript. Partner with the wider Codex team to deliver a high-performance, stable, and secure application platform for client development. This includes API design and implementation (mostly in C++) and the infrastructure that supports deploying it (in Python, TypeScript, and agentic skills). Work with a small, experienced team of engineers on this critical and rapidly growing product. You might thrive in this role if you: Have significant experience building technically complex features end-to-end. Are a strong C++ developer, especially with experience in browser environments like Chromium and Electron. Since this role is more backend focused, general knowledge of web development and TypeScript is helpful but not required. Thrive in a fast-paced, ambiguous environment. Communicate clearly and concisely across many different roles in the organization. Are self-directed, identifying important work and executing it end-to-end. About OpenAI OpenAI is an AI
Other cities to consider
More places hiring for this role
Get new agentic risk analyst jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime