About the Team OpenAI’s Application Engineering team builds the internal products and platforms that help OpenAI operate securely and at scale. We engineer, own, and evolve OpenAI’s core productivity ecosystem, creating secure applications, integrations, automation, and reusable tooling where off-the-shelf software is not enough. Our work spans employee-facing experiences and the services, APIs, control planes, and governance that make them reliable, permission-aware, and scalable. We also act as a customer zero for OpenAI’s technology, building the enterprise foundations that let employees and agents safely access the context, tools, and actions they need. We partner closely with IT, Security, product teams, and platform providers to turn company-wide problems into durable systems, learn from real internal workflows, and help shape the products we deploy. We create paved paths that let teams move quickly without compromising security or operational quality. About the Role As a Staff Software Engineer on Agent Productivity, you will shape the foundation that enables teams to build agents with secure access to the context and capabilities they need. Slack will be the first and deepest implementation surface—and where you spend most of your time—owning its application architecture, integrations, APIs, governance, and administration automation while building patterns that extend to internal systems, identity platforms, and other enterprise applications. This is a hands-on engineering role with broad organizational impact as agents support more employee workflows. You will define platform architecture, build reusable foundations, and establish secure patterns for identity, permissions, connectivity, and operations that make agents easier to develop, deploy, and manage. In this role, you will: Own the technical strategy and architecture that enable teams to build, connect, and deploy agents quickly and safely, using Slack as the primary implementation surface. Design and
Jobs in United States
Team Lead in San Francisco
1,331 active opportunities · Updated October 2026
Showing
15 jobs
Explore current team lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Recursive Self-Improvement (RSI) team works across research, engineering, product, and infrastructure to build AI systems that accelerate and ultimately conduct high-quality research at OpenAI. We work to automate real research workflows and improve research productivity by building systems and feedback loops, designing evaluations, and training models to develop missing capabilities. Our work spans the full lifecycle of model training, evaluation, and deployment to help researchers move faster and tackle increasingly ambitious problems. About the Role We’re hiring research scientists , research engineers , and AI systems engineers to work on automating research at OpenAI. This role is based in San Francisco, CA. In this role, you will: Design evaluations for research judgment, hypothesis generation and testing, and long-horizon experiment execution. Turn real research workflows and model failures into data and evaluation flywheels. Improve model research capabilities through agent harnesses, synthetic data, RL environments, and model training. Build and maintain safe, reliable integrations between our models and OpenAI’s research infrastructure. Develop research agents, experiment-orchestration systems, and sandboxed runtimes that support real research workflows. Create metrics and economic models to understand RSI’s current and future effects on research productivity, model capabilities, and the safety of internal deployments. This is a high-ownership role for researchers and engineers who thrive in ambiguity, move fluidly between research and implementation, and turn emerging opportunities into rigorous, reliable, scalable results. You might thrive in this role if you: Have research or engineering experience across LLM training, model evaluations, agent systems, synthetic data, research infrastructure, or large-scale distributed systems. Are a strong generalist who can move between open-ended research and practical implementation, turning ambig
About the Team The Strategic Initiatives & Operations team is in need of a Technical Program Manager (TPM) to streamline our processes, including full safety governance and integration of various safety research and mitigations into our ChatGPT, API, and any frontier models. This role is critical for driving safe deployment of our new models, synthesizing inputs from multiple stakeholders, ranging across research, product, engineering, legal and policy, and ensuring all the risks are effectively and properly monitored, mitigated or resolved. About the Role As a TPM, you will be responsible for critical tasks ranging from tracking safety research progress and risk tables to overseeing the quality of human data campaigns – acting as the connective tissue to enhance the deployment of OpenAI’s safety system. Additionally, you will create and execute a compute roadmap for your team to ensure that our top priorities are resourced while taking advantage of new opportunities to make key safety research discoveries. Your primary focus will be to ensure our models are qualified for safe deployment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Manage key risk areas and corresponding stakeholders. Keep track of a stack of existing and future mitigations for every major product and model deployment. Standardize the lifecycle of risk assessment, setting safety bars, consolidating inputs from multiple stakeholders across research, product, engineering, legal and policy, pre-launch safety reviews and post-launch followup. Manage pre-launch safety reviews. Share launch calendars and key safety practices and evaluations with our key parter (i.e. Microsoft). Develop comprehensive documentation for all the safety work, including metrics, evaluations, and progress tracking across multiple teams within OpenAI. Help with publishing and open sourcing safety
ABOUT THE TEAM Critical Harm Operations sits within User Safety & Risk Operations and builds enforcement systems for Frontier Risk and Material Harm that are accurate, fast, defensible, and built to scale. The Cyber vertical turns policy into reviewer standards, calibrated judgment, quality systems, escalation paths, and automation guardrails. ABOUT THE ROLE We are looking for a senior cybersecurity practitioner and operations strategist to raise the quality, scalability, and technical rigor of our Cyber Operations. You will combine hands-on cyber judgment with systems-level operating design: resolve the hardest dual-use questions, evolve SOPs, uplift reviewers and vendors, and build practical tools and automations. This is a senior IC role. Success is not primarily cases closed; it is durable improvement in the operating model and the reviewers who run it. IN THIS ROLE, YOU WILL: Drive the Cyber Operations operating model across domain priorities, SOPs, escalation paths, quality health, vendor capability, roadmap inputs and help inform trusted access strategies. Serve as the senior cyber expert for complex or high-risk decisions across ChatGPT, API, Codex, agents, and emerging product surfaces. Translate policy ambiguity, quality misses, appeals, and reviewer disagreement into clear decision rules, calibration examples, training, and tooling requirements. Build durable operating systems and quality loops: golden sets, holdouts, double-labeling, adjudication, error taxonomies, reviewer calibration, and automation evaluations. Raise FTE and BPO capability through onboarding, certification, coaching, recurring calibration, and vendor-performance partnership. Use quality, appeals, SLA, backlog, and disagreement signals to diagnose root causes and prioritize high-leverage fixes. Build hands-on solutions—SQL analyses, scripts, dashboards, LLM eval workflows, evidence enrichment, routing logic, and lightweight automations—that improve decision quality and reduce manua
About the Team API Agents builds the shared agent harness, tools, and infrastructure that turn OpenAI’s frontier models into systems that can reliably complete real work. We carry the capabilities behind Codex into a much broader set of products and workflows across software engineering, research, finance, healthcare, enterprise operations, and more. Our work spans search and connected context, computer use, memory, delegation and multi-agent coordination, and safe execution. Sitting at the intersection of Research, Codex, infrastructure, and applied product teams, we build reusable agent capabilities that compound across the ecosystem. About the Role We are looking for an experienced backend software engineer to build the core systems behind the next generation of agents. You will design reliable services and abstractions that help agents find the right context, use tools and computers, retain knowledge, coordinate over long-running workflows, and take action safely. The role combines deep backend and infrastructure work with strong product judgment, with opportunities to work across agent runtimes, orchestration, search, execution environments, identity and permissions, observability, and evaluations. This is software and systems engineering rather than model training: success comes from strong backend fundamentals, high agency, and the ability to turn fast-moving research capabilities into dependable production primitives. In this role, you will: Design, build, and operate the shared agent harness and backend infrastructure that power long-running, high-value workflows across OpenAI and third-party products. Build reusable capabilities across search and connected context, computer use, memory, tool execution, delegation, subagents, and multi-agent orchestration. Establish the foundations agents need to operate safely in production, including secure execution environments, identity and permissions, observability, evaluations, reliability, and cost and latency effi
About the Team Our economics team is continuously working to improve our understanding of an AI-driven economy. About the Role We are seeking a highly technical Economist to join the OpenAI Economic Research team studying the real-world economic impacts of AI. This role is designed for economists with up to 5 years of professional experience post-Ph.D. who are interested in using novel, large-scale datasets to study how AI is reshaping economic systems. We are looking for candidates with deep expertise in at least one core domain relevant to AI’s economic impact, and an interest in contributing to a broader research agenda spanning labor markets, firm behavior, market dynamics, and macroeconomic change. This is an individual contributor role where the candidate will organize and execute on their own data-oriented projects. You will work at the intersection of economic research, data science, and public policy to produce rigorous empirical work that informs decision-makers across the public, industry, and government. Research Areas of Interest We are particularly interested in candidates with demonstrated expertise in one or more of the following areas: Economic Measurement of AI Impact (e.g., adoption trajectories, labor market transitions, productivity growth, and forecasting/scenario modeling for AI-driven economic change) Macroeconomic Implications of AI (e.g., productivity, technology diffusion, economic growth) AI and the Labor Market (e.g., employment, wages, job search, task-level impacts, skill acquisition) Applicants are not expected to have experience across all domains. We aim to build a team with complementary strengths across these areas. In this role, you will: Design and execute empirical research using large-scale observational or experimental data. Apply causal inference and/or structural modeling techniques to study AI-driven economic change. Collaborate with cross-functional teams to translate research questions into testable frameworks and applic
About the Team We're building the foundation for a new kind of AI coworker: persistent agents that have their own environments, can meet people wherever they work, and continue making progress for as long as a task requires. Our goal is to help individuals, teams, and organizations delegate meaningful work to AI—not just ask questions or complete a single turn. Our work brings together product, agent, and infrastructure capabilities across OpenAI. Together, we are creating always-on virtual coworkers that can carry context across tasks, operate through the right tools and channels, and deliver reliable results in real workplace environments. About the Role We’re looking for a full stack product engineer to shape how people discover, direct, and collaborate with AI coworkers. You’ll own product experiences end to end, spending roughly equal time building intuitive frontend surfaces and the backend systems that make agentic workflows persistent, reliable, and useful. You’ll work at the intersection of product engineering, design, agent capabilities, and enterprise readiness—turning rapidly evolving model and platform capabilities into experiences customers can understand, trust, and use every day. In this role, you will: Design, build, and ship full stack product experiences that help people and teams delegate meaningful work to persistent AI coworkers. Create intuitive frontend interfaces for directing agents, reviewing their progress, understanding their actions, and collaborating on the work they produce. Build backend APIs, services, and data models that support persistent agent state, asynchronous execution, orchestration, and progress reporting. Develop workflows that help agents access relevant context, use tools, and collaborate with people across workplace surfaces and channels. Partner with product, design, research, and engineering teams to turn shared capabilities into cohesive products. Build trust into the product through clear user controls, understanda
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You’ll own the hands-on and automation work that brings WAN, fiber, carrier, and cloud-interconnect circuits into service. Partner with network engineers, fiber providers, cloud service providers, colocation teams, and data-center technicians to move each connection from ordered and patched to verified, stable, and ready for handoff. You’ll own Layer 1 troubleshooting and circuit bring-up while building workflows that translate reliable system or model output into precise, approved technician actions, capture field feedback, and drive each connection to a green-port handoff. The right person combines strong physical-networking judgment with practical automation skills: patch-panel and port mappings, optics and light levels, provider coordination, structured operational data, API or scripting workflows, and human-in-the-loop LLM tooling. Responsibilities Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence. Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports. Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching
About the Team API Multimodal builds the developer-facing products and infrastructure that bring OpenAI’s image, audio, and real-time model capabilities into the world. We are responsible for high-scale APIs for image generation, speech transcription, speech generation, and low-latency voice interactions. We partner closely with Research and Inference to bring frontier model capabilities to developers and use customer feedback to improve our models. About the Role As a software engineer on API Multimodal, you will build and operate the products and distributed systems behind OpenAI’s image, audio, and real-time APIs. You will work across model integration, API design, and production infrastructure to turn new research capabilities into reliable developer experiences. This hands-on role combines backend and systems depth with product judgment: you will own projects end to end, partner with Research, Inference, and Safety, and help make multimodal AI useful at scale. Model training experience is not required. In this role, you will: Design, build, and ship developer-facing APIs and backend services that serve frontier models. Architect low-latency streaming, request, session, and model integration systems that make complex multimodal interactions reliable and intuitive at scale. Work directly with Research to bring new model capabilities into production, shape the systems around them, and incorporate feedback from real-world developers and customers. Own the availability, latency, scalability, and cost efficiency of the services you build. Own projects from technical design and implementation through launch and ongoing iteration, while raising the team’s engineering standards. Your background might look something like: 7+ years of professional experience, excluding internships, in backend, infrastructure, platform, or product engineering roles. A track record of designing, building, and operating production backend services, developer-facing APIs, or distributed syste
About the Team At OpenAI, our User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We operate at the front line of real-world safety and risk management, translating user and operational signals into timely decisions, effective interventions, and improvements to our systems. This role sits on a team focused on building operational capacity for new, ambiguous, and fast-moving areas of work. The team defines what needs to be built, creates the operating model to support it, and works with partner teams to make the work scalable and durable over time. About the Role We are seeking a Device Safety & Risk Operations Specialist to build the safety operating model for a new category of consumer hardware. This is a senior individual-contributor role for someone who can turn emerging product risks and incomplete requirements into practical workflows, controls, launch plans, and durable systems. You will define how product-safety incidents, critical escalations, regulated cases, and privacy-sensitive issues should be identified, investigated, escalated, resolved, and learned from. You will also establish operational requirements for case management, data access, decision logging, quality assurance, monitoring, and cross-functional response. You will stand up priority workflows through launch and early operations, then help transition them into durable homes across USRO and partner teams. The right person combines deep operational judgment with strong technical and hardware product fluency. They can move from executive-level risk framing to detailed workflow design, tabletop exercises, launch readiness, frontline guidance, and post-launch improvement. Location / work model: San Francisco, CA; hybrid, 3 days/week in-office. Please note: This role may involve exposure to sensitive or concerning material. Strong discretion, judgment, and resilience are essential. In This Role, You Will:
About the Team The Legal team is building the next generation of AI-powered products and experiences for the legal industry. We are exploring how advanced AI systems can transform legal workflows, improve access to information, and enable legal professionals and organizations to work more effectively. As a founding member of the Legal engineering team, you will help define the technical foundation for this new product area from the earliest stages. You’ll operate at the intersection of AI, product, and real-world legal workflows—identifying opportunities, building prototypes, and turning emerging ideas into scalable products that can create meaningful impact. We operate with a startup-like mindset inside OpenAI: small teams, rapid iteration cycles, and a willingness to explore bold ideas, learn quickly, and adapt based on user feedback. Our goal is to build products that meaningfully improve how legal professionals work while leveraging OpenAI’s cutting-edge models and infrastructure. About the Role As a Founding Full-Stack Software Engineer on the Legal team, you will help imagine, build, and scale new AI-powered products for the legal industry. You’ll work across the stack to design intuitive user experiences, build robust backend systems, and create the foundations for products used by legal professionals and organizations around the world. You’ll have significant ownership from the earliest stages—working closely with product, design, research, and go-to-market partners to understand customer needs, shape product direction, and deliver high-impact solutions. This includes rapidly prototyping new concepts, building production-quality applications on top of OpenAI’s platforms, and developing new technical approaches when existing systems are not sufficient. We’re looking for engineers who thrive in ambiguity, have strong product instincts, and enjoy building from 0→1. You should be comfortable moving quickly, making thoughtful technical decisions, and taking owner
About the Team The Core Models team helps shape how OpenAI’s frontier models are built, measured, and launched. We work across Research, Engineering, Model Design, Data Science, and Product to turn advances in model capabilities into reliable, useful experiences for people. Our scope includes model planning and launches as well as building data flywheels, evaluations and measurement systems to ensure our models have strong capabilities and behavior. About the Role As a Product Manager for the Core Models team, you'll be at the forefront of defining and guiding the future of how our AI models work in real-world applications. You will connect user needs to model and systems decisions: how prompts are understood; how information is aggregated and made useful for training and evaluation data; and how capabilities move from research prototypes into the mainline model and launch stack. You will operate comfortably across research, infrastructure, and consumer product surfaces, creating clarity where ownership and technical boundaries are still emerging. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate user and product goals into clear model requirements, system architecture choices, and research priorities across query understanding, indexing, retrieval, ranking, tool boundaries, data, training, inference, and evaluation. Build closed learning loops that turn product usage, explicit feedback, and other user signals into datasets, evaluations, experiments, training priorities, and launch decisions. Define success across offline evaluations and online product metrics, balancing model quality, usefulness, latency, safety, reliability, and cost. Partner closely with post-training research, applied product engineering, Model Design, and Data Science to integrate capabilities into the mainline model stack. Create reusable platforms and operatin
About the Team The Core Models team shapes how our models interact with people. We view the model as the product itself, aiming for intuitive experiences that exceed user expectations and feel like magic. About the Role As a Model Designer, you’ll have an outsized impact on how our models interact and resonate with users. You’ll strike a delicate balance between maximizing the model’s capabilities, reading in between the lines in user queries to understand how best to help, and upholding user trust. We’re looking for people who are passionate about the intersection of design, technology, and user experience — and are up for the challenge of defining new human-AI interaction paradigms. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you’ll join a small team in evolving and expanding the model design function. You will: Collaborate closely with researchers to understand, predict, and design model behavior. Partner with product managers and designers across the company to ensure a cohesive voice and approach. Proactively identify ways to improve our models based on product sense, user feedback, quantitative insights, and the research roadmap. Come up with creative strategies for collecting high-quality data. Do whatever needs to be done to make our models better. You might thrive in this role if you: Possess exceptional taste, creativity, and writing skills, allowing you to craft responses that delight users. Don’t mind ambiguity — you’re happy to throw yourself into a new, unfamiliar environment, build relationships, define a problem, and make progress. Love experimentation, and are willing to test and reject new ideas when the results don’t pan out. Exhibit high levels of empathy and self-awareness required to serve everyone in the world. Enjoy tackling profound and often philosophical questions while always driving towards clarity. Demonstrate technic
About the Team This team builds and operates the systems that enable OpenAI researchers to run reliable, scalable, and efficient research workflows. The team sits close to research and works across infrastructure, systems, and automation to make sure researchers have the tools and environments they need to move quickly. The work spans software engineering, infrastructure, systems administration, cluster operations, and reliability engineering. As OpenAI’s infrastructure evolves from bespoke bare-metal systems toward more standard, scalable platforms, the team needs engineers who can understand how systems work end-to-end and build the right abstractions without reinventing the wheel. About the Role As a Software Engineer on this team, you will build and operate the infrastructure that supports frontier research and critical research-facing systems. You will work on systems that sit close to the metal, but the role is not limited to classic operations or sysadmin work. We are looking for someone who can reason about networking, bootstrapping, Kubernetes, scalability, automation, and reliability - while also writing software to make these systems better over time. This role is a strong fit for an independent, high-ownership engineer who enjoys reliability-heavy infrastructure work but still wants to build. You do not need to come in as a kernel expert or highly algorithmic optimization engineer, but you should be deeply curious about infrastructure, comfortable debugging complex systems, and excited to support researchers doing novel work. We expect you to: Build and operate reliable infrastructure for research workloads and research-facing services. Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services. Improve cluster bootstrapping, provisioning, automation, and deployment workflows. Debug issues across networking, compute, storage, orchestration, and service reliability layers.
About the Team The CoT Monitorability team at OpenAI studies whether and when the chain-of-thought of frontier reasoning models is monitorable enough to support scalable oversight. We study how to measure monitorability , which training mechanisms affect monitorability, and speculative methods to improve monitorability. While we mostly focus on CoT monitorability at the moment, we care more generally about any form of monitorability, auditing methods, and improving alignment. We were the first to show that chain-of-thought monitoring can be a practical additional safety mechanism, and today our monitoring systems are actively used on OpenAI’s largest RL training runs to detect misbehavior. The issues we surface are then used to help improve our reward functions, environments, etc (without directly training against a CoT monitor). Our work sits in Alignment and intersects with model training, alignment evaluations, monitoring, and frontier-risk research.We care most about monitorability where the stakes are high, and about preserving useful oversight signals as models become more capable. About the Role We’re looking for a researcher with strong empirical ML expertise and a deep interest in model behavior, alignment, or interpretability. Direct chain-of-thought interpretability experience is welcome but not required; strong candidates may come from broader interpretability, alignment, model training, or investigative model-behavior work. As a researcher on the Alignment team, you will design and run experiments that improve our understanding of model monitorability. You will investigate how training interventions across the model-development pipeline influence whether reasoning remains legible, build evaluations that make those questions measurable, and help translate findings into practical oversight and training recommendations. You may also help develop new monitoring models or methods and apply them to OpenAI’s largest training runs. This role is especially well
Other cities to consider
More places hiring for this role
Get new team lead jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime