Jobs in United States

Deployment Lead in San Francisco

313 active opportunities · Updated October 2026

Explore current deployment lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Recursive Self-Improvement (RSI) team works across research, engineering, product, and infrastructure to build AI systems that accelerate and ultimately conduct high-quality research at OpenAI. We work to automate real research workflows and improve research productivity by building systems and feedback loops, designing evaluations, and training models to develop missing capabilities. Our work spans the full lifecycle of model training, evaluation, and deployment to help researchers move faster and tackle increasingly ambitious problems. About the Role We’re hiring research scientists , research engineers , and AI systems engineers to work on automating research at OpenAI. This role is based in San Francisco, CA. In this role, you will: Design evaluations for research judgment, hypothesis generation and testing, and long-horizon experiment execution. Turn real research workflows and model failures into data and evaluation flywheels. Improve model research capabilities through agent harnesses, synthetic data, RL environments, and model training. Build and maintain safe, reliable integrations between our models and OpenAI’s research infrastructure. Develop research agents, experiment-orchestration systems, and sandboxed runtimes that support real research workflows. Create metrics and economic models to understand RSI’s current and future effects on research productivity, model capabilities, and the safety of internal deployments. This is a high-ownership role for researchers and engineers who thrive in ambiguity, move fluidly between research and implementation, and turn emerging opportunities into rigorous, reliable, scalable results. You might thrive in this role if you: Have research or engineering experience across LLM training, model evaluations, agent systems, synthetic data, research infrastructure, or large-scale distributed systems. Are a strong generalist who can move between open-ended research and practical implementation, turning ambig

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You’ll own the hands-on and automation work that brings WAN, fiber, carrier, and cloud-interconnect circuits into service. Partner with network engineers, fiber providers, cloud service providers, colocation teams, and data-center technicians to move each connection from ordered and patched to verified, stable, and ready for handoff. You’ll own Layer 1 troubleshooting and circuit bring-up while building workflows that translate reliable system or model output into precise, approved technician actions, capture field feedback, and drive each connection to a green-port handoff. The right person combines strong physical-networking judgment with practical automation skills: patch-panel and port mappings, optics and light levels, provider coordination, structured operational data, API or scripting workflows, and human-in-the-loop LLM tooling. Responsibilities Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence. Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports. Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching

AWSAzureRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Our Cyber team builds AI systems and products that help trusted defenders understand and respond to cyber threats while improving the safety and reliability of frontier models in security-sensitive settings. The team works across product engineering, model training, evaluations, safeguards, and deployment to make advanced cyber capabilities useful to defenders and responsibly managed. We collaborate closely with Safety/Preparedness, Research, Security, Legal, Communications, GTM, and external partners across OpenAI’s broader cyber work. About the Role We’re looking for research and software engineers to join Codex Cyber. You’ll help define and ship security products, work with trusted defenders and customers, shape model training and access patterns, and build research and evaluation systems for assessing cyber capabilities, validating safeguards, and improving training data. This role is hands-on and cross-functional, connecting product launches, model development, safety work, and real-world security use cases. In this role, you will: Help define and execute the technical roadmap for Codex Cyber’s security products, including evaluations, safeguards, trusted-defender workflows, and deployment decisions. Work with trusted defenders, customers, and partner teams to understand cyber use cases, evaluate risk, and turn feedback into product and research priorities. Shape cyber-specific model training and access patterns, including data, evaluations, validation, and deployment criteria. Build and validate systems for measuring cyber capabilities, monitoring misuse risk, and proving safeguards work in practice. Collaborate with Safety/Preparedness, Research, Security, Legal, Communications, Go-to-Market, and external partners on company-wide cyber priorities. Translate frontier cyber research into launch-ready tools, operational playbooks, and durable infrastructure for Codex and security products. You might thrive in this role if you: Enjoy 0 -> 1 envi

JavaScriptTypeScriptPythonJava
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Codex is OpenAI's software engineering agent. Codex Security extends that work into one of the most important product areas in AI: helping organizations find, validate, prioritize, and fix real vulnerabilities in the software they build and depend on. The Codex Cyber team is building the product and platform foundations for AI-native application security. This includes Codex Security product experiences, cloud-based security analysis, platform controls across Codex, customer deployment and support tooling, and infrastructure that helps security researchers and cyber models improve over time. The team is early, small, and growing quickly, with a mandate to move fast and hire exceptional builders. About the Role We are looking for software engineers first: strong full-stack or product-minded generalists who can own ambiguous product and platform problems end to end. Security experience is helpful, and security curiosity is important, but this is not a role for security specialists who only occasionally write code. The right person is an excellent builder who is excited to work in security and can turn complex research, product, and customer needs into reliable systems. You will work across user-facing product surfaces, developer workflows, backend services, security analysis pipelines, cloud infrastructure, and internal tooling. You may build features that make Codex Security more useful for application security teams, systems that scale cloud-based security analysis, platform controls that make agentic coding safer, or infrastructure that helps security researchers and models become more effective. You will collaborate closely with engineering, product, security research, infrastructure, and customer-facing partners as Codex Cyber becomes a major product and platform investment for OpenAI. In this role, you will: Build end-to-end product features for Codex Security, from developer-facing interfaces to APIs, backend services, and workflow tooling. Own a

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Monetization team is a cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products, including next-generation ads experiences that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to build measurement systems that connect ad interactions to meaningful advertiser outcomes while protecting user privacy. In this foundational role, you’ll design infrastructure for conversion signals, attribution, reporting, and feedback loops across OpenAI’s ads products. This role is ideal for engineers who have built large-scale ads measurement, data, experimentation, marketplace, or distributed systems and want to apply that experience in a highly ambiguous 0→1 environment. You’ll work across event collection and normalization, deduplication and matching, attribution and modeled measurement, privacy-safe aggregation, reporting, and high-quality labels for ads optimization. We are hiring engineers who can independently own complex systems, make sound technical tradeoffs, and help define what should be built. You’ll work closely with Ads Delivery, Ads ML, Product, Research, Privacy, Data Sc

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

PythonAWSAzureGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside our cloud partners, infrastructure providers, and internal engineering teams, we operate hyperscale AI campuses that support the training and deployment of frontier AI models. The Site Operations team serves as OpenAI's on-site operational presence, helping ensure campuses operate safely, efficiently, and in alignment with Industrial Compute standards. We work closely with Hardware Operations, Infrastructure Delivery, Network Operations, Security, Facilities, Construction, and our infrastructure partners to support day-to-day site execution and maintain operational readiness. As Industrial Compute continues to expand globally, Site Operations plays a critical role in ensuring each campus is prepared to support reliable AI infrastructure at scale. About the Role We are seeking a Site Operations Technician to support the daily operation of Industrial Compute campuses. This role acts as OpenAI's on-site technical representative, helping coordinate activities across hardware operations, facilities, construction, logistics, security, and external service providers. You will perform routine site inspections, support asset tracking, coordinate vendor activities, assist with operational readiness, document site conditions, and help ensure infrastructure issues are identified and resolved quickly. The ideal candidate enjoys working in highly technical environments, is detail-oriented, and thrives in fast-paced operational settings where no two days are the same. Key Responsibilities Perform routine walkthroughs of Industrial Compute facilities to verify operational readiness and identify potential issues. Monitor site conditions and report abnormalities involving hardware spaces, network rooms, utilities, logistics areas, and common infrastructure. Support coordination of vendors, contractors, and partner organizations performing work o

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Applied AI team safely brings OpenAI's technology to the world. We released ChatGPT, Plugins, DALL·E, and the APIs for GPT-4, GPT-3, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate horizon. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. We serve end-users directly through ChatGPT, and serve developers through our APIs, which power product features that were never before possible. About the Role The Engineering Acceleration team designs, builds and maintains the foundational systems that engineers use to build ChatGPT and the API. This is a fast-growing team and you will get a chance to own and define the strategy, vision, and plan for how to increase developer productivity. In this role, you will: Drive the design, development, and implementation of tools, systems, and processes that accelerate engineering velocity, reduce manual effort, and increase the quality of output. Use our latest AI tools to re-think how we can be the most productive team in the industry. Work closely with various teams within OpenAI to understand their workflows, challenges, and needs, and ensure the tools and systems built by the Engineering Acceleration team address these requirements. Bring new features and research capabilities to the world by partnering with product engineers to lay the necessary technical foundations. Guide and advise product engineering teams on best practices for ensuring observable, scalable systems. Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years of experience in engineering, including 3+ years of experience in infrastructure building tooling for developers. Have experi

PythonAWSKubernetesRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team The Applied team safely brings OpenAI's technology to the world. We released ChatGPT; Plugins; DALL·E; and the APIs for GPT-5, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate horizon. Our customers build fast-growing businesses around our APIs, which power product features that were never before possible. ChatGPT is a prime example of what is currently possible. We simultaneously ensure that our powerful tools are used responsibly. Safe deployment is more important to us than unfettered growth. The Fraud Engineering team works within our Applied Engineering organization identifying and responding to fraudsters on our platform. We are looking for a software engineer with anti fraud & abuse experience to help architect and build our next-generation anti-fraud systems. About the role The Scaled Abuse team protects OpenAI’s products and customers by detecting, preventing, and responding to fraudulent and abusive behavior at scale. We build and operate the backend and data systems that power real-time detection, investigation workflows, and enforcement — balancing strong protections with a great user experience as the platform grows. Our work sits at the intersection of engineering and abuse expertise: we partner closely with Trust & Safety, Security, and Product to understand emerging attack patterns, translate messy signals into clear system behavior, and continuously harden our defenses. The problems are dynamic and ambiguous by default, so we value engineers who can quickly dive into an unfamiliar codebase, develop strong intuition about how it works end-to-end, and propose pragmatic improvements that make the entire stack more resilient. In this role, you will: Design and build systems for fraud detection and remediation while balancing fraud loss, cost of implementation, and customer experience Work closely with finance, security, product, research, and trust & safety ope

PythonAWSAzureKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Safety Systems team is responsible for various safety work to ensure our best models can be safely deployed to the real world to benefit the society and is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Model Safety Research team aims to fundamentally advance our capabilities for precisely implementing robust, safe behavior in AI models, and to leverage these advances to make OpenAI’s deployed models safe and beneficial. This requires a breadth of new ML research to address the growing set of safety challenges as AI becomes more powerful and used in more settings. Key focus areas include how to enforce nuanced safety policies without trading off helpfulness and capabilities, how to make the model robust to adversaries, how to address privacy and security risks, and how to make the model trustworthy in safety-critical domains. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. About the Role OpenAI is seeking a senior researcher with passion for AI safety and experience in safety research. Your role will set directions for research to enable and empower safe AGI and work on research projects to make our AI systems safer, more aligned and more robust to adversarial or malicious use cases. You will play a critical role in shaping how a safe AI system should look like in the future at OpenAI, making a significant impact on our mission to build and deploy safe AGI. In this role, you will: Conduct state-of-the-art research on AI safety topics such as RLHF, adversarial training, robustness, and more. Implement new methods in OpenAI’s core model training and launch safety improvements in OpenAI’s products. Set the research directions and strategies to make our AI systems safer, more aligned and more robust. Coordinate and collaborate with cross-functional team

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team With Codex we’re building an AI software engineer. One that you can pair with, delegate to, or even ask to take on future tasks proactively. Our team is a fast-moving group within OpenAI, bringing together research, engineering, design, and product. We iteratively build the Codex agent harness and product to get the most out of the model, and we iteratively train the model to be great at complex software engineering tasks. The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. We operate across research, engineering, product, and infrastructure; owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. Codex Enterprise builds the ecosystem, governance, and enterprise capabilities that help Codex spread across developers, teams, and organizations worldwide. The User Activation team owns the product experiences that help developers discover Codex, understand its capabilities, connect it to their workflows, and turn initial usage into sustained adoption across teams. About the Role As Codex adoption grows, our challenge is no longer just building powerful AI capabilities. It is helping developers and teams quickly understand how Codex fits into their work, connect it to the tools and codebases they already use, and unlock workflows that make Codex feel like a true teammate. This role will help build the full-stack product surfaces that drive activation and adoption across Codex Enterprise. You will work across onboarding, workspace setup, integrations, discovery, collaboration, usage insights, and ecosystem capabilities that help Codex spread naturally through organizations. You will partner closely with product, design, research, infrastructure, GTM, and customers to identify where users get stuck, where teams fail to adopt Codex, and what product experiences can turn curiosity int

TypeScriptPythonReactAWS
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role We’re building the observability product for OpenAI—from scalable infrastructure to a rich, AI-powered UI. Our systems ingest over petabytes of logs and billions of time series metrics across our fleet. We're now layering intelligence on top—think agents that summarize SEVs, auto-generate dashboards, or help engineers debug through notebook-like UIs. We’re hiring software engineers across the stack—infra, backend, and product. You’ll join a small, gritty team building both foundational infra and novel internal tools to make OpenAI's production systems reliable, performant, and observable. What You’ll Do Own core observability infrastructure, including distributed logging, time series, and trace storage Build AI-native tools that help engineers detect, understand, and resolve issues autonomously. Contribute to UI experiences like dashboards, notebooking, or interactive debugging Collaborate closely with engineers, researchers, user ops, and other teams across the company to build the next generation observability product You Might Be a Fit If You: Have operated large-scale distributed systems in production. ( especially logging systems or some other time series databases) Thrive in ambiguous environments and roll up your sleeves to solve unscoped problems. Have full-stack chops or product sensibilities—you're excited to build real tools people use. Have strong fundamentals in systems, networking, and cloud infra (Kubernetes, AWS, etc). Bonus : built or contributed to observability systems (e.g. Prometheus, OpenTelemetry, etc). Why This Team We’re b

AWSKubernetesRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We're seeking a Security Engineer to join our First-Party Hardware team. In this role, you will own the end-to-end security foundation for OpenAI's first-party AI hardware systems, working across hardware security, embedded security, system security, and practical deployment at data center scale. You will partner with silicon, hardware, firmware, infrastructure, manufacturing, operations, and security teams to define and deliver system-level device trust. This includes boot integrity, device identity, provisioning, attestation, management-plane security, storage encryption, debug controls, firmware update and recovery, RMA, and decommissioning. You will be accountable for turning threat models into requirements, requirements into implementation, and implementation into validation evidence that can support launch decisions. Location: San Francisco, CA (Hybrid: 3 days/week onsite) Relocation assistance available. In this role, you will: Own security requirements, threat models, validation strategy, and launch-readiness evidence for first-party hardware platforms from early design through production deployment. Design and review secure boot, measured boot, roots of trust, platform firmware resilience, firmware signing, recovery, and anti-rollback strategies across heterogeneous devices. Own device identity, provisioning, enrollment, attestation, certificate lifecycle, and key-management requirements across manufacturing and data center bring-up. Harden management

AWSRestAIC++
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

By applying to this role, you will be considered for Research Scientist roles across all teams at OpenAI. About the Role As a Research Scientist here, you will develop innovative machine learning techniques and advance the research agenda of the team you work on, while also collaborating with peers across the organization. We are looking for people who want to discover simple, generalizable ideas that work well even at large scale, and form part of a broader research vision that unifies the entire company. We expect you to: Have a track record of coming up with new ideas or improving upon existing ideas in machine learning, demonstrated by accomplishments such as first author publications or projects Possess the ability to own and pursue a research agenda, including choosing impactful research problems and autonomously carrying out long-running projects Be excited about OpenAI’s approach to research Nice to have: Interested in and thoughtful about the impacts of AI technology Past experience in creating high-performance implementations of deep learning algorithms About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement . Background checks for

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,

AWSKubernetesCI/CDLinux
🔔

Get new deployment lead jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime