Jobs in United States

Infrastructure Sourcing Operations Lead in San Francisco

522 active opportunities · Updated October 2026

Explore current infrastructure sourcing operations lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The ChatGPT organization at OpenAI supports our mission by bringing advanced AI capabilities to hundreds of millions of users worldwide. The Image Generation team is responsible for one of the fastest-growing experiences in ChatGPT, enabling users to create, edit, and transform images through natural language. Recent breakthroughs in multimodal AI have dramatically improved image quality, instruction following, editing precision, consistency, and text rendering. We're building the systems and experiences that turn these research advances into products used daily by creators, professionals, businesses, and consumers around the world. Our team sits at the intersection of research, product, design, and infrastructure. We work closely with model researchers, mobile engineers, frontend engineers, and platform teams to build intuitive experiences and scalable systems that power image generation at global scale. Whether users are creating marketing assets, visualizing ideas, editing photos, designing products, or simply exploring their creativity, our goal is to make visual creation feel as natural as having a conversation. About the Role We are looking for an experienced Full Stack Engineer to join the Image Generation team and help shape the future of AI-powered visual creation. In this role, you'll own features end-to-end across both frontend and backend systems, building the experiences that enable users to generate, edit, organize, and interact with images inside ChatGPT. You'll work across the entire stack—from highly interactive user interfaces and real-time workflows to backend services, APIs, orchestration systems, and data infrastructure. This role is ideal for engineers who enjoy moving fluidly between product development and systems engineering, collaborating closely with design, product, and research teams to rapidly bring new AI capabilities to users. You'll help define entirely new interaction paradigms as multimodal AI continues to evolve. In

TypeScriptReactAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Finance Platform & Technology team at OpenAI builds and scales the future-proof systems and data architecture that power our core financial operations. We enable business agility, compliance, and operational excellence across quote-to-cash, procure-to-pay, inventory, and asset management for both B2B and B2C. Our focus is on modernizing workflows through strategic integrations, scalable automation, and seamless data flows empowering smarter decisions, reliable reporting, and sustainable growth as OpenAI evolves About the Role We are looking for a Supply Chain Transformation Architect to redesign and modernize our end-to-end supply chain operations supporting robotics, consumer hardware, and data center infrastructure. This role sits at the intersection of process, systems, and data. Your primary focus will be transforming supply chain processes across planning, procurement, manufacturing, logistics, and fulfillment—then enabling those processes with the right systems architecture, data foundation, and AI-driven automation. You will help move the organization from manual, reactive operations to intelligent, data-driven supply chain execution. In this role you will: Lead End-to-End Supply Chain Transformation Evaluate and redesign core supply chain processes across demand planning, supply planning, procurement, manufacturing operations, logistics, and fulfillment. Identify operational bottlenecks, fragmented workflows, and manual processes that limit scalability. Build standardized process frameworks and operating models that support rapid scaling of hardware programs. Drive Operational Excellence Implement structured supply chain practices such as: S&OP / Integrated Business Planning Supply risk management Inventory optimization Supplier collaboration frameworks Logistics visibility and execution models Establish operational KPIs and governance to improve predictability, responsiveness, and resilience. Architect the Digital Supply Chain Tra

ReactAWSGitRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Applied Foundations team at OpenAI is dedicated to ensuring that our cutting-edge technology is not only revolutionary but also secure from a myriad of adversarial threats. We strive to maintain the integrity of our platforms as they scale. The Applied Foundations team is at the front lines of defending against financial abuse, scaled attacks, and other forms of misuse that could undermine the user experience or harm our operational stability. Integrity Foundations provides the core building blocks and infrastructure for this work. About the Role At OpenAI, our mission is to advance AI in a way that is safe, reliable, and aligned with broad societal values. The applied foundations role is crucial for maintaining the trustworthiness of our platforms. You will be pivotal in developing robust defenses against a spectrum of adversarial behaviors that threaten our ecosystem. In this role, you'll work with our entire engineering team to design and implement systems that detect and prevent abuse, promote user safety, and reduce risk across our platform. You'll be at the forefront of our efforts to ensure that the immense potential of AI is harnessed in a responsible and sustainable manner. In this role, you will: Develop and enhance systems to detect and prevent various forms of abuse including financial fraud, botting, and scripting. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. Assist with response to active incidents on the platform and build new tooling and infrastructure that address the fundamental problems. You might thrive in this role if you: Have at least 3 years of professional software engineering experience. Have experience setting up and maintaining production backend services and data pipelines. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Are self-directed

PythonAWSAzureKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role We are seeking a Software Engineer, Security Observability to join our Security team. In this role, you will be responsible for building secure, scalable systems that enhance our security observability infrastructure. Leveraging your strong engineering skills, you will collaborate with cross-functional teams to develop, deploy, and maintain robust software solutions that support our security and detection capabilities. This role is open to remote employees, or relocation assistance is available to one of our OpenAI offices in San Francisco, Seattle, or New York City. Due to requirements associated with work this role may support, applicants for this position must be U.S. citizens. In this role, you will: Design and develop scalable software systems that facilitate security observability across our infrastructure. Build and maintain data pipelines that centralize and store security-relevant data from diverse sources. Proactively improve the resilience and reliability of data systems to ensure high platform availability Collaborate closely with Detection & Response (D&R) and other security teams to reduce the company’s security risk. Contribute to data engineering in support of forensic investigations and compliance efforts. You might thrive in this role if you have: Strong software engineering experience, with proficiency in programming languages such as Python, Golang, or similar. A background in infrastructure as code, with exp

PythonAWSAzureRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Codex Web Layer team provides the web-based systems and user experiences for Codex across the entire stack, from the Electron-like application framework that powers the application, to the user-facing in-app browser. About the Role In this role, you will be responsible for designing and implementing infrastructure and features end-to-end for the Codex desktop client application. You will help define what it means to be a hybrid agentic/interactive web browser. The team embodies “full stack” development from the lowest-level OS integration to the highest-level interaction design. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role you will: Partner closely with product and design to conceive, design, and build features for Codex web browsing features on macOS and Windows. This role will focus mostly on the backend C++ layer and Chromium, but many features cross the full stack including some TypeScript. Partner with the wider Codex team to deliver a high-performance, stable, and secure application platform for client development. This includes API design and implementation (mostly in C++) and the infrastructure that supports deploying it (in Python, TypeScript, and agentic skills). Work with a small, experienced team of engineers on this critical and rapidly growing product. You might thrive in this role if you: Have significant experience building technically complex features end-to-end. Are a strong C++ developer, especially with experience in browser environments like Chromium and Electron. Since this role is more backend focused, general knowledge of web development and TypeScript is helpful but not required. Thrive in a fast-paced, ambiguous environment. Communicate clearly and concisely across many different roles in the organization. Are self-directed, identifying important work and executing it end-to-end. About OpenAI OpenAI is an AI

TypeScriptPythonAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man

PythonSQLAWSGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Applications Engineering organization builds and operates the products (such as ChatGPT & Codex) that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role We’re hiring Backend Software Engineers to design and implement safe services and infrastructure that power our core products. What You’ll Do Architect, build, and improve scalable backend systems and APIs. Drive performance, reliability, and safety across distributed services. Implement data storage, retrieval, compute, and integration solutions. Participate in long-term architectural planning and technical design reviews. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. You Might Thrive Here If You: Have strong experience with distributed systems, APIs, and backend languages (e.g., Go, Python, Rust, C++). Have experience setting up and maintaining production backend services and data pipelines. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Enjoy building resilient services that handle large scale and complexity. Are self-directed and enjoy figuring out the best way to solve a particular problem Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i

PythonReactNode.jsAngular
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,

AWSKubernetesCI/CDLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Applications Engineering organization builds and operates the products (such as ChatGPT & Codex) that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role We’re hiring Full-Stack Software Engineers to design and implement safe services, systems and infrastructure that power our core products. In this role, you will: Architect, build, and improve frontend and backend systems. Work across the full stack to build products and systems from initial exploration through launch readiness. Participate in long-term architectural planning and technical design reviews. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. You might thrive in this role if you: Have built and shipped full-stack apps or systems end-to-end — in fast-moving, startup-like environments. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Enjoy building resilient products and services that handle large scale and complexity. Are self-directed and enjoy figuring out the best way to solve a particular problem Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool t

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

Overview: The Data Acquisition team within the Foundations organization at OpenAI is responsible for all aspects of data collection to support our model training operations. Our team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams. We are looking for a skilled Full-Stack Engineer to join our Data Acquisition team to build and optimize the interfaces and tools that power our data infrastructure. Responsibilities: Develop and maintain full-stack applications that support data acquisition, including internal tools and dashboards. Collaborate closely with cross-functional teams, including Data Processing, Architecture, and Scaling, to ensure seamless data ingestion and workflow management. Design and implement APIs to facilitate data interactions between internal services and external data sources. Enhance user experience by developing intuitive web-based interfaces for managing and monitoring data pipelines. Optimize backend services for performance, scalability, and security in a distributed computing environment. Work with legal and compliance teams to ensure our data acquisition processes adhere to privacy regulations and best practices. Deploy and maintain infrastructure using Kubernetes and Infrastructure-as-Code (IaC) methodologies. Analyze system performance, conduct experiments, and improve data workflows to maximize efficiency. Qualifications: BS/MS/PhD in Computer Science or a related field. 4+ years of industry experience in full-stack development. Proficiency in frontend frameworks (React, Vue, or similar) and backend technologies such as Python, Node.js, or Go. Strong expertise in RESTful APIs, GraphQL, and database design (SQL and NoSQL). Experience building data-intensive applications that handle large-scale datasets. Familiarity with cloud platforms (AWS, GCP, or Azure) and container orchestration (Kubernetes, Docker). Prior experience with web crawling and large-scale data processing is a

PythonReactNode.jsVue
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p

PythonAWSCI/CDRest
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%

$295K – $380K/yr

Quick readStrong listing-quality and freshness signals

About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Senior Software Engineer, ML Systems & Training Infrastructure, you will be a deeply hands-on engineering force multiplier for the robotics team. You will help keep the training framework and surrounding infrastructure healthy, review and improve code quickly, debug failures across ML systems and infrastructure, and unblock researchers and engineers when the path from idea to working training job gets rough. We’re looking for people who love writing, reading, reviewing, and fixing code; who can get productive quickly in unfamiliar systems; and who bring strong practical judgment without a lot of ego or process overhead. This role will be based in San Francisco, CA and be expected in office 5 days per week and offer relocation assistance to new employees. In this role, you will: Review, improve, and clean up code across training frameworks and adjacent infrastructure. Identify risky or low-quality changes before they land, and raise the code quality bar without slowing the team down. Debug issues across ML training systems, GPUs, clusters, networking, and related infrastructure. Help researchers and engineers unblock broken training jobs, flaky workflows, and brittle internal tooling. Improve the reliability, maintainability, and usability of the robotics team’s training framework. Move quickly on practical engineering problems that directly affect team velocity. You might thrive in this role if you: Have strong software engineering fundamentals and excellent code review judgment. Have experience with ML systems, training fr

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system

PythonSQLAWSLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Employee Technology & Experience (ETX) team is responsible for delivering a world-class internal technology experience that enables employees to do their best work. We support and operate the employee-facing systems that keep the company moving quickly and efficiently. ETX spans support, logistics, AV, identity, endpoints, SaaS administration, automation, enterprise tooling, and internal infrastructure operations. We partner closely with Security, Engineering, Workplace, Finance, People, and other teams to keep OpenAI’s internal technology reliable, scalable, and moving at the pace of the company. About the Role We are hiring a Program Manager to help scale how IT operates across OpenAI. This role will lead complex cross-functional programs that improve operational maturity, streamline how teams work together, and turn high-impact initiatives into durable operational capabilities. You will work across IT, Security, Engineering, Workplace, and other functions to drive alignment, remove friction, and help build the operational foundation needed to support OpenAI’s rapid growth. You’ll be responsible for: Lead cross-functional operational programs that improve scalability, consistency, and operational maturity. Drive operational excellence initiatives across IT Support, employee lifecycle operations, meeting room and calendaring services, onsite support, vending, and research support environments. Build operating models, readiness plans, escalation paths, governance cadences, and success metrics for complex operational programs. Partner with technical teams to ensure new deployments, infrastructure investments, and internal platforms are operationally ready and sustainably supported at scale. Drive high-priority operational programs supporting company growth, including infrastructure expansion, operational integrations, and other emerging initiatives. Improve operational visibility, stakeholder alignment, and coordination across long-running cros

AWSRestAIGo
🔔

Get new infrastructure sourcing operations lead jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime