Jobs in United States

Cloud Operations Lead in San Francisco

131 active opportunities · Updated October 2026

Explore current cloud operations lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads. As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI’s rapidly expanding compute environment. About the Role We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI’s third-party infrastructure partners, including major cloud service providers and strategic compute vendors. In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership. Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure. Key Responsibilities Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations. Develop operationa

AWSAzureGCPRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

AWSLinuxRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team At OpenAI, our User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We translate real-world user and operational signals into timely decisions, practical interventions, and improvements to our products and systems. This role will take on new, ambiguous, or underdeveloped operational risks and help mature them into scalable capabilities. We work across USRO and partner closely with Product, Engineering, Data Science, Product Policy, Legal, Safety, Support, and external vendors or partnership stakeholders. About the Role We are seeking a Senior Operations Analyst to take on complex, ambiguous safety and risk problems and turn them into practical operational solutions that can scale. This is a senior individual-contributor role for a versatile operator who is comfortable moving between queues, investigation, analysis, workflow design, hands-on execution, and cross-functional leadership. Depending on team needs, the role may focus on emerging-risk incubation, cloud deployment partnerships, or other new operational areas. You will be expected to move quickly, work hands-on, and create structure without waiting for perfect requirements or a large support team. The work starts with the problem, not a prescribed process. You may investigate unstructured user signals, stand up a lightweight workflow, build an AI-assisted tool, improve an existing operation, or help a new launch become operationally ready. The goal is to produce durable systems that other people can run, not simply complete a series of individual tasks. The portfolio will change with company priorities and may span established harm areas, emerging-risk incubation, cloud deployments and partnerships, device safety, or new product launches. Some hires may focus primarily on cloud deployment operations, including launch readiness, partner coordination, safety workflows, and operational monitoring. You will ty

SQLAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Order to Cash (OTC) team oversees the complete flow of commercial transactions from order intake and provisioning through billing, collections, and cash application — ensuring accuracy, compliance, and operational excellence in support of OpenAI’s mission to ensure artificial general intelligence benefits all of humanity. About the Role We are looking for a senior, hands-on operator to own Order Management and Billing execution across OpenAI’s cloud marketplace and partner ecosystem, including platforms such as AWS, GCP, Oracle Cloud, GovCloud, and future channels. This senior individual contributor role will translate partner requirements into scalable workflows and ensure launch readiness, accurate billing, and reliable daily execution. As a senior individual contributor within the Cloud Marketplaces team, you will own the end-to-end order-to-invoice lifecycle for your assigned portfolio. You will ensure that private offers, commercial terms, provisioning, pricing, usage, billing data, credits, settlements, and partner-specific reporting flow through our systems accurately, on time, and with audit-ready controls. You will implement and continuously improve the common cloud marketplace operating model, lead cross-functional execution for your assigned portfolio, and surface risks, requirements, and improvement opportunities. You will partner across Revenue Systems, Product, Engineering, GTM, Finance, Partner Operations, and external marketplace stakeholders. This role is critical to building the operational backbone for OpenAI’s expansion across cloud marketplaces and government-cloud channels. You will combine deep operational judgment with process and control execution, automation, clear communication, and hands-on problem solving to improve billing reliability, partner experience, customer outcomes, and financial integrity at scale. This role

AWSGCPRestAI
M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -67.9%

AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. About the Role We're seeking a Revenue Operations Manager with a strong track record, a builder's mindset, and a bias for action to join our in-person team in New York or SF. This is a high-impact, hands-on role. You'll own the entire revenue operations function, from top-of-funnel lead routing through deal close and commission administration. You'll work closely with our Head of Finance & People Ops and sales leadership to build the systems, dashboards, and processes that scale our go-to-market motion. What You'll Do: Own the lead routing process from inbound and partnering with marketing to ensure proper attribution Run effective territory management & strategy for Geo based decisioning Support & strategise every aspect of revenue operations in your territory Own the strategy for capacity forecasting, inputs, throughputs & outputs being the conduit back to finance in

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team We are a small and fast-moving partnerships team that shapes and executes OpenAI’s most important collaborations. Your mission is to build and grow partnerships across the cybersecurity ecosystem. Reporting to the Cybersecurity Partnerships Lead, you will own a portfolio of cybersecurity technology partners and help them validate, launch, and scale solutions powered by OpenAI models and platforms. About the Role You are an experienced partnerships operator who combines business development, partner management, technical curiosity, and strong execution. You can manage external relationships while driving detailed cross-functional work across product, engineering, technical success, sales, marketing, legal, security, and operations. You move with urgency, follow through consistently, and are comfortable managing a portfolio in a fast-changing market. In this role, you will: Source, close and manage a portfolio of cybersecurity technology partners. Identify high-value use cases across security operations, identity, cloud security, application security, threat intelligence, governance, risk, and compliance. Develop partner plans covering integration, launch, enablement, co-marketing, co-sell, and growth. Support partnership structuring and coordinate product, technical, commercial, legal, and security workstreams. Help partners move from concept and technical validation to production launch and scaled customer adoption. Run regular partner reviews, track commitments, resolve blockers, and identify expansion opportunities. Coordinate closely with product, engineering, technical success, sales, marketing, legal, security, and operations. Measure partner pipeline, launches, model adoption, consumption, customer outcomes, and partner health. You might thrive in this role if you have: 8+ years of experience in partnerships, business development, alliances, partner success, or ecosystem roles. Experience in cybersecurity, enterprise software, cloud platforms, o

AWSRestAIGo
B
📍 San Francisco, California, United States· Full-time
✓ High-confidence listing

$192K – $240K/yr

Quick readStrong listing-quality and freshness signals

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Security Operations Engineer at Brex, you will focus on preventing, detecting and responding to security threats across Brex's corporate and cloud environments. You will use existing systems and develop tools to improve our security capabilities. Our team is responsible for functions across corporate security, detection & response and infrastructure security domains; and we perform systems engineering and automation to support those functions. Security Operations is part of our wider Trust & IT o

PythonAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. In partnership with leading cloud providers, hardware manufacturers, utilities, construction partners, and internal engineering organizations, we are delivering hyperscale AI campuses that power the next generation of frontier AI models. Infrastructure Delivery Operations sits at the center of this effort. Our team develops the operating model that connects infrastructure strategy, supply planning, manufacturing operations, and delivery into a single, integrated system that enables OpenAI to deploy AI infrastructure predictably at scale. We partner across Hardware Engineering, Network Engineering, Capacity Delivery, Hardware Operations, Security, Finance, Strategic Sourcing, and external infrastructure partners to create a single, integrated view of program health. Through governance, operational analytics, executive reporting, and scalable operating mechanisms, we enable leaders to proactively manage risk, optimize capacity, and deliver infrastructure predictably at Industrial Compute speed. About the Role We are seeking a Technical Program Manager, Infrastructure Delivery Operations to drive integrated strategy and delivery across OpenAI's rapidly expanding AI infrastructure portfolio. This role sits at the intersection of infrastructure strategy, New Product Introduction (NPI), supply planning, manufacturing operations, and infrastructure delivery. You will lead highly cross-functional programs spanning engineering, supply planning, manufacturing, logistics, construction, commissioning, and operations, ensuring technical and operational dependencies remain synchronized from planning through production readiness. Beyond driving program execution, you will leverage operational insights to improve capacity planning, infrastructure strategy, and deployment readiness. You will also help operationalize new technologies and suppliers by partnering w

AWSRestAgileAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Order to Cash (OTC) team oversees the complete flow of commercial transactions from order intake and provisioning through billing, collections, and cash application — ensuring accuracy, compliance, and operational excellence in support of OpenAI’s mission to ensure artificial general intelligence benefits all of humanity. About the Role We are seeking a Senior Manager to build and scale Order to Cash across OpenAI’s API and ChatGPT businesses, cloud marketplaces, and strategic partnership channels. This role will own end-to-end Order Management and Billing operations across API and ChatGPT while leading OTC readiness and execution for AWS Marketplace, Google Cloud Marketplace, Oracle Cloud Marketplace, GovCloud, and future partner channels. The role will also own the end-to-end Order Management and Billing close, setting the close calendar, readiness standards, review and sign-off expectations, while leading the team responsible for execution. You will oversee the Order Management and Billing lifecycle across API, ChatGPT, marketplace, and partner transactions, spanning commercial readiness, order intake, provisioning, usage and transaction data, pricing validation, invoicing, credits, settlements, and product and partner reporting. Your work will ensure transactions are accurate, timely, complete, and supported by audit-ready controls. This is a leadership role that combines strategic ownership, cross-functional leadership, and hands-on operational execution. You will define the target operating model, lead first-of-kind launches, shape product and systems roadmaps, and oversee the resolution of complex contract modifications, non-standard pricing, usage disputes, reconciliation breaks, settlement variances, and customer- or partner-impacting escalations. You will collaborate with teams across Finance, GTM, Product, Engineering, Legal, Tax, Reven

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The User Operations team (Support) is central to ensuring that our customers' experience with our products is nothing short of exceptional. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. About the Role We’re seeking a Program Manager to lead support readiness and operational programs for OpenAI’s cloud and strategic partnerships. You’ll work closely with Engineering, Product, Support Delivery, and external partners to translate complex technical and business requirements into scalable customer support experiences. This role will help define how we support customers across partner ecosystems, from launch planning and issue-routing workflows to escalation management and ongoing operational improvements. In this role, you will: Lead support readiness for new and existing cloud and strategic partnerships, including launch planning, operational design, and ongoing program execution. Partner closely with Engineering, Product, go-to-market teams, and external partners to develop support models for technically complex products and integrations. Define partner-specific customer journeys, support workflows, escalation paths, ownership models, and cross-company handoffs. Identify operational and technical risks, align stakeholders on solutions, and drive improvements that strengthen the customer experience. Establish clear success metrics and operating rhythms to monitor partnership health, launch readiness, and support performance. Use AI and automation to improve partner-related support workflows and scale operations effectively. You might thrive in this role if you: Have 8+ years of experience

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Governance, Risk, and Compliance team helps ensure security and privacy are grounded in how our products and systems actually operate. Assurance Operations partners with Security, Engineering, Infrastructure, Product, Privacy, and Legal to make controls provable, risk decisions explicit, and audit readiness a result of well-designed systems. About the Role We are hiring a technical, product-minded GRC builder who can own consequential audits while improving the control and evidence systems behind them. You will build a reusable common control framework, use Codex to automate assurance work, validate changing system scope, and turn repeated audit friction into measurable improvements. We are looking for someone who questions inherited assumptions, solves novel problems creatively, works closely with engineers, and makes the next audit easier by improving the underlying system. You’ll be responsible for: Lead external, internal, customer, and certification audit work from scoping through evidence review, fieldwork, remediation, and closeout. Build a common control framework linking risk, control intent, implementation, owner, system, environment, evidence, and applicable frameworks. Validate actual scope and ownership instead of assuming last year's controls, product boundaries, or evidence remain accurate. Use Codex to build and test evidence checks, control mappings, request triage, owner workflows, monitoring, and remediation reporting. Partner with engineers on cloud architecture, identity, logging, data flows, software changes, vulnerabilities, and control effectiveness. Design maintainable, permission-aware tools that preserve source provenance, human review, and evidence integrity. Reduce repeated requests and operational burden for control owners through measurable workflow improvements. Define roadmaps, decision rights, milestones, success metrics, and clear cross-functional escalations. We’re looking for someone with: Direct ownership

SQLAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside leading cloud providers, engineering firms, construction partners, utilities, and equipment manufacturers, we are delivering hyperscale AI campuses that enable the next generation of frontier AI models. The Strategic Sourcing team develops and executes the commercial strategies that ensure our infrastructure programs have reliable access to the equipment, materials, and strategic partners needed to deliver at unprecedented scale. We partner closely with Infrastructure Delivery, Capacity Planning, Design Engineering, Hardware Operations, Finance, Legal, and our external suppliers to build a resilient global supply network capable of supporting Industrial Compute's long-term growth. As we continue expanding globally, strategic sourcing becomes a critical competitive advantage, ensuring our infrastructure programs remain cost-effective, resilient, and capable of executing against aggressive deployment timelines. About the Role We are seeking a Strategic Sourcing Manager, Data Center Infrastructure: Owner Furnished Equipment to lead sourcing strategy for the critical infrastructure systems that power Industrial Compute campuses. This role will develop commercial strategies, negotiate strategic supplier agreements, and manage relationships across engineering, construction, manufacturing, and infrastructure partners responsible for delivering mission-critical facilities. You will work closely with Infrastructure Delivery, Capacity Planning, Engineering, Finance, Construction, and external suppliers to ensure Industrial Compute has the capacity, supplier relationships, and commercial frameworks required to support rapid global expansion. The ideal candidate has experience sourcing major infrastructure systems for hyperscale data centers, mission-critical facilities, industrial construction, semiconductor manufacturing, energy infrastr

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside leading cloud providers, engineering firms, construction partners, utilities, and equipment manufacturers, we are delivering hyperscale AI campuses that enable the next generation of frontier AI models. The Strategic Sourcing team develops and executes the commercial strategies that ensure our infrastructure programs have reliable access to the equipment, materials, and strategic partners needed to deliver at unprecedented scale. We partner closely with Infrastructure Delivery, Capacity Planning, Design Engineering, Hardware Operations, Finance, Legal, and our external suppliers to build a resilient global supply network capable of supporting Industrial Compute's long-term growth. As we continue expanding globally, strategic sourcing becomes a critical competitive advantage, ensuring our infrastructure programs remain cost-effective, resilient, and capable of executing against aggressive deployment timelines. About the Role We are seeking a Strategic Sourcing Manager, Data Center Infrastructure to lead sourcing strategy for the critical infrastructure systems that power Industrial Compute campuses. This role will develop commercial strategies, negotiate strategic supplier agreements, and manage relationships across engineering, construction, manufacturing, and infrastructure partners responsible for delivering mission-critical facilities. You will work closely with Infrastructure Delivery, Capacity Planning, Engineering, Finance, Construction, and external suppliers to ensure Industrial Compute has the capacity, supplier relationships, and commercial frameworks required to support rapid global expansion. The ideal candidate has experience sourcing major infrastructure systems for hyperscale data centers, mission-critical facilities, industrial construction, semiconductor manufacturing, energy infrastructure, or similarly comple

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside leading cloud providers, engineering firms, construction partners, utilities, and equipment manufacturers, we are delivering hyperscale AI campuses that enable the next generation of frontier AI models. The Strategic Sourcing team develops and executes the commercial strategies that ensure our infrastructure programs have reliable access to the equipment, materials, and strategic partners needed to deliver at unprecedented scale. We partner closely with Infrastructure Delivery, Capacity Planning, Design Engineering, Hardware Operations, Finance, Legal, and our external suppliers to build a resilient global supply network capable of supporting Industrial Compute's long-term growth. As we continue expanding globally, strategic sourcing becomes a critical competitive advantage, ensuring our infrastructure programs remain cost-effective, resilient, and capable of executing against aggressive deployment timelines. About the Role We are seeking a Strategic Sourcing Manager, Data Center Infrastructure to lead sourcing strategy for the critical infrastructure systems that power Industrial Compute campuses. This role will develop commercial strategies, negotiate strategic supplier agreements, and manage relationships across engineering, construction, manufacturing, and infrastructure partners responsible for delivering mission-critical facilities. You will work closely with Infrastructure Delivery, Capacity Planning, Engineering, Finance, Construction, and external suppliers to ensure Industrial Compute has the capacity, supplier relationships, and commercial frameworks required to support rapid global expansion. The ideal candidate has experience sourcing major infrastructure systems for hyperscale data centers, mission-critical facilities, industrial construction, semiconductor manufacturing, energy infrastructure, or similarly comple

AWSRestAIGo
P
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. About the Team The Information Security organization at Postman operates across three pillars: Governance Risk & Compliance (GRC), Product Security, and Security Operations. We are a team of builders, not checkbox-checkers. We hold active SOC 2 Type II, ISO 27001, ISO 42001, and HIPAA compliance postures, and we are pursuing FedRAMP High and CMMC Level 2 authorization. Our security stack includes Wiz, SentinelOne, Okta, Jamf, and 1Password, and we operate across a multi-cloud environment. The Offensive Security team is the "red" pulse of this organization. We don't just find bugs — we simulate the adversary to ensure our defenses hold up under real-world pressure. We focus on continuous security validation, AI-augmented adversary emulation, and offensive AI security research at Postman's scale. The Opportunity We are looking for a Principal Offensive Security Engineer who is as much a strategist as they are a hacker. You will own the strategic direction of Postman's offensive security program — including building out a dedicated Offensive AI Security capability from the ground up — and operat

AWSKubernetesCI/CDGraphql
🔔

Get new cloud operations lead jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime