Jobs in United States

Cloud Operations Lead in San Francisco

131 active opportunities · Updated October 2026

Explore current cloud operations lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. Our Go-to-Market team helps organizations understand, adopt, and deploy OpenAI’s technology to solve meaningful business challenges and create lasting value. The Technology team works with leading software, internet, cloud, infrastructure, cybersecurity, semiconductor, and digital-native companies as they build new AI-powered products, transform internal operations, and rethink how they serve their customers. We partner with executives, product leaders, engineers, and go-to-market teams to help organizations integrate OpenAI’s capabilities into their products and businesses responsibly and at scale. The team collaborates closely with Solutions Engineering, Customer Success, Product, Research, Partnerships, Marketing, and Operations to turn customer priorities into successful, durable deployments. About the Role We are looking for an experienced Account Director, Tech to help build and grow OpenAI’s business across the technology industry. You will own relationships with a portfolio of strategic technology companies, helping executive, product, and technical leaders understand how OpenAI’s products can accelerate innovation, improve productivity, and create differentiated customer experiences. You will be responsible for developing account strategies, creating qualified pipeline, navigating complex enterprise sales cycles, and expanding adoption across products, teams, and use cases. This role requires a combination of enterprise sales leadership, technical fluency, commercial judgment, and the ability to operate credibly with both business and engineering stakeholders. You should be comfortable engaging with customers that have sophisticated technical environments, rapidly evolving AI strategies, and high expectations for product performance, security, reliability, and scale. Success in this role will be measured by revenue growth, depth of customer adoption,

AWSGitRestMachine Learning
S
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.4%
Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About this role At Sentry, Finance plays a critical role in helping the company understand where we're investing, what we're getting in return, and where we can operate more effectively. We are looking for an FP&A Manager who will help build Sentry's FP&A function while focusing on establishing best-in-class FinOps practices. You'll partner across the business on forecasting and financial analysis, with a significant focus on cloud infrastructure, AI costs, vendor spend, and other areas where better financial visibility can translate directly into better decisions. You will be expected to understand the economics behind our infrastructure, build the financial models and reporting needed to manage it, identify meaningful opportunities, and work with Engineering and other partners to turn those insights into action. You'll also operate as a generalist FP&A partner and play an important role in building the processes, models, and operating rhythms that Finance will use as Sentry grows. You will report to the Director of FP&A. In this role you will Own forecasting, reporting, and financial analysis for significant areas of Sentry's operations, with particular emphasis on cloud infrastructure, AI-related spend, software, vendors, and other major cost categories. Partner closely with Engineering and infrastructure leaders to understand the drivers of cloud and AI spend, translate technical consumption into financial forecasts, and identify opportunities to improve unit economics and gross margin. Develop reporting that makes cloud and infrastructure costs understandable and actionable, including trends, cost alloca

RestAIGoRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. Safely delivering increasingly capable AI systems requires scalable technical safeguards, clear ownership of emerging risks, rigorous deployment readiness, and close coordination across research, engineering, product, operations, legal, policy, and external partners. Our Technical Program Managers lead complex, high-stakes initiatives that turn safety commitments into deployed systems and measurable outcomes. We work across model development, infrastructure, product, and operational response to help ensure our technology is deployed responsibly and cannot be used to cause serious real-world harm. About the Role We’re seeking Technical Program Managers to drive complex product, platform, and safety initiatives across ChatGPT, API, enterprise, and related deployment environments. These roles operate at the intersection of technical strategy and execution: you will turn safety and product priorities into actionable plans, influence architectural and operational decisions, and deliver durable capabilities across model, infrastructure, application, and platform layers. Depending on the role, you may enable sensitive or high-impact model deployments, integrate safeguards into cloud and API platforms, prevent violent misuse and other serious harms, improve detection and enforcement systems, create platform solutions for safety or establish new programs as risks evolve. You will partner deeply with engineers, researchers, product managers, and operational teams while communicating technical tradeoffs and program decisions to senior leadership. You bring technical fluency, product judgment, and a strong execution record. You’re comfortable navigating ambiguity, advocating for users and developers, balancing safety with model usefulness, and leading cross-functional work with urgency, rigor, and empathy. Specific focus areas and scope will vary by opening and level. Thi

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team Security is at the foundation of OpenAI's mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect model weights, customer data, and critical systems across multiple cloud environments. The team partners across OpenAI, including Applied Engineering, Research, IT, Security, Infrastructure, and Engineering, to provide secure and scalable platforms for identity, access management, permissioning, orchestration, and safe AI research. About the Role We’re looking for an engineering leader to lead Identity Infrastructure Engineering, the team building the systems that govern and scale access across OpenAI’s research, engineering, and internal platforms. This role sits at the center of cloud infrastructure, identity, software engineering, and security-critical operations. You’ll lead engineers building control planes, policy systems, workload and agent authorization patterns, infrastructure-as-code, and operational foundations that help OpenAI move quickly while keeping access reliable, auditable, least-privileged, and safe under failure. The ideal candidate has led teams responsible for large-scale, mission-critical infrastructure. They can go deep into code and architecture when needed, while giving engineers and technical leads the clarity and ownership to do their best work. They set technical direction, grow strong teams, make durable architecture decisions, and turn ambiguous 0-to-1 problems into platforms OpenAI can trust and build on for years. In this role, you will: Build and lead a high-performing Identity Infrastructure team, going deep enough technically to set direction while empowering the team to own delivery. Define the strategy for identity platform as the policy plane for access across people, agents, workloads, services, clouds, and internal systems. Scale Acc

AWSGitRestAI
P
📍 San Francisco, CA, United States· Remote
✓ High-confidence listingCompany trend -84.3%
Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is seeking a Sr. Manager to lead our Capacity Engineering team. The team ensures that Pinterest’s cloud infrastructure has the capacity it needs while operating reliably, efficiently and with clear financial accountability. You’ll lead the full portfolio across forecasting and supply, capacity-management systems, compute and GPU efficiency, infrastructure data and governance and capacity operations. What you’ll do: Lead the Capacity Engineering team and establish its 12–18 month functional and technical strategy, roadmap and success measures tied to Infrastructure and company goals. Develop CPU and GPU forecasts and supply plans that account for workload demand, delivery constraints, cost and reliability requirements. Guide the design and delivery of capacity requests, reservations, entitlements, allocation policy and infra

KubernetesAIFinance
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -73.6%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking an experienced and proactive Security Engineer to help us build, maintain, and continuously improve the security posture of our rapidly growing ML infrastructure platform. As one of the first dedicated security hires at Baseten, you will work cross-functionally with engineering and operations teams to ensure we’re meeting the highest standards of confidentiality, integrity, and availability. You’ll have an opportunity to shape our security strategy and best practices from the ground up, influencing the way our platform handles sensitive data for both internal and external stakeholders. RESPONSIBILITIES Security architecture and design: Collaborate with engineering teams to design and implement secure systems and infrastructure, including cloud (AWS/GCP) environments and container orchestration platforms. Vulnerability management: Lead proactive vulnerability assessments, pen tests, and remediation efforts to ensure our products and infrastructure remain secure. Incident response: Develop and maintain incident response processes, including detection, analysis, containment, eradication, and post-incident reviews. Identity and access management (IAM): Oversee IAM strategies and tools to ensure the right people have the right level of access to our systems and data. Security compliance and audits: Work closely with operations to ensure compliance with relevant standards (e.g., SOC 2, ISO 27001) and

AWSGCPCI/CDMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, cloud services, mechanical engineering, electrical engineering, and product design to deliver reliable, production-ready devices at scale. Within Consumer Devices, Hardware Engineering eXperience, or HEX, is a new bootstrapped team building the environments, applications, compute, product-data systems, and workflows that let hardware engineers do their work without needing to troubleshoot the machinery underneath. HEX owns virtual engineering environments, HPC/GPU compute, storage, networking, licensing, MCAD/ECAD/CAE applications, PLM, product data, automation, validation, and support as one connected system. About the Role As a Staff PLM & Engineering Applications Engineer, you will be one of the first technical builders of HEX and the primary counterpart to the HEX lead. You will own the engineering-application and product-data side of the hardware engineering experience, with an initial focus on NX, Teamcenter, licensing, parts import, integrations, packaging, validation, and user workflows. This is not a traditional Teamcenter administration role and not a Corporate IT application-support role. You will take complex, fragile workflows and turn them into reliable engineering systems. This role is highly hands-on and systems-oriented. You will not inherit a mature environment and support queue. You will help build a fresh one, replacing manual setup guides, tribal knowledge, repeated support issues, and team handoffs with tested automation and reliable workflows. In This Role, You Will Own the technical architecture, deployment, configuration, integration, validation, and long-term operation of NX and Teamcenter. Build reliable workflows for parts import, product-data migration, metadata quality, BOMs, revisions, lifecycle states, and releas

PythonAWSRestAgile
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op

PythonAWSRestAI
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -73.6%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE This role owns Baseten's relationships and market intelligence across hyperscalers and strategic neoclouds, including NVIDIA cloud partners. This is a technical and commercial role in equal measure: you'll evaluate capacity from the GPU to the data center, negotiate cost and terms with suppliers, and stay close enough to the market to develop and defend a real point of view on where it's heading. Given current market conditions, Baseten needs a much stronger pulse on this part of the market so we can track pricing, stay close to the right relationships, and move fast the moment more capacity is needed. This is a senior, experienced hire who will also help pair with and develop 1-2 junior to mid-level teammates covering the same space. WHAT YOU'LL DO Build and maintain deep relationships across hyperscalers and strategic neoclouds (including NVIDIA cloud partners), working each organization from top to bottom rather than a single point of contact Maintain a consistent, "top of mind" presence with key accounts so Baseten is positioned to move quickly when capacity needs arise Evaluate capacity from the GPU to the data center — hardware generation, rack and node configuration, interconnect, power density, and cooling — so you know what a configuration will actually deliver, not just what the spec sheet claims Live in compute pricing daily: track rates by GPU generation, region, and contract term to keep Baseten inf

Machine LearningAIGoHR
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -73.6%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. ABOUT THE TEAM Supply is responsible for knowing everything happening in the compute market: who's building, who's buying, and on what terms. This role owns a specific and fast-moving slice of that map — emerging clouds and international markets — and owns the full relationship lifecycle in that space, from first outreach through to closed terms. RESPONSIBILITIES Build and maintain a real-time picture of the emerging cloud and international compute landscape — who's active, what they're building, and what terms are available Own the full partnership lifecycle in this space — from identifying and sourcing new providers, to negotiating terms, to ongoing relationship management Develop and manage relationships across a broad set of emerging and international providers, from account reps up through leadership Identify, structure, and help close opportunities where Baseten can move quickly to secure favorable capacity terms Define compelling value propositions tailored to different types of providers, rather than a one-size-fits-all pitch Partner closely with others in the team already covering this space to build out a durable, well-organized intelligence and relationship function Collaborate with the broader Supply and Deals functions to bring opportunities to the table and support negotiation when it's time to close WHAT WE’RE LOOKING FOR Equal parts relationship-builder and operator — you can open a door and also drive it

Machine LearningAIGoHR
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -86%

$185K – $220K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: We are seeking a strategic and technically fluent Lead, IT Audit to join our Finance team reporting to the Head of Internal Audit. This is a broad, high-impact role spanning both IT SOX compliance and operational IT audits. You will help establish and elevate our technology controls program end to end — owning the IT SOX lifecycle, designing the IT general and application controls framework, embedding AI and automation into how we test and monitor controls, and delivering value-added operational IT and cybersecurity audits that strengthen how the company builds and runs its systems. You will partner with leaders across Engineering, Security, IT, Finance, and the business to ensure sound technology controls are built into how the company operates as we scale. This role is ideal for someone who thinks like a builder, not just an auditor — someone who can translate complex control and security requirements into practical, scalable processes in a fast-moving SaaS environment with modern cloud architecture and complex data flows. This role can be based in either San Francisco or New York City. We work from our offices on M

AWSAzureGCPCI/CD
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -73.6%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As the Engineering Manager for Baseten's Cloud Platform team, you will directly manage a team of cloud platform engineers responsible for building the systems and processes that keep our infrastructure scalable, reliable, and efficient — from automated deployments and monitoring to performance optimization and incident response. You are a people-first leader with a strong cloud infrastructure background. You set a high bar for reliability and operational excellence, engage credibly in technical discussions and code reviews, and know how to build a culture of ownership and accountability. You'll spend most of your time close to the work: unblocking your team, shaping technical direction on day-to-day decisions, and developing your engineers. At Baseten, we work closely with our users to understand their struggles operationalizing ML — you'll keep your team connected to that mission and translate user learnings into better infrastructure. RESPONSIBILITIES Recruit, hire, and grow a high-performing team of cloud platform engineers; provide ongoing coaching, feedback, and career development through regular 1:1s. Set clear performance expectations, hold a high bar, and create an environment where engineers do their best work. Foster a culture of ownership, accountability, and continuous improvement. Drive day-to-day technical decisions through design reviews, code reviews, and architectural discussions; translate th

KubernetesCI/CDGitMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. Technical Program Managers at OpenAI play a key leadership role in scaling these efforts, partnering deeply with product, engineering, design, and go-to-market teams to bring ambitious ideas to life and ensure clarity and discipline in execution. About the Role We are hiring a Technical Program Manager to support OpenAI's critical AI deployments across strategic cloud partners. This role is designed for a candidate who can operate as an end-to-end owner across internal engineering teams and external partner organizations. This role will drive the technical strategy and execution required to bring OpenAI models and platform capabilities into partner environments responsibly and at scale. The work spans engineering deliverables, shared roadmaps, model launch pipelines, technical integration, launch readiness, and post-launch follow-through. You will work closely with senior leaders across OpenAI engineering, infrastructure, product, safety, security, legal, finance, and go-to-market, as well as technical counterparts at our partners. The job is to turn broad partnership commitments into concrete execution plans, align both sides on what must land, and build repeatable mechanisms for launching OpenAI capabilities on third-party platforms. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end execution for major cloud partner programs spanning model deployment, product integration, operational readiness, launch follow-through, and partner-platform adoption. Own integrated technical roadma

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun

AWSAzureGCPKubernetes
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc

AWSAzureGCPKubernetes
🔔

Get new cloud operations lead jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime