Jobs in United States

Cloud Operations Lead in United States

698 active opportunities · Updated October 2026

Explore current cloud operations lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -73.6%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. ABOUT THE TEAM Supply is responsible for knowing everything happening in the compute market: who's building, who's buying, and on what terms. This role owns a specific and fast-moving slice of that map — emerging clouds and international markets — and owns the full relationship lifecycle in that space, from first outreach through to closed terms. RESPONSIBILITIES Build and maintain a real-time picture of the emerging cloud and international compute landscape — who's active, what they're building, and what terms are available Own the full partnership lifecycle in this space — from identifying and sourcing new providers, to negotiating terms, to ongoing relationship management Develop and manage relationships across a broad set of emerging and international providers, from account reps up through leadership Identify, structure, and help close opportunities where Baseten can move quickly to secure favorable capacity terms Define compelling value propositions tailored to different types of providers, rather than a one-size-fits-all pitch Partner closely with others in the team already covering this space to build out a durable, well-organized intelligence and relationship function Collaborate with the broader Supply and Deals functions to bring opportunities to the table and support negotiation when it's time to close WHAT WE’RE LOOKING FOR Equal parts relationship-builder and operator — you can open a door and also drive it

Machine LearningAIGoHR
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -83.5%

From $154K/yr

Quick readStrong listing-quality and freshness signals

Datadog’s Implementation Services team helps customers implement and deploy Datadog quickly and successfully. Our team of architects leads the discovery, design, build, and launch of the Datadog platform to help customers accelerate time to value and get the most out of their investment. As a Senior Services Architect focused on Security and Cloud SIEM, you will help customers design, implement, and operationalize Datadog’s security capabilities across cloud, infrastructure, application, and log data sources. You will lead structured, outcome-driven professional services engagements delivered through a day-based professional services delivery model, partnering directly with customers through co-development working sessions, architecture workshops, implementation planning, and operational handoff. This role is ideal for someone who combines customer-facing consulting experience with strong cybersecurity knowledge, hands-on Cloud SIEM implementation skills, and an understanding of security control frameworks such as NIST 800-53, the NIST Cybersecurity Framework, CIS Controls, MITRE ATT&CK, SOC 2, PCI, HIPAA, ISO 27001, or similar standards. At Datadog, we place value in our office culture — the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Design and guide execution of Datadog implementations, focusing on Security and Cloud SIEM deployment, including discovery, requirements gathering, technical architecture, deployment planning, and launch. Partner with customers to map security requirements, controls, and monitoring objectives to Datadog capabilities, including frameworks such as NIST 800-53, NIST Cybersecurity Framework, CIS Controls, MITRE ATT&CK, SOC 2, PCI, HIPAA, ISO 27001, or similar standards. Advise customers on security data strategy, including log source prioritization, parsing, normalizat

KubernetesAIGoRust
P
📍 New York, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. What You’ll Do Regional & Executive Leadership Own the overall Americas Channel Sales strategy, operating model, and results across Enterprise West, Enterprise East, LATAM, and Federal/SLED. Build, lead, and scale the Americas channel team hiring, developing, and retaining 4-5 regional channel ICs plus evolving Americas-wide scope. Establish clear priorities, coverage models, KPIs, and operating rhythms aligned to global Postman objectives. Act as the senior leader and voice for Americas Channel within Postman. Partner Ecosystem Strategy Define and execute the Americas partner strategy across SIs, resellers, distributors, and technology alliances. Build a scalable, services-capable partner ecosystem (WWT, SHI, CDW, Insight, SoftwareONE, Carahsoft, plus FSIs BAH, Deloitte Federal, Accenture Federal Services, Leidos, SAIC, CGI Federal, GDIT, CACI plus LATAM partners including Caylent, Mission Cloud, and regional SIs). Ensure partners are enabled, certified, and accountable for sourcing, selling, and delivering Postman Enterprise. Revenue & Pipeline Ownership Own partner-sourced and partner-influenced pipeline and revenue acro

AWSAIGoRust
C
📍 Work At Home Texas, United States
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary Health100 is an AI‑native health technology platform that unifies pharmacies, providers, insurers, PBMs, and digital health solutions into a single, consumer‑focused ecosystem. Powered by Google Cloud AI, we’re reimagining personalized and connected health experiences. As a Senior Software Engineer for Health100, you will play a crucial role within a collaborative team — designing, developing, and maintaining backend services and APIs while ensuring releases are well-coordinated, fully prepared, and successfully deployed to production. The ideal candidate brings strong technical expertise in modern backend development, excellent problem-solving skills, and a proactive approach to production monitoring, issue triage, and cross-team coordination. This position is critical in maintaining high engineering standards, ensuring smooth release cycles, and driving operational excellence across the development lifecycle. *This role can be based anywhere in the US; hybrid or remote with preference for candidates to work out of our corporate headquarters in Woonsocket, RI. Responsibilities: Partner with technical leaders and the open-source community to contribute to technical designs, frameworks, roadmap definition, and requirements-gathering. Provide domain knowledge and engineering insight to guide early designs, ac

JavaAzureGCPAI
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -86%

$185K – $220K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: We are seeking a strategic and technically fluent Lead, IT Audit to join our Finance team reporting to the Head of Internal Audit. This is a broad, high-impact role spanning both IT SOX compliance and operational IT audits. You will help establish and elevate our technology controls program end to end — owning the IT SOX lifecycle, designing the IT general and application controls framework, embedding AI and automation into how we test and monitor controls, and delivering value-added operational IT and cybersecurity audits that strengthen how the company builds and runs its systems. You will partner with leaders across Engineering, Security, IT, Finance, and the business to ensure sound technology controls are built into how the company operates as we scale. This role is ideal for someone who thinks like a builder, not just an auditor — someone who can translate complex control and security requirements into practical, scalable processes in a fast-moving SaaS environment with modern cloud architecture and complex data flows. This role can be based in either San Francisco or New York City. We work from our offices on M

AWSAzureGCPCI/CD
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -73.6%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As the Engineering Manager for Baseten's Cloud Platform team, you will directly manage a team of cloud platform engineers responsible for building the systems and processes that keep our infrastructure scalable, reliable, and efficient — from automated deployments and monitoring to performance optimization and incident response. You are a people-first leader with a strong cloud infrastructure background. You set a high bar for reliability and operational excellence, engage credibly in technical discussions and code reviews, and know how to build a culture of ownership and accountability. You'll spend most of your time close to the work: unblocking your team, shaping technical direction on day-to-day decisions, and developing your engineers. At Baseten, we work closely with our users to understand their struggles operationalizing ML — you'll keep your team connected to that mission and translate user learnings into better infrastructure. RESPONSIBILITIES Recruit, hire, and grow a high-performing team of cloud platform engineers; provide ongoing coaching, feedback, and career development through regular 1:1s. Set clear performance expectations, hold a high bar, and create an environment where engineers do their best work. Foster a culture of ownership, accountability, and continuous improvement. Drive day-to-day technical decisions through design reviews, code reviews, and architectural discussions; translate th

KubernetesCI/CDGitMachine Learning
R
📍 New York City, NY, United States· Full-time
✓ High-confidence listingCompany trend -99.2%

From $10K/yr

Quick readStrong listing-quality and freshness signals

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role The Security Engineering team helps make Ramp the most secure place for our customers to collect, manage, and put to work their business’ financial information Our work centers in three areas: Ramp builds products with an eye for security Ramp detects and responds to threats before they cause harm Security powers Ramp’s growth Check out our Engineering Blog for more on our tech stack, mission and values! What You’ll Do Drive our cloud security roadmap: review our cloud deployments to identify opportunities for improvement Design and build security-focused infrastructure primitives and integrate them into our existing products and development processes Lead remediation of prioritized issues across our technology stack Partner with infrastructure, data, and devops teams to design and deploy solutions that are inherently secure What You Need Minimum 5 years of experience building software Minimum 3 years of experience building in AWS (with Terraform) A strong sense of ownership: you need to drive projects from inception to scaling it in

PythonAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. Technical Program Managers at OpenAI play a key leadership role in scaling these efforts, partnering deeply with product, engineering, design, and go-to-market teams to bring ambitious ideas to life and ensure clarity and discipline in execution. About the Role We are hiring a Technical Program Manager to support OpenAI's critical AI deployments across strategic cloud partners. This role is designed for a candidate who can operate as an end-to-end owner across internal engineering teams and external partner organizations. This role will drive the technical strategy and execution required to bring OpenAI models and platform capabilities into partner environments responsibly and at scale. The work spans engineering deliverables, shared roadmaps, model launch pipelines, technical integration, launch readiness, and post-launch follow-through. You will work closely with senior leaders across OpenAI engineering, infrastructure, product, safety, security, legal, finance, and go-to-market, as well as technical counterparts at our partners. The job is to turn broad partnership commitments into concrete execution plans, align both sides on what must land, and build repeatable mechanisms for launching OpenAI capabilities on third-party platforms. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end execution for major cloud partner programs spanning model deployment, product integration, operational readiness, launch follow-through, and partner-platform adoption. Own integrated technical roadma

AWSRestAIGo
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. NVIDIA has a rapidly expanding ecosystem of data center platform & node designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each bringing together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We're searching for a highly technical, motivated manager to lead & manage the team responsible for rack-scale system software architecture. From firmware, kernel drivers, operating systems, networking, fabrics and associated user mode drivers + manageability software. You will work with component leads internally and engage with industry leading hyperscalar / cloud service providers on taking these products to market. What you’ll be doing: Drive the software end-to-end architecture for NVIDIA's rack-scale products Maintain deep understanding of the product portfolio and roadmap; translate forward-looking plans into clear, formal software requirements that anchor execution across the organization. Ensure high quality & reliable

S
📍 United States· Full-time
✓ Quality checkedCompany trend -81%

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. Remote (US East Coast preferred, for timezone coverage) About the team Cloud Infrastructure owns the platform every Synthesia product runs on — AWS, Kubernetes, MongoDB, Temporal, our observability stack, and the vendor and cost relationships underneath them. We're a small, high-leverage team scaling toward a domain-ownership model: small groups that both build and operate the systems they're accountable for. The role We're hiring a dedicated SRE to take real ownership of operational excellence across Cloud Infrastructure. Today, too much critical operational knowledge — vendor relationships, cost management, and incident response — lives with one or two people. Your mission is to take genuine ownership of those domains, make them resilient to any single person, and raise the bar on how reliably we run. This is not simply a ticket-queue or keep-the-lights-on role. You'll own domains end to end: understand them deeply, operate them well, and build the automation and tooling that make them boring . We deliberately pair operational and engineering work so the role grows rather than narrows. What you'll own Incident management & operational excellence — take custody of the incident process: on-call quality, resp

PythonMongoDBAWSKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun

AWSAzureGCPKubernetes
P
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity At Postman, we are revolutionizing the way developers build, trace, and automate API workflows with Postman Flows , a powerful visual programming tool designed to simplify the development and sharing of API-powered applications. With an intuitive drag-and-drop interface, Postman Flows enables teams to collaborate and showcase their APIs regardless of technical expertise. We are looking for a Software and Systems Engineer to help scale and maintain the Flows runtime system. This system runs mission-critical automations in the cloud, with a focus on low latency, high throughput, and high availability. You’ll play a vital role in developing, deploying, and operating our backend services and infrastructure in a Kubernetes-based cloud environment. We’re looking for an experienced engineer who is excited not only about hands-on building as we ship and iterate on a weekly basis to get our product ready for GA, but who can also serve as a role model and mentor to other engineers. This role involves making key technical decisions and improvements to the system, as well as effectively making impact through influence wit

Node.jsAWSAzureGCP
C
📍 Hartford, United States
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary The Executive Director, Digital Engineering- Aetna Member Care and Journey Services is a senior technology leader responsible for setting the technical vision, architectural direction, and engineering execution for member centric services. This role leads large-scale engineering teams that build high-performance backend APIs, microservices, and cloud-native systems that power member experiences across digital, agent, provider, and partner channels. The leader ensures exceptional service stability, resiliency, innovation velocity, and alignment with enterprise user experience and operational goals. Key Responsibilities 1. Backend API & Microservices Engineering Leadership • Lead the design, development, and delivery of scalable backend systems, APIs, and microservices powering member-facing capabilities. • Define API contract standards, and integration patterns used across Member Services platforms. • Drive service modernization by adopting cloud‑native architectures, containerization, and event-driven patterns. 2. Service Stability, Observability & Resiliency • Establish standards for availability, resiliency, performance, and disaster recovery across all services. • Implement SLO/SLI/error budget frameworks, health checks, and high‑availability architectures. <p

C
📍 United States· Remote
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary: As a Senior Software Development Engineer at Aetna, you will play a critical leadership role in the design, development, and continuous enhancement of enterprise-scale Provider Applications. You will drive technical solutions for complex business problems, ensure application stability, and lead cross-functional initiatives as a Project Owner. This role requires a balance of hands-on engineering expertise, technical leadership, and delivery ownership, including overseeing vendor/contractor teams, ensuring alignment with enterprise architecture, and delivering high-impact solutions that improve provider data systems and operational efficiency. Required Qualifications: 5&#43; years of hands-on application development experience with Python and Google Cloud Platform (GCP) 2&#43; years of experience leading or contributing to large-scale application development initiatives Preferred Qualifications: Experience working in Agile/SCRUM environments Proven experience in project/program management, including planning, execution tracking, and delivery management Strong organizational, leadership, and planning skills with the ability to manage multiple priorities Experience working with distributed teams and cross-functional stakeholders Prior exposure to

C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc

AWSAzureGCPKubernetes
🔔

Get new cloud operations lead jobs in United States by email

Daily job updates · Unsubscribe anytime