We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary We are seeking an accomplished Principal Cloud Storage Engineer to lead the design, engineering, and evolution of our private cloud storage platforms. This role will focus on large-scale storage architecture, data protection, cyber recovery, and resiliency technologies across complex enterprise environments. The ideal candidate will combine deep technical expertise in storage systems with strong leadership, architectural vision, and the ability to influence technical direction across the organization. Key Responsibilities Architect and engineer enterprise storage platforms that ensure data integrity, availability, security, and disaster recovery readiness Design and implement end-to-end storage solutions, including Software Defined Storage, SAN, NAS, and object storage across private cloud and data center environments Drive strategic technology decisions by evaluating emerging products, tools, and standards supporting storage, data protection, cloud, and compute platforms Lead infrastructure initiatives involving storage modernization, data protection, cyber recovery, data migration, and resilience engineering Develop and execute enterprise strategies for backup, recovery, cyber vaulting, and business continuity Create and maintain comprehensive documentation of storage architectures, configurations, policies, and operation
Jobs in United States
Cloud Operations Lead in United States
698 active opportunities · Updated October 2026
Showing
15 jobs
Explore current cloud operations lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives, spanning AI research specialists, silicon designers, software engineers and systems architects. Job Summary We are looking for an experienced Principal Engineer to join our System Management team and help lead the development of critical interfaces used by internal and external customers to manage system state. You will provide technical leadership within assigned areas of System Management, guide architecture and implementation choices, mentor engineers and translate broader technical direction into effective execution. This is a hands-on engineering role for someone who can lead complex technical work, improve reliability and operational readiness, and collaborate effectively across multiple engineering disciplines. The Team The System Management team sits within the Software Platform group and helps build Graphcore products into large-scale AI solutions for our customers. The team is responsible for developing the interfaces between hardware, AI software and frameworks, as well as providing interfaces for public and private cloud environments. This includes system management capabilities that abstract complex hardware administration and enable reliable deployment and operation at scale. As one of the first teams to work with new hardware and software, we regularly solve complex system-level problems
From $244K/yr
Role Summary: Datadog is seeking a Staff Software Engineer to help shape the future of our Bring Your Own Cloud (BYOC) Logs offering by unifying observability pipelines with log management software that customers deploy and manage in their own infrastructure. This role will focus on building and scaling systems that process, route, and store high-volume observability data within customer-managed infrastructure. You will operate as a hands-on technical leader, driving architecture, cross-team delivery, and product direction across a complex and evolving space. This is a high-impact opportunity to influence product strategy, mentor engineers, and solve deeply technical challenges at scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Make customer-controlled deployments feel like a managed Datadog product: deployment, upgrades, configuration, observability, diagnostics, reliability, and secure operation across diverse customer cloud environments Build and scale high-throughput systems for log processing, routing, and transformation across distributed environments Lead cross-team initiatives, aligning engineers, product managers, and stakeholders to deliver complex, multi-team projects Design and implement software that runs reliably that customers deploy and operate within their own cloud infrastructure. Improve system performance, scalability, and cost efficiency through thoughtful trade-off analysis and capacity planning Contribute hands-on to critical code paths, debugging, and deployment challenges in customer environments Who You Are: You have significant experience building software that is installed, deployed, and operated in customer environments rather than only as a fully managed SaaS service. You have strong expertise in distributed systems,
About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, cloud services, mechanical engineering, electrical engineering, and product design to deliver reliable, production-ready devices at scale. Within Consumer Devices, Hardware Engineering eXperience, or HEX, is a new bootstrapped team building the environments, applications, compute, product-data systems, and workflows that let hardware engineers do their work without needing to troubleshoot the machinery underneath. HEX owns virtual engineering environments, HPC/GPU compute, storage, networking, licensing, MCAD/ECAD/CAE applications, PLM, product data, automation, validation, and support as one connected system. About the Role As a Staff PLM & Engineering Applications Engineer, you will be one of the first technical builders of HEX and the primary counterpart to the HEX lead. You will own the engineering-application and product-data side of the hardware engineering experience, with an initial focus on NX, Teamcenter, licensing, parts import, integrations, packaging, validation, and user workflows. This is not a traditional Teamcenter administration role and not a Corporate IT application-support role. You will take complex, fragile workflows and turn them into reliable engineering systems. This role is highly hands-on and systems-oriented. You will not inherit a mature environment and support queue. You will help build a fresh one, replacing manual setup guides, tribal knowledge, repeated support issues, and team handoffs with tested automation and reliable workflows. In This Role, You Will Own the technical architecture, deployment, configuration, integration, validation, and long-term operation of NX and Teamcenter. Build reliable workflows for parts import, product-data migration, metadata quality, BOMs, revisions, lifecycle states, and releas
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Network Engineer Location: MPK/Bellevue/Dublin Snowflake's Enterprise Technology Network Services team is looking for a Senior Network Engineer to lead the design, operation, and optimization of our Zero Trust and secure-access platform. This role is Zscaler-centric — you will own the health, performance, and roadmap of our ZIA/ZPA deployment — while working across a modern, multi-cloud network stack that supports a global workforce of 10,000+ users. You'll be the escalation point for the most complex connectivity issues and a driver of automation and observability across the environment. What You'll Do Own and operate the Zscaler platform (ZIA, ZPA, ZDX, ZCC) end-to-end, including policy frameworks, app-segmentation models, PAC/traffic-forwarding standards, App Connector topology, and NSS/log-streaming design. Troubleshoot secure-access incidents like tunnel flapping, broker/connector health, SSL inspection edge cases, DNS/DTLS failures, and lead root-cause analysis for systemic issues. Manage Palo Alto firewalls, Panorama and GlobalProtect VPN, while planning migration toward Zscaler solutions. Support Aruba (Central) Wireless, Ekahau and Cisco Catalyst switches globally. Design, install, and configure network devices and ISP circuits at new offices. Build automatio
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Corporate Systems Engineering builds and operates the software platforms, integrations, and automations that power Smartsheet’s core business functions across Finance, Sales/GTM, and People & Culture. Our team owns mission-critical systems and workflows that enable how the company hires, sells, bills, pays, reports, and scales. We operate at the intersection of software engineering, enterprise platforms, and business-critical data, treating internal systems with the same rigor, reliability, and product mindset as customer-facing software. The Automation team builds human-to-system and system-to-system automations that reduce manual effort and friction across the business. We combine cloud-native services, agentic AI, and workflow orchestration to enable employees to interact with enterprise systems through intelligent, secure, and auditable automation. As a Senior Software Engineer I (Automation), you will lead the design, build, and operation of systems and workflows that directly support business execution at scale. You will own complex technical initiatives, partner with Product Managers and stakeholders on technical roadmaps, and mentor junior engineers. This full-time position reports to the Sr. Director, Development and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Architect AI Agents: Take a leading role in designing Agentic Workflows using AWS Step Functions and Bedrock Agents that reason
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches
About Datadog: Datadog is the essential monitoring and security platform for cloud applications. We bring together end-to-end traces, metrics, and logs to make your applications, infrastructure, and third-party services entirely observable. These capabilities help businesses secure their systems, avoid downtime, and ensure customers are getting the best user experience. The Team: The Financial Planning & Analysis (FP&A) team analyzes company financial data (revenue, customers, headcount, expenses, etc.) in order to support the business’ growth and success. Within FP&A, the R&D Finance team enables the financial strategy behind Datadog’s Engineering and Product organizations. We partner directly with technical leadership to drive decision-making around our most critical investments. Your work will be highly cross-functional and play a pivotal role in connecting the dots across the organization through a financial lens, ensuring operational alignment and informing decision making. The Opportunity: As a part of the R&D Finance team, this person will be key in supporting our product and engineering leadership team. Reporting to the Senior Manager, you will help create our annual budget and financial targets and work with operational leaders to support execution against our goals. Your work will be highly cross-functional and strategic, and you will play a pivotal role in connecting the dots across the organization through a financial lens, in order to help ensure operational alignment and inform decision making. What You’ll Do: Partner with senior business leaders to manage their departmental budgets on a regular basis, including but not limited to the company’s co-founder and CTO Manage financial forecasts and analytics, which include data across revenue, customers, product, company expenses, and headcount / workforce. Help manage significant portions of the company’s spend, including but not limited to AI and GPU spend across the compan
From $150.5K/yr
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. We are seeking a Global Senior Manager of Leave of Absence to join our team in a hybrid capacity—based out of our Santa Clara, CA or Bellevue, WA offices (in-office Tuesday–Thursday)—or as a Remote (US) team member. Reporting directly to the AMS Director of Benefits within the Total Rewards department, you will serve as the strategic architect and hands-on operational lead for our global leave, time off, and disability programs. In this high-impact role, you will design, implement, and scale leave policies and processes tailored for a diverse, rapidly growing global workforce. Your mission is to ensure our programs remain highly market-competitive and legally compliant, while championing an "employee-first" philosophy that deeply reflects our core values. What you’ll do (Role Expectations) Design and standardize global leave frameworks that balance statutory requirements with company benefits, serving as the lead SME for
Become a part of our caring community Humana is seeking a Lead Cloud Architect – NoSQL Databases to provide strategic leadership, architecture direction, and engineering oversight for enterprise NoSQL database platforms across Humana’s cloud environments. This role will focus on the design, implementation, modernization, and governance of NoSQL database solutions, including MongoDB, Azure Cosmos DB, Neo4j, and vector database technologies. The successful candidate will help define and advance Humana's enterprise NoSQL strategy, support platform rationalization initiatives, and ensure database solutions are secure, scalable, resilient, automated, and aligned with enterprise architecture standards. This role requires hands-on technical depth, strong cloud architecture experience, and the ability to collaborate across security, engineering, quality, application, and business teams. Key Responsibilities Develop, document, and maintain enterprise-wide NoSQL database standards, reference architectures, design patterns, and best practices. Lead architecture and engineering efforts for NoSQL platforms including MongoDB, Azure Cosmos DB, Neo4j, and vector databases. Support NoSQL platform rationalization and modernization initiatives, including migration planning and execution from Cosmos DB to MongoDB where appropriate. Architect secure, highly available, scalable, and performant NoSQL database solutions across cloud environments, including Azure and/or Google Cloud Platform. Define database architecture patterns for document databases, graph databases, key-value workloads, and vector search use cases. Guide application teams on NoSQL data modeling, partitioning, indexing, query patterns, performance optimization, and operational readiness. Oversee automation of NoSQL provisioning, configuration, monit
From $296K/yr
Datadog’s Cloud Observability group is one of the core data retrieval and processing groups powering our foundational product, Infrastructure Monitoring. The group’s scope includes integration with all major hyperscalers (AWS, Azure, GCP, OCI), as well as both regional and GPU-specific cloud providers. As Director, you will own engineering for all clouds, generating more than 10 million metric points per second, managing ~40 engineers through a team of Engineering Managers. You’ll partner with Senior Directors and product leadership to shape the roadmap, not just execute against it, managing the growth of one of Datadog’s foundational teams. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You'll Do: Own engineering for all of Cloud Observability Manage ~40 engineers through a layer of Engineering Managers; this is a manager-of-managers role Shape the roadmap alongside product leadership rather than simply executing against it — push back on, iterate on, and help author the strategy for your area Drive AI adoption across the engineering org, from tooling and workflows to product features and team practices Navigate cross-team dependencies across the Agent, Telemetry Onboarding, Integrations, Action Platform, and Infrastructure Monitoring. Build and retain engineering talent in NYC, Boston, and Paris, mentor Engineering Managers toward Director readiness, and participate in the on-call rotation Who You Are: You have directly managed Engineering Managers, not just individual contributors You have deep experience with one or more cloud providers, ideally with experience operating large-scale systems in the cloud. You have a solid understanding of cloud economics, as well as how to balance performance and cos
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE This role owns Baseten's relationships and market intelligence across hyperscalers and strategic neoclouds, including NVIDIA cloud partners. This is a technical and commercial role in equal measure: you'll evaluate capacity from the GPU to the data center, negotiate cost and terms with suppliers, and stay close enough to the market to develop and defend a real point of view on where it's heading. Given current market conditions, Baseten needs a much stronger pulse on this part of the market so we can track pricing, stay close to the right relationships, and move fast the moment more capacity is needed. This is a senior, experienced hire who will also help pair with and develop 1-2 junior to mid-level teammates covering the same space. WHAT YOU'LL DO Build and maintain deep relationships across hyperscalers and strategic neoclouds (including NVIDIA cloud partners), working each organization from top to bottom rather than a single point of contact Maintain a consistent, "top of mind" presence with key accounts so Baseten is positioned to move quickly when capacity needs arise Evaluate capacity from the GPU to the data center — hardware generation, rack and node configuration, interconnect, power density, and cooling — so you know what a configuration will actually deliver, not just what the spec sheet claims Live in compute pricing daily: track rates by GPU generation, region, and contract term to keep Baseten inf
The Engineering Lead Analyst – SonarQube & Code Quality Engineering is a senior-level engineering role responsible for leading static code analysis, automated code quality governance, security vulnerability remediation, and AI-augmented developer enablement across enterprise software delivery pipelines. In this role, you will champion software reliability, maintainability, clean-coding standards, and automated quality gates. You will partner with development teams, system architects, and platform engineering to integrate and manage enterprise-scale code quality platforms (such as SonarQube) both on-premises and in cloud/SaaS environments. Additionally, you will drive modern engineering practices by embedding Behavior-Driven Development (BDD) within your own software delivery and leveraging Agentic AI workers and Model Context Protocol (MCP) architectures to optimize developer experience, streamline code governance, and boost engineering velocity. Key Responsibilities 1. Code Quality & Static Analysis Platform Ownership Lead the architecture, deployment, administration, and continuous enhancement of enterprise Static Application Security Testing (SAST) and Code Quality platforms (e.g., SonarQube , DeepSource, Codacy, Semgrep). Configure, calibrate, and enforce automated Quality Gates, code rulesets, technical debt calculation models, and code-coverage baselines across multi-language enterprise repositories. Oversee version upgrades, patching, high availability, and operational maintenance for on-premises and SaaS/cloud-hosted code quality infrastructure. 2. CI/CD & Pipeline Integration <li style=
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. The Principal Product Marketing Manager for the Cloud Memory Business Unit is responsible for defining and communicating the value of Micron's memory portfolio for AI, cloud, and data center markets. This role serves as a strategic link between product management and global marketing to shape market perception, drive demand, and enable competitive differentiation. At the Principal level, this role operates with significant independence and serves as a subject matter authority, influencing cross-functional decisions and shaping business outcomes at the portfolio level. This Principal PMM operates as a subject matter authority, influencing both near-term go-to-market execution and long-term category leadership positioning in high-growth markets. The role is expected to harness advanced AI technologies, including the creation, deployment, and management of AI agents and automated workflows, to optimize market analysis, customer insights, content development, campaign execution, competitive intelligence, and operational effectiveness. The successful candidate will demonstrate proficiency in finding opportunities where AI can accelerate business outcomes, increase productivity, improve decision-making, and scale marketing impact across the organization. Responsibilities: Product Positioning & Value Definition - Define and articulate clear product value propositions, differentiated positioning, and customer-facing messaging; translate complex technical capabilities into business outcomes Go-to-Ma
Other cities to consider
More places hiring for this role
Get new cloud operations lead jobs in United States by email
Daily job updates · Unsubscribe anytime