Jobs in United States

Cloud Operations Lead in United States

698 active opportunities · Updated October 2026

Explore current cloud operations lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

Z
📍 California, United States· Full-time· Remote
✓ High-confidence listing

From $164.5K/yr

Quick readStrong listing-quality and freshness signals

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Principal Production Engineer to join our team. This role is available as a hybrid opportunity 3 days a week in San Jose, CA or Remote reporting to Production Engineering in the Cloud Infrastructure & Operations department. Join Zscaler to be a force multiplier for the reliability of a global platform processing 200+ billion transactions daily across tens of millions of enterprise users. In this role, you will provide the technical vision and hands-on execution to drive an "automation-first" culture across the company. By maturing our observability and architectural standards, you will directly reduce our Mean Time to Mitigate (MTTM) and shape the scalability of our globally distributed, multi-cloud infrastructure. What you’ll do (Role Expectations) Design and implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments Drive an "automati

PythonAWSAzureGCP
Z
📍 Bellevue, Washington, United States· Full-time· Remote
✓ High-confidence listing

From $96.6K/yr

Quick readStrong listing-quality and freshness signals

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Detection Analyst to join our team. This is a Remote (USA Only) role, reporting to the Detection Analyst Manager in the Detection Operations Team. You will utilize our advanced detection platform to analyze EDR telemetry, alerts, and log sources across key detection domains—including Endpoint, Identity, Network, and Cloud—to identify and publish threats for our Managed Detection and Response customers. By driving projects that improve our workflow through orchestration and automation, you will ensure our customers receive high-fidelity threat analysis and concise, actionable communication regarding emerging threats at scale. What you’ll do (Role Expectations) Analyze EDR telemetry, alerts, and log sources across several detection domains including Endpoint, Identity, Network, and Cloud/SaaS Publish threats for customers using concisely written communication to effectively convey key indica

SQLAWSGitAgile
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel

PythonKubernetesLinuxArtificial Intelligence
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys

PythonJavaKubernetesLinux
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $131K/yr

Quick readStrong listing-quality and freshness signals

Datadog is a world-class monitoring and security platform for cloud applications. We're dedicated to creating, developing, and supporting our product and customers, fostering seamless collaboration and problem-solving among Dev, Ops, and Security teams globally. Built by engineers for engineers, our SaaS product is trusted by organizations of all sizes across various industries, driving digital transformation, cloud migration, and infrastructure monitoring for our customers' entire technology stacks. As cloud technologies continue to evolve and digital operations become ever more crucial, Datadog remains at the forefront of innovation and is well-positioned for sustained growth. Datadog's Finance team partners with stakeholders across the organization, providing commercial, operational, and analytical support to ensure that Datadog's business continues its rapid and efficient growth. The Equity Administration team is responsible for overseeing the full cycle of stock administration on a global basis, collaborating with groups across Accounting, Tax, People, Payroll, and Legal to manage Datadog's equity plans worldwide. Our work spans the processing of equity transactions including RSU releases, stock option exercises, and ESPP purchases; management of equity-related inquiries from employees, executives, investors, and other key partners; day-to-day management of the E*TRADE Equity Edge Online platform; global equity reporting and tax compliance; oversight of equity procedures, policies, and controls; and employee education on equity programs. We are seeking a hands-on Equity Administration Assistant Manager to own our complex and high-risk equity workstreams. Reporting to the Senior Manager, Equity Administration, the Assistant Manager will serve as a first level preparer or reviewer on equity transactions, assist with process documentation, and work closely with various internal teams across Accounting, Tax, People Operations, Payroll, and Legal. This is a ro

PythonSQLGitAI
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $195K/yr

Quick readStrong listing-quality and freshness signals

Here at Datadog, we think about offensive security a little bit differently. We embrace automation and AI to run adversary simulations continuously across a massive cloud-native environment, and we expect our offensive engineers to build the tooling that makes that possible. We're looking for a Senior Security Engineer who can execute sophisticated red team operations, write the code that scales them, and take an AI-first approach to offensive security engineering. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Plan and execute red team engagements end-to-end, simulating real-world threat actors across cloud infrastructure (AWS, GCP), Kubernetes, CI/CD pipelines, and corporate environments Build and maintain custom offensive tooling, automation frameworks, and engagement infrastructure, treating offensive operations as a software engineering problem Develop custom payloads and evasion capabilities tailored to Datadog's environment and modern defensive controls (EDR, SIEM, network monitoring) Improve the efficiency of offensive operations through thoughtful use of automation and AI, accelerating reconnaissance, vulnerability analysis, and reporting workflows Partner with the Detection & Response team on purple team exercises to validate detection logic, improve alert fidelity, and influence threat models Translate offensive findings into concrete improvements by working directly with defensive security and engineering teams to close gaps Who You Are: You have 5+ years of hands-on experience in offensive security (red teaming, penetration testing, or adversary simulation) with a track record of operating against mature, well-defended environments You write production-quality code (Python, Go, or similar), can build your own tools, and automate your w

PythonAWSAzureGCP
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -93.7%

From $140K/yr

Quick readStrong listing-quality and freshness signals

MongoDB's Partner Program is how we formalize value for every partner type in our ecosystem — Technology/ISV partners, Services and Solutions partners (including our Systems Integrators), Cloud partners (AWS, Azure, Google Cloud), Built with MongoDB partners, and Certified by MongoDB DBaaS partners. We're looking for a Director, Partner Programs to own this program end to end: the tiering and benefits structure, the incentive framework, certification (including our SI Associate and SI Architect tracks), and the operational rigor that keeps thousands of partners engaged and productive. You'll report to the VP of Partner Strategic Operations and oversee all MongoDB Partner Programs globally, managing a team of Partner Program Managers. You'll be the connective tissue between the field-facing work being done in Sales Plays & Offerings, the platform work being built by our Business Product Manager (PRM), and the automation work coming out of Digital Process & AI Operations — making sure the underlying program itself stays coherent, well-governed, and aligned to company strategy as all of that activity happens around it. This role can be hybrid near any of our MongoDB hub offices in the United States or remote in the United States. What You'll Do Own Partner Program strategy and structure Own and evolve the overall architecture of the MongoDB Partner Program across all partner categories — Technology, Services and Solutions, Cloud, Built with MongoDB, and Certified by MongoDB DBaaS — including tiering, entitlements, and benefits at each level Ensure the program structure stays aligned to MongoDB's broader GTM strategy, including emerging priorities like AI-driven partnerships and cloud marketplace growth Own partner certification pathways (e.g., SI Associate and SI Architect certifications) and ensure they reflect current MongoDB technology and sales priorities Design and govern incentives Own the global partner incentive framework — MDF, SPIFs, rebates, and co-s

MongoDBAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

Overview: The Data Acquisition team within the Foundations organization at OpenAI is responsible for all aspects of data collection to support our model training operations. Our team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams. We are looking for a skilled Full-Stack Engineer to join our Data Acquisition team to build and optimize the interfaces and tools that power our data infrastructure. Responsibilities: Develop and maintain full-stack applications that support data acquisition, including internal tools and dashboards. Collaborate closely with cross-functional teams, including Data Processing, Architecture, and Scaling, to ensure seamless data ingestion and workflow management. Design and implement APIs to facilitate data interactions between internal services and external data sources. Enhance user experience by developing intuitive web-based interfaces for managing and monitoring data pipelines. Optimize backend services for performance, scalability, and security in a distributed computing environment. Work with legal and compliance teams to ensure our data acquisition processes adhere to privacy regulations and best practices. Deploy and maintain infrastructure using Kubernetes and Infrastructure-as-Code (IaC) methodologies. Analyze system performance, conduct experiments, and improve data workflows to maximize efficiency. Qualifications: BS/MS/PhD in Computer Science or a related field. 4+ years of industry experience in full-stack development. Proficiency in frontend frameworks (React, Vue, or similar) and backend technologies such as Python, Node.js, or Go. Strong expertise in RESTful APIs, GraphQL, and database design (SQL and NoSQL). Experience building data-intensive applications that handle large-scale datasets. Familiarity with cloud platforms (AWS, GCP, or Azure) and container orchestration (Kubernetes, Docker). Prior experience with web crawling and large-scale data processing is a

PythonReactNode.jsVue
L
📍 Bethesda, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Leidos has an exciting opportunity for a Principal DevOps Engineer in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary This Principal DevOps Engineer role provides mission critical system support to our customer. You will closely work with the Development team as well as other technology stakeholders to maintain, develop and support IC enterprise products – legacy and new products – in an Agile SAFe environment. The role will also work collaboratively with software engineering to deploy and operate systems. Additionally, this role will help automate and streamline operations and processes; as well as build and maintain tools for deployment, monitoring and operations, and troubleshoot and resolve issues in dev, test, and production environments. Primary Responsibilities: Supports software deployments, cloud infrastructure baselines, and operational availability of production systems. Managing, building, configuring, administering, operating and maintaining all components that comprise the DevOps environment. Defining enterprise Continuous Integration/Continuous Deployment processes and best practices Codifying DevOps best practices across the enterprise Developing and maintaining scripts to automate tool deployment to an AWS cloud environment and other tasks. <l

JavaScriptPythonJavaAWS
L
📍 Bethesda, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Leidos has an exciting opportunity for a Principal Software Engineer in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary As a Software Engineer on this program, you will have the opportunity to build strong systems, software, and cloud environments while providing operations and maintenance for critical systems. This role will provide technical expertise in the design, development, implementation and testing of customer tools and applications. Based in a DevOps framework, this role participates in and/or directs major deliverables of projects through all aspects of the software development lifecycle including scope and work estimation, architecture and design, coding and unit testing. Primary Responsibilities: Participates in and/or directs software programming initiatives using Java, JavaScript, Python, SpringBoot, and Hibernate. Develops software system validation and testing methods using Junit and Katalon and uses integrated custom developed software solutions to leverage automated deployment technologies Develop, prototype and deploy solutions within Commercial Cloud Solutions leveraging infrastructure platform services Coordinate closely with team members, Product Owners and Scrum Masters to ensure User Story alignment and implementation to customer use cases Supp

JavaScriptPythonJavaAWS
L
📍 Bethesda, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Leidos has an exciting opportunity for a Principal Java Developer in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary A Java developer on this program provides Agile development and operations and maintenance for mission critical systems. Based in DevOps framework, this role participates in major deliverables of projects through all aspects of the software development lifecycle including scope and work estimation, architecture and design, coding and unit testing. Primary Responsibilities: Design, develop, test and maintain high-performance, scalable backend microservices using Java and the Spring Framework to meet customer information technology needs. Participates in software programming initiatives, shaping backend architecture, mentoring team members, and conducting code reviews. Develops and directs software system validation and testing methods using Junit and Katalon and uses integrated custom developed software solutions to leverage automated deployment technologies Develop, prototype and deploy solutions within a cloud-based platform leveraging platform services. Analyze (though proof of concept, performance, and end-to-end testing) and effectively coordinate Infrastructure needs driven by developed software to meet customer mission needs Support the Agil

PythonJavaPostgreSQLAWS
L
📍 Bethesda, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Leidos has an exciting opportunity for a Sr. DevOps Engineer in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary This DevOps Engineer role provides mission critical system support to our customer. You will closely work with the Development team as well as other technology stakeholders to maintain, develop and support IC enterprise products – legacy and new products – in an Agile SAFe environment. The role will also work collaboratively with software engineering to deploy and operate systems. Additionally, this role will help automate and streamline operations and processes; as well as build and maintain tools for deployment, monitoring and operations, and troubleshoot and resolve issues in dev, test, and production environments. Primary Responsibilities: Supports software deployments, cloud infrastructure baselines, and operational availability of production systems. Managing, building, configuring, administering, operating and maintaining all components that comprise the DevOps environment. Defining enterprise Continuous Integration/Continuous Deployment processes and best practices Codifying DevOps best practices across the enterprise Developing and maintaining scripts to automate tool deployment to an AWS cloud environment and other tasks. Scripting and

JavaScriptPythonJavaAWS
L
📍 Bethesda, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Leidos has an exciting opportunity for a DevOps Engineer (SME) in our Intel Security Sector's Analysis Solutions Business Are) . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary This DevOps Engineer role provides mission critical system support to our customer. You will closely work with the Development team as well as other technology stakeholders to maintain, develop and support IC enterprise products – legacy and new products – in an Agile SAFe environment. The role will also work collaboratively with software engineering to deploy and operate systems. Additionally, this role will help automate and streamline operations and processes; as well as build and maintain tools for deployment, monitoring and operations, and troubleshoot and resolve issues in dev, test, and production environments. Primary Responsibilities: Supports software deployments, cloud infrastructure baselines, and operational availability of production systems. Managing, building, configuring, administering, operating and maintaining all components that comprise the DevOps environment. Defining enterprise Continuous Integration/Continuous Deployment processes and best practices Codifying DevOps best practices across the enterprise Developing and maintaining scripts to automate tool deployment to an AWS cloud environment and other tasks. Scripting an

JavaScriptPythonJavaAWS
L
📍 Bethesda, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Leidos has an exciting opportunity for a Sr. Software Engineer in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary As a Software Engineer on this program, you will have the opportunity to build strong systems, software, and cloud environments while providing operations and maintenance for critical systems. This role will provide technical expertise in the design, development, implementation and testing of customer tools and applications. Based in a DevOps framework, this role participates in and/or directs major deliverables of projects through all aspects of the software development lifecycle including scope and work estimation, architecture and design, coding and unit testing. Primary Responsibilities: Participates in and/or directs software programming initiatives using Java, JavaScript, Python, SpringBoot, and Hibernate. Develops software system validation and testing methods using Junit and Katalon and uses integrated custom developed software solutions to leverage automated deployment technologies Develop, prototype and deploy solutions within Commercial Cloud Solutions leveraging infrastructure platform services Coordinate closely with team members, Product Owners and Scrum Masters to ensure User Story alignment and implementation to customer use cases Support th

JavaScriptPythonJavaAWS
L
📍 Bethesda, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Leidos has an exciting opportunity for a Software Engineer (SME) in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary As a Software Engineer on this program, you will have the opportunity to build strong systems, software, and cloud environments while providing operations and maintenance for critical systems. This role will provide technical expertise in the design, development, implementation and testing of customer tools and applications. Based in a DevOps framework, this role participates in and/or directs major deliverables of projects through all aspects of the software development lifecycle including scope and work estimation, architecture and design, coding and unit testing. Primary Responsibilities: Participates in and/or directs software programming initiatives using Java, JavaScript, Python, SpringBoot, and Hibernate. Develops and directs software system validation and testing methods using Junit and Katalon and uses integrated custom developed software solutions to leverage automated deployment technologies Develop, prototype and deploy solutions within Commercial Cloud Solutions leveraging infrastructure platform services Coordinate closely with team members, Product Owners and Scrum Masters to ensure User Story alignment and implementation to customer use cases

JavaScriptPythonJavaAWS
🔔

Get new cloud operations lead jobs in United States by email

Daily job updates · Unsubscribe anytime