Jobs in United States

Lead Cloud Operations Engineer in United States

2,434 active opportunities · Updated October 2026

Explore current lead cloud operations engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op

PythonAWSRestAI
O
📍 United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches

AWSAzureGCPKubernetes
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives, spanning AI research specialists, silicon designers, software engineers and systems architects. Job Summary We are looking for an experienced Principal Engineer to join our System Management team and help lead the development of critical interfaces used by internal and external customers to manage system state. You will provide technical leadership within assigned areas of System Management, guide architecture and implementation choices, mentor engineers and translate broader technical direction into effective execution. This is a hands-on engineering role for someone who can lead complex technical work, improve reliability and operational readiness, and collaborate effectively across multiple engineering disciplines. The Team The System Management team sits within the Software Platform group and helps build Graphcore products into large-scale AI solutions for our customers. The team is responsible for developing the interfaces between hardware, AI software and frameworks, as well as providing interfaces for public and private cloud environments. This includes system management capabilities that abstract complex hardware administration and enable reliable deployment and operation at scale. As one of the first teams to work with new hardware and software, we regularly solve complex system-level problems

PythonKubernetesCI/CDGit
S
📍 Bellevue, WA, United States· Full-time
✓ High-confidence listingCompany trend -91.7%
Quick readStrong listing-quality and freshness signals

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Corporate Systems Engineering builds and operates the software platforms, integrations, and automations that power Smartsheet’s core business functions across Finance, Sales/GTM, and People & Culture. Our team owns mission-critical systems and workflows that enable how the company hires, sells, bills, pays, reports, and scales. We operate at the intersection of software engineering, enterprise platforms, and business-critical data, treating internal systems with the same rigor, reliability, and product mindset as customer-facing software. The Automation team builds human-to-system and system-to-system automations that reduce manual effort and friction across the business. We combine cloud-native services, agentic AI, and workflow orchestration to enable employees to interact with enterprise systems through intelligent, secure, and auditable automation. As a Senior Software Engineer I (Automation), you will lead the design, build, and operation of systems and workflows that directly support business execution at scale. You will own complex technical initiatives, partner with Product Managers and stakeholders on technical roadmaps, and mentor junior engineers. This full-time position reports to the Sr. Director, Development and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Architect AI Agents: Take a leading role in designing Agentic Workflows using AWS Step Functions and Bedrock Agents that reason

JavaScriptTypeScriptPythonJava
H
📍 Louisville, United States
✓ Quality checkedCompany trend +310%

Become a part of our caring community Humana is seeking a Lead Cloud Architect – NoSQL Databases to provide strategic leadership, architecture direction, and engineering oversight for enterprise NoSQL database platforms across Humana’s cloud environments. This role will focus on the design, implementation, modernization, and governance of NoSQL database solutions, including MongoDB, Azure Cosmos DB, Neo4j, and vector database technologies. The successful candidate will help define and advance Humana's enterprise NoSQL strategy, support platform rationalization initiatives, and ensure database solutions are secure, scalable, resilient, automated, and aligned with enterprise architecture standards. This role requires hands-on technical depth, strong cloud architecture experience, and the ability to collaborate across security, engineering, quality, application, and business teams. Key Responsibilities Develop, document, and maintain enterprise-wide NoSQL database standards, reference architectures, design patterns, and best practices. Lead architecture and engineering efforts for NoSQL platforms including MongoDB, Azure Cosmos DB, Neo4j, and vector databases. Support NoSQL platform rationalization and modernization initiatives, including migration planning and execution from Cosmos DB to MongoDB where appropriate. Architect secure, highly available, scalable, and performant NoSQL database solutions across cloud environments, including Azure and/or Google Cloud Platform. Define database architecture patterns for document databases, graph databases, key-value workloads, and vector search use cases. Guide application teams on NoSQL data modeling, partitioning, indexing, query patterns, performance optimization, and operational readiness. Oversee automation of NoSQL provisioning, configuration, monit

PythonMongoDBAzureKubernetes
C
📍 Tampa Florida United States, United States
✓ High-confidence listingCompany trend +800%
Quick readStrong listing-quality and freshness signals

The Engineering Lead Analyst – SonarQube & Code Quality Engineering is a senior-level engineering role responsible for leading static code analysis, automated code quality governance, security vulnerability remediation, and AI-augmented developer enablement across enterprise software delivery pipelines. In this role, you will champion software reliability, maintainability, clean-coding standards, and automated quality gates. You will partner with development teams, system architects, and platform engineering to integrate and manage enterprise-scale code quality platforms (such as SonarQube) both on-premises and in cloud/SaaS environments. Additionally, you will drive modern engineering practices by embedding Behavior-Driven Development (BDD) within your own software delivery and leveraging Agentic AI workers and Model Context Protocol (MCP) architectures to optimize developer experience, streamline code governance, and boost engineering velocity. Key Responsibilities 1. Code Quality & Static Analysis Platform Ownership Lead the architecture, deployment, administration, and continuous enhancement of enterprise Static Application Security Testing (SAST) and Code Quality platforms (e.g., SonarQube , DeepSource, Codacy, Semgrep). Configure, calibrate, and enforce automated Quality Gates, code rulesets, technical debt calculation models, and code-coverage baselines across multi-language enterprise repositories. Oversee version upgrades, patching, high availability, and operational maintenance for on-premises and SaaS/cloud-hosted code quality infrastructure. 2. CI/CD & Pipeline Integration <li style=

JavaScriptTypeScriptPythonJava
R
📍 New York City, NY, United States· Full-time
✓ High-confidence listingCompany trend -99.2%

From $10K/yr

Quick readStrong listing-quality and freshness signals

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role The Security Engineering team helps make Ramp the most secure place for our customers to collect, manage, and put to work their business’ financial information Our work centers in three areas: Ramp builds products with an eye for security Ramp detects and responds to threats before they cause harm Security powers Ramp’s growth Check out our Engineering Blog for more on our tech stack, mission and values! What You’ll Do Drive our cloud security roadmap: review our cloud deployments to identify opportunities for improvement Design and build security-focused infrastructure primitives and integrate them into our existing products and development processes Lead remediation of prioritized issues across our technology stack Partner with infrastructure, data, and devops teams to design and deploy solutions that are inherently secure What You Need Minimum 5 years of experience building software Minimum 3 years of experience building in AWS (with Terraform) A strong sense of ownership: you need to drive projects from inception to scaling it in

PythonAWSAzureGCP
DC
📍 New York, New York, United States· Full-time
✓ High-confidence listing

From $131K/yr

Quick readStrong listing-quality and freshness signals

Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production

PythonAWSAzureCI/CD
S
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.4%
Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role A strong and reliable platform is essential to scaling Sentry for the future. Our Platform organization is responsible for everything that powers Sentry—from cloud infrastructure and streaming systems to storage, deployment, and security. We own the core services and technical foundations that enable every product and engineering team at Sentry to move fast and build with confidence. We're looking for a passionate and pragmatic Senior Staff Software Engineer to help lead this evolution. In this role, you’ll report directly to the VP of Engineering and collaborate with teams across the company to shape the future of Sentry’s platform. What You’ll Do Architect the future of Sentry by translating business needs and product strategy into clear, scalable technical blueprints. Partner with product and engineering leaders to align technical roadmaps with company goals. Lead cross-cutting initiatives across the Platform org—owning them end-to-end and driving meaningful outcomes. Promote engineering excellence by mentoring platform engineers, sharing best practices, and setting high standards for system design, scalability, and operational quality. Review major architectural proposals and help ensure consistency, maintainability, and long-term technical health across the company. You’ll Love This Job If You... Enjoy designing and building platforms that help teams move faster and scale safely. Thrive on solving complex, multi-dimensional problems across product, infrastructure, and organizational layers. Want to make architectural decisions that shape Sentry’s long-term success. Bring new ideas, tools, and frameworks t

AWSGCPKubernetesAI
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are looking for a highly skilled PSIRT Engineer to lead the vulnerability response program for Replit’s cloud-native AI platform. You will own the lifecycle of security vulnerabilities affecting our products and services—from intake to validation, remediation coordination, and public disclosure. This role requires strong technical ability to reproduce vulnerabilities , deep understanding of web/app/cloud exploit classes, and experience operating bug bounty and coordinated disclosure programs. You will work closely with Engineering, Cloud Security, SecOps, SRE, and IT teams to ensure vulnerabilities are fixed quickly and communicated responsibly. What You’ll Do Vulnerability Intake, Triage & Validation Manage intake from bug bounty platforms (HackerOne preferred), customer reports, automated scanners, pentest reports, and coordinated disclosure channels. Independently validate, reproduce, severity-score, and document findings. Identify duplicates and maintain a clean vulnerability records pipeline. Assess relevance and exploitability using OWASP, cloud misconfiguration patterns, and identity/authentication/authorization risks (Oauth, OIDC). Remediation Coordination & SLA Management Work with Engineering, SecOps, IT, SRE, and Cloud Security to confirm product impact and drive remediation. Provide detailed reproduction steps, proof-of-concepts, and technical analyses. Track SLAs, remediation progress, regression testing, and systemic improvements. Support SOC 2, ISO 27001, and pentest evidence needs as part of vulnerability lifecycle governance. Bug Bounty & Vulnerability Disclosure Program Management Design and evolve the bug bounty program, including scope, rules, and reward structures. Man

PythonGCPCI/CDAI
G
📍 United Kingdom; Remote, United States· Full-time· Remote
✓ Quality checkedCompany trend -97.9%

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a member of the Infrastructure Security Team within the Product Security Department , you will work with teams across GitLab to ensure that the components that comprise our public cloud infrastructure are built from the beginning with resiliency and set security expectations that our customers rely on to power their DevSecOps goals. As a Staff Security Engineer, you will serve as a technical lead across the topics the Infrastructure Security team owns, including our SaaS Platforms (e.g. GitLab Dedicated, Cells) and Self-Managed offerings. You will define the technical direction for how the team approac

PythonAWSAzureGCP
C
📍 United States· Remote
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary: As a Senior Software Development Engineer at Aetna, you will play a critical leadership role in the design, development, and continuous enhancement of enterprise-scale Provider Applications. You will drive technical solutions for complex business problems, ensure application stability, and lead cross-functional initiatives as a Project Owner. This role requires a balance of hands-on engineering expertise, technical leadership, and delivery ownership, including overseeing vendor/contractor teams, ensuring alignment with enterprise architecture, and delivering high-impact solutions that improve provider data systems and operational efficiency. Required Qualifications: 5&#43; years of hands-on application development experience with Python and Google Cloud Platform (GCP) 2&#43; years of experience leading or contributing to large-scale application development initiatives Preferred Qualifications: Experience working in Agile/SCRUM environments Proven experience in project/program management, including planning, execution tracking, and delivery management Strong organizational, leadership, and planning skills with the ability to manage multiple priorities Experience working with distributed teams and cross-functional stakeholders Prior exposure to

B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.1%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE This role owns Baseten's relationships and market intelligence across hyperscalers and strategic neoclouds, including NVIDIA cloud partners. This is a technical and commercial role in equal measure: you'll evaluate capacity from the GPU to the data center, negotiate cost and terms with suppliers, and stay close enough to the market to develop and defend a real point of view on where it's heading. Given current market conditions, Baseten needs a much stronger pulse on this part of the market so we can track pricing, stay close to the right relationships, and move fast the moment more capacity is needed. This is a senior, experienced hire who will also help pair with and develop 1-2 junior to mid-level teammates covering the same space. WHAT YOU'LL DO Build and maintain deep relationships across hyperscalers and strategic neoclouds (including NVIDIA cloud partners), working each organization from top to bottom rather than a single point of contact Maintain a consistent, "top of mind" presence with key accounts so Baseten is positioned to move quickly when capacity needs arise Evaluate capacity from the GPU to the data center — hardware generation, rack and node configuration, interconnect, power density, and cooling — so you know what a configuration will actually deliver, not just what the spec sheet claims Live in compute pricing daily: track rates by GPU generation, region, and contract term to keep Baseten inf

Machine LearningAIGoHR
M
📍 O Fallon, Missouri, United States
✓ Quality checkedCompany trend +212.5%

Our Purpose Mastercard powers economies and empowers people in 200&#43; countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview Join a team focused on transforming how Mastercard's payment systems are built, scaled, and operated. As a Senior Software Engineer, you will lead the design and development of cloud-ready applications, microservices, and distributed systems that support large-scale payment processing platforms while helping advance modernization, automation, and engineering excellence across the organization. In this role, you will contribute to software architecture decisions, drive technical design discussions, and partner with engineers to deliver scalable, resilient, and maintainable software solutions. You'll have the opportunity to solve complex technical challenges, mentor other engineers, and influence how software is designed, developed, tested, and supported across critical technology platforms. What You Will Do •Design software solutions and contribute to software architecture decisions that support scalability, maintainability, and operational excellence. •Translate complex product requirements into technical designs and implementation plans. •Lead development of modular, extensible, high-performance applications. •Design and implement comprehensive unit, functional, and integration testing strategies. •Analyze, optimize, and improve application performance, scal

JavaAIRecruitment
C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 ClickUp is looking for an experienced Engineering Manager to lead our fullstack team responsible for building and scaling our flagship products. As the leader of the team that owns the APIs and core experiences powering ClickUp, you will play a pivotal role in shaping the future of our platform. You will guide engineers working across the stack, from frontend experiences to backend infrastructure. Your focus will be on driving the development of new features, addressing performance and reliability challenges, and ensuring operational excellence as we continue to grow. This is an opportunity to make a significant impact on our core product while fostering a culture of technical excellence and collaboration. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our produ

Node.jsAWSMachine LearningAI
🔔

Get new lead cloud operations engineer jobs in United States by email

Daily job updates · Unsubscribe anytime