Jobs in United States

Lead Cloud Operations Engineer in United States

2,434 active opportunities · Updated October 2026

Explore current lead cloud operations engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

AWSLinuxRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads. As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI’s rapidly expanding compute environment. About the Role We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI’s third-party infrastructure partners, including major cloud service providers and strategic compute vendors. In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership. Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure. Key Responsibilities Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations. Develop operationa

AWSAzureGCPRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Order to Cash (OTC) team oversees the complete flow of commercial transactions from order intake and provisioning through billing, collections, and cash application — ensuring accuracy, compliance, and operational excellence in support of OpenAI’s mission to ensure artificial general intelligence benefits all of humanity. About the Role We are looking for a senior, hands-on operator to own Order Management and Billing execution across OpenAI’s cloud marketplace and partner ecosystem, including platforms such as AWS, GCP, Oracle Cloud, GovCloud, and future channels. This senior individual contributor role will translate partner requirements into scalable workflows and ensure launch readiness, accurate billing, and reliable daily execution. As a senior individual contributor within the Cloud Marketplaces team, you will own the end-to-end order-to-invoice lifecycle for your assigned portfolio. You will ensure that private offers, commercial terms, provisioning, pricing, usage, billing data, credits, settlements, and partner-specific reporting flow through our systems accurately, on time, and with audit-ready controls. You will implement and continuously improve the common cloud marketplace operating model, lead cross-functional execution for your assigned portfolio, and surface risks, requirements, and improvement opportunities. You will partner across Revenue Systems, Product, Engineering, GTM, Finance, Partner Operations, and external marketplace stakeholders. This role is critical to building the operational backbone for OpenAI’s expansion across cloud marketplaces and government-cloud channels. You will combine deep operational judgment with process and control execution, automation, clear communication, and hands-on problem solving to improve billing reliability, partner experience, customer outcomes, and financial integrity at scale. This role

AWSGCPRestAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking an experienced and proactive Security Engineer to help us build, maintain, and continuously improve the security posture of our rapidly growing ML infrastructure platform. As one of the first dedicated security hires at Baseten, you will work cross-functionally with engineering and operations teams to ensure we’re meeting the highest standards of confidentiality, integrity, and availability. You’ll have an opportunity to shape our security strategy and best practices from the ground up, influencing the way our platform handles sensitive data for both internal and external stakeholders. RESPONSIBILITIES Security architecture and design: Collaborate with engineering teams to design and implement secure systems and infrastructure, including cloud (AWS/GCP) environments and container orchestration platforms. Vulnerability management: Lead proactive vulnerability assessments, pen tests, and remediation efforts to ensure our products and infrastructure remain secure. Incident response: Develop and maintain incident response processes, including detection, analysis, containment, eradication, and post-incident reviews. Identity and access management (IAM): Oversee IAM strategies and tools to ensure the right people have the right level of access to our systems and data. Security compliance and audits: Work closely with operations to ensure compliance with relevant standards (e.g., SOC 2, ISO 27001) and

AWSGCPCI/CDMachine Learning
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. We are looking for a Security Operations Lead (SOC Lead) to build, mature, and operate our 24/7 detection and response capabilities across a modern cloud-native and AI-driven environment. This role leads the global SOC function—monitoring, SIEM ownership, detection engineering, alert triage, and operational readiness—while also evaluating and integrating emerging AI-based SOC products and autonomous response platforms . You will oversee monitoring across multi-cloud environments (GCP primary, AWS/Azure secondary), Kubernetes, SaaS services, endpoints, developer tools, and AI workloads . You’ll collaborate closely with Cloud Security, Compliance/GRC, SRE, Platform Engineering, IT/Endpoint teams, and AI Infrastructure to ensure our detection strategy scales and stays ahead of evolving threats. This is a hands-on leadership role perfect for someone who wants to shape the SOC of the future while solving complex challenges in a high-scale AI setting. What You’ll Do SOC Leadership & 24/7 Monitoring Lead, mentor, and scale a global SOC team responsible for 24/7 monitoring, alert intake, triage, correlation, and escalation. Build operational rigor: processes, runbooks, SLAs, metrics, and quality standards for high-scale environments. Cover monitoring across: Cloud infrastructure (GCP, AWS, Azure) Kubernetes/GKE/EKS/AKS clusters SaaS platforms (Google Workspace, GitHub, Slack, Okta, etc.) Endpoints (macOS, Linux, Windows) including EDR/XDR telemetry Developer platforms + CI/CD pipelines AI/ML systems and model-serving workflows AI-Based SOC Integration & Innovation Evaluate, adopt, and integrate AI-native SOC technologies for triaging, detection, and correlation Identify opportunities to automate triage, investigations,

PythonAWSAzureGCP
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're hiring a Compute Strategy and Operations lead to own how Modal plans for and acquires GPU and CPU capacity. You'll size our infrastructure needs ahead of demand, source supply across hyperscalers, neoclouds, and datacenter operators, and negotiate and close the contracts to secure it. The compute you secure directly determines what Modal can sell and build. In this role, you will: Own end-to-end procurement of GPU and CPU capacity across hyperscalers, neoclouds, and datacenter operators Build and maintain a strong pipeline of supplier relationships Evaluate supply options on price, availability, hardware specs, networking capabilities, and SLA terms Negotiate and close contracts: reserved capacity agreements, spot arrangements, MSAs, DPAs, and order forms Work closely with our engineering teams to translate technical requirements into procurement specs Track

H
📍 Louisville, United States
✓ Quality checkedCompany trend +310%

Become a part of our caring community Humana is seeking a self-driven and collaborative Lead Engineer to join our Interactive Voice Response (IVR) team. In this role, you will deliver innovative IVR solutions and develop robust omnichannel APIs for our enterprise platforms. You will have the opportunity to drive the success of a high-impact, customer-facing application within a Fortune 50 company, working closely with multiple teams throughout the software development lifecycle (SDLC). Lead Engineer –Omnichannel Humana is seeking a self-driven and collaborative Lead Engineer to join our Omnichannel team. In this role, you will design, develop, secure, and enhance enterprise APIs that support high-impact, member facing, applications across Humana's digital and voice channels. This role offers the opportunity to modernize and strengthen existing API capabilities while helping deliver resilient, scalable, and secure omnichannel solutions within a Fortune 50 organization. Key Responsibilities Design, develop, and maintain scalable Omnichannel APIs that support enterprise applications and customer-facing capabilities. Enhance the security, resiliency, performance, and reliability of existing APIs through modernization, improved architecture, observability, testing, and operational controls. Apply AI and AI-assisted engineering practices to accelerate development, improve quality, automate testing, enhance documentation, and identify opportunities for optimization. Partner with architecture, security, cloud, product, engineering, and operations teams to deliver secure, resilient, and enterprise-aligned API solutions. Collaborate with agile teams to plan, track, and deliver API enhancements, platform improvements, and cloud-based capabilities. Develop proofs of

Machine LearningAIRecruitment
B
📍 Berkeley, United States
✓ Quality checkedCompany trend +515.8%

Cloud Platform Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations is currently seeking a Cloud Platform Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking a skilled platform engineer to help build and operate the cloud platform services that host critical enterprise applications and software toolchains. In this role, the selected candidate will focus on the shared platform capabilities that enable teams to deploy, run, and maintain containerized and cloud-hosted solutions in a consistent and supportable manner. As both an individual contributor and technical leader, this position will help define and implement platform standards for Kubernetes, container hosting, deployment automation, configuration management, and operational support. This role is focused on platform reliability, repeatability, scalability, and service enablement, rather than custom application software development. Position Responsibilities: Design, implement, and maintain cloud platform services supporting Kubernetes, containers, ingress, storage integration, secrets management, and service connectivity Build and sustain reusable deployment patterns for Commercial-Off-The-Shelf (COTS), Open Source Software (OSS), and internally customized applications Develop and maintain automation for platform provisioning, upgrades, patching, and lifecycle support Manage cluster lifecycle activities including: Cluster upgrades Node management <

AWSAzureDockerKubernetes
B
📍 Berkeley, United States
✓ Quality checkedCompany trend +515.8%

Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations organization is currently seeking a Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking an experienced cloud infrastructure professional to help design, build, and sustain the foundational cloud environment supporting critical program needs. In this role, the selected candidate will help establish and operate secure, scalable, and resilient cloud infrastructure environments in Microsoft Azure to enable enterprise applications, software toolchains, and digital engineering workloads. As both an individual contributor and technical leader, this position will work across network, computer, storage, identity, security, and automation domains to deliver repeatable cloud infrastructure patterns and operational excellence. This role is focused on infrastructure operations, sustainment, automation, and reliability, rather than application software development. Position Responsibilities: Design, implement, and maintain Microsoft Azure-based infrastructure solutions including networking, compute, storage, identity integration, and supporting services Develop and maintain Infrastructure as Code (IaC) and configuration automation solutions using Terraform, Ansible, PowerShell, and Bash Implement cloud policies to enforce security, ensure regulatory compliance, and manage user access Build repeatable landing zones and cloud infrastructure patterns that support mul

AzureTerraformAnsibleSap
H
📍 Boston, Massachusetts, United States· Full-time
✓ High-confidence listing

From $119.6K/yr

Quick readStrong listing-quality and freshness signals

We take play seriously. We’re looking for curious adventurers ready to find their party, fueled by imagination and drive to build what’s never been built before. At Hasbro and Wizards of the Coast, you’ll collaborate with passionate teams to reimagine our iconic brands and create experiences that spark joy, connection, and community through the magic of play. This is your chance to shape legendary play that lasts a lifetime. The Senior Network Engineer leads Wizards of the Coast's enterprise network — datacenters, corporate offices, studios, and AWS cloud. Our stack runs on a Juniper/Mist campus fabric, a Palo Alto Networks security edge, and cloud-native AWS connectivity. You'll set technical direction, drive complex initiatives end to end, and mentor the broader team. Come help us build the future of network operations. What You'll Do: Own the architecture, build, and roadmap for our Juniper/Mist campus and branch infrastructure, and lead its evolution across sites and business units. Lead the Palo Alto Networks security stack, from policy architecture to secure-by-design standards across teams. Own end-to-end AWS cloud network implementation — VPC, Transit Gateway, Direct Connect/VPN, Route 53 — and hybrid connectivity, including BGP, OSPF, and SD-WAN traffic engineering at scale. Drive automation and AI adoption across network operations, from config deployment to AI-powered monitoring, observability, and root-cause analysis. Be the go-to for critical issues, and mentor less-experienced engineers through code review, troubleshooting, and skill-building. What You'll Bring: 10+ years in enterprise network engineering — routing, switching, wireless — with deep expertise in Juniper (EX/QFX/SRX, Mist) and/or Cisco (Catalyst, Nexus), plus hands-on work with intelligent ops tools like Mist/Marvis at scale. Expert-level BGP, OSPF, SD-WAN, and load balancing chops. You've built these solutions from scratch, not just maintained them! Extensive Palo Alto Net

AWSAIGoTerraform
G
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -97.9%

From $126.4K/yr

Quick readStrong listing-quality and freshness signals

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Lead People Systems Engineer, you'll drive how GitLab's People systems evolve through an AI-first transformation mindset, putting AI to work across everything we build. You won't just support our people systems, you'll build them: owning the code, pipelines, and integrations that keep GitLab's People automation ecosystem running reliably at scale across Workato, Google Cloud, Claude, Workday, and the tools that power our global workforce. This is a strong fit if you think like a software engineer first, write clean and maintainable code, and thrive at the intersection of people systems, AI, innovatio

PythonAWSGCPDocker
N
📍 Remote, United States· Remote
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data

PythonSQLAWSAzure
C
📍 Hartford, United States
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary The Senior Software Engineer / Technical Lead (AI & Automation) will provide technical leadership for the design, development, modernization, and support of critical applications supporting Prior Authorization Operations (PAOps) under PBM line of business. This role will be responsible for building and maintaining scalable, cloud-native solutions that enable intelligent workflow automation, AI-driven decisioning, document processing, and business process optimization. The ideal candidate is a hands-on technical leader with strong software engineering and cloud architecture expertise, coupled with practical experience implementing Generative AI, Agentic AI, and Large Language Model (LLM) solutions in production environments. This individual will collaborate closely with Data Engineering, Data Science, Product, and Business teams to deliver highly available, secure, and scalable applications while driving innovation through AI-powered solutions. Key areas of focus include: Application architecture, development, and production support Cloud-native engineering and platform modernization Microservices and distributed systems AI/GenAI, Agentic AI, and LLM-based solutions Event-driven and streaming architectures Engineering best practices, mentoring, and technical leadership R

PythonSQLGCPDocker
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Governance, Risk, and Compliance team helps ensure security and privacy are grounded in how our products and systems actually operate. Assurance Operations partners with Security, Engineering, Infrastructure, Product, Privacy, and Legal to make controls provable, risk decisions explicit, and audit readiness a result of well-designed systems. About the Role We are hiring a technical, product-minded GRC builder who can own consequential audits while improving the control and evidence systems behind them. You will build a reusable common control framework, use Codex to automate assurance work, validate changing system scope, and turn repeated audit friction into measurable improvements. We are looking for someone who questions inherited assumptions, solves novel problems creatively, works closely with engineers, and makes the next audit easier by improving the underlying system. You’ll be responsible for: Lead external, internal, customer, and certification audit work from scoping through evidence review, fieldwork, remediation, and closeout. Build a common control framework linking risk, control intent, implementation, owner, system, environment, evidence, and applicable frameworks. Validate actual scope and ownership instead of assuming last year's controls, product boundaries, or evidence remain accurate. Use Codex to build and test evidence checks, control mappings, request triage, owner workflows, monitoring, and remediation reporting. Partner with engineers on cloud architecture, identity, logging, data flows, software changes, vulnerabilities, and control effectiveness. Design maintainable, permission-aware tools that preserve source provenance, human review, and evidence integrity. Reduce repeated requests and operational burden for control owners through measurable workflow improvements. Define roadmaps, decision rights, milestones, success metrics, and clear cross-functional escalations. We’re looking for someone with: Direct ownership

SQLAWSRestAI
H
📍 Louisville, United States
✓ High-confidence listingCompany trend +310%
Quick readStrong listing-quality and freshness signals

Become a part of our caring community Job Description Summary The Lead Solutions Architect provides architecture leadership for CenterWell Home Health programs and platforms, shaping conceptual and reference architectures, governing solution designs, and aligning delivery teams to cloud and data strategies. The scope includes high-priority initiatives as well as interoperability and provider-data integrations that span CenterWell and Humana Insurance. As Lead Solution Architect, you'll be the senior individual contributor on a team with broad accountability across CenterWell's dispensing pharmacy portfolio — mail order, specialty, retail, and associated platforms. You'll own the architectural vision for complex, multi-system initiatives, shape how technology decisions get made, and act as a connective force between business strategy, engineering execution, and enterprise standards. You will operate within CenterWell IT – Cross-CenterWell Architecture. You will collaborate with product, engineering, EA Activation, security, data, and operations. You will engage governance forums to enable Integrated Health across CenterWell and Humana Insurance. Key Activities Quickly conduct structured knowledge transfer with existing architects and relevant stakeholders to capture critical in-flight designs and decisions. Review current initiatives and establish an architectural roadmap aligned with organizational priorities. Develop or refine reference architectures and design patterns for core platforms and solutions. Collaborate with governance and compliance teams to validate designs against enterprise standards and regulatory requirements. Define integration strategies and solution blueprints for key systems and data flows. Establish architecture review processes and decision forums to support de

AWSAzureAIRecruitment
🔔

Get new lead cloud operations engineer jobs in United States by email

Daily job updates · Unsubscribe anytime