About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op
Jobs in United States
Lead Cloud Infrastructure Engineer in United States
2,434 active opportunities · Updated October 2026
Showing
15 jobs
Explore current lead cloud infrastructure engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
From $131K/yr
Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 We're looking for a Staff Data Engineer to own the architecture and technical vision of our data platform. This is a high-leverage, high-autonomy role where you'll set the technical bar for the team, drive cross-functional alignment on data infrastructure strategy, and solve our hardest engineering problems. You'll operate across AWS serverless technologies, Snowflake, dbt, and Terraform, but your impact goes well beyond any single tool: you'll shape how we think about reliability, scalability, cost, and developer experience at the platform level. This role is for someone who doesn't just build great systems, but makes the engineers around them better. The Role: Own the technical architecture of ClickUp's data platform, making design decisions that balance scalability, cost, reliability, and velocity. Define and drive the technical roadmap for data infrastructure in partnership with leadership. Design systems at scale : build frameworks, abstractions, and patterns that other engineers use daily. Lead complex, cross-team technical initiatives spanning data engineering, analytics engineering, data science, and data analytics. Drive cost optimization across cloud infrastructure and compute, turning efficiency into a competitive advantage. Build and evolve our data pipelines using AWS serverless (Lambda, Fargate, Step Functions, Kinesis, S3, DynamoDB, Aurora), Snowflake, and dbt. Establish and champion engineering standards : observability, testing, CI/CD, code review, and documentation practices. Design and maintain infrastructure for AI/ML workloads , including LLM frameworks, feature pipelines, training
About the Team Security is at the foundation of OpenAI's mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect model weights, customer data, and critical systems across multiple cloud environments. The team partners across OpenAI, including Applied Engineering, Research, IT, Security, Infrastructure, and Engineering, to provide secure and scalable platforms for identity, access management, permissioning, orchestration, and safe AI research. About the Role We’re looking for an engineering leader to lead Identity Infrastructure Engineering, the team building the systems that govern and scale access across OpenAI’s research, engineering, and internal platforms. This role sits at the center of cloud infrastructure, identity, software engineering, and security-critical operations. You’ll lead engineers building control planes, policy systems, workload and agent authorization patterns, infrastructure-as-code, and operational foundations that help OpenAI move quickly while keeping access reliable, auditable, least-privileged, and safe under failure. The ideal candidate has led teams responsible for large-scale, mission-critical infrastructure. They can go deep into code and architecture when needed, while giving engineers and technical leads the clarity and ownership to do their best work. They set technical direction, grow strong teams, make durable architecture decisions, and turn ambiguous 0-to-1 problems into platforms OpenAI can trust and build on for years. In this role, you will: Build and lead a high-performing Identity Infrastructure team, going deep enough technically to set direction while empowering the team to own delivery. Define the strategy for identity platform as the policy plane for access across people, agents, workloads, services, clouds, and internal systems. Scale Acc
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a member of the Infrastructure Security Team within the Product Security Department , you will work with teams across GitLab to ensure that the components that comprise our public cloud infrastructure are built from the beginning with resiliency and set security expectations that our customers rely on to power their DevSecOps goals. As a Staff Security Engineer, you will serve as a technical lead across the topics the Infrastructure Security team owns, including our SaaS Platforms (e.g. GitLab Dedicated, Cells) and Self-Managed offerings. You will define the technical direction for how the team approac
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Engineering Manager, Cloud Efficiency Snowflake runs large scale cloud infrastructure to deliver its own service — production and internal deployments, Kubernetes fleets, CI/CD, etc. Our cloud spend is in billions of dollars per year. We are looking for an experienced Engineering Manager to lead the Cloud Efficiency engineering team. In this role, you will own the technical vision and execution for building a unified, self-serve cloud efficiency platform along with AI skills and agents that makes resource usage and spend attributable and governable while driving insights and optimization of our cloud spend. AS AN ENGINEERING MANAGER IN CLOUD EFFICIENCY, YOU WILL: Lead and grow our talented team of software engineers, fostering a culture of technical excellence, ownership, and continuous learning. Drive the roadmap for Cloud Efficiency — translating company-level spend objectives into engineering systems: authoritative cost data, resource ownership registry, attribution pipelines, cost and unit economics modeling, observability, governance policies, and optimization workflows — in partnership with Product, Engineering, Finance, and Data Science. Set technical strategy for backend systems, data pipelines, and APIs that measure, attribute and surface cost and usage insights at
$275K – $350K/yr
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We are seeking a strategically-minded, technology-focused, and customer-centric Engineering Manager to lead one of our Infrastructure teams here. You will lead a team responsible for building, operating, and scaling the cloud infrastructure and platform systems that underpin CLEAR’s services, ensuring reliability, performance, and security across our environments. A successful candidate brings strong experience in cloud infrastructure, distributed systems, and operational excellence, along with a solid foundation in software engineering. You are an effective communicator who can lead complex infrastructure initiatives from inception through delivery, and thrive in fast-paced environments. This role requires a focus on building resilient, scalable systems, driving automation, and leading and developing high-performing engineering teams. What you'll do: Hire, develop, and grow engineering talent through coaching, mentorship, performance management, and career development planning Set clear goals and expectations, provide regular feedback, and foster accountability across the team Own and execute the roadmap for cloud infrastructure and platform engineering, and reliability initiatives Design, build, and operate a scalable, secure, and highly available cloud platform infrastructure Drive automation across infrastructure provisioning, deployment, and operations to improve efficiency and reduce manual overhead Establish and enforce best practices for system reliability, observability, incident response, and disaster recovery Partner with eng
Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations organization is currently seeking a Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking an experienced cloud infrastructure professional to help design, build, and sustain the foundational cloud environment supporting critical program needs. In this role, the selected candidate will help establish and operate secure, scalable, and resilient cloud infrastructure environments in Microsoft Azure to enable enterprise applications, software toolchains, and digital engineering workloads. As both an individual contributor and technical leader, this position will work across network, computer, storage, identity, security, and automation domains to deliver repeatable cloud infrastructure patterns and operational excellence. This role is focused on infrastructure operations, sustainment, automation, and reliability, rather than application software development. Position Responsibilities: Design, implement, and maintain Microsoft Azure-based infrastructure solutions including networking, compute, storage, identity integration, and supporting services Develop and maintain Infrastructure as Code (IaC) and configuration automation solutions using Terraform, Ansible, PowerShell, and Bash Implement cloud policies to enforce security, ensure regulatory compliance, and manage user access Build repeatable landing zones and cloud infrastructure patterns that support mul
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary We are seeking an accomplished Principal Cloud Storage Engineer to lead the design, engineering, and evolution of our private cloud storage platforms. This role will focus on large-scale storage architecture, data protection, cyber recovery, and resiliency technologies across complex enterprise environments. The ideal candidate will combine deep technical expertise in storage systems with strong leadership, architectural vision, and the ability to influence technical direction across the organization. Key Responsibilities Architect and engineer enterprise storage platforms that ensure data integrity, availability, security, and disaster recovery readiness Design and implement end-to-end storage solutions, including Software Defined Storage, SAN, NAS, and object storage across private cloud and data center environments Drive strategic technology decisions by evaluating emerging products, tools, and standards supporting storage, data protection, cloud, and compute platforms Lead infrastructure initiatives involving storage modernization, data protection, cyber recovery, data migration, and resilience engineering Develop and execute enterprise strategies for backup, recovery, cyber vaulting, and business continuity Create and maintain comprehensive documentation of storage architectures, configurations, policies, and operation
About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability Team Focus Areas Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely. Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads. Databases: Building high performance, distributed database systems that power all of OpenAI's product stack. In this
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Principal Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. As a principal engineer, you will set technical direction and drive execution on high-impact infrastructure security programs, partnering across various orgs at OpenAI to deliver durable controls that raise the security bar at OpenAI scale. In this role, you will: Own end-to-end security outcomes for one or more critical infrastructure areas, including multi-quarter strategy, roadmap, and delivery. Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Lead cross-functional programs to deploy security enhancements and control changes across broad-scale infrastructure, balancing security guarantees with reliability and velocity. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. Replit is building the world’s most accessible AI coding agent. Replit Agent can be used by anybody to bring their ideas to life. Whether it’s an app for yourself, the next great startup idea, or a tool to make you more productive at work, Replit Agent can help build it. Replit builds complete apps better than anybody thanks to our full suite of services that handle app integrations, storage, hosting, analytics, and more. We don’t just build apps in development, we handle the full lifecycle into production and beyond. About the role: Help power the development of Replit Agent as a technical leader for the Replit Cloud organization. You will report to the Vice President of Engineering. The Replit Cloud team builds Replit’s first party cloud infrastructure so users can build, scale, and succeed entirely on Replit. They manage databases, application storage, app publishing and hosting, development/production environment splitting, custom domains, and more. By having a set of first party services that integrate seamlessly, you will power one of Replit’s key product differentiators. You will: Help lead major projects, either by taking new products from 0->1 or doubling down on our first party primitives to keep winning users. Work closely with designers and product managers, to quickly iterate on Replit Cloud to continually grow and improve the product. Identify the hardest technical and/or quality problems holding us back, and then build solutions. Mentor and develop new senior engineers to help grow the team. Ship product and build infrastructure as a true full stack builder using: TypeScript, React, CSS, Postgres, Go, and Terraform. Examples of what you could do: Leverage our unique cloud infrastructure to build diffe
From $212K/yr
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Airbnb's Security Engineering organization protects a global community of millions of Hosts and guests. The Cloud & Data Security team is responsible for the security of the infrastructure and data platforms that Airbnb runs on - spanning cloud environments, data infrastructure, identity and access, and the paved roads that engineering teams build on every day. We partner closely with other Information Security teams, Cloud Infrastructure, Data Platform, and Enterprise teams to make secure the easiest path for engineers to take. The Difference You Will Make: As the Engineering Manager for Cloud & Data Security, you will lead a team of security engineers responsible for securing Airbnb's cloud infrastructure, data platforms, and the controls that govern how sensitive data is accessed and moved. You will set the team's technical direction, coach engineers through complex architectural work, and partner across the company to raise the bar on how Airbnb builds and operates its infrastructure. You will own the team's roadmap, its people, and its outcomes, building a durable, high-trust function that shifts security left through paved roads, automation, and deep partnership with the teams you protect. A Typical Day: Lead and grow a team of security engineers focused on cloud infrastructure security, data security, and identity and access controls. Set and drive the team's roadmap in alignment with organizational security priorities, balancing embedded partnership with high-leverage automation and paved-road investment. Partner with Infrastructure, Data Platform, and Application Se
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an
Other cities to consider
More places hiring for this role
Get new lead cloud infrastructure engineer jobs in United States by email
Daily job updates · Unsubscribe anytime