Jobiba hiring network

Lead Cloud Infrastructure Engineer Jobs

6,876 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead cloud infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale is seeking a Staff Software Engineer to lead the technical vision for our Infrastructure team. As a Staff Engineer, you will be responsible for the architectural evolution of our control plane and data plane, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. You will act as a force multiplier, setting the standards for Kubernetes-based cloud-native infrastructure while mentoring engineers and driving cross-functional alignment across the Ray open-source community and our proprietary product teams. Key Responsibilities Architectural Leadership: Define and drive the multi-year technical roadmap for services that orchestrate Ray clusters across diverse cloud and on-premises environments. Systemic Optimization: Lead the design and optimization of high-performance control plane components specifically tailored for large-scale, heterogeneous AI/ML workloads. Platform Reliability: Establish the organization-wide standards for the reliability, scalability, and observability of Anyscale-managed infrastructure. Strategic Integration: Direct the long-term strategy for accelerator integration (GPUs, TPUs) and container management to ens

pythonawsazure
View job →
O
1mo ago

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Principal Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. As a principal engineer, you will set technical direction and drive execution on high-impact infrastructure security programs, partnering across various orgs at OpenAI to deliver durable controls that raise the security bar at OpenAI scale. In this role, you will: Own end-to-end security outcomes for one or more critical infrastructure areas, including multi-quarter strategy, roadmap, and delivery. Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Lead cross-functional programs to deploy security enhancements and control changes across broad-scale infrastructure, balancing security guarantees with reliability and velocity. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets

awsazurekubernetes
View job →
R
Replit
📍 Foster City• Full-time
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. Replit is building the world’s most accessible AI coding agent. Replit Agent can be used by anybody to bring their ideas to life. Whether it’s an app for yourself, the next great startup idea, or a tool to make you more productive at work, Replit Agent can help build it. Replit builds complete apps better than anybody thanks to our full suite of services that handle app integrations, storage, hosting, analytics, and more. We don’t just build apps in development, we handle the full lifecycle into production and beyond. About the role: Help power the development of Replit Agent as a technical leader for the Replit Cloud organization. You will report to the Vice President of Engineering. The Replit Cloud team builds Replit’s first party cloud infrastructure so users can build, scale, and succeed entirely on Replit. They manage databases, application storage, app publishing and hosting, development/production environment splitting, custom domains, and more. By having a set of first party services that integrate seamlessly, you will power one of Replit’s key product differentiators. You will: Help lead major projects, either by taking new products from 0->1 or doubling down on our first party primitives to keep winning users. Work closely with designers and product managers, to quickly iterate on Replit Cloud to continually grow and improve the product. Identify the hardest technical and/or quality problems holding us back, and then build solutions. Mentor and develop new senior engineers to help grow the team. Ship product and build infrastructure as a true full stack builder using: TypeScript, React, CSS, Postgres, Go, and Terraform. Examples of what you could do: Leverage our unique cloud infrastructure to build diffe

typescriptreactai
View job →
A
Airbnb
📍 United States• Full-time• From $212K/yr
1mo ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Airbnb's Security Engineering organization protects a global community of millions of Hosts and guests. The Cloud & Data Security team is responsible for the security of the infrastructure and data platforms that Airbnb runs on - spanning cloud environments, data infrastructure, identity and access, and the paved roads that engineering teams build on every day. We partner closely with other Information Security teams, Cloud Infrastructure, Data Platform, and Enterprise teams to make secure the easiest path for engineers to take. The Difference You Will Make: As the Engineering Manager for Cloud & Data Security, you will lead a team of security engineers responsible for securing Airbnb's cloud infrastructure, data platforms, and the controls that govern how sensitive data is accessed and moved. You will set the team's technical direction, coach engineers through complex architectural work, and partner across the company to raise the bar on how Airbnb builds and operates its infrastructure. You will own the team's roadmap, its people, and its outcomes, building a durable, high-trust function that shifts security left through paved roads, automation, and deep partnership with the teams you protect. A Typical Day: Lead and grow a team of security engineers focused on cloud infrastructure security, data security, and identity and access controls. Set and drive the team's roadmap in alignment with organizational security priorities, balancing embedded partnership with high-leverage automation and paved-road investment. Partner with Infrastructure, Data Platform, and Application Se

awsgcpkubernetes
View job →
C
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the Role We're looking for an Engineering Manager to lead our Deployment Engineering team in EMEA. This isn't a typical management role — we need someone who leads from the front, gets their hands dirty, and drives impact. You'll manage a team of Forward Deployed Engineers who are on the front lines of deploying Cohere's North platform into customer environments. You should be ready to be a force to be reckoned with. Location: UK (can be remote but based in UK) What You'll Do Lead and mentor a team of Forward Deployed Engineers across EMEA Drive end-to-end deployment of North in private cloud and on-premises environments Take ownership of customer success from technical implementation through delivery Collaborate closely with Product, Engineering, and Sales to shape how we deliver AI to enterprises Mentor your team on cloud infrastructure, Kubernetes, and enterprise-grade deployments Optimize performance for OpenSearch, databases, and other K8s services Define scaling guidelines for GPU and CPU compute resources Build processes & technology that scales — we're growing fast. What We're Looking For 5+ years of experience

awsazuregcp
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! We're looking for an Engineering Manager to lead our Deployment Engineering team. This isn't a typical management role — we need someone who leads from the front, gets their hands dirty, and drives impact. You'll manage a team of Forward Deployed Engineers who are on the front lines of deploying Cohere's North platform into customer environments. You should be ready to be a force to be reckoned with. Location: North America (remote-first) What You'll Do Lead and mentor a team of Forward Deployed Engineers Drive end-to-end deployment of North in private cloud and on-premises environments Take ownership of customer success from technical implementation through delivery Collaborate closely with Product, Engineering, and Sales to shape how we deliver AI to enterprises Mentor your team on cloud infrastructure, Kubernetes, and enterprise-grade deployments Optimize performance for OpenSearch, databases, and other K8s services Define scaling guidelines for GPU and CPU compute resources Build processes & technology that scales — we're growing fast. What We're Looking For 5+ years of experience in software engineering with demonstrate

awsazuregcp
View job →
C
Cohere
📍 San Francisco• Full-time
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an

kubernetesgitai
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t

pythonsqlpostgresql
View job →
N
Nuro
📍 Mountain View• Full-time• From $132K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Team We're looking for a Fullstack Software Engineer to join the Fleet Platform & Operations Tooling team. You'll build the internal tools and platform systems that Nuro's operations, engineering, and partner teams rely on to manage, monitor, and maintain our autonomous vehicle fleet. This is hands-on product engineering work at the intersection of fleet operations and software infrastructure. You'll own features end-to-end — from building responsive frontends that surface real-time vehicle and fleet data, to designing backend services that integrate with onboard systems, cloud infrastructure, and partner APIs. The systems you build will directly impact how reliably and efficiently Nuro operates vehicles on public roads. You'll report to the Technical Lead Manager for On-Road Experience. What You'll Do Design, build, and ship full-stack applications using React, TypeScript, and modern web development practices. Develop

typescriptreactsql
View job →

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Product Security Data Platforms team is a newly established engineering team within Stripe Security. Our mission is to build the foundational infrastructure that provides our users with unprecedented visibility into the security posture of their Stripe integration. While Stripe is renowned for industry-leading payment protection, we are expanding our focus to provide a comprehensive security telemetry platform that helps businesses protect their entire digital ecosystem on Stripe. As a founding member of this team, you'll architect a large-scale customer-facing security data pipeline and presentation layer. Much like modern security observability platforms and data lakes that have transformed cloud infrastructure, we're building an API-first service that transforms massive streams of behavioral data into actionable security intelligence. This team operates at the intersection of high-throughput data engineering and cybersecurity, creating the systems that will allow the world’s most sophisticated companies to monitor, detect, and respond to threats in real time. What you’ll do As a Senior Software Engineer on this founding team, you'll lead the technical design and implementation of our core security data pipelines. You'll define how we capture security signals, process them at scale, and deliver them to our users through robust, developer-friendly interfaces. If you have security domain knowledge, you'll have opportunities to help shape

typescriptjavareact
View job →
R
Replit
📍 Foster City• Full-time
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. We are looking for a Security Operations Lead (SOC Lead) to build, mature, and operate our 24/7 detection and response capabilities across a modern cloud-native and AI-driven environment. This role leads the global SOC function—monitoring, SIEM ownership, detection engineering, alert triage, and operational readiness—while also evaluating and integrating emerging AI-based SOC products and autonomous response platforms . You will oversee monitoring across multi-cloud environments (GCP primary, AWS/Azure secondary), Kubernetes, SaaS services, endpoints, developer tools, and AI workloads . You’ll collaborate closely with Cloud Security, Compliance/GRC, SRE, Platform Engineering, IT/Endpoint teams, and AI Infrastructure to ensure our detection strategy scales and stays ahead of evolving threats. This is a hands-on leadership role perfect for someone who wants to shape the SOC of the future while solving complex challenges in a high-scale AI setting. What You’ll Do SOC Leadership & 24/7 Monitoring Lead, mentor, and scale a global SOC team responsible for 24/7 monitoring, alert intake, triage, correlation, and escalation. Build operational rigor: processes, runbooks, SLAs, metrics, and quality standards for high-scale environments. Cover monitoring across: Cloud infrastructure (GCP, AWS, Azure) Kubernetes/GKE/EKS/AKS clusters SaaS platforms (Google Workspace, GitHub, Slack, Okta, etc.) Endpoints (macOS, Linux, Windows) including EDR/XDR telemetry Developer platforms + CI/CD pipelines AI/ML systems and model-serving workflows AI-Based SOC Integration & Innovation Evaluate, adopt, and integrate AI-native SOC technologies for triaging, detection, and correlation Identify opportunities to automate triage, investigations,

pythonawsazure
View job →
RS
10 days ago

About the Role Redwood is scaling public cloud infrastructure and AI features across multiple product lines, and we need a FinOps Lead to bring rigor, visibility, and accountability to that spend. This is a senior individual-contributor role with the authority to drive cross-functional cost governance directly. This role owns the translation of raw cloud cost data and AI spent into the models, forecasts, and governance mechanisms that let engineering, product, and executive leadership make informed decisions. This is not a bill-monitoring role. You will build the cost attribution infrastructure that ties cloud/AI spend to specific product lines and, ultimately, to ROI involving architecture and engineering teams to make the right trade offs and decisions in line with the strategic roadmap. The role will report to the Senior Director, Product Engineering Operations, and work closely with Cloud Engineering, the CPO's org, and engineering leadership to create cost visibility and defensibility informing leadership on efficiency strategies. Your ability to be an effective communicator, collaborative team player, and analytical thinker will be keys to success in this role. Responsibilities Cost Visibility, Attribution & Optimization Own and continuously improve cost models that attribute AWS (and other public cloud) spend by product line, team, and environment Drive tagging governance and hygiene, define standards, audit compliance, and close attribution gaps that prevent accurate cost-per-product reporting Build shared cost allocation models to drive transparency into per-team cost drivers where there are prevalent savings plans, reserved instances and network charges Build toward feature-level cost attribution that connects infrastructure spend to product ROI, not just aggregate bill totals Create and manage optimization programs with achievable savings targets, including working across engineering and finance to rightsize, clean up and modernize infrastructur

sqlawskubernetes
View job →
A
Amplitude
📍 Remote• Full-time• $198K – $299K/yr
1mo ago

Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Staff Platform Engineer, you'll set technical direction for the platform across teams, lead our highest-complexity and highest-leverage initiatives, and shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll operate across team boundaries — partnering with product engineering, fellow Staff+ engineers, and engineering leadership to make Kubernetes and cloud infrastructure effortless across the entire engineering org. You'll build the self-service automation, shared standards, and scalable AWS and GCP infrastructure that let dozens of product teams ship faster, safer, and with less cognitive load — and you'll multiply the engineers around you while you do it. Key Responsibilities Set technical direction — shape platform and domain-level technical strategy that improves developer experience, reliability, security, and cost, and lead the high-complexity, cross-cutting initiatives that deliver it with measurable impact for the organization. Drive clarity through ambiguity. Take on the most loosely-defined problems, validate the critical assumptions early, and create alignment with stakeholders across teams so others can move quickly and confidently — driving cross-team decisions to a timely close and escalating when needed. Build the AI-augmented platform. Design org-wide tooling, guardrails, and policy-as-code that help every engineer get more out of AI-assisted development — infra primitives an LLM can safely reason about and PR against, automated review, and standards that hold as AI changes how code gets written. Own Infrastructure-as-Code standards for Kubernetes, AWS, and GCP using Terraform, Helm, Kustomize, and emerging tooling — setti

pythonawsazure
View job →

You’ll shape the future of a business‑critical platform as the technical lead across both product engineering and cloud infrastructure. You’ll modernize a mature .NET application running on AWS today, while steering its evolution toward a cloud‑native, React/Node.js, AI‑enabled architecture. If you enjoy owning architecture end‑to‑end, from backend and frontend through CI/CD, DevOps, and AWS infrastructure, this role gives you real influence at Staff Engineer level and the opportunity to set engineering standards that others follow. You’ll spend your time leading complex .NET and React features, designing scalable AWS infrastructure with Infrastructure as Code, and building automation that makes releases fast, safe, and repeatable. You’ll work on performance, reliability, and modernization in equal measure—fixing what’s slowing the platform down today and designing what it will look like in the next generation. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture and development of enterprise .NET services and APIs that power a business‑critical platform. Design and operate AWS infrastructure (using AWS CDK in TypeScript) to support secure, scalable, multi‑environment deployments. Build and optimize CI/CD pipelines (AWS CodePipeline, CodeBuild, Windows build agents) to make shipping .NET and React changes fast and reliable. Drive modernization initiatives across the stack, including clean architecture, refactoring legacy components, and reducing technical debt. Design and tune PostgreSQL and MSSQL database solutions for performance, scalability, and reliability. Mentor engineers and influence engineering practices across teams, raising the bar on cloud, DevOps, and software design. These are the essentials you’ll need to get an interview Significant experience (typically 8+ years) delivering and operating scalable enterprise software, owning both application code and cloud infrastructure. Deep hands‑on expertise with C

typescriptreactnode.js
View job →
M
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: Modal is seeking an experienced Forward Deployed Engineer (FDE) to partner with our sales team and drive technical sales success. As an FDE, you will be the technical voice in our sales process, working directly with Account Executives to help enterprise customers understand how Modal can transform their AI/ML infrastructure. You will: Partner with Account Executives to identify, qualify, and close strategic enterprise opportunities Lead technical discovery sessions with prospective customers to understand their current infrastructure, pain points, and requirements Design and present compelling technical solutions that demonstrate how Modal addresses customer needs Architect migration paths from existing cloud infrastructure (AWS, GCP, Azure) to Modal's serverless platform Conduct technical demos, experiments, and proof-of-concepts that showcase Modal's capabilitie

sqlawsazure
View job →
🔔

Get new lead cloud infrastructure engineer jobs by email

Daily job updates · Unsubscribe anytime