Jobiba hiring network

Lead Cloud Infrastructure Engineer Jobs

6,876 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead cloud infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. As a Principal Engineer on the Data Platform team, you will be a technical anchor — setting direction, solving the hardest engineering problems, and elevating the entire team. You will work at the intersection of systems design, long-term architecture, and real-world engineering execution. In this specific role, you will lead the evolution of our declarative data pipelines, focusing on streaming ingestion and transformations. What You'll Do Drive the technical strategy and architecture for key Data Platform initiatives, focusing heavily on real-time data movement and incremental processing. Lead design reviews and set engineering standards across teams. Identify and resolve high-impact technical challenges and systemic bottlenecks to reduce latency and improve the efficiency of dynamic tables at cloud scale. Mentor and grow senior engineers; raise the engineering bar. Partner closely with product and infrastructure leadership on roadmap and direction. Continuous hands-on technical deliverables in the most critical areas. What We're Looking For 14+ years of software engineering experience with deep expertise in distributed systems. Demonstrated track record of defining and delivering platform-scale technical initiatives. Expert-level knowledge of streaming and/or batch data

aigorust
View job →
W
12 days ago

WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: Responsible for leading the Cloud Automation Engineering function. Primary focus will be leading a team of other engineers in designing and implementing automation solutions to improve customer experience and increase productivity in our cloud estates. Responsible for maintaining and delivering automation solutions through infrastructure as code, ensuring security best practice, evangelising automation practice and tools, and supporting customer needs, both internal and external. What you'll be doing: Identify opportunities for improvement and automation of operations Design, build, test and implement use cases to drive automation adoption and improve operational efficiency Work closely with the IT Operations team to develop automated incident detection and response mechanisms. Implement proactive monitoring and alerting systems to quickly respond to and resolve critical issues, minimizing downtime and service disruptions Responsible for driving CSI initiatives to improve operations (processes/tools) working with various stakeholders Responsible for providing feedback at various leve

pythonawsazure
View job →
R
Roblox
📍 San Mateo• Full-time• From $345K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's engineering team is expanding rapidly, and we're looking for a seasoned engineering leader to help us scale our AI-powered creation platform and team. As a Senior Engineering Manager within the Creator Group, you will lead the Studio Assistant team — driving the strategy and execution of our agentic AI system that powers creation for millions of Roblox creators. Your team owns everything end-end from the user experience in Studio, to the infrastructure underneath backend by a single cloud-native architecture: one harness, one tool set, one eval framework. Your work will empower millions of creators to plan, generate, and ship content faster than ever before. At Roblox we move fast and ship code to production daily; your leadership will increase productivity by removing obstacles and keeping processes lean. You will identify and mitigate risk and make sure the technology outpaces our growth. Roblox teams are inclusive, helpful, and have a strong sense of ownership over the things they build. If you have a desire to grow and learn, you will fit right in with our highly-skilled and ever-expanding engineering team. You Are: Battle ready: You have experience defining architecture for AI

awsgitai
View job →
R
Roblox
📍 San Mateo• Full-time• From $295.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Join Roblox as an Engineering Manager of Application Security and lead a team responsible for improving the security of our products, services, and development ecosystem. In this role, you will drive security across the software lifecycle, partnering with engineering teams to identify risks, improve secure development practices, and build scalable solutions that protect Roblox at scale. You will balance hands-on security work with longer-term investments in automation, tooling, and developer enablement. You will work closely with engineering, infrastructure, and security teams to reduce risk while enabling teams to move quickly and safely. This role reports to the Senior Manager of Application Security and is based in San Mateo with a hybrid schedule. You Have: 8+ years of experience in Information Security 2+ years of experience managing engineers Strong background in Application Security or Product Security Experience driving security programs across the software development lifecycle Solid understanding of common vulnerabilities (e.g., OWASP Top 10) and secure coding practices Experience working closely with engineering teams in modern environments (cloud, microservices, CI/CD) Prov

awsci/cdgit
View job →
O
1mo ago

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. Technical Program Managers at OpenAI play a key leadership role in scaling these efforts, partnering deeply with product, engineering, design, and go-to-market teams to bring ambitious ideas to life and ensure clarity and discipline in execution. About the Role We are hiring a Technical Program Manager to support OpenAI's critical AI deployments across strategic cloud partners. This role is designed for a candidate who can operate as an end-to-end owner across internal engineering teams and external partner organizations. This role will drive the technical strategy and execution required to bring OpenAI models and platform capabilities into partner environments responsibly and at scale. The work spans engineering deliverables, shared roadmaps, model launch pipelines, technical integration, launch readiness, and post-launch follow-through. You will work closely with senior leaders across OpenAI engineering, infrastructure, product, safety, security, legal, finance, and go-to-market, as well as technical counterparts at our partners. The job is to turn broad partnership commitments into concrete execution plans, align both sides on what must land, and build repeatable mechanisms for launching OpenAI capabilities on third-party platforms. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end execution for major cloud partner programs spanning model deployment, product integration, operational readiness, launch follow-through, and partner-platform adoption. Own integrated technical roadma

awsrestai
View job →

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. Technical Program Managers at OpenAI play a key leadership role in scaling these efforts, partnering deeply with product, engineering, design, and go-to-market teams to bring ambitious ideas to life and ensure clarity and discipline in execution. About the Role We are hiring a Technical Program Manager to support OpenAI's critical AI deployments across strategic cloud partners. This role is designed for a candidate who can operate as an end-to-end owner across internal engineering teams and external partner organizations. This role will drive the technical strategy and execution required to bring OpenAI models and platform capabilities into partner environments responsibly and at scale. The work spans engineering deliverables, shared roadmaps, model launch pipelines, technical integration, launch readiness, and post-launch follow-through. You will work closely with senior leaders across OpenAI engineering, infrastructure, product, safety, security, legal, finance, and go-to-market, as well as technical counterparts at our partners. The job is to turn broad partnership commitments into concrete execution plans, align both sides on what must land, and build repeatable mechanisms for launching OpenAI capabilities on third-party platforms. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end execution for major cloud partner programs spanning model deployment, product integration, operational readiness, launch follow-through, and partner-platform adoption. Own integrated technical roadma

awsrestai
View job →

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Security Operations (SecOps) team protects Robinhood and its customers by detecting, investigating, and responding to security threats across our production systems, endpoints, and cloud environments. We combine detection engineering, incident response, and threat intelligence to identify emerging risks, strengthen our defenses, and reduce the impact of attacks before they reach our customers. As we continue to evolve our security platform, we're embracing AI and automation to help our teams move faster, uncover threats more effectively, and scale our defenses against increasingly sophisticated adversaries. As the Manager of Response, Automation, Intelligence and Detection Engineering (RAID) within SecOps, you will lead teams across North America, shaping the strategy, execution, and long term evolution of our defensive security capabilities. You'll be responsible for building high performing teams, maturing our detection and incident response programs, and ensuring we stay ahead of an increasingly sophisticated threat landscape. Working closely with Security, Engineering, Infrastructure, and Trust & Safety, you'll translate emerging threats into scalable defenses wh

vueawsai
View job →

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: Most of the value of owning a model shows up at serving time. We're building a platform that covers the whole life of an LLM -- train it, deploy it, observe it -- and inference is where teams feel the difference every day. We already run elastic inference, sandboxes, distributed volumes, and multi-node training, and we control the infrastructure underneath, so the serving stack is ours to shape rather than something we resell. You will do hands-on inference research at Modal, working with the research lead to pick high-impact bets and owning them end to end. The bets that matter most are the ones that move cost per token and tail latency on the workloads our customers actually run. What you'll do: Own end-to-end inference research bets: speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, autoscaling for spik

restaigo
View job →
E
Everpure
📍 Bengaluru• Full-time
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... We’re in an unbelievably exciting area of tech, fundamentally reshaping the cloud-native, modern virtualization, and AI infrastructure landscape. Here, you’ll lead with innovative thinking, grow alongside us, and work with some of the smartest minds in the industry. We’re looking for engineers passionate about system testing, distributed systems, Kubernetes, storage, modern virtualization, AI workloads, and automation. You’ll work on complex, real-world scenarios involving HA, resiliency, disaster recovery, scalability, and failure testing across large-scale Kubernetes environments. As enterprises modernize their infrastructure, containers, virtual machines, and AI/ML workloads are increasingly converging on Kubernetes. From traditional enterprise applications and VMs to GPU-accelerated AI training, inference, and data-intensive workloads, Kubernetes is rapidly becoming the common platform for running the next generation of applications. This role gives you the opportunity to test and influence how Portworx delivers enterprise-grade storage, data protection, availability, and resiliency across these workloads—including containerized applications, KubeVirt/OpenShift Virtualization VMs, and demanding AI/ML workloads running on Kubernetes. WHAT YOU WILL DO: Own system-level quality for Portworx Enterprise across Kubernetes, storage, and modern virtualization environments. Develop comprehens

pythonawsazure
View job →
E
Everpure
📍 Bengaluru• Full-time
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... We’re in an unbelievably exciting area of tech, fundamentally reshaping the cloud-native, modern virtualization, and AI infrastructure landscape. Here, you’ll lead with innovative thinking, grow alongside us, and work with some of the smartest minds in the industry. We’re looking for engineers passionate about system testing, distributed systems, Kubernetes, storage, modern virtualization, AI workloads, and automation. You’ll work on complex, real-world scenarios involving HA, resiliency, disaster recovery, scalability, and failure testing across large-scale Kubernetes environments. As enterprises modernize their infrastructure, containers, virtual machines, and AI/ML workloads are increasingly converging on Kubernetes. From traditional enterprise applications and VMs to GPU-accelerated AI training, inference, and data-intensive workloads, Kubernetes is rapidly becoming the common platform for running the next generation of applications. This role gives you the opportunity to test and influence how Portworx delivers enterprise-grade storage, data protection, availability, and resiliency across these workloads—including containerized applications, KubeVirt/OpenShift Virtualization VMs, and demanding AI/ML workloads running on Kubernetes. WHAT YOU WILL DO: Own system-level quality for Portworx Enterprise across Kubernetes, storage, and modern virtualization environments. Develop comprehens

pythonawsazure
View job →
E
Everpure
📍 Bengaluru• Full-time
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Exa team and lead the charge in redefining enterprise storage by unifying block, file, and object protocols across hybrid-cloud environments. You will combine deep technical expertise in distributed systems with hands-on people leadership to guide architectural decisions and mentor high-impact engineers. This is a unique opportunity to build new engineering teams from the ground up and drive industry-leading innovation alongside Product and Architecture partners. Your work will directly impact how customers consume, scale, and operate mission-critical storage infrastructure. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management, validation, and support teams to align technical roadmaps, establish architectural standards, and drive ent

pythonjavaaws
View job →
E
Everpure
📍 Bengaluru• Full-time
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Exa team and lead the charge in redefining enterprise storage by unifying block, file, and object protocols across hybrid-cloud environments. You will combine deep technical expertise in distributed systems with hands-on people leadership to guide architectural decisions and mentor high-impact engineers. This is a unique opportunity to build new engineering teams from the ground up and drive industry-leading innovation alongside Product and Architecture partners. Your work will directly impact how customers consume, scale, and operate mission-critical storage infrastructure. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management, validation, and support teams to align technical roadmaps, establish architectural standards, and drive ent

pythonjavaaws
View job →
E
Everpure
📍 Bengaluru• Full-time
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Exa team and lead the charge in redefining enterprise storage by unifying block, file, and object protocols across hybrid-cloud environments. You will combine deep technical expertise in distributed systems with hands-on people leadership to guide architectural decisions and mentor high-impact engineers. This is a unique opportunity to build new engineering teams from the ground up and drive industry-leading innovation alongside Product and Architecture partners. Your work will directly impact how customers consume, scale, and operate mission-critical storage infrastructure. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management, validation, and support teams to align technical roadmaps, establish architectural standards, and drive ent

pythonjavaaws
View job →
E
Everpure
📍 Bengaluru• Full-time
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Exa team and lead the charge in redefining enterprise storage by unifying block, file, and object protocols across hybrid-cloud environments. You will combine deep technical expertise in distributed systems with hands-on people leadership to guide architectural decisions and mentor high-impact engineers. This is a unique opportunity to build new engineering teams from the ground up and drive industry-leading innovation alongside Product and Architecture partners. Your work will directly impact how customers consume, scale, and operate mission-critical storage infrastructure. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management, validation, and support teams to align technical roadmaps, establish architectural standards, and drive enterprise

pythonjavaaws
View job →
E
Everpure
📍 Bengaluru• Full-time
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Exa team and lead the charge in redefining enterprise storage by unifying block, file, and object protocols across hybrid-cloud environments. You will combine deep technical expertise in distributed systems with hands-on people leadership to guide architectural decisions and mentor high-impact engineers. This is a unique opportunity to build new engineering teams from the ground up and drive industry-leading innovation alongside Product and Architecture partners. Your work will directly impact how customers consume, scale, and operate mission-critical storage infrastructure. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management, validation, and support teams to align technical roadmaps, establish architectural standards, and drive ent

pythonjavaaws
View job →
🔔

Get new lead cloud infrastructure engineer jobs by email

Daily job updates · Unsubscribe anytime