At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. About the team The Apps & Experiences Platform team powers the systems and services behind all of our user-facing applications, including Snowsight, Snowflake Intelligence, and new mobile experiences. Our mission is to build innovative backend services, developer tooling, platform infrastructure, and AI-powered capabilities that enable exceptional product experiences at scale. As part of our team, you’ll work across feature development, platform engineering, infrastructure, and internal tooling to support both end users and developers. We care deeply about building systems that are reliable, scalable, maintainable, and performant. Snowflake is a high-growth AI Data Cloud company, and we’re looking for exceptional engineers to help us scale the next generation of our platform. A key part of this is our work on our internal AI developer agent, which is designed to democratize end-to-end app development across all engineering teams by translating product specs and designs into fully functional, production-ready features. AS A SENIOR SOFTWARE ENGINEER FOR THE APPS & EXPERIENCES PLATFORM TEAM, YOU WILL: Design, build, and operate scalable backend services and platform infrastructure that power Snowflake’s user-facing applications. Contribute across the full development l
Jobiba hiring network
Senior Cloud Infrastructure Engineer Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior cloud infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! About the Team The NR Lens team builds New Relic's Federated Data Platform — a distributed SQL query engine that lets customers connect external data sources (Snowflake, PostgreSQL, Google Sheets, AWS CloudWatch) and query them directly from New Relic. You'll work on the distributed query execution layer, connector architecture, and API gateway that powers cross-source JOINs and unified analytics across customer data stores. What You'll Do Design, build, and maintain cloud-native Java microservices in the NR Lens query path: SQL Gateway, Query Gateway, and data source connector plugins Develop and harden JDBC connector integrations — including connection lifecycle management, credential handling, query pushdown optimization, and security validation Improve query reliability and performance across a multi-tenant distributed SQL deployment serving 50+ customers Build and operate services on AWS (EKS, IAM/STS, S3) with infrastructure-as-code practices Implement
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's engineering team is expanding rapidly, and we're looking for a seasoned engineering leader to help us scale our AI-powered creation platform and team. As a Senior Engineering Manager within the Creator Group, you will lead the Studio Assistant team — driving the strategy and execution of our agentic AI system that powers creation for millions of Roblox creators. Your team owns everything end-end from the user experience in Studio, to the infrastructure underneath backend by a single cloud-native architecture: one harness, one tool set, one eval framework. Your work will empower millions of creators to plan, generate, and ship content faster than ever before. At Roblox we move fast and ship code to production daily; your leadership will increase productivity by removing obstacles and keeping processes lean. You will identify and mitigate risk and make sure the technology outpaces our growth. Roblox teams are inclusive, helpful, and have a strong sense of ownership over the things they build. If you have a desire to grow and learn, you will fit right in with our highly-skilled and ever-expanding engineering team. You Are: Battle ready: You have experience defining architecture for AI
Cloud Platform Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations is currently seeking a Cloud Platform Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking a skilled platform engineer to help build and operate the cloud platform services that host critical enterprise applications and software toolchains. In this role, the selected candidate will focus on the shared platform capabilities that enable teams to deploy, run, and maintain containerized and cloud-hosted solutions in a consistent and supportable manner. As both an individual contributor and technical leader, this position will help define and implement platform standards for Kubernetes, container hosting, deployment automation, configuration management, and operational support. This role is focused on platform reliability, repeatability, scalability, and service enablement, rather than custom application software development. Position Responsibilities: Design, implement, and maintain cloud platform services supporting Kubernetes, containers, ingress, storage integration, secrets management, and service connectivity Build and sustain reusable deployment patterns for Commercial-Off-The-Shelf (COTS), Open Source Software (OSS), and internally customized applications Develop and maintain automation for platform provisioning, upgrades, patching, and lifecycle support Manage cluster lifecycle activities including: Cluster upgrades Node management <
The Team This role can sit in our NYC HQ on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas platform. As a senior SRE, you will be expected to be able to design & build complex systems, operate with autonomy and act as owner for everything you do. The SRE Atlas team works alongside the various Atlas software engineering teams to provide expertise about running systems at scale, build new tooling and automation and perform essential maintenance of the Atlas fleet. This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions that have the ability to impact our customer’s most crucial workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This role requires engineers to have a customer-first mindset to ensure that everything we do results in a stronger product and a better experience for all Atlas customers. The ideal candidate should Have 5+ years of experience running critical systems at scale Value efficiency in processes and operations, and display a preference for automation over manual processes (“allergic to ops work”) Be familiar with a major cloud provider (AWS, Azure, or GCP) and possess the ability to build and operate systems in a multi-cloud environment A strong understanding of how to run a large scale Linux environment, including low level fundamentals Firm grasp of at least one modern programming language, beyond basic scripting (Go, Ruby, Python) Solid understanding of web and network protocols and standards (HTTP, TLS, DNS, etc) Special Requirements: Be a US Citizen Expectations Participate in the development of a reliable and resilient multi-cloud platform that hosts business critic
Artefact is a next-generation data and AI consulting firm dedicated to accelerating the adoption of data and AI to create measurable business impact across the full enterprise value chain. We sit at the intersection of consulting, data science, AI technologies, data engineering, and digital transformation. We do not just advise — we build, implement, and deliver results our clients can measure. Our teams bring together consultants, data scientists, data engineers, AI engineers, analysts, and digital experts to solve complex business challenges with pragmatic, production-ready solutions. As Artefact continues to grow globally, we are building a team of entrepreneurial data and AI talent who can help clients move beyond experimentation and into scalable, governed, value-generating AI adoption. The Role Artefact is looking for a Senior AI & Agentic Engineer: a full-stack engineer who takes AI features from idea to production. You will design and build the interfaces, services, and agentic systems at the heart of our client work, such as conversational applications over enterprise data, multi-step agents that automate business workflows, and the retrieval and data pipelines that support them. You will own your components end to end: the React front end, the Python or Node service behind it, the RAG pipeline feeding it, and the evaluations proving it works. This role combines breadth and depth: the ability to take a feature from front end to cloud deployment, together with strong expertise in at least one major AI platform — Google (Gemini), Anthropic (Claude), or OpenAI. You will work with direct client exposure, and you will support the professional development of the junior engineers around you. What You'll Do Build Full-Stack AI Applications, End to End You will build AI products across the entire stack, from interface to infrastructure. Develop user-facing interfaces in TypeScript/React and the backend services and APIs behind them in Python or Node. Implement a
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As Manager / Senior Manager, Cell Infrastructure , you'll help shape the foundation that lets GitLab run reliably across multiple cloud providers in the agentic era. As more enterprise customers adopt AI-driven development workflows, GitLab needs a cell-based infrastructure that can scale horizontally, support strong reliability, and operate with clear cost efficiency. In this role, you'll report to the VP Engineering, Platform Scale & Architecture and lead a 0 to 1 effort to build the control layer that determines how GitLab cells are provisioned, placed, and operated across clouds. This is a high-im
NVIDIA DGX Cloud is an AI Factory designed to power the next generation of AI and industrial-scale breakthroughs. As a Principal Engineer for Security Architecture, within our Security Engineering organization, you will own a core security domain of the AI factory: the architecture, the paved road that delivers it, and much of the code underneath. You will hold the security design bar across DGX Cloud from inside the teams doing the building, and this is a founding seat on a new team. Security Engineering is a new organization at DGX Cloud, accountable for the security outcome of the platform, and Security Architecture is the function inside it that holds the design bar. Security here is fleet horizontal and stack vertical, so your work will cross every DGX Cloud engineering organization: you will embed with the teams building GPU clusters, control planes, and services, join their designs as a participant rather than an approver, and leave behind systems in which an entire class of risk is no longer possible. There is no architecture review board here and no approval queue. You are a senior IC with deep security domain knowledge, and the security bar holds because you helped set it and then helped ship it. What You Will Be Doing: Own a Security Domain End to End: Take architectural ownership of a core domain of DGX Cloud security, from the design through the system running in production. That could be tenant and GPU workload isolation, workload identity, infrastructure and network, supply-chain provenance, hardened baselines and patching, or deploy-time policy and admission control. Embed with the Teams Building It: Join the design early, write the code, and help land it. The posture is not "you did this wrong." It is "here are the considerations we need to meet, I will help, let's go to work." Build Paved Roads, Not
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are seeking an experienced and technically influential Senior Software Development Engineer to join our Cloud Tooling and Pipelines team. This pivotal team is responsible for the design, development, and maintenance of our core Continuous Delivery (CD) platform (leveraging Spinnaker and custom tooling), Infrastructure as Code (IaC) execution engines (primarily Terraform), and a suite of supporting microservices. These systems are critical for enabling and managing our extensive resource footprint across AWS ECS and EKS . As a Senior Software Development Engineer, you will be a key contributor, driving the implementation of scalable, reliable, and secure software solutions that automate infrastructure provisioning and application deployments. Your deep expertise in software engineering principles and cloud-native development will be essential in building and enhancing our critical tooling for infrastructure provisioning, vulnerability management, and IaC deployments. You will also play a vital role in mentoring other engineers and influencing the team's technical roadmap. If you have a strong passion for building robust software systems that empower operational efficiency at scale, we encourage you to apply. Key Responsibilities Design and Develop Core Platform Components: Lead the design and development of scalable and reliable microservices and tools that form the backbone of Okta's Continuous Delivery (CD) platform (including components for Spinnaker,
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We’re at the forefront of the data revolution, committed to building the world’s greatest data and applications platform. Snowflake is striving to become the cloud data platform of choice for everyone who needs affordable, reliable, scalable and elastic data management. Armed with great technical talent, our expanding product offering is designed for the cloud and automatically handles infrastructure, optimization, availability, data protection, and more so our users can focus on using their data, not managing it. And we’re doing this across all the major cloud providers, breaking new ground in what it means to be multi-cloud and cloud agnostic. We are looking to hire an experienced Senior Technical Program Manager who can be part of a world-class team that builds Snowflake’s product and cloud platform. This role is a unique opportunity to make a significant impact by building a flexible, large scale, high-performance, and resilient platform for Snowlake features and services. Be part of the vision to create the best data cloud platform. Our ‘get it done’ culture allows everyone at Snowflake to have an equal opportunity to innovate on new ideas, create work with a lasting impact, and excel in a culture of collaboration. AS A SENIOR TECHNICAL PROGRAM MANAGER, YOU WILL: Be re
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . What We're Looking For: 2–4 years of program or project management experience (or equivalent), ideally touching infrastructure, platform, or cloud environments. Working knowledge of cloud platforms (AWS or similar) — compute, storage, and basic usage/cost concepts. Interest in, and some exposure to, capacity management and infrastructure efficiency (right-sizing, utilization, waste reduction); deep FinOps expertise is not required. Strong organizational and execution skills: able to track a program's moving pieces, follow up, and keep things on schedule. Clear written and verbal communication; comfortable presenting status and asks to engineering partners. Collaborative and coachable — works well with engineers and more senior TPMs, seeks input, and takes feedback well. Experience at a large-scale consumer tech or infrastructure organization is
JOB TITLE Senior Storage Engineer A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO The Senior Storage Engineer will be responsible for monitoring, implementation, and management of the organization's File/Block storage infrastructure. This role requires to cover one of the 2 possible shifts in India as well as a deep understanding of File/Block storage technologies and cloud-based solutions, with a focus on ensuring data availability, scalability, and security. The ideal candidate will have extensive experience in storage engineering and a proven ability to manage complex storage environments. • Provide advanced troubleshooting and support for complex storage issues, minimizing downtime and ensuring seamless operations. • Oversee and optimize File/Block storage systems, ensuring high availability and performance. • Conduct capacity planning and forecasting to ensure adequate storage resources are available to meet future demands. • Monitor storage performance and work with US resource to find recommendations for improvements to enhance efficiency and reduce costs. • Implement integration and automation solutions and provide feedback for solution to US Team to streamline operations and improve storage management. • Implement robust security measures to protect data integrity and ensure compliance with industry regulations and standards. • Maintain comprehensive documentation of storage configurations and processes and provide training and guidance to junior team members. WHAT’S
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our
About Datadog: Datadog is the world-class monitoring and security platform for cloud applications. We’re dedicated to creating, developing, and supporting our product and customers, allowing for seamless collaboration and problem-solving among Dev, Ops and Security teams globally. Built by engineers, for engineers, our SaaS product is used by organizations of all sizes across a wide range of industries to enable digital transformation, cloud migration, and infrastructure monitoring of our customers’ entire technology stack. Given the resilience of cloud technologies and importance placed today in digital operations and agility, Datadog continues to innovate and is well positioned for the long term. The Team: Datadog's Finance team collaborates with teams across the organization, providing commercial, operational, and analytical support to ensure that Datadog's business continues to grow as rapidly and efficiently as possible. The Opportunity: We are seeking a Senior Revenue Accountant to join our growing Finance team at Datadog. As a member of the finance team, the Senior Revenue Accountant will be a key member of the team in developing more efficient revenue / billing data flow and close processes and be on the front lines of supporting the rapid growth of the Company. You Will: Work closely with the broader finance team to support the finance operations, accounting and compliance function for our business. Complete month-end close responsibilities including preparing journal entries, balance sheet reconciliations, and supporting schedules. Prepare memos and analyses surrounding revenue recognition under ASC 606 for more complicated billing arrangements Assist in calculation of key internal revenue metrics and invoicing during month-end close. Generate and issue prepaid and monthly invoices issued in arrears. Help customers understand their bills, which vary with usage and plan type. Effectively detect, communicate and work to resolve customer bi
MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workflows Pr
Get new senior cloud infrastructure engineer jobs by email
Daily job updates · Unsubscribe anytime