Jobiba hiring network

Cluster Lead Facilities Services Jobs

315 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cluster lead facilities services jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We're entering a world where AI agents don't just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents' ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova's throughput, correctness, and operational rigor grows dramatically. We're looking for a Senior Software Engineer who wants to go deep on the engine internals and the infrastructure underneath. You'll own significant components of a modern OLAP system — across query execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — and drive meaningful improvements to performance, cost-efficiency, and reliability. You'll grow your technical influence through the quality of your code, your design contributions, and your collaboration with other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of enterprise customers. What You'll Do Build and improve core query engine components Contribute across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Help ensure Nova's components support

pythonjavaredis
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! We’re looking for a senior engineer to help build, maintain and evolve the training framework that powers our frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure. You will design and maintain the core components that enable fast, reliable, and scalable model training — and build the tooling that connects research ideas to thousands of GPUs. If you enjoy working across the full stack of ML systems, this role gives you the opportunity and autonomy to have massive impact. What You’ll Work On Build and own the training framework responsible for large-scale LLM training. Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing). Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100). Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics. Collaborate closely with infra teams to ensure our cluster, container environments, and hardware configurations support high-performance training. Investigate and res

dockerkubernetesgit
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Security Clearance: Active Secret+ clearance strongly preferred; candidates eligible and willing to obtain clearance will also be considered. More information about Canadian Security Clearance is available here . As an Infrastructure Security Engineer, your key responsibilities include: Deploy, and manage infrastructure for Protected B classified environments, ensuring compliance with ITSG-33 and Canadian government standards Design and implement security controls for cloud (AWS, GCP, Azure) and hybrid/multi-cloud deployments Evaluate, implement, and manage security tools and technologies for training cluster and inference infrastructure hardening Implement security best practices including IAM, encryption, logging, and monitoring Participate in security incident response activities, including detection, analysis, containment, and remediation Conduct regular vulnerability assessments and penetration testing of infrastructure components Maintain comprehensive security documentation, procedures, and configurations for classified environments Maintain active Secret+ security clearance and adhere to all Canadian government security

awsazuregcp
View job →
PE
Private Employer
📍 Seattle• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Mission Manager is Palantir’s PaaS for enabling US Government customers and vendors to run software securely and compliantly in the most sensitive environments, but without the overhead — whether connected, disconnected, cloud, or edge. Built on the strength of Palantir’s Apollo platform, it provides the critical infrastructure needed to rapidly onboard and deploy applications into a secure Kubernetes-based ecosystem, freeing our customers to focus on building and powering mission-critical systems. The Mission Manager offering is still in its earliest days, and by joining us now, you’ll define the strategy for how we develop and scale it — witnessing firsthand the impact of your work on critical missions and the new capabilities you unlock. You’ll drive this by building elegant, robust APIs powered by Kubernetes controllers, bridging the gap between a raw Kubernetes cluster and a fully featured, infrastructure-agnostic runtime that can meet the operational demands of hundreds of specialized microservices. You’ll undertake this challenge alongside an energized team with a wide array of backgrounds and skillsets, all united by an ambitious vision for what’s possible.

kubernetesmicroservicesai
View job →
PE
Private Employer
📍 Washington• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Mission Manager is Palantir’s PaaS for enabling US Government customers and vendors to run software securely and compliantly in the most sensitive environments, but without the overhead — whether connected, disconnected, cloud, or edge. Built on the strength of Palantir’s Apollo platform, it provides the critical infrastructure needed to rapidly onboard and deploy applications into a secure Kubernetes-based ecosystem, freeing our customers to focus on building and powering mission-critical systems. The Mission Manager offering is still in its earliest days, and by joining us now, you’ll define the strategy for how we develop and scale it — witnessing firsthand the impact of your work on critical missions and the new capabilities you unlock. You’ll drive this by building elegant, robust APIs powered by Kubernetes controllers, bridging the gap between a raw Kubernetes cluster and a fully featured, infrastructure-agnostic runtime that can meet the operational demands of hundreds of specialized microservices. You’ll undertake this challenge alongside an energized team with a wide array of backgrounds and skillsets, all united by an ambitious vision for what’s possible.

kubernetesmicroservicesai
View job →
PE
Private Employer
📍 New York• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Mission Manager is Palantir’s PaaS for enabling US Government customers and vendors to run software securely and compliantly in the most sensitive environments, but without the overhead — whether connected, disconnected, cloud, or edge. Built on the strength of Palantir’s Apollo platform, it provides the critical infrastructure needed to rapidly onboard and deploy applications into a secure Kubernetes-based ecosystem, freeing our customers to focus on building and powering mission-critical systems. The Mission Manager offering is still in its earliest days, and by joining us now, you’ll define the strategy for how we develop and scale it — witnessing firsthand the impact of your work on critical missions and the new capabilities you unlock. You’ll drive this by building elegant, robust APIs powered by Kubernetes controllers, bridging the gap between a raw Kubernetes cluster and a fully featured, infrastructure-agnostic runtime that can meet the operational demands of hundreds of specialized microservices. You’ll undertake this challenge alongside an energized team with a wide array of backgrounds and skillsets, all united by an ambitious vision for what’s possible.

kubernetesmicroservicesai
View job →

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. Overview of the Role GitLab.com is being rebuilt as a horizontally scalable cluster. Two primitives make that possible: Organizations , the tenant boundary that makes a customer's data a single addressable, portable unit, and Cells , independent GitLab instances that host those Organizations and can be provisioned on demand across clouds and regions. The target is a 100x increase in headroom for GitLab.com, delivered without customers ever seeing a cell. Same domain, same URLs, same product — a substrate underneath that scales by adding capacity rather than by negotiating with a single database. The boundary we are building

gitrestai
View job →
NR
New Relic
📍 Spain• Full-time
1mo ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity As a Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with a proven track record of building and scaling resilient infrastructure. What you'll do Platform Orchestration: Work on a large-scale K8s infrastructure platform, ensuring high availability and performance. Automation: Drive the evolution of internal tooling to streamline platform delivery. This role requires Experience: Solid hands-on background in DevOps, Site Reliability, or Infrastructure Engineering. Kubernetes Mastery: Deep internal knowledge of K8s primitives (Deployments, StatefulSets, Services) and hands-on experience writing custom Kubernetes Operators. Golang Proficiency: Proficiency in Go, specifically for infrastructure automation and systems programming. Operations-Heavy Mindset: A proven track record of managing production environments and handling high-severity incidents. Cloud Infrastructure: Hands-on experience with cloud-native scaling tools (e.g., Karpenter, Cluster API) and Day 1/Day 2 operations of K8s clusters. Tooling: Familiarity with Helm and GitOps workflows (e.g., ArgoCD or Flux). Please note that visa sponsorship is not available for this position. Fostering a diverse, welcoming and inclusive environment is important to us. We work hard to make everyone feel comfortable bringing their best, most authentic selves to work every day. We cele

kubernetesgitrest
View job →
N
Nuro
📍 Mountain View• Full-time• From $160.4K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro takes a machine-learning-first approach to autonomous driving, and the ML Infrastructure team builds and operates the infrastructure that makes that possible. We own the systems that train the models at the core of the Nuro Driver™ - from distributed GPU training and closed-loop reinforcement learning, to the workflows, orchestration, observability, and cost management that keep the fleet running efficiently. Our work sits directly on the critical path of autonomy development. When a training run stalls, when a pipeline silently regresses, or when GPU utilization slips, it shows up in how fast the rest of the company can ship. We care as much about reliability and operational maturity as we do about raw scale. About the Work Contribute to Nuro’s training infrastructure, spanning multi-generation accelerators, and multi-cluster scheduling and orchestration. Design and operate large-scale data pipelines - batch and strea

pythongcpkubernetes
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro takes a machine-learning-first approach to autonomous driving, and the ML Infrastructure team builds and operates the infrastructure that makes that possible. We own the systems that train the models at the core of the Nuro Driver™ - from distributed GPU training and closed-loop reinforcement learning, to the workflows, orchestration, observability, and cost management that keep the fleet running efficiently. Our work sits directly on the critical path of autonomy development. When a training run stalls, when a pipeline silently regresses, or when GPU utilization slips, it shows up in how fast the rest of the company can ship. We care as much about reliability and operational maturity as we do about raw scale. About the Work Contribute to Nuro’s training infrastructure, spanning multi-generation accelerators, and multi-cluster scheduling and orchestration. Design and operate large-scale data pipelines - batch and strea

pythongcpkubernetes
View job →
R
Roblox
📍 San Mateo• Full-time• From $196.8K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a senior software engineer on the Cell Platform team at Roblox, you will build systems that Roblox engineers use to create and deploy resources onto Kubernetes. Our engineers deploy their services in a complex, hybrid, multiple-cluster, K8s environment. The Cell Platform manages this complexity for our users, with tools, APIs, K8s controllers, and UX, simplifying infrastructure for our internal customers. You Have: A desire to work on critical, large-scale distributed systems An appreciation of observability and instrumentation and tooling to make your life easier 3+ years of experience as software engineer Bachelor's degree in Computer Science or an equivalent field You will: Build our Roblox-wide control plane using Kubernetes primitives (and plenty of custom resources) Work on the interface of the few hundred person infrastructure organization to the thousands of Roblox engineers Write and review high quality code and tests (largely Golang) Work on a team that cares about inclusivity and shipping For roles that are based at our headquarters in San Mateo, CA: The starting base pay for this position is as shown below. The actual base pay is dependent upon a variety of job-related

awskubernetesgit
View job →
M
Mongodb
📍 Bengaluru• Full-time
1mo ago

We are seeking a Staff Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Staff Site Reliability Engineer on this new team, you will be responsible for providing technical leadership for the operational foundations that enable deployment at scale of AI applications. You will own the reliability architecture of the platform as it expands across regions and cloud providers, and set the technical direction for how the platform is operated, including capacity planning, multi-cloud expansion, incident response, and SLO discipline. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Bengaluru for our hybrid working model. Position Expectations Own the reliability architecture of the platform across regions and cloud providers Collaborate with the teams building the platform, providing internal support and guidance on operability, capacity, and best practices Set operational standards for the team: on-call quality, incident response, SLO discipline Mentor and technically develop the SRE team Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure Qualifications 10+ years of experience working on software and operating distributed systems, with deep Kubernetes expertise, including designing or evolving multi-cluster platforms Proficiency in Python, Go, or a similar programming language Understand workload isolati

pythonmongodbaws
View job →
M
Mongodb
📍 Gurugram• Full-time
1mo ago

We are seeking a Senior Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Senior Site Reliability Engineer on this new team, you will be responsible for enabling deployment at scale of AI applications and improving the performance, scalability, and reliability of the distributed systems infrastructure for this new product. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Position Expectations Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing Identify and configure key metrics to detect incidents and quantify service health, availability, and performance Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure Mentor early-career SREs and contribute to the team’s operational practices as it grows Qualifications Strong background in software development and operating distributed systems 6+ years of experience building and operating distributed systems, with proficiency in Python, Go, or a similar programming language Experience operating Kubernetes in production and debugging below the abstraction layer, including scheduling, cluster networking, and node-level issues Expertise in cloud infrastructure platforms, in

pythonmongodbaws
View job →

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold builders and sharp problem-solvers who are wired to deliver great outcomes. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. The DevX team’s mission is to build and operate the core developer infrastructure at Robinhood. Our team owns and scales the systems that thousands of engineers rely on daily, partnering with software developers across the company to make development fast, reliable, and cost-efficient! As a Senior Software Developer, you will focus on building and scaling our developer infrastructure. You will manage and optimize monorepo builds with Bazel and own the remote build execution cluster that powers them, and you will scale CI pipelines to be fast, safe, and reliable across thousands of daily builds. You will also improve the developer onboarding experience and provide remote development environments to accelerate development workflows. In this role, you will ensure developers can code, test, and deploy with minimal friction. This is an incredible opportunity to make a massive difference for the entire engineering organization! This role is based in our Toronto, ON office, with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. What you’ll do Manage and

pythonawskubernetes
View job →
C
Coinbase
📍 - Canada• Full-time• Remote• From C$191.1K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Software Engineer on the Compute Platform team within the Platform group, you'll own the primary compute orchestration infrastructure that every service at Coinbase runs on. Built largely on CNCF technologies including Kubernetes and Istio, this platform underpins the scalability, reliability, and efficiency of our entire product suite. You'll design and ship tooling, automation, and net-new capabilities that make it easy for hundreds of engineers to deploy and operate critical services, while partnering closely with Security, Reliability, and Observability teams to raise the bar across the stack. What you'll do: Own the design, build, and operation of Kubernetes cluster management tooling and automation that keeps our compute platform reliable and self-healing at scale. Build developer-facing tooling and workflows that improve how engineers across Coinbase interact with Kubernetes, with a heavy emphasis on integrating AI-driven processes and support. Deliver net-new compute capabilities for service owners, such as one-off jobs, cron scheduling, deployment strategies, EFS support, and automated right-sizing. Drive operational excellence by automating toil, reducing on-call burden, and continuously improving platform observability and incident response. Partner with Security, Reliability, and Observability teams to ensure the compute platform meets C

REMOTEawsgcpkubernetes
View job →
🔔

Get new cluster lead facilities services jobs by email

Daily job updates · Unsubscribe anytime