Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As a Datacenter Liquid Cooling Architect, you will define, design, and architect next-generation liquid cooling infrastructure for Tenstorrent’s large-scale AI training and inference clusters. You will partner with systems engineering, mechanical engineering, software, and cross-functional design teams to develop chassis-, rack-, and cluster-scale cooling solutions, including CDU integration, telemetry and control, leak detection, and resilient operating strategies. This role will help shape reliable AI datacenter architectures and deployments for both internal and external customers. This role is on-site, based out of Toronto, Canada, Austin, Texas or Santa Clara, California. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A datacenter and system thermal design professional with 10+ years of experience architecting cooling infrastructure for complex computing environments. An experienced liquid cooling architect who can design chassis- and rack-scale solutions for large AI training and inference clusters. A systems thinker who understands how mechanical, electrical, software, facility, and systems engineering decisions come toge
Jobiba hiring network
Cluster Lead Facilities Services Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster lead facilities services jobs. Use filters to narrow by work mode, employment type, experience and date posted.
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The Sandbox service team at SpaceXAI builds and maintains a secure, scalable system that gives our models safe, controlled access to computational environments. This infrastructure powers critical workloads across training and product, enabling models to run code, build software, interact with tools, and even control applications with user interfaces. We provision containers and virtual machines on large-scale clusters, granting models interactive control over these remote environments. Our work spans the full stack: from orchestrating massive jobs and resource scheduling at the cluster level, to fine-tuning filesystem performance on nodes. The Sandbox service enables Grok to safely run and test code in real-time for user queries, and supports reinforcement learning in training, where models interactively explore tools ranging from compilers to productivity apps. BASIC QUALIFICATIONS: Expert knowledge of Rust, C++ or Go Familiarity with Python Deep experience with either Linux or Windows systems (familiarity with both is a strong plus) Experience with virtualisation and containerisation technologies (e.g., cgroups, KVM, gVisor, QEMU) Solid knowledge of the networking stack COMPENSATION AND BENEFITS: £107,000 -
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE The CPBU Manageability team is responsible for Middleware of FlashArray and FlashBlade products. We are building a scale-out all-flash file and object store, designed for the modern world. To really understand how our customers work with data, we are deeply immersed in AI, modern backup, log analytics with Splunk and Elastic, data pipeline with Kafka, cluster computing with Spark, and many more use cases. You will love it on the CPBU Manageability team if you: want to understand how modern applications work with data and how we can make it better. are ready to dive into a complex problem and be the one who will drive it to a resolution. enjoy working with distributed systems, algorithms, operating systems, Linux kernel, database internals, hypervisors, containers, compilers and hardware… or at least some of those. want to work with other great engineers and develop or refine skills that will serve your entire career. enjoy working in a collaborative team environment in an open office. If this describes you, let's talk! You can take a part in changing how the world works with data. WHAT YOU'LL DO Own and deliver innovation end-to-end, from concept to shipped product Design, develop and maintain customer-facing and internal-facing API and command line interface for end to end configuration and management of FlashArray and FlashBlade products using Java, Python and beyond Experimenting with new technologies and archi
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... This is an opportunity to be one of the seed members of a growing product team in the FlashBlade BU, one of the fastest growing at Pure. With the FlashBlade product, we are building a scale-out all-flash file and object store, designed for the modern world. To really understand how our customers work with data, we are deeply immersed in AI, modern backup, log analytics with Splunk and Elastic, data pipeline with Kafka, cluster computing with Spark, and many more use cases. You will love it on the FlashBlade team if you: Want to understand how modern applications - like AI or Splunk - work with data and how we can make it better. Enjoy working with distributed systems, algorithms, operating systems, Linux kernel, database internals, hypervisors, containers, compilers and hardware... or at least some of those. Are ready to dive into a complex problem and be the one who will drive it to a resolution. Want to work with other great engineers and develop or refine skills that will serve your entire career. If you, like us, say “bring it on” to exciting challenges that change the world, we have endless opportunities where you can make your mark. WHAT YOU’LL NEED TO BRING TO THIS ROLE... Design, collaborate and implement creative new algorithms and technologies for high-performance, highly reliable systems (think six 9’s). Own and deliver innovation end-to-end, from concept to shipped prod
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... This is an opportunity to be one of the seed members of a growing product team in the FlashBlade BU, one of the fastest growing at Pure. With the FlashBlade product, we are building a scale-out all-flash file and object store, designed for the modern world. To really understand how our customers work with data, we are deeply immersed in AI, modern backup, log analytics with Splunk and Elastic, data pipeline with Kafka, cluster computing with Spark, and many more use cases. You will love it on the FlashBlade team if you: Want to understand how modern applications - like AI or Splunk - work with data and how we can make it better. Enjoy working with distributed systems, algorithms, operating systems, Linux kernel, database internals, hypervisors, containers, compilers and hardware... or at least some of those. Are ready to dive into a complex problem and be the one who will drive it to a resolution. Want to work with other great engineers and develop or refine skills that will serve your entire career. If you, like us, say “bring it on” to exciting challenges that change the world, we have endless opportunities where you can make your mark. WHAT YOU’LL NEED TO BRING TO THIS ROLE... Design, collaborate and implement creative new algorithms and technologies for high-performance, highly reliable systems (think six 9’s). Own and deliver innovation end-to-end, from concept to shipped prod
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... This is an opportunity to be one of the seed members of a growing product team in the FlashBlade BU, one of the fastest growing at Pure. With the FlashBlade product, we are building a scale-out all-flash file and object store, designed for the modern world. To really understand how our customers work with data, we are deeply immersed in AI, modern backup, log analytics with Splunk and Elastic, data pipeline with Kafka, cluster computing with Spark, and many more use cases. You will love it on the FlashBlade team if you: Want to understand how modern applications - like AI or Splunk - work with data and how we can make it better. Enjoy working with distributed systems, algorithms, operating systems, Linux kernel, database internals, hypervisors, containers, compilers and hardware... or at least some of those. Are ready to dive into a complex problem and be the one who will drive it to a resolution. Want to work with other great engineers and develop or refine skills that will serve your entire career. If you, like us, say “bring it on” to exciting challenges that change the world, we have endless opportunities where you can make your mark. WHAT YOU’LL NEED TO BRING TO THIS ROLE... Design, collaborate and implement creative new algorithms and technologies for high-performance, highly reliable systems (think six 9’s). Own and deliver innovation end-to-end, from concept to shipped prod
JOB TITLE Cloud Compute Engineer A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. WHAT YOU'LL DO Design, build, and operate Kubernetes clusters on Amazon EKS, including cluster lifecycle management, networking, autoscaling, and workload scheduling. Manage and optimize EC2-based compute infrastructure, including instance selection, placement strategies, capacity planning, and utilization analysis. Operate and improve ECS-based services where applicable, ensuring consistency across our container runtime environments. Develop and maintain Infrastructure as Code (IaC) using Terraform to provision and manage compute resources at scale. Collaborate with development and platform teams to define compute patterns, containerization standards, and deployment best practices. Monitor compute environments for availability, performance, and cost, driving continuous optimization across the fleet. Contribute to architectural decisions around workload placement, multi-tenancy, OS image and container li
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As our TT-Distributed Software Engineer, you will develop and optimize distributed software systems that power the most efficient and highest-performing AI and HPC clusters. In this role, you'll work on distributed programming across multiple nodes, utilizing systems programming, inter-node communication, and Tenstorrent’s scalable architectures to advance the state-of-the-art distributed inference and training infrastructure. This role is hybrid, based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Strong C or C++ engineer with solid foundations in systems programming, operating systems, and distributed systems principles. Enthusiastic about distributed computing, including IPC, socket programming, and cluster resource coordination. Comfortable reasoning about scalability, fault tolerance, and performance across multi-node environments. Curious and first-principles thinker who challenges conventional approaches to distributed system design. Motivated to grow into a deep technical expert in large-scale distributed AI infrastructure. What We Need Architect, implement, and optim
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Developer Infrastructure org is the engine behind Robinhood's entire engineering organization — a collection of tightly integrated teams whose collective mission is to make every engineer at Robinhood faster, more reliable, and exponentially more productive. The org spans four major teams: DevX (Developer Experience), TestX (Test Infrastructure), Backend Platform, and Mobile Platform. DevX owns Robinhood's monorepo and Bazel-based build infrastructure — the critical layer between a developer writing code and that code being ready to ship — along with the company's full CI/CD pipeline and remote build execution cluster. TestX owns the infrastructure behind Robinhood's entire test experience: the integration test environments, and personal development environments that serve as miniature simulations of the full Robinhood system, giving engineers a safe, isolated space to test their code end-to-end before it ever touches production. Backend Platform and Mobile Platform own the core language runtimes, libraries, IDEs, and developer toolchains across Python, Go, TypeScript, Swift, and Android. Together, these teams share a single north star: leveraging AI and agentic systems
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments. We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment. This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment.. RESPONSIBILITIES Core Responsibilities: Drive suppliers to keep the maximum amount of the GPU fleet online and healthy. Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs. Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier. SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short. Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability. Scope and
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're hiring a Product Data Scientist to establish how product decisions at Baseten are made with data. You'll work directly with Product and Engineering, alongside GTM to determine measurement, strategy, experimentation and implementation. This is a foundational, hands-on role. You'll define what success looks like across a technical, usage-based platform and turn ambiguous questions into analyses, forecasts, and experiments that shape product strategy. You'll work from clickstream and product events through inference telemetry and observability data, helping Baseten make faster decisions about reliability, performance, adoption and developer experience. RESPONSIBILITIES Partner directly with Product and Engineering: frame the questions that matter, define success criteria, and turn analysis into roadmap, launch, and prioritization decisions. Define how product success is measured: establish metrics across activation, adoption, retention, expansion, reliability and user experience. Support experimentation and launches: design measurement plans, analyze A/B experiments and controlled rollouts, and translate results into product decisions. Diagnose reliability and scaling behavior: join customer signals with request, replica, deployment, and cluster telemetry to find patterns in release bottlenecks, unhealthy replicas, and models without traffic. Define the enterprise customer journey and measure feature adoption
MongoDB Technical Services Engineers use their exceptional problem solving and customer service skills, along with their deep technical experience, to advise customers and to solve their complex MongoDB problems. Technical Service Engineers are experts in the entire MongoDB ecosystem - database server, drivers, cloud, and infrastructure. This also includes services such as Atlas (database as a service), or Cloud Manager (which helps customers with automation, backup and monitoring of their MongoDB systems). Our engineers combine their MongoDB expertise with passion, initiative, teamwork, and a great sense of humor to help our customers be successful with MongoDB. The position will be based in our Gurugram office. The standard work week will be Monday to Friday, every week. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Cool things you’ll do You'll be working alongside our largest customers, solving their complex challenges - resolving questions on architecture, performance, recovery, security, and everything in between. You'll be an expert resource on best practices in running MongoDB at scale, whatever that scale may be. You'll be an advocate for customers' needs - interfacing with our product management and development teams on their behalf. And you'll contribute to internal projects, including software development of support tools for performance, benchmarking, and diagnostics. What you need You should have 5+ years of proven experience, we consider all candidates with an eye for those who are self-taught, insatiably curious, and multi-faceted. The ideal candidates should have strong technical experience in more than one of the following areas Systems engineering experience, including Linux performance, memory management, I/O tuning, configuration, security, networking, clusters, and troubleshooting Understand core Kubernetes concepts, including containers, namespaces, custom resources and multi-cluster dep
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Staff DB SRE to build the runtime foundation for NVIDIA’s enterprise AI platforms — with a strong emphasis on database infrastructure at scale. This role blends large-scale database transformation with the building and development of GPU-accelerated platforms. You'll develop the software systems, automation frameworks, and high-performance database services that power NVIDIA’s AI workloads at scale. What you'll be doing: Design and operate highly available database clusters (MySQL, MSSQL, Oracle) with automated replication, failover, point-in-time recovery, and disaster-recovery strategies at enterprise scale. Drive database performance engineering — own query optimization, indexing strategies, connection pooling, lock-contention analysis, and storage-engine tuning for production systems handling millions of transactions. Build self-service database lifecycle automation — from one-click cluster provisioning and schema migrations to zero-downtime upgrades, blue-green deployments, and automated capacity scaling. Bridge relational and AI-native data infrastructure — extend traditional database exper
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an OS / K8s Systems Engineer at Baseten, you’ll build the automation and systems that turn raw GPU hardware into production-ready compute. From provisioning to orchestration, you’ll own the software layer that makes our infrastructure reproducible, scalable, and reliable across data centers. This is a senior, hands-on role focused on building systems not operating them. You’ll work close to the metal designing OS images, building provisioning pipelines, and automating cluster bring-up from scratch. Your work will define how quickly we can turn new capacity into usable compute. EXAMPLE INITIATIVES Zero-to-cluster automation Build workflows that take new hardware from unprovisioned to fully operational cluster. Provisioning systems Design PXE-based or equivalent systems for imaging and lifecycle management. Reproducible infrastructure — Ensure clusters deploy consistently across data centers. RESPONSIBILITIES Own the end-to-end automation of cluster bring-up and lifecycle management. Build and maintain OS images, provisioning systems, and configuration pipelines. Deploy and operate cluster orchestration platforms (Kubernetes, Slurm, or similar). Design systems for reproducibility across sites and hardware generations. Automate upgrades, rollouts, and failure recovery. Optimize system performance, including GPU utilization and networking. Partner with hardware and network teams to validate and improve system b
About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We’re entering a world where AI agents don’t just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents’ ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova’s throughput, correctness, and operational rigor grows dramatically. We’re looking for a Staff Software Engineer who wants to go deep on both the engine internals and the infrastructure underneath it. You’ll work across the full stack of a modern OLAP system — query planning and execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — while driving meaningful improvements to performance, cost-efficiency, and reliability at scale. You’ll influence technical direction through your work, your design reviews, and your mentorship of other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of customers. What You’ll Do Build and evolve core query engine infrastructure Work across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Design and implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Design for high-throughput automated quer
Get new cluster lead facilities services jobs by email
Daily job updates · Unsubscribe anytime