The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 55616
Jobiba hiring network
Cluster Hr Head Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster hr head jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 55614
The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 54048
The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 51581
The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 51579
The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 42490
The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 45606
The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 42492
About the Role: We're hiring Senior and Staff Data Platform Engineers to join the Data Infrastructure teams in Toronto. Together these teams own the infrastructure that processes billions of events per day: Spark-on-Kubernetes, Flink and Kinesis pipelines, a multi-petabyte Delta Lake, a large-scale MemoryDB feature store, Databricks multi-environment operations, and the catalog and lifecycle systems that govern it. The team is small and senior. Each engineer owns major platform components: you design it, build it, and support it in production. This is a hybrid-role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: Spark-on-Kubernetes — EKS-based compute platform for Spark workloads: cluster configuration, Pod Identity IAM, job environment setup, Kustomize overlays, and shadow canary validation Event ingestion — Rust services and Flink jobs processing billions of events per day over Kinesis; throughput, reliability, on-call response, and AI-assisted operational tooling to reduce toil Platform infrastructure — Terraform modules for environment provisioning, cross-account AWS IAM, ARC runner infrastructure, and CI/CD for data platform changes Feature store and ML compute — Flink-based real-time feature pipelines feeding a large-scale MemoryDB cluster; GPU capacity governance and Databricks multi-environment operations for ML training workloads Workflow orchestration and CDC — Airflow-based DAG deployment, change data capture pipeline operations, and data quality monitoring Your Background: 3+ years building and operating production data platform infrastructure at the cluster or platform level, across Spark, Flink, Kinesis, Kubernetes, or equivalent Deep experience in at least one of: Spark-on-K8s cluster operations, Rust-based data or systems engineering, Kubernetes platform engineering and IaC, or data catalog and governance tooling Production AWS experience or equivalent: EKS, S3, Kinesis, and mu
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The Sandbox service team at SpaceXAI builds and maintains a secure, scalable system that gives our models safe, controlled access to computational environments. This infrastructure powers critical workloads across training and product, enabling models to run code, build software, interact with tools, and even control applications with user interfaces. We provision containers and virtual machines on large-scale clusters, granting models interactive control over these remote environments. Our work spans the full stack: from orchestrating massive jobs and resource scheduling at the cluster level, to fine-tuning filesystem performance on nodes. The Sandbox service enables Grok to safely run and test code in real-time for user queries, and supports reinforcement learning in training, where models interactively explore tools ranging from compilers to productivity apps. BASIC QUALIFICATIONS: Expert knowledge of Rust, C++ or Go Familiarity with Python Deep experience with either Linux or Windows systems (familiarity with both is a strong plus) Experience with virtualisation and containerisation technologies (e.g., cgroups, KVM, gVisor, QEMU) Solid knowledge of the networking stack COMPENSATION AND BENEFITS: £107,000 -
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Powering the METCA AI & Sovereign Cloud Revolution As Everpure , we are transcending traditional storage to deliver the Enterprise Data Cloud - a unified architecture engineered to fuel the world’s most ambitious AI, Deep Learning, and HPC projects. The METCA region is a global epicenter for AI infrastructure investment, characterised by massive capital flowing into Tier-1 sovereign GPU clouds (CSPs, Neo-scalers) and enterprise AI factories. We are seeking a Battle-Trained, High-Conviction Hunting Systems Engineer (SE) to serve as our technical tip of the spear. This is not a passive, box-pushing relationship management role. You will partner aggressively with an Enterprise Account Executive to target, break into, and land the largest AI infrastructure projects in the market, displacing legacy architectures and securing net-new footprints. WHAT YOU'LL DO Execute High-Impact Hunting: Partner closely with Account Executives to actively map out and break into net-new enterprise accounts, sovereign GPU clouds, and high-performance computing clusters. Architect the AI Factory: Design high-performance, multi-tenant data pipelines. Move beyond basic storage architecture to design full-stack environments, optimising how data nodes interact within massive GPU fabrics. Drive Technical Consensus: Lead deep-dive architectural workshops with customer GPU cluster architects while simultaneously translating complex en
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As our TT-Distributed Software Engineer, you will develop and optimize distributed software systems that power the most efficient and highest-performing AI and HPC clusters. In this role, you'll work on distributed programming across multiple nodes, utilizing systems programming, inter-node communication, and Tenstorrent’s scalable architectures to advance the state-of-the-art distributed inference and training infrastructure. This role is hybrid, based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Strong C or C++ engineer with solid foundations in systems programming, operating systems, and distributed systems principles. Enthusiastic about distributed computing, including IPC, socket programming, and cluster resource coordination. Comfortable reasoning about scalability, fault tolerance, and performance across multi-node environments. Curious and first-principles thinker who challenges conventional approaches to distributed system design. Motivated to grow into a deep technical expert in large-scale distributed AI infrastructure. What We Need Architect, implement, and optim
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments. We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment. This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment.. RESPONSIBILITIES Core Responsibilities: Drive suppliers to keep the maximum amount of the GPU fleet online and healthy. Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs. Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier. SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short. Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability. Scope and
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're hiring a Product Data Scientist to establish how product decisions at Baseten are made with data. You'll work directly with Product and Engineering, alongside GTM to determine measurement, strategy, experimentation and implementation. This is a foundational, hands-on role. You'll define what success looks like across a technical, usage-based platform and turn ambiguous questions into analyses, forecasts, and experiments that shape product strategy. You'll work from clickstream and product events through inference telemetry and observability data, helping Baseten make faster decisions about reliability, performance, adoption and developer experience. RESPONSIBILITIES Partner directly with Product and Engineering: frame the questions that matter, define success criteria, and turn analysis into roadmap, launch, and prioritization decisions. Define how product success is measured: establish metrics across activation, adoption, retention, expansion, reliability and user experience. Support experimentation and launches: design measurement plans, analyze A/B experiments and controlled rollouts, and translate results into product decisions. Diagnose reliability and scaling behavior: join customer signals with request, replica, deployment, and cluster telemetry to find patterns in release bottlenecks, unhealthy replicas, and models without traffic. Define the enterprise customer journey and measure feature adoption
About the Team The Core Services organization builds and runs the mission-critical online services that product teams rely on in production. We own foundational distributed systems and platform capabilities that enable reliable execution, high-performance services, and large-scale file/data needs across our products. This team is distinct from developer infrastructure and data infrastructure—our focus is production service foundations and core runtime services. About the Role We’re hiring an Engineering Manager, Core Services to help lead teams responsible for highly reliable, high-scale distributed systems that sit on the critical path for OpenAI products. Your team will own foundational production systems that OpenAI’s product engineering teams build on. You’ll collaborate closely with product and infrastructure partners to ship reliable services quickly, and help scale systems and teams as OpenAI grows. You’ll partner closely with senior engineering leaders to scale the org, mature operations, and drive major platform initiatives. This role requires strong technical ability. You’ll be responsible for: Managing and growing a high-performing team of infrastructure engineers. Leading teams building and operating large, critical production platforms, including cluster reliability, scaling, and rollout safety. Building and operating mission-critical distributed systems with strong operational rigor (SLOs, incident response, capacity planning, reliability). Setting technical direction for platform foundations such as workflow/orchestration capabilities, large-scale file/blob/storage services, and core service foundations. Partnering with a broad set of stakeholders, including product engineering, adjacent infrastructure teams, and (where relevant) finance/cost partners. Coaching, mentoring, and developing engineers and emerging leaders. You might thrive in this role if you: Have significant experience leading teams that run mission-critical infrastructure in production
Get new cluster hr head jobs by email
Daily job updates · Unsubscribe anytime