Jobiba hiring network

Cluster Lead Facilities Services Jobs

315 active opportunities Β· Updated for October 2026

Fresh results

15 shown

Explore current cluster lead facilities services jobs. Use filters to narrow by work mode, employment type, experience and date posted.

AG
Adani Group
πŸ“ Mundraβ€’ Full-time
1mo ago

The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 55614

aiexcelsap
View job β†’
AG
Adani Group
πŸ“ Mundraβ€’ Full-time
1mo ago

The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 54048

aiexcelsap
View job β†’
AG
Adani Group
πŸ“ Navi Mumbaiβ€’ Full-time
1mo ago

Asset Head - Techno Commercial is responsible for translating BU-wide Techno-Commercial strategies into effective execution across airport sites by leading sourcing, contracting, logistics, and vendor performance activities. The role ensures timely and cost-efficient procurement of airport-specific CapEx, OpEx, and services, supports operational continuity, enhances commercial controls, and drives process improvements and digital adoption while managing teams and enabling capability development at the site or cluster level. Source: Adani Group | Job ID: 52862

gitailogistics
View job β†’
AG
Adani Group
πŸ“ Mundraβ€’ Full-time
1mo ago

The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 51581

aiexcelsap
View job β†’
AG
1mo ago

The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 51579

aiexcelsap
View job β†’
AG
1mo ago

The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 42490

aiexcelsap
View job β†’
AG
Adani Group
πŸ“ Mundraβ€’ Full-time
1mo ago

The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 45606

aiexcelsap
View job β†’
AG
Adani Group
πŸ“ Mundraβ€’ Full-time
1mo ago

The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 42492

aiexcelsap
View job β†’
B
Biohub
πŸ“ Redwood Cityβ€’ Full-timeβ€’ Hybridβ€’ $241K – $331K/yr
1mo ago

Biohub is the first large-scale initiative bringing frontier AI models, massive compute, and frontier experimental capabilities under one roof. We're building a general-purpose system to accelerate scientific discovery, integrating frontier AI models, biological foundation models, and lab capabilities, with the ultimate goal of curing disease. Our technology powers scientists around the world, translating AI capabilities into tools that accelerate research everywhere. The Team The AI Cluster Production Engineering team is part of the AI Compute Platform organization at Biohub, a non-profit research lab committed to open science and open-source AI. We own the design, operation, and reliability of large-scale multi-GPU AI clusters that power frontier AI biology research: protein language models, genomic foundation models, and scientific reasoning systems built to be shared, not monetized. Our clusters run Slurm on Kubernetes infrastructure and support everything from day-to-day AI researcher workflows to multi-node hero training runs at thousands of GPUs. The team works at the intersection of AI tooling, distributed systems, HPC, and frontier AI, debugging deep AI infrastructure problems and building AI systems critical to the entire AI organization. The Opportunity CZ Biohub's mission is to cure or prevent all human disease. Achieving that requires training frontier-scale AI biology models, and that demands reliable, high-performance compute infrastructure. This is production engineering work at a frontier AI lab, with the twist that the mission is biology and the science is open. You'll keep GPU clusters running at high utilization, debug the toughest distributed systems failures, and build the operational foundations for scaling to multi-thousand GPU hero runs. The technical problems are genuinely hard (e.g., multi-node distributed training, InfiniBand fabrics, large-scale storage, Slurm at scale) inside an organization where the work is aimed at helping peop

pythonkubernetesgit
View job β†’
A
Anyscale
πŸ“ Remoteβ€’ Full-time
1mo ago

About Anyscale: At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we're building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Anyscale is looking for a Software Engineer to join the Platform and Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the team, we build the scalable, secure, and robust backbone that enables this vision, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads. We are seeking a talented Software Engineer with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale's cloud platform. You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impa

pythonawsazure
View job β†’
A
Anyscale
πŸ“ Indiaβ€’ Full-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About Ray Data Team: Ray Data is Python-native data processing engine that is a one stop shop for all AI data processing needs. Ray Data provides performant, first-class integration with cutting edge AI frameworks using both multi-modal and structured data. The Ray Data team currently develops and maintains Ray Data . We are a team of engineers passionate about building a Data processing engine which is a one-stop shop for all of your ML/AI needs. We are looking for exceptional engineers to build, optimize, and scale Ray for modern and increasingly complex AI workloads. As part of this role, you will: Improve the performance of Ray Data and multi-modal batch inference use cases. Ensure efficient scaling across different stages of the Data pipeline in a heterogeneous environment. Building data loading solutions for production training workloads. Focus on stability and fault tolerance at high scale Working with customers and new age AI native companies in scaling their AI workloads. We'd love to hear from you if have: At least 3-4 years of relevant work experience Solid background in building scalable and fault-tolerant distributed systems Experience with data processing, database internals. Passionate about large

pythonmachine learningai
View job β†’
A
Anyscale
πŸ“ Remoteβ€’ Full-time
1mo ago

About Anyscale At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure. As part of this role, you will Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices We'd love to hear from you if you have Familiarity with running ML inference at large scale with high throughput and low latency Familiarity with deep learning and deep learning frameworks (e.g. PyTorch) Solid understanding of distributed systems, ML inference challenges Bonus points

machine learningai
View job β†’
A
Anyscale
πŸ“ Remoteβ€’ Full-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend. About the Ray Core Team The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray. A snapshot of projects you can work on: Optimizing performance of large-scale workloads on Ray Stability and stress testing infrastructure Improving fault tolerance (HA) As part of this role, you will: Leading cross-team projects while mentoring junior team members Develop high quality open source software to simplify distributed programming (Ray) Identify, implement, and evaluate architectural improvements

restmachine learningai
View job β†’
A
Anyscale
πŸ“ Remoteβ€’ Full-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Ray aims to provide a universal API for building distributed applications (e.g. a machine learning pipeline of feature engineering, model training, and evaluation). Data is usually a core element connecting these different stages, and therefore plays a critical role in Ray’s usability, performance, and stability. We are looking for strong engineers to build, optimize, and scale Ray’s Datasets library and data processing capabilities in general. About the Ray Data team: The Ray Data team currently develops and maintains the Ray Datasets library, which is already powering critical production use cases (e.g. large scale data compaction at Amazon , and ML pipeline at Alibaba ). Ray Datasets is a Python library built on top of Apache Arrow and Ray Core (Ray’s C++ backend), and the Ray Data team interacts closely with Ray Core components including the scheduler and the memory & I/O subsystems. The Ray Data team also works closely with Ray’s ML libraries including Train, RLlib, and Serve. A snapshot of projects you will work on: - Performance of Ray Datasets at large scale (leveraging Arrow primitives, optimizing Ray object manager, etc.) - Integration with ML training and data sources - Stability an

pythonmachine learningai
View job β†’
A
Anyscale
πŸ“ Remoteβ€’ Full-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend. About the Ray Core Team The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray. A snapshot of projects you can work on: - Optimizing performance of large-scale workloads on Ray - Stability and stress testing infrastructure - Improving fault tolerance (HA) As part of this role, you will: Develop high quality open source software to simplify distributed programming (Ray) Identify, implement, and evaluate architectural improvements to Ray core Improve the testing process for Ray to make re

restmachine learningai
View job β†’
πŸ””

Get new cluster lead facilities services jobs by email

Daily job updates Β· Unsubscribe anytime