Jobiba hiring network

Cluster Lead Facilities Services Jobs

315 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cluster lead facilities services jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Business Systems drives efficiency across Datadog through business process analysis, systems automation and integrations, AI agent and MCP development, and vendor/software review. The team is increasingly embedded in cross-functional initiatives across People, Finance, GTM , Legal, Recruiting, and Technical Solutions — translating ambiguous business problems into scoped, buildable solutions and owning delivery end-to-end. This is not a generalist BSA role. Each Senior BSA will own a cluster of business functions end-to-end, acting as an internal product manager for their domain rather than processing inbound requests reactively. There are two openings, each covering a different domain: People, Recruiting, Legal or Finance, GTM, Procurement As the role evolves alongside AI, Senior BSAs leverage agentic tools for data aggregation and context gathering while remaining the critical human-in-the-loop layer — owning business context and product outlook, validating use cases, managing stakeholder relationships, and making the judgment calls agents can't. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Discovery and scoping: Translate ambiguous asks from business stakeholders into well-defined requirements that Business Systems Engineers can build against. Many stakeholders don't know what they want or the cross-functional impact of what they're asking for — you surface both before engineering begins. Cross-functional visibility: Identify dependencies, downstream impacts, and integration considerations that requesting teams miss. Proactive opportunity identification: Develop deep domain knowledge to identify automation and AI opportunities before they become inbound requests, shifting the team from reactive in

aisalesforcefinance
View job →
N
Nvidia
📍 Santa Clara, United States
10 days ago

At NVIDIA, we push the boundaries of computing innovation. Our ASIC Verification Engineers focus on developing the world’s top SoCs and GPUs. Joining us as a Senior ASIC Verification Engineer - GPU means working on modern technology powering consumer graphics and AI applications. This position is ideal for those passionate about technology and eager to impact computing’s future. What you'll be doing: As a key member of our ASIC Verification team, you will verify the design and implementation of the industry's leading GPUs. You will be responsible for verifying the ASIC build, architecture, golden models, and micro-architecture using advanced verification methodologies such as UVM or equivalent. Understand the design and implementation of your unit/cluster/chip, define the verification scope, develop the verification infrastructure, and verify the correctness of the design. Collaborate with architects, designers, and pre- and post-silicon verification teams to accomplish your task. What we need to see: Bachelor's Degree in EE, CS, or CE or equivalent experience. 5+ years of relevant experience. Experience in verification using random stimulus along with functional coverage and assertion-based verification methodologies. Experience with design and verification tools (VCS or equivalent simulation tools, debug tools like Verdi, Indago, GDB). Expertise in System Verilog or similar HVL. Strong debugging and analytical skills. Perl and C/C++ programming language experience desirable. Strong communication skills and the ability & desire to work as a great teammate are huge pluses. Experience in crafting test bench environments for unit and system level verification. #LI-Hybrid Your base salary will be det

E
11 days ago

We build and operate the compute infrastructure our researchers run on, supporting large-scale processing of historical market data and model training on our own hardware across multiple data centers. Our environment includes bare-metal Linux, virtualization, storage, and GPU clusters, where performance, reliability, and predictable system behavior are critical. Our Infrastructure team covers monitoring and automation, distributed storage, hardware and OS provisioning, GPU clusters and workload scheduling, high-speed networking, L2/L3 Linux support, and security engineering. Engineers here own their tasks end to end, so there's room to go deeper in your area and pick up the parts you haven't touched yet. We’re looking for a Linux Infrastructure Engineer who can work hands-on with server and cluster environments, from deployment and configuration to performance tuning, troubleshooting, and ongoing improvement What You’ll Be Doing: Deploying, configuring, and maintaining Linux-based bare-metal servers across our data centers Building and operating clustered environments, including virtualization, storage, GPU compute, and database clusters Troubleshooting complex Linux, hardware, networking, and cluster-level issues Performance tuning for throughput, latency, stability, and resource utilization Monitoring infrastructure health and performance, identifying bottlenecks, and preventing recurring issues Supporting the full server lifecycle: provisioning, setup, upgrades, and maintenance Improving reliability and predictability during failures, maintenance, and scaling Automating provisioning, configuration, and operational tasks, primarily using Ansible and scripting What We Look For In You: Strong hands-on Linux administration and troubleshooting experience Production experience with on-premise, bare-metal infrastructure Good understanding of Linux performance and bottleneck analysis Experience with: infrastructure monitoring and troubleshooting production issues,

REMOTElinuxansible
View job →
T-
Tubi - Canada
📍 Toronto• Full-time• From C$1.4M/yr
15 days ago

About the Role: We're hiring Senior and Staff Data Platform Engineers to join the Data Infrastructure teams in Toronto. Together these teams own the infrastructure that processes billions of events per day: Spark-on-Kubernetes, Flink and Kinesis pipelines, a multi-petabyte Delta Lake, a large-scale MemoryDB feature store, Databricks multi-environment operations, and the catalog and lifecycle systems that govern it. The team is small and senior. Each engineer owns major platform components: you design it, build it, and support it in production. This is a hybrid-role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: Spark-on-Kubernetes — EKS-based compute platform for Spark workloads: cluster configuration, Pod Identity IAM, job environment setup, Kustomize overlays, and shadow canary validation Event ingestion — Rust services and Flink jobs processing billions of events per day over Kinesis; throughput, reliability, on-call response, and AI-assisted operational tooling to reduce toil Platform infrastructure — Terraform modules for environment provisioning, cross-account AWS IAM, ARC runner infrastructure, and CI/CD for data platform changes Feature store and ML compute — Flink-based real-time feature pipelines feeding a large-scale MemoryDB cluster; GPU capacity governance and Databricks multi-environment operations for ML training workloads Workflow orchestration and CDC — Airflow-based DAG deployment, change data capture pipeline operations, and data quality monitoring Your Background: 3+ years building and operating production data platform infrastructure at the cluster or platform level, across Spark, Flink, Kinesis, Kubernetes, or equivalent Deep experience in at least one of: Spark-on-K8s cluster operations, Rust-based data or systems engineering, Kubernetes platform engineering and IaC, or data catalog and governance tooling Production AWS experience or equivalent: EKS, S3, Kinesis, and mu

pythonjavaaws
View job →
A
Anyscale
📍 Remote• Full-time• $171K – $211K/yr
1mo ago

At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale runs on a small, high-leverage IT function, and we're looking for an IT Engineer to own it. This is a hands-on engineering role, not a help desk or tier-1 support role. Roughly half the job is software: you will own and extend a set of internal automation and notification services wired into our directory and hardware systems. You should be comfortable owning a codebase and a GitHub repository, not only vendor admin consoles. The other half is the identity, endpoint, hardware, and vendor backbone that keeps a company of roughly 200 people, close to half of them engineers, productive across three public clouds. Reporting to the Head of Security and IT, you will work closely with security, engineering, and the people team. This role is based in the San Francisco Bay Area on a hybrid schedule. What You'll Do Own identity and access: Okta and Google Workspace administration, group-based authorization, 1Password Enterprise, passkeys, and automation driven provisioning and deprovisioning. Own and extend our internal automation services, and run them on scoped service accounts with sound security and hygiene. Own the endpoint fleet through mobile device management, including device trust posture checks, endpoint

azuregitmachine learning
View job →

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for a Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the Infra team, we build the scalable, secure, and robust backbone that enables this vision. Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads. We are seeking a talented engineers with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale’s cloud platform. You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers. A snapshot of projects you may work on Design, build, and scale services that orches

REMOTEpythonawsazure
View job →
A
Anyscale
📍 Remote• Full-time• Remote
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we're building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Anyscale is looking for a Software Engineer to join the ML Developer Experience (MLDevX) team. MLDevX owns the experience layer of the Anyscale platform: the interfaces through which users and coding agents discover, configure, run, observe, debug, and productionize AI workloads. Every user journey crosses this layer through the CLI, SDKs, APIs, UI, Workspaces, MCP, or the workflows and integrations built on top of them. Together, these form the user’s primary interface into Anyscale, turning distributed computing from a systems problem back into a coding problem. We build the common contracts, tools, control-plane services, and architecture that power these surfaces. You will work across the stack from developer tooling to cloud infrastructure and the Ray runtime. Manage long-running operations and make failures across jobs, tasks, actors, nodes, and GPUs easier to diagnose. The systems you build must scale with the platform, remain predictable through failures, and be intuitive for developers, programmable for applications, and operable by coding agents. This is a high impact individual-contributor role with end-to-end ownership. You will work directly with users and field teams to identify high

REMOTEmachine learningaigo
View job →
A
Anyscale
📍 Remote• Full-time• Remote
1mo ago

Senior Cloud Security Engineer At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale's security needs are growing as we operate more production and cloud infrastructure for larger and more demanding customers. We're looking for a Senior Cloud Security Engineer to own the security of that infrastructure. This is a hands-on, high-ownership role: you will own how our production and cloud environments are hardened, isolated, and monitored. You will set and drive the direction for infrastructure and production security, reporting to the Head of Security and partnering closely with the wider engineering organization. This role is based in the San Francisco, Bay Area. In your first year, success looks like hardened and well-segmented production environments, strong runtime security coverage across our container footprint, and a clear, defensible story for how we secure the infrastructure our customers rely on. What You'll Do Own the security posture of Anyscale's production and cloud infrastructure across AWS and Azure, including hardening, network segmentation, and tenant isolation. Own runtime security coverage across our Kubernetes environments, from deployment through detection of anomalous activity. Partner with engineering on secure infrastructure architectur

REMOTEawsazurekubernetes
View job →
A
1mo ago

At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale's product security needs are growing as we ship to larger and more demanding customers. We're looking for a Senior Product Security Engineer to own our secure software development lifecycle and to be engineering's partner on building security into the product. Reporting to the Head of Security, you will work in close partnership with engineering. This is a senior, high-ownership role. You will own and operate a scalable SSDL, partner with engineering on security features and secure design, review the security of existing systems and new initiatives, and own how we find, track, drive to resolution and report on vulnerabilities in what we ship. This role is based in India. In your first year, success looks like an SSDL that scales with engineering rather than gating it, security review embedded in how new initiatives ship, and accurate, on-demand vulnerability reporting backed by a working path to resolution. What You'll Do Own and operate a scalable secure software development lifecycle: threat modeling, security requirements, secure design practices, and scanning that engineering can readily adopt. Partner with engineering on security features and secure-by-design architecture, from early design through i

machine learningaisupply chain
View job →
A
Anyscale
📍 Remote• Full-time
1mo ago

At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale's security and compliance needs are growing as we work with larger and more demanding customers. Compliance is increasingly a customer-facing, contractual function rather than an internal exercise, and we are looking for someone to own it. This role owns that function end to end: our audits, our evidence base, our risk register, and the security diligence that customers put us through before and during a contract. You will work directly with the Head of Security and across engineering, IT, legal, and sales. This is a program-ownership role with the autonomy and accountability that implies. You will not have a senior compliance function above you to defer to; you are that function. In your first year, success looks like a complete and defensible evidence base with clean audit outcomes, a repeatable way to answer customer security diligence, and a risk register that leadership actually uses. What You'll Do Own our SOC 2 Type II and ISO 27001 programs, and future frameworks as we take them on, including scope, evidence, control operation, and the relationship with our external auditors. Own and complete the control evidence base in our compliance automation platform, moving controls from partially substantia

machine learningaigo
View job →
A
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: As a Site Reliability Engineer, you will play a crucial role in ensuring the smooth operation of all user-facing services and other Anyscale production systems. Anyscale values diversity and inclusion, and we encourage applications from individuals of all backgrounds. This includes processes for provisioning, negotiating prices, managing costs, seeing opportunities for teams to reduce wastage by finding applications across the company. You will apply sound engineering principles, operational discipline, and mature automation to our environments and the Anyscale codebase as we scale. As part of this role, you will: Develop a unified perspective on how cloud components are utilized across the company, taking into account diverse needs and requirements. Ensure that deployment methodologies align with the company's reliability goals. Build systems that promote understanding of production environments, facilitating quick identification of issues through robust observability infrastructure for metrics, logging, and tracing. Create monitoring and alerting systems at different levels, enabling teams to easily contribute and enhance the overall monitoring capabilities. Establish testing infrastructure to s

machine learningai
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
1mo ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI’s AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production. In this role, you will: Design and implement the LLM inference runtime for frontier models running on custom silicon. Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization. Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads. Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack. Enable new

REMOTEpythonawsrest
View job →
N
Nvidia
📍 Santa Clara, United States
1mo ago

NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel

pythonkuberneteslinux
View job →
A
Anyscale
📍 Remote• Full-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About Ray Data Team: Ray Data is Python-native data processing engine that is a one stop shop for all AI data processing needs. Ray Data provides performant, first-class integration with cutting edge AI frameworks using both multi-modal and structured data. The Ray Data team currently develops and maintains Ray Data . We are a team of engineers passionate about building a Data processing engine which is a one-stop shop for all of your ML/AI needs. We are looking for exceptional engineers to build, optimize, and scale Ray for modern and increasingly complex AI workloads. As part of this role, you will: Improve the performance of Ray Data and multi-modal batch inference use cases. Ensure efficient scaling across different stages of the Data pipeline in a heterogeneous environment. Building data loading solutions for production training workloads. Focus on stability and fault tolerance at high scale Working with customers and new age AI native companies in scaling their AI workloads. We'd love to hear from you if have: At least 3-4 years of relevant work experience Solid background in building scalable and fault-tolerant distributed systems Experience with data processing, database internals. Passionate about large

pythonmachine learningai
View job →
AG
Adani Group
📍 Mundra• Full-time
1mo ago

The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 55616

aiexcelsap
View job →
🔔

Get new cluster lead facilities services jobs by email

Daily job updates · Unsubscribe anytime