About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for a Software Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the Infra team, we build the scalable, secure, and robust backbone that enables this vision. Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads. We are seeking a talented Software Engineer with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale’s cloud platform. You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers. A snapshot of projects you may work on Design, build, and scale services that orches
Jobiba hiring network
Cluster Lead Facilities Services Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster lead facilities services jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is growing its Sales Team! We're looking for a Bay Area based Enterprise Account Executive to build out our enterprise-level client relationships. This is a Hunter role, so you will be prospecting, developing, and closing new business while focusing on the clients’ requirements. The Enterprise AE’s must have the confidence and ability to negotiate and close agreements with clients and support new customers. This role is a unique opportunity to contribute in a meaningful way to high visibility, high impact projects at a very exciting time for the company. Anyscale is an innovative, high-growth, customer-focused company in a large and growing market. If you are an energetic, self-managed professional with experience managing a complex sales process and possess excellent presentation and listening skills, organization and contact management capabilities, we’d love to hear from you. As part of this role, you will: Achieve sales quotas for named accounts on a quarterly and annual basis by developing a sales strategy in the allocated territory with a target prospect list, and a regional sales plan Develop marketing plans with the marketing team to drive revenue growth Be the trusted advisor to
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role If you joined us today, you get to author the strategy by which each of Palantir’s software platforms - Foundry, Gotham, Apollo - achieves full containerization across an intimidating diversity of infrastructure types. Tomorrow, you get to do the same for Palantir’s ever expanding customer community. You will drive those goals by building elegant, robust APIs powered by K8s controllers which bridge the gap between a raw Kubernetes cluster and a fully-featured, infrastructure agnostic runtime that can scale to the operational needs of 100s of specialized microservices. Joining you on that journey is a highly motivated team with a diverse group of backgrounds and skillsets, brimming with ambition.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role If you joined us today, you get to author the strategy by which each of Palantir’s software platforms - Foundry, Gotham, Apollo - achieves full containerization across an intimidating diversity of infrastructure types. Tomorrow, you get to do the same for Palantir’s ever expanding customer community. You will drive those goals by building elegant, robust APIs powered by K8s controllers which bridge the gap between a raw Kubernetes cluster and a fully-featured, infrastructure agnostic runtime that can scale to the operational needs of 100s of specialized microservices. Joining you on that journey is a highly motivated team with a diverse group of backgrounds and skillsets, brimming with ambition.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Apollo is Palantir’s autonomous software management and deployment platform. It enables seamless, continuous delivery of mission-critical software (Foundry, Gotham, AIP) across a vast range of environments: on-prem, public cloud, disconnected (air-gapped) networks, and highly regulated settings (including IL-5 and FedRAMP). As a Software Engineer on the Apollo team, you’ll build and operate a large-scale distributed system to allow the remote operation and maintenance of Kubernetes clusters. Our mission is to extract the entire state of a cluster into a portable, high-performance artifact within minutes, enabling full and almost instant cluster reconstruction from the ground up—all while pushing the limits of speed, reliability, and scale. You’ll design and implement backup and restore solutions for Kubernetes, leveraging proprietary compression infrastructure tailored to Palantir’s unique deployment models. You’ll also build and optimize our container artifact store, which is based on the OCI (Open Container Initiative) distribution spec—the industry standard for storing and distributing container images and artifacts. You’ll own the backbone of every environment Apollo supports, from hyperscalers to Army trucks. If you’re excited by challenges at the intersection of container technologies like OCI and docker, storage, and distributed systems, you’ll find opportunities here to dive deep into storage formats and low-level optimizations, where milliseconds matter. As we increasingly automate cluster creation and management on diverse hardware, you’ll play a key role in scaling Palantir’s presence at the edge and solving tough distributed systems proble
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Apollo is Palantir’s autonomous software management and deployment platform. It enables seamless, continuous delivery of mission-critical software (Foundry, Gotham, AIP) across a vast range of environments: on-prem, public cloud, disconnected (air-gapped) networks, and highly regulated settings (including IL-5 and FedRAMP). As a Software Engineer on the Apollo team, you’ll build and operate a large-scale distributed system to allow the remote operation and maintenance of Kubernetes clusters. Our mission is to extract the entire state of a cluster into a portable, high-performance artifact within minutes, enabling full and almost instant cluster reconstruction from the ground up—all while pushing the limits of speed, reliability, and scale. You’ll design and implement backup and restore solutions for Kubernetes, leveraging proprietary compression infrastructure tailored to Palantir’s unique deployment models. You’ll also build and optimize our container artifact store, which is based on the OCI (Open Container Initiative) distribution spec—the industry standard for storing and distributing container images and artifacts. You’ll own the backbone of every environment Apollo supports, from hyperscalers to Army trucks. If you’re excited by challenges at the intersection of container technologies like OCI and docker, storage, and distributed systems, you’ll find opportunities here to dive deep into storage formats and low-level optimizations, where milliseconds matter. As we increasingly automate cluster creation and management on diverse hardware, you’ll play a key role in scaling Palantir’s presence at the edge and solving tough distributed systems proble
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Fleet Management team provides the core runtime environment that empowers our developers to build and ship products to delight our customers. We manage the end-to-end lifecycle of our Kubernetes fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As our infrastructure scales to support new use cases and products, we are spearheading a migration from Terraform-based Infrastructure as Code (IaC) to an Operator-driven lifecycle management model. This role can be based out of our Austin, Boston, Los Angeles, New York City, Raleigh, or San Francisco offices, remotely in the United States region, or our European office in Dublin. Responsibilities Contribute to developing and maintaining a scalable and secure runtime environment on top of Kubernetes that supports product needs across MongoDB Provide internal support for our Kubernetes ecosystem, partnering with engineering teams to help them solve domain-specific problems Participate in a 24/7 on-call rotation to resolve critical issues Prioritize blameless post-mortems and dedicate engineering time to systemic fixes, ensuring you aren’t paged for the same issue twice You may be a good fit if you Have 6+ years of experience in software development and operating distributed systems Are proficient in Go, Python, or a similar language, with a strong commitment to code quality and testing practices (writing unit, integration, and E2E tests) Have deep experience using and extending containerization technologies, preferably Kubernetes Have a solid understanding
About the Team This team builds and operates the systems that enable OpenAI researchers to run reliable, scalable, and efficient research workflows. The team sits close to research and works across infrastructure, systems, and automation to make sure researchers have the tools and environments they need to move quickly. The work spans software engineering, infrastructure, systems administration, cluster operations, and reliability engineering. As OpenAI’s infrastructure evolves from bespoke bare-metal systems toward more standard, scalable platforms, the team needs engineers who can understand how systems work end-to-end and build the right abstractions without reinventing the wheel. About the Role As a Software Engineer on this team, you will build and operate the infrastructure that supports frontier research and critical research-facing systems. You will work on systems that sit close to the metal, but the role is not limited to classic operations or sysadmin work. We are looking for someone who can reason about networking, bootstrapping, Kubernetes, scalability, automation, and reliability - while also writing software to make these systems better over time. This role is a strong fit for an independent, high-ownership engineer who enjoys reliability-heavy infrastructure work but still wants to build. You do not need to come in as a kernel expert or highly algorithmic optimization engineer, but you should be deeply curious about infrastructure, comfortable debugging complex systems, and excited to support researchers doing novel work. We expect you to: Build and operate reliable infrastructure for research workloads and research-facing services. Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services. Improve cluster bootstrapping, provisioning, automation, and deployment workflows. Debug issues across networking, compute, storage, orchestration, and service reliability layers.
This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions tailored to the demands of advanced AI workloads. We work across the full stack—from silicon to system integration—partnering closely with internal teams and external vendors to define and deliver next-generation AI infrastructure. Our team focuses on defining scalable, high-performance system architectures and reference designs that balance performance, cost, and operational efficiency across rapidly evolving technologies. About the Role We are seeking a 3P Architect to define and drive rack- and cluster-level reference designs in collaboration with external partners. This role is responsible for translating workload requirements and system-level goals into concrete architectures, aligning partners on critical design attributes, and ensuring vendor roadmaps meet our infrastructure needs. You will work closely with performance modeling and internal architecture teams to evaluate tradeoffs, while owning the end-to-end definition and execution of third-party system designs. This includes identifying gaps in current technologies, driving vendor development, and shaping future infrastructure capabilities. This role requires strong system intuition, cross-functional leadership, and the ability to operate effectively across internal teams and external ecosystems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Define rack- and cluster-level reference architectures for AI infrastructure deployments. Translate workload requirements into clear system design specifications and partner deliverables. Collaborate with performance modeling teams to evaluate architectural tradeoffs and system behaviors. Align internal stakeholders and external partners on critical system attributes (performance, cost, power, reliability, scalability). Identify gaps in current technology offerings and dr
About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t
NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world. NVIDIA GH200 superchip provides performance and productivity required for strong scaling for HPC and generative AI workload. Scale out is inherent to design of this massive superchip. We are looking for expert engineers to come and help design rack level solutions for next generation scaling AI supercomputing platforms. We are looking for a strong technical architect to own end to end manageability architecture for these products in data centers. You will work with various component leads internally and externally, drive customer use cases, align architecture with customer requirements and release best products to market. Join us at the forefront of technological advancement. What you’ll be doing: Drive server management for large clusters and data centers deploying GPUs and Grace solution from Nvidia. Work with data center architects and cloud customers to narrow down on requirements for implementation to ensure speed of light product development. Work with internal teams to make sure requirements are designed and implemented in right way with each firmware and software module Collaborate with other leads to design & build data center health management workflow. Drive reliability and optimization in firmware architecture from a data center view point. Work closely with cluster bring up team and resolve is
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer (Java) Job Description Summary Who is Mastercard? Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital economy that benefits everyone, everywhere by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. With connections across more than 210 countries and territories, we are building a sustainable world that unlocks priceless possibilities for all. Overview • The Decision Management program enables intelligent decision-based products through streaming analytics with the ability to govern these decisions and manage their outcomes with business agility. • This program leverages business rules & AI engines, a streaming big data cluster, an in-memory data grids, APIs, & UIs to deliver real time decisions at global scale • This person will be responsible for mentoring the team as well as stay hands on. We are looking for a Senior Software Developer to join our DMP team in Vancouver office.<
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Department: Treasury Reporting To: Manager - Treasury Job Location: Mumbai Education Qualification Required: B Com/M Com Number of Years of Experience Required: 2-4Yrs Summary: The Treasury Analyst is responsible for the day-to-day management of treasury operations, ensuring efficient cash flow, accurate financial reporting, and compliance with internal controls and regulations. Key Responsibilities: Daily: Cash Management: Prepare and share daily cash position reports with OpCo and regional teams. Report cluster and cash balances to GL and Treasury Manager. Manage fund flow and cash flow, including loan disbursements, payments, rollovers, and follow-ups. Share bank debits with the Payments team and credits with the Collection team. Share bank statements (collections) with relevant teams. Accounting and Reporting: Account for and approve treasury-related entries and share with the GL team for posting. Prepare Bank Reconciliation Statements (BRSs). Follow up with Collection and Payment teams to resolve open items in BRS. Arrange entries for Conversion of EEFC balances to INR as needed. Tran
Get new cluster lead facilities services jobs by email
Daily job updates · Unsubscribe anytime