About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation
Jobiba hiring network
Cluster Lead Facilities Services Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster lead facilities services jobs. Use filters to narrow by work mode, employment type, experience and date posted.
MongoDB is seeking an Engineering Manager to join the Atlas Organization. The organization is responsible for building MongoDB Atlas, our database-as-a-service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. This includes developing software to interface with the three major cloud providers (AWS, Azure, and GCP) in order to bring security, durability, availability, and performance to all deployments of MongoDB. The Atlas Data Federation & Archiving team is an engineering team responsible for the Atlas capabilities that allow customers to move data from hot to cold storage and run federated queries over that data. The team builds Atlas Data Federation, a distributed query engine that lets users query data across Atlas Clusters and cloud object storage through a unified service. The team also builds Atlas Online Archive which allows customers to move data from Atlas Clusters into fully managed cloud object storage while preserving a seamless query experience across hot and cold datasets. We are forming a new Atlas Data Federation & Archiving team in the Dublin area. The Engineering Manager who fills this position will be pivotal in growing that team. We are looking to speak to candidates who are based in Cork and would like a hybrid or in-office working model. What you’ll do Lead a team of motivated individual contributors who are eager to learn and grow Contribute to the code, design, and architecture of the systems your team develops Work with stakeholders throughout MongoDB to build our roadmap and product offerings Work with customers and support engineers to fix issues and become part of our on-call rotation Collaborate with team members to develop our codebase, best practices, and design principles Foster an inclusive and respectful work environment according to MongoDB's Core Values We’re looking for someone who Has at least 6 years of professi
MongoDB is seeking an Engineering Manager to join the Atlas Organization. The organization is responsible for building MongoDB Atlas, our database-as-a-service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. This includes developing software to interface with the three major cloud providers (AWS, Azure, and GCP) in order to bring security, durability, availability, and performance to all deployments of MongoDB. The Atlas Data Federation & Archiving team is an engineering team responsible for the Atlas capabilities that allow customers to move data from hot to cold storage and run federated queries over that data. The team builds Atlas Data Federation, a distributed query engine that lets users query data across Atlas Clusters and cloud object storage through a unified service. The team also builds Atlas Online Archive which allows customers to move data from Atlas Clusters into fully managed cloud object storage while preserving a seamless query experience across hot and cold datasets. We are forming a new Atlas Data Federation & Archiving team in the Dublin area. The Engineering Manager who fills this position will be pivotal in growing that team. We are looking to speak to candidates who are based in Dublin and would like a hybrid or in-office working model. What you’ll do Lead a team of motivated individual contributors who are eager to learn and grow Contribute to the code, design, and architecture of the systems your team develops Work with stakeholders throughout MongoDB to build our roadmap and product offerings Work with customers and support engineers to fix issues and become part of our on-call rotation Collaborate with team members to develop our codebase, best practices, and design principles Foster an inclusive and respectful work environment according to MongoDB's Core Values We’re looking for someone who Has at least 6 years of profes
About the Team Training Runtime builds the distributed systems that power OpenAI's largest model training runs - most recently GPT-5.5! The Data Movement area owns the infrastructure that keeps training jobs supplied with the right data at the right time, and keeps model state moving safely and efficiently across large clusters. Our work spans machine learning systems, distributed storage, high-throughput data loading, reliability engineering, and developer experience. Success means researchers can move quickly while training runs remain fast, reproducible, debuggable, and resilient at scale. About the Role We are looking for a deeply hands-on Technical Lead Manager to own datasets throughout our training infrastructure. This person will set the direction for how training jobs read data: the APIs, storage contracts, versioning model, benchmarks, debugging tools, and reliability guarantees that make data access consistent across current and future training frameworks. You will begin as the primary technical owner for dataset reads, working directly in the code while aligning researchers, training framework owners, storage teams, and infrastructure partners around a durable platform. The problem is deceptively hard at frontier scale: make enormous, heterogeneous datasets easy to consume, correct across distributed workers, observable when something goes wrong, and flexible enough to support pretraining, reinforcement learning, and multimodal training. In this role, you will Design and build a unified dataset read platform for multiple current and future training frameworks. Define dataset APIs, storage-format expectations, registration/versioning, and migration paths that make data access reproducible and maintainable. Build reliability into the read path, including stateful iteration, caching, fast restart, recovery, and clear operational contracts. Build terminal and web-based visualizers that let teams inspect text, multimodal, and reinforcement learning data late
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the world, delivering the object, block, and file storage platforms that power GoDaddy's hosting infrastructure, internal services, OpenStack environments, and next-generation AI/HPC workloads. If you're passionate about distributed systems, storage architecture, and solving failure scenarios at massive scale, this is an opportunity to work on infrastructure few engineers will experience in their careers. Ceph is a strategic platform at GoDaddy — not an ancillary service. Our global footprint includes 80+ production clusters, 20,000+ OSDs, 1,830 storage nodes, 300 PB of raw capacity, and 69 billion objects spanning five datacenters across three continents. The platform supports RBD, RGW (S3/Swift), and CephFS workloads through more than 1,550 pools, 574,000 placement groups, and 900+ MDS daemons, creating engineering challenges that demand deep expertise in storage architecture, data durability, performance optimization, automation, and observability. As a Lead Senior Site Reliability Engineer, you'll serve as one of the principal technical leaders for GoDaddy's Ceph platform. You'll design the next generation of storage clusters, lead major platform upgrades, drive capacity and hardware strategy, and establish the standards that govern how the platform scales. You'll be the engineer the team turns to for the most complex s
About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. We are looking for a Security Operations Lead (SOC Lead) to build, mature, and operate our 24/7 detection and response capabilities across a modern cloud-native and AI-driven environment. This role leads the global SOC function—monitoring, SIEM ownership, detection engineering, alert triage, and operational readiness—while also evaluating and integrating emerging AI-based SOC products and autonomous response platforms . You will oversee monitoring across multi-cloud environments (GCP primary, AWS/Azure secondary), Kubernetes, SaaS services, endpoints, developer tools, and AI workloads . You’ll collaborate closely with Cloud Security, Compliance/GRC, SRE, Platform Engineering, IT/Endpoint teams, and AI Infrastructure to ensure our detection strategy scales and stays ahead of evolving threats. This is a hands-on leadership role perfect for someone who wants to shape the SOC of the future while solving complex challenges in a high-scale AI setting. What You’ll Do SOC Leadership & 24/7 Monitoring Lead, mentor, and scale a global SOC team responsible for 24/7 monitoring, alert intake, triage, correlation, and escalation. Build operational rigor: processes, runbooks, SLAs, metrics, and quality standards for high-scale environments. Cover monitoring across: Cloud infrastructure (GCP, AWS, Azure) Kubernetes/GKE/EKS/AKS clusters SaaS platforms (Google Workspace, GitHub, Slack, Okta, etc.) Endpoints (macOS, Linux, Windows) including EDR/XDR telemetry Developer platforms + CI/CD pipelines AI/ML systems and model-serving workflows AI-Based SOC Integration & Innovation Evaluate, adopt, and integrate AI-native SOC technologies for triaging, detection, and correlation Identify opportunities to automate triage, investigations,
The Infrastructure Engineering team is responsible for building and maintaining a self-service internal development platform that enables MongoDB engineering teams to reliably deploy and operate their own production services and products. We work with numerous engineering teams across the company to understand their infrastructure requirements and development workflows, develop broadly applicable self-service platform services and tooling, continuously monitor how platform services are being utilized, and look for ways to improve developer productivity through automation and education. We are big open source enthusiasts and use a number of open source tools in our stack (contributing upstream whenever possible). Some of the tools we use regularly include Go, AWS, Kubernetes, Crossplane, Terraform, Helm, Drone, Prometheus, and Grafana. However, technology is nothing without a stellar team of engineers that are focused on doing high quality work and working as a team to solve complex distributed computing and platform engineering problems. This is where you come in! We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Our ideal candidate 2+ years of experience managing and mentoring a team of 3+ engineers Has 5+ years of experience owning the design and implementation of large software/infrastructure projects Has built and operated large-scale distributed systems in cloud providers (AWS strongly preferred) Has a strong backend programming background. Fluency in Go is strongly preferred; deep experience with another compiled or strongly-typed backend language is acceptable Pragmatic, detail-oriented, self-motivated, and understands the benefits of collaboration Strong experience operating production Kubernetes clusters, not just deployed to it Has practical experience defining and operating against SLI/SLOs for services they owned Strong experience with observability tooling: metrics, logging, traces, Prometheus, Grafana, OpenTe
About us Paytm is India's leading mobile payments and financial services distribution company. Pioneer of the mobile QR payments revolution in India, Paytm builds technologies that help small businesses with payments and commerce. Paytm’s mission is to serve half a billion Indians and bring them to the mainstream economy with the help of technology. Role Overview We are seeking a Database STL (Individual Contributor) with deep expertise in MySQL and strong working knowledge of MongoDB, PostgreSQL, and Cassandra. This role combines hands-on database administration and optimization with strategic ownership of database reliability, automation, and cloud adoption. The candidate will lead by example—driving technical excellence, influencing best practices, and partnering cross-functionally with DevOps, SRE, and product engineering teams to deliver highly available, secure, and scalable database platforms. Key Responsibilities 1. End-to-End Ownership of MySQL databases in production & staging—availability, performance, and reliability. 2. Architect, manage, and support MongoDB, PostgreSQL, and Cassandra clusters for scale and resilience. 3. Define and enforce backup, recovery, HA, and DR strategies across all critical database platforms. 4. Drive database performance engineering—tuning queries, optimizing schemas, indexing, and partitioning for high-volume workloads. 5. Own replication, clustering, and failover architectures ensuring business continuity. 6. Champion automation & AI-driven operations—design self-healing scripts, predictive scaling, and proactive monitoring solutions. Collaborate with Cloud/DevOps teams on AWS database services (RDS, Aurora, DynamoDB, EC2, S3) to optimize cost, security, and performance. 7. Establish monitoring dashboards & alerting mechanisms for slow queries, replication lag, deadlocks, and capacity planning. Ensure compliance & security standards—encryption, auditing, and regulatory requirements. 8. Lea
Role Summary This role serves as the single point of accountability between USMAPPS and the Specialty Care Business Unit, owning access strategy, payer marketing, and brand-level contracting and pricing strategy across the full Specialty Care portfolio. The role drives and owns outcomes, integrating all USMAPPS capabilities into one coherent access plan that directly supports the Specialty Care BU President and franchise leads. The VP is expected to operate at full strategic weight, lead a team aligned to Specialty Care brands and franchise clusters, and be the single person the Specialty Care BU holds accountable for market access results. The Vice President US Market Access Lead – Specialty Care reports directly to the Senior Vice President of US Market Access & Pfizer Patient Services (USMAPPS) with a dotted line to the Specialty Care BU President. This position requires close partnership with the Specialty Care BU President, franchise leads, Strategic Contracting & Analytics, Strategic Account Management, the Patient Services, and USMAPPS leadership. The role sits on the Specialty Care BU leadership team and on the USMAPPS Leadership Team. Role Responsibilities 1. Specialty Care Access Strategy Ownership </spa
Role Summary This role serves as the single point of accountability between USMAPPS and the Primary Care Business Unit, owning access strategy, payer marketing, and brand-level contracting and pricing strategy across the full Primary Care portfolio. The role drives and owns outcomes, integrating all USMAPPS capabilities into one coherent access plan that directly supports the Primary Care BU President and franchise leads. The VP is expected to operate at full strategic weight, lead a team aligned to Primary Care brands and franchise clusters, and be the single person accountable the Primary Care BU holds accountable for market access results. The Vice President, US Market Access Lead – Primary Care reports directly to the Senior Vice President of US Market Access & Pfizer Patient Services (USMAPPS) with a dotted line to the Primary Care BU President . This position requires close partnership with the Primary Care BU President, franchise leads, Strategic Contracting & Analytics, Strategic Account Management, the Patient Services, and USMAPPS leadership. The role sits on the Primary Care BU leadership team and on the USMAPPS Leadership Team. Role Responsibilities 1. Primary Care Access Strategy Ownership </
Role Summary This role serves as the single point of accountability between USMAPPS and the Oncology Business Unit, owning access strategy, payer marketing, and brand-level contracting and pricing strategy across the full oncology portfolio. The role drives and owns outcomes, integrating all USMAPPS capabilities into one coherent access plan that directly supports the Oncology BU President and franchise leads. The Vice President is expected to operate at full strategic weight, lead a team aligned to oncology brands and franchise clusters, and be the single person the Oncology BU holds accountable for market access results. The Vice President, US Market Access Lead – Oncology reports directly to the Senior Vice President of US Market Access & Pfizer Patient Services (USMAPPS) with a dotted line to the Oncology BU President . This position requires close partnership with the Oncology BU President, franchise leads, Strategic Contracting & Analytic s , Strategic Account Management, the Patient Services, and USMAPPS leadership. The role sits on the Oncology BU leadership team and on the USMAPPS Leadership Team. Role Responsibilities 1. Oncology Access Strategy Ownership Develop and maintain a fully integrated brand market access plan for all oncology brands, covering payer marketing, contracting strategy, and pricing strategy in one coherent plan </
MongoDB is seeking a Senior Software Engineer to join the Atlas Clusters Organization. The organization is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. This includes developing software to interface with the three major cloud providers (AWS, Azure, and GCP) in order to bring security, durability, availability, and performance to all deployments of MongoDB. We are forming a new Atlas Clusters team in the Dublin area. We are looking to speak to candidates who are based in Dublin for our hybrid working model. What you’ll do Build and design new features for MongoDB Atlas Contribute to and lead complex technical projects Work with stakeholders throughout MongoDB to build our roadmap and product offerings Work with customers and support engineers to fix issues and become part of our on-call rotation Collaborate with team members to develop our codebase, best practices, and design principles Foster an inclusive and respectful work environment according to MongoDB's Core Values We’re looking for someone who Has at least 6 years of professional software development experience Is skilled at writing large-scale, distributed backend systems in a compiled language (Go, Java, C#, etc.). Has experience with at least one major cloud provider technology (AWS, Azure, GCP) Has led the launch of a new module and maintained it in production Is eager to solve tough problems Has excellent communication skills Is curious, collaborative, and motivated Success Measures In 3 months, you'll have shipped code into production and collaborated with the team to solve tough problems In 6 months, you'll have contributed to a large project and joined our on-call rotation In 12 months, you'll have designed new features, led development work, and become a go-to expert on parts of the system About MongoDB MongoDB is built for
MongoDB is seeking an Engineering Manager to join the Atlas Clusters Organization. The organization is responsible for building MongoDB Atlas, our database-as-a-service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. This includes developing software to interface with the three major cloud providers (AWS, Azure, and GCP) in order to bring security, durability, availability, and performance to all deployments of MongoDB. The Atlas Clusters Fleet Signal Management team is a product engineering team that builds machinery to rollout hardware and software to the Atlas data plane, detect regressions at scale, and automate remediative actions. The core mission is to establish platform stability by owning the rollout and monitoring systems designed to swiftly detect and mitigate potential issues. We are forming a new Atlas Clusters Fleet Signal Management team in the Dublin area. The Engineering Manager who fills this position will be pivotal in growing that team. We are looking to speak to candidates who are based in Dublin for our hybrid working model. What you’ll do Lead a team of motivated individual contributors who are eager to learn and grow Contribute to the code, design, and architecture of the systems your team develops Work with stakeholders throughout MongoDB to build our roadmap and product offerings Work with customers and support engineers to fix issues and become part of our on-call rotation Collaborate with team members to develop our codebase, best practices, and design principles Foster an inclusive and respectful work environment according to MongoDB's Core V
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE You will join the Portworx team in Everpure, which is responsible for delivering the highest quality Portworx Enterprise products. You will be contributing clean & robust code, be customer oriented and put quality first.A WHAT YOU'LL DO Designing and developing cloud native microservices and integrating new features to Portworx products Bringing a focus on design, development, unit/functional testing, code reviews, documentation, continuous integration and continuous deployment Debug product and performance issues in large scale clusters using AI tooling Collaborating with peers and stake-holders to take solutions from initial design to production Take full ownership of design and development activity by adapting to customer feedback and handling issues found in unit testing, system testing and customer deployments Experimenting with new technologies in order to push the state-of-the-art and innovate new solutions. We are primarily an in-office environment and therefore, you will be expected to work from the Bangalore office in compliance with Everpure's policies, unless you are on PTO, or work travel, or other approved leave. WHAT YOU BRING BS in Computer Science 7+ years of experience in Designing, Development and Testing of Enterprise products (Golang preferred). Good understanding of Microservice Architectures and Cloud Native platforms Designing and owning micro services to operate and scale in a di
Get new cluster lead facilities services jobs by email
Daily job updates · Unsubscribe anytime