Jobiba hiring network

Cluster Lead Facilities Services Jobs

315 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cluster lead facilities services jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. Our team builds the software, tooling, and operational systems that help manage this fleet at scale. We work across production engineering, distributed systems, capacity management, and operational automation to improve reliability, reduce manual work, and make better use of available compute. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will develop the systems that help manage the GPU fleet powering ChatGPT, including tooling for fleet health, capacity planning, operational automation, and incident response. You will work closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and compute utilization. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference. Develop systems for capacity planning, fleet health monitoring, and resource utilization. Automate operational workflows, including incident detection, diagnosis, and response. Identify and address bottlenecks affecting fleet reliability, scalability, and performance. Partner with infrastructure, research, and product engineering teams to improve the compute platform. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems. Have a background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering. Have built software that automates operational workflows and reduces manual work. Have worked with distributed infrastructure, cluster orchestration, or large-scale int

pythonawsrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,

awskubernetesci/cd
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system

pythonsqlaws
View job →
S
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer, Snowflake Natsec Running Snowflake in public sectors in different countries and regions, even in different industry verticals, requires us to build a compliant, secure, and auditable infrastructure. Many key design decisions are deeply rooted in the Snowflake product architecture. As a Senior Software Engineer, you will be responsible for leading several key areas and collaborating with various engineering groups in addition to the Public Sector team. To be successful in the area, you will need to have (and continue to build) a broad and in-depth knowledge base on cloud infrastructure, privacy, and governance, compliance controls, data security and data residency in various aspects of Snowflake. AS A SENIOR SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Solve real business needs at large scale by applying your software engineering and analytical problem solving skills. Design, implement and maintain scalable distributed systems for our cloud automation platform that include cloud control plane, Kubernetes container platform and traffic and networking. Work directly with customers to quickly understand their critical problems and design and implement solutions Deploy and maintain availability of cloud compute servers and Kubernetes cluster that power the

pythonjavaaws
View job →
S
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer, Snowflake Natsec Running Snowflake in public sectors in different countries and regions, even in different industry verticals, requires us to build a compliant, secure, and auditable infrastructure. Many key design decisions are deeply rooted in the Snowflake product architecture. As a Senior Software Engineer, you will be responsible for leading several key areas and collaborating with various engineering groups in addition to the Public Sector team. To be successful in the area, you will need to have (and continue to build) a broad and in-depth knowledge base on cloud infrastructure, privacy, and governance, compliance controls, data security and data residency in various aspects of Snowflake. AS A SENIOR SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Solve real business needs at large scale by applying your software engineering and analytical problem solving skills. Design, implement and maintain scalable distributed systems for our cloud automation platform that include cloud control plane, Kubernetes container platform and traffic and networking. Work directly with customers to quickly understand their critical problems and design and implement solutions Deploy and maintain availability of cloud compute servers and Kubernetes cluster that power the

pythonjavaaws
View job →
S
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - Container Platform (Menlo Park) About the Role We build the foundational container platform that runs Snowflake's production, AI/ML, and CI workloads across AWS, Azure, and GCP, including a rapidly growing AI/ML footprint. Hundreds of large Kubernetes clusters under management and growing. The work is to make that fleet reliable, automated, and invisible to the thousands of engineers building on top of it. This is a staff-level role on a senior, high-performing platform team. You'll own hard problems end to end, drive technical direction across teams, and build the automation and platform abstractions that make operating at this scale sustainable. There is significant unsolved work ahead: improving the developer experience for thousands of internal engineers and continuing to scale the platform to meet Snowflake's growth. What You'll Do Own the design and delivery of large, complex platform initiatives spanning cluster lifecycle management, multi-cloud automation, and internal developer tooling. Identify and drive cross-team technical improvements across the platform, from architecture through adoption. Make and defend architectural trade-offs grounded in reliability, scalability, and operational reality. Act as a technical anchor for the team, dev

awsazuregcp
View job →

As a Sr. Staff Technical Program Manager, you will partner with key Engineering, Product, Product Design, Marketing, and Analytics stakeholders to conduct data-driven experiments and deliver features for MongoDB. As a seasoned program leader, you will be responsible for one of our most mission-critical programs this fiscal year, and own executive level communication related to the program. We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model, or remote within Ireland. The right candidate for this role will be: Experienced with 15+ years of working in an engineering organization leading complex cross-functional technical programs Experienced with 10 years of Software development background, with Cloud storage and compute products Experience with Service Oriented Architecture and Cloud Infrastructure Able to leverage their knowledge and experience to influence technical discussions, summarizing outcomes and next steps Skilled at communicating across a diverse set of engineers and stakeholders Hyper-organized and capable of coordinating across multiple independent work streams and organizations A role model for effective execution practices, driven by an attuned sense of priority and urgency Able to leverage their experience in program delivery to influence improvements to our tools, operations, and architecture Trained in working with project tracking software (e.g. Jira, Rally, MS Project) Familiar with MongoDB or a comparable technology Interested in business automation work such as scripting in Python, Google Apps Script and Slack Position Expectations: Leverage technical acumen and analytical skills to drive engineering programs forward and maximize business value delivery Recognize patterns in a sea of information and take action accordingly Design, maintain, and improve the processes and tools that power program delivery Build strategic partnerships with Product and Engineering stakeholders Expand knowledge into new

pythonmongodbaws
View job →
M
Mongodb
📍 New York City; United States• Full-time• From $106K/yr
1mo ago

Join the MongoDB Networking & Observability team and help build the core of a distributed database! Our team focuses on creating and enhancing components which facilitate communication between distributed processes and make these processes, and their communication, easily observable. Networking Observability’s responsibilities include improving MongoDB networking, improving the efficiency of resource utilization, and building low-overhead observability features. Our team includes engineers located in New York City and fully remote engineers. We operate close to the bottom of the stack, and have a lot of influence over the availability, performance, and robustness of our open source database. Recently, we’ve improved connection handling, explored new networking architectures, and integrated OpenTelemetry to make issues easier to diagnose and connect MongoDB to modern observability tools. We are planning to further improve our networking’s stack performance, availability and scalability as well as further enhance our observability stack using open observability frameworks. Are you excited to help the MongoDB engineering team build a better database? We are! Join us today, and we can build a faster, more reliable, exceptionally observable, database system together. This role can be based out of our New York City office or remotely within the United States and Canada. Candidate Profile 3+ years of experience building distributed systems Passionate about delivering and deploying a product with cross-team stakeholders Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Hands-on experience with building production-level code. Experience in C++ is required Interest in furthering their knowledge of networking, observability and how computer architecture and internals impact the availability of SaaS Solid verbal and written communication skills and highly motivated to collaborate with colleagues Po

mongodbawsazure
View job →
M
Mongodb
📍 Alberta• Full-time• From C$122K/yr
1mo ago

Join the MongoDB Networking & Observability team and help build the core of a distributed database! Our team focuses on creating and enhancing components which facilitate communication between distributed processes and make these processes, and their communication, easily observable. Networking Observability’s responsibilities include improving MongoDB networking, improving the efficiency of resource utilization, and building low-overhead observability features. Our team includes engineers located in New York City and fully remote engineers. We operate close to the bottom of the stack, and have a lot of influence over the availability, performance, and robustness of our open source database. Recently, we’ve improved connection handling, explored new networking architectures, and integrated OpenTelemetry to make issues easier to diagnose and connect MongoDB to modern observability tools. We are planning to further improve our networking’s stack performance, availability and scalability as well as further enhance our observability stack using open observability frameworks. Are you excited to help the MongoDB engineering team build a better database? We are! Join us today, and we can build a faster, more reliable, exceptionally observable, database system together. This role will be based remotely in Canada. Candidate Profile 3+ years of experience building distributed systems Passionate about delivering and deploying a product with cross-team stakeholders Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Hands-on experience with building production-level code. Experience in C++ is required Interest in furthering their knowledge of networking, observability and how computer architecture and internals impact the availability of SaaS Solid verbal and written communication skills and highly motivated to collaborate with colleagues Position Expectations Understand and improve the current funct

mongodbawsazure
View job →

Come join the Server Ingress Security team, where we are rearchitecting MongoDB Server’s ingress networking to make MongoDB clusters even more secure. This new team is building the Atlas Network Protection layer, a set of performant, security-critical services that harden MongoDB's pre-authentication attack surface and provides the ability to respond rapidly to emergent threats. We are looking for talented Staff Engineers to join the team and be founding members, where you will play a crucial role in our multi-year roadmap. Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to be a key technical leader on a collaborative team that applies security and systems engineering fundamentals to protect a popular database at scale, join us! We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. Candidate Profile 10+ years of experience building production-quality systems software with a large user base, robust design structure, and rigorous code quality Experience with large backend/compiled codebases and performance-sensitive software, preferably in Rust Bonus points for experience working hands-on in security-sensitive or networking-adjacent domains Strong systems fundamentals, including multi-threaded programming and performance profiling. Bonus points for: Understanding of network protocols, TLS, and connection lifecycle management. Familiarity with security concepts such as attack surface reduction, input validation, memory safety, and defense-in-depth architectures. Excellent verbal and written technical communication skills for communicating to a wide variety of audiences ranging from junior engineers to executive stakeholder Strong mentorship skills, and excitement about leveling up your peers and teammates through coaching, feedback, and enablement Strong time management skills and the ability to realistically assess project complexity B.Sc. in Computer Science or a related

mongodbawsazure
View job →

Come join the Server Ingress Security team, where we are rearchitecting MongoDB Server’s ingress networking to make MongoDB clusters even more secure. This new team is building the Atlas Network Protection layer, a set of performant, security-critical services that harden MongoDB's pre-authentication attack surface and provides the ability to respond rapidly to emergent threats. We are looking for talented Senior Engineers to join the team and be founding members, where you will play a crucial role in our multi-year roadmap. Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies security and systems engineering fundamentals to protect a popular database at scale, join us! We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. Candidate Profile 5+ years of experience building production-quality systems software Experience with large backend/compiled codebases and performance-sensitive software, preferably in Rust Bonus points for experience working hands-on in security-sensitive or networking-adjacent domains Strong systems fundamentals, including multi-threaded programming and performance profiling. Bonus points for: Understanding of network protocols, TLS, and connection lifecycle management Familiarity with security concepts such as attack surface reduction, input validation, memory safety, and defense-in-depth architectures Excellent verbal and written technical communication skills, with a strong desire to collaborate with colleagues Strong time management skills and the ability to realistically assess project complexity B.Sc. in Computer Science or a related field, or equivalent practical experience, with strong competencies in data structures, algorithms, and software design/architecture. Interest in the theory and practice of high-availability, security-critical systems Position Expectations Design, implement, and operate production

mongodbawsazure
View job →
M
Mongodb
📍 Toronto• Full-time• From C$137K/yr
1mo ago

We are hiring a Senior Software Engineer to join our Server Security team. The Server Security team is a development-focused group within MongoDB's core engineering organization. Operating "close to the bottom of the stack," the team builds features that enable database users to secure their data globally. You will work on critical components including: Cryptography: Queryable Encryption , at-rest data encryption, and fundamental cryptographic principles. Identity & Access: Authentication and authorization systems, TLS, and X.509 certificate management Network Security: High-performance, low-latency networking protocols (PKI, Hashing, CRLs) System Integrity: Resilience, observability, and compliance assurance within a large-scale distributed database Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies distributed systems fundamentals to deliver core features of a popular database, join us! Let’s change what’s possible for application developers, system architects, and database operators. The role As a Senior Engineer, you will apply distributed systems fundamentals to deliver core security features. You will be a leader in improving MongoDB's security posture by owning features and leading investigations into complex areas of the codebase. What you’ll do: Build and test new security features in a large, feature-rich C++ codebase Work across engineering, cloud services, and support teams to coordinate feature rollouts and changes Stand for code quality and security best practices, assisting fellow engineers in writing well-reasoned, secure code Use strong diagnostic intuition to solve thorny technical issues related to distributed systems, concurrency, and OS internals This role can be remote or hybrid anywhere in the USA or Canada. We will prioritize candidates who are already located in one of these countries. Candidate Profile We are looking for a highly technical engineer w

javamongodbaws
View job →
M
Mongodb
📍 New York City• Full-time• From $126K/yr
1mo ago

We are hiring a Senior Software Engineer to join our Server Security team. The Server Security team is a development-focused group within MongoDB's core engineering organization. Operating "close to the bottom of the stack," the team builds features that enable database users to secure their data globally. You will work on critical components including: Cryptography: Queryable Encryption , at-rest data encryption, and fundamental cryptographic principles. Identity & Access: Authentication and authorization systems, TLS, and X.509 certificate management Network Security: High-performance, low-latency networking protocols (PKI, Hashing, CRLs) System Integrity: Resilience, observability, and compliance assurance within a large-scale distributed database Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies distributed systems fundamentals to deliver core features of a popular database, join us! Let’s change what’s possible for application developers, system architects, and database operators. The role As a Senior Engineer, you will apply distributed systems fundamentals to deliver core security features. You will be a leader in improving MongoDB's security posture by owning features and leading investigations into complex areas of the codebase. What you’ll do: Build and test new security features in a large, feature-rich C++ codebase Work across engineering, cloud services, and support teams to coordinate feature rollouts and changes Stand for code quality and security best practices, assisting fellow engineers in writing well-reasoned, secure code Use strong diagnostic intuition to solve thorny technical issues related to distributed systems, concurrency, and OS internals This role can be remote or hybrid anywhere in the USA or Canada. We will prioritize candidates who are already located in one of these countries. Candidate Profile We are looking for a highly technical engineer w

javamongodbaws
View job →
M
23 days ago

The MongoDB Atlas Clusters Team is seeking a Senior Staff Product Manager to provide division-wide, product leadership for MongoDB Atlas which is our best-in-class, fully managed, multi-cloud database service that simplifies the complexities of deploying and managing highly available database deployments across multiple regions and cloud providers. As the most senior product management contributor, you will shape the multi-year strategy, product architecture, and market trajectory for the entire Atlas Clusters business unit. You will be the primary product authority on the most complex trade-offs, anticipating systemic risks and cutting through structural friction that blocks multiple teams and initiatives. This role requires collaboration with distinguished engineers to rigorously debate strategy and pressure-test assumptions and advise executive leadership on options that balance long-term monetization, strategy, and business outcomes. And, where relevant, owns the pricing and monetization strategy for the product line. We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. Responsibilities Product Vision and Roadmap: Spearhead strategic unity across multiple product areas over a multi-year horizon to drive a coherent product experience Business Leadership: Originate and de-risk the division’s most consequential strategic bets, owning rigorous business cases and advising leadership on multi-year economic outcomes Customer Advocacy: Engage with strategically significant customers at the senior level (e.g., advisory boards and executive briefings) to surface cross-portfolio opportunities and inform the roadmap PM Excellence: Elevate the global standard of product management by mentoring Staff and Senior PMs and pioneering new data-driven playbooks across the organization. Create frameworks and standards for customer and market research that others use for conducting high-quality customer and market research Cross-Fu

mongodbawsazure
View job →
🔔

Get new cluster lead facilities services jobs by email

Daily job updates · Unsubscribe anytime