Jobs in United States

Cluster Lead Facilities Services in United States

87 active opportunities · Updated October 2026

Explore current cluster lead facilities services jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You’ll own the hands-on and automation work that brings WAN, fiber, carrier, and cloud-interconnect circuits into service. Partner with network engineers, fiber providers, cloud service providers, colocation teams, and data-center technicians to move each connection from ordered and patched to verified, stable, and ready for handoff. You’ll own Layer 1 troubleshooting and circuit bring-up while building workflows that translate reliable system or model output into precise, approved technician actions, capture field feedback, and drive each connection to a green-port handoff. The right person combines strong physical-networking judgment with practical automation skills: patch-panel and port mappings, optics and light levels, provider coordination, structured operational data, API or scripting workflows, and human-in-the-loop LLM tooling. Responsibilities Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence. Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports. Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching

AWSAzureRestAI
N
📍 Remote, United States· Remote
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA DGX Cloud is an AI Factory designed to power the next generation of AI and industrial-scale breakthroughs. As a Principal Engineer for Security Architecture, within our Security Engineering organization, you will own a core security domain of the AI factory: the architecture, the paved road that delivers it, and much of the code underneath. You will hold the security design bar across DGX Cloud from inside the teams doing the building, and this is a founding seat on a new team. Security Engineering is a new organization at DGX Cloud, accountable for the security outcome of the platform, and Security Architecture is the function inside it that holds the design bar. Security here is fleet horizontal and stack vertical, so your work will cross every DGX Cloud engineering organization: you will embed with the teams building GPU clusters, control planes, and services, join their designs as a participant rather than an approver, and leave behind systems in which an entire class of risk is no longer possible. There is no architecture review board here and no approval queue. You are a senior IC with deep security domain knowledge, and the security bar holds because you helped set it and then helped ship it. What You Will Be Doing: Own a Security Domain End to End: Take architectural ownership of a core domain of DGX Cloud security, from the design through the system running in production. That could be tenant and GPU workload isolation, workload identity, infrastructure and network, supply-chain provenance, hardened baselines and patching, or deploy-time policy and admission control. Embed with the Teams Building It: Join the design early, write the code, and help land it. The posture is not "you did this wrong." It is "here are the considerations we need to meet, I will help, let's go to work." Build Paved Roads, Not

KubernetesLinuxArtificial IntelligenceAI
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated

PythonKubernetesLinuxMachine Learning
T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is building the world’s fastest, most efficient AI compute clusters. TT-Fabric is the high-performance nervous system of this platform: the low-level networking layer that lets thousands of RISC-V and AI processors snap together into a single, massively parallel distributed supercomputer. If you love squeezing nanoseconds out of hot paths, designing protocols that move data at absurd scale, and turning messy hardware constraints into elegant distributed systems, this is an opportunity to shape the fabric that future AI models will run on This role is hybrid based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who We Are Strong systems engineer with deep C or C++ experience and comfort working in low-level or bare-metal environments. Passionate about hardware-software interaction, performance tuning, and eliminating inefficiencies at the protocol level. Curious about networking, synchronization, and communication across large clusters. Comfortable reasoning from first principles and challenging industry conventions. Motivated by building infrastructure that directly impacts large-scale

AWSAIC++SEM
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,

AWSKubernetesLinuxRest
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

We are developing advanced multi-rack, multi-tenant AI/ML datacenters with NVIDIA GB200, and upcoming GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team to focus on the cloud-native stack for datacenter products like GB200. In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex scheduling challenges across racks, tenants, and clouds as part of the CSP engagements team. What you’ll be doing: Perform deep-dive debugging of multi-rack, multi-tenant clusters: scheduler behavior, container runtime issues, device-plugin crashes, RDMA/IB fabric anomalies, etc. Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom micro-services that expose new GPU capabilities. Drive joint architecture reviews and “whiteboard” sessions with CSP and internal platform teams; convert findings into RFCs and upstream pull requests. Create reproducible testbeds (Helm/Ansible/Terraform) that mirror customer environments; automate validation and benchmark suites. Deliver technical collateral-design docs, how-to guides, demo scripts-and present at customer on-sites, KubeCon, and SlurmUG. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Strong source-level expertise in Kubernetes internals (scheduler, CRI/CNI/CSI, operators) and Slurm (federation, power-save, plugins). Hands-on experience integrating next-gen GPUs (Blackwell/GB200/GB300) or comparable accelerators into containerized clusters. Proven track record debugging large-scale, cloud-native stacks across ne

PythonKubernetesArtificial IntelligenceAI
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

We are looking for a Senior Software Engineer to become part of our storage management plane team. The management plane is a web-based application crafted to provide our storage customers the capabilities to handle and supervise our distributed storage infrastructure. Our team is continually dedicated to acquiring and implementing ground breaking technologies to overcome obstacles and innovate solutions for improving our ability to handle large clusters of machines efficiently. What You Will Be Doing: Maintain and develop Kubernetes operators and our Container Storage Interface (CSI) plugin. Develop a web-based solution that manages, operates and monitors our distributed storage. Work closely with other teams to define and implement new APIs. What We Need to See: B.Sc., M.Sc. or Ph.D. in Computer Science, or related discipline, or equivalent experience. 8+ years of experience in web development ( both client and server ) Proven experience with Kubernetes (K8s), including developing or maintaining operators and/or CSI plugins. Experience scripting with Python, Bash or similar. Experience with nodejs is a must At least 5 years of experience working in a Linux OS environment You’re smart and a quick learner You do what it takes to get the job done Passionate about coding and big challenges Ways to stand out from the crowd: NodeJS for the server side: dominant modules are async & express . Kafka, MongoDB, K8s JavaScript frameworks: React, jQuery, c3j

JavaScriptPythonReactNode.js
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

PythonAWSAzureGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i

PythonReactNode.jsAngular
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training. Our mission is to make some of the world's largest GPU clusters work reliably for frontier training. We bring new platforms and clusters online, safely maintain installed fleets, and partner with hardware, infrastructure, and research teams to resolve the system-level issues that keep jobs from running. That means building and maintaining the software closest to the machine: Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation. We make these components reproducible, compatible, and safe to operate across heterogeneous fleets. About the Role We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate. You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models

AWSLinuxRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a software engineer on the Scaling team, you’ll help build and optimize the low-level stack that orchestrates computation and data movement across OpenAI’s supercomputing clusters. Your work will involve designing high-performance runtimes, building custom kernels, contributing to compiler infrastructure, and developing scalable simulation systems to validate and optimize distributed training workloads. You will work at the intersection of systems programming, ML infrastructure, and high-performance computing, helping to create both ergonomic developer APIs and highly efficient runtime systems. This means balancing ease of use and introspection with the need for stability and performance on our evolving hardware fleet. This role is based in San Francisco, CA, with a hybrid work model (3 days/week in-office). Relocation assistance is available. In this role, you will: Design and build APIs and runtime components to orchestrate computation and data movement across heterogeneous ML workloads. Contribute to compiler infrastructure, including the development of optimizations and compiler passes to support evolving hardware. Engineer and optimize compute and data kernels, ensuring correctness, high performance, and portability across simulation and production environments. Profile and optimize system bottlenecks, especially around I/O, memory hierarchy, and interconnects, at both local and distributed scales. Develop simulation infrastructure to validate runtime b

PythonAWSRestAI
O
📍 United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Security Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments—GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Security Software Engineer, you will design and build critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Architect and implement production-grade security services (e.g., auth services, access brokers, secure proxies, key-management infrastructure) that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD. Partner with infrastructure and research engineers to embed security into high-performance compute clusters, enabling rapid model training and deployment without compromising protection. Develop automation and detection tooling to continuously identif

PythonAWSAzureGCP
O
📍 United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. In this role, you will: Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Collaborate with engineering and security teams to drive deployment of security enhancements and control changes across broad-scale infrastructure. Tackle high-impact projects such as checkpoint encryption, network isolation, secret management, and machine identity, while continuously raising the security bar for emerging AI workloads. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets to adapt to evolving challenges. You will thrive in this role if you have: Deep understanding of security principles, best practices, and common vulnerabilities. A proactive mindset, with the ability to identify and address secu

AWSAzureKubernetesCI/CD
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. This role is based in San Francisco, CA. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change This role is exclusively based in our San Francisco HQ. We offer relocation assistance to new employees. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely depl

AWSKubernetesRestAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%

From $230K/yr

Quick readStrong listing-quality and freshness signals

About the Role The Engineering Acceleration Delivery / Continuous Deployment team builds and operates the systems that safely ship OpenAI’s infrastructure and product code to production. We own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across OpenAI to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale infrastructure orchestration. In This Role, You Will Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces, and alerts to detect regressions and trigger safe rollbacks. Build systems that support agent-assisted or autonomous deployment workflows using modern AI tooling. Technologies commonly used in this environment include: Kubernetes for large-scale container orchestration and runtime infrastructure Python and FastAPI for internal services Terraform for infrastructure as code GitOps-based deployment workflows (e.g., ArgoCD, Flux, or similar systems) Buildkite for CI orchestration You may be a strong fit if you: Have worked with Kubernetes-based deployment systems at scale Have experience building or operating continuous deployment platforms Are familiar with GitOps tooling such as

PythonAWSKubernetesGit
🔔

Get new cluster lead facilities services jobs in United States by email

Daily job updates · Unsubscribe anytime