Jobs in United States

Cluster Hr Head in San Francisco

45 active opportunities · Updated October 2026

Explore current cluster hr head jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

C
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? The Data Infrastructure team at Cohere is responsible for the storage and data movement layer underlying every model training run. We're building the unified storage layer that feeds our training workloads. It needs to serve petabytes of training data and model checkpoints fast enough to keep thousands of GPUs busy across several training clusters. In this role, you’d have an opportunity to build this system from the ground up. You’d be a key contributor, working on a problem few teams have had to solve at this scale. In this role, you will: Design, build, and operate the distributed storage system that feeds model training and evaluation. Run this system multiple on Kubernetes clusters at petabyte scale. Work with researchers and training-infra teams on how jobs actually read and write data, and turn that into throughput, latency, and durability requirements Work through the networking, I/O, and consistency problems of moving large datasets and checkpoints across regions and backends, with GPU idle time and time-to-insight as the measures of success You may be a good fit if you have: Strong storage fundamentals,

PythonKubernetesGitRest
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI’s AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production. In this role, you will: Design and implement the LLM inference runtime for frontier models running on custom silicon. Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization. Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads. Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack. Enable new

PythonAWSRestAI
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.1%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments. We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment. This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment.. RESPONSIBILITIES Core Responsibilities: Drive suppliers to keep the maximum amount of the GPU fleet online and healthy. Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs. Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier. SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short. Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability. Scope and

Machine LearningAIGoFinance
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.1%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're hiring a Product Data Scientist to establish how product decisions at Baseten are made with data. You'll work directly with Product and Engineering, alongside GTM to determine measurement, strategy, experimentation and implementation. This is a foundational, hands-on role. You'll define what success looks like across a technical, usage-based platform and turn ambiguous questions into analyses, forecasts, and experiments that shape product strategy. You'll work from clickstream and product events through inference telemetry and observability data, helping Baseten make faster decisions about reliability, performance, adoption and developer experience. RESPONSIBILITIES Partner directly with Product and Engineering: frame the questions that matter, define success criteria, and turn analysis into roadmap, launch, and prioritization decisions. Define how product success is measured: establish metrics across activation, adoption, retention, expansion, reliability and user experience. Support experimentation and launches: design measurement plans, analyze A/B experiments and controlled rollouts, and translate results into product decisions. Diagnose reliability and scaling behavior: join customer signals with request, replica, deployment, and cluster telemetry to find patterns in release bottlenecks, unhealthy replicas, and models without traffic. Define the enterprise customer journey and measure feature adoption

PythonSQLMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Core Services organization builds and runs the mission-critical online services that product teams rely on in production. We own foundational distributed systems and platform capabilities that enable reliable execution, high-performance services, and large-scale file/data needs across our products. This team is distinct from developer infrastructure and data infrastructure—our focus is production service foundations and core runtime services. About the Role We’re hiring an Engineering Manager, Core Services to help lead teams responsible for highly reliable, high-scale distributed systems that sit on the critical path for OpenAI products. Your team will own foundational production systems that OpenAI’s product engineering teams build on. You’ll collaborate closely with product and infrastructure partners to ship reliable services quickly, and help scale systems and teams as OpenAI grows. You’ll partner closely with senior engineering leaders to scale the org, mature operations, and drive major platform initiatives. This role requires strong technical ability. You’ll be responsible for: Managing and growing a high-performing team of infrastructure engineers. Leading teams building and operating large, critical production platforms, including cluster reliability, scaling, and rollout safety. Building and operating mission-critical distributed systems with strong operational rigor (SLOs, incident response, capacity planning, reliability). Setting technical direction for platform foundations such as workflow/orchestration capabilities, large-scale file/blob/storage services, and core service foundations. Partnering with a broad set of stakeholders, including product engineering, adjacent infrastructure teams, and (where relevant) finance/cost partners. Coaching, mentoring, and developing engineers and emerging leaders. You might thrive in this role if you: Have significant experience leading teams that run mission-critical infrastructure in production

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI's Industrial Compute organization builds and operates the infrastructure required to train and serve frontier AI models. The Capacity Planning team connects rapidly changing research and product demand with the compute, networking, storage, power, data center, hardware, and operational resources required to make that demand executable. About the Role We are seeking a Technical Program Manager to build and lead capacity planning across OpenAI's large-scale AI infrastructure. You will translate uncertain workload demand into clear infrastructure requirements, allocation decisions, supply commitments, activation priorities, and long-range capacity strategies. This role sits at the intersection of research, engineering, infrastructure, finance, sourcing, deployment, and operations. You will create the planning models, operating cadences, governance mechanisms, and source-of-truth systems that allow teams to understand what capacity is required, what is available, what is at risk, and what decisions must be made. This is not a finance-only forecasting or reporting role. Success requires technical fluency across the infrastructure stack, strong analytical judgment, and the ability to move consequential decisions forward when requirements, timelines, and supply conditions change quickly. Key Responsibilities Own capacity-planning processes across near-term workload allocation, quarterly execution, and longer-range infrastructure horizons. Translate research, training, inference, and product demand into compute, accelerator, cluster, networking, storage, rack, power, and site requirements. Develop scenarios that make assumptions, confidence levels, constraints, sensitivities, and decision points explicit. Reconcile requested demand against contracted, delivered, installed, activated, and workload-usable capacity. Partner with research and engineering teams to understand workload priorities, technical dependencies, utilization patterns, and changing req

PythonSQLAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,

AWSKubernetesLinuxRest
W
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend +8.1%
Quick readStrong listing-quality and freshness signals

🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Join WRITER's security team as a staff detection and response engineer and help protect the AI infrastructure that's transforming how the world works. You'll build sophisticated detection systems that identify attacks targeting our AI platform, training data, and model deployments while creating automated response capabilities that scale with our explosive growth. This isn't just traditional security work – you're defending cutting-edge AI/AGI systems against adversaries who are evolving their tactics as fast as AI itself advances. This role combines hands-on security engineering with strategic thinking to stay ahead of novel threats that don't exist in textbooks yet. You'll be the operational arm of our security function, translating threat intelligence into real-time detections, coordinating incident response across multiple teams, and hunting for sophisticated attacks across GPU clusters and distributed training environments. If you're excited by the challen

PythonRestAIGo
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.1%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Manager focused on TPUs at Baseten, you will lead the "engine room" for our non-NVIDIA accelerator fleet, architecting, securing, and optimizing the Google Cloud TPU (and broader emerging accelerator) capacity that powers our customers' AI workloads. You'll own the end-to-end journey of capacity management for this fleet, from securing large-scale TPU pod allocations to building the automation that ensures reliable uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering, with a specific focus on the TPU ecosystem. You will act as the fleet orchestrator for Google's TPU architecture, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics as we diversify beyond NVIDIA. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the latest generation of TPU hardware, like Google's Trillium (v6e) architecture, and partnering closely with the Model Performance (MP) team to ensure workloads are tuned for TPU-specific execution. EXAMPLE INITIATIVES The TPU Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's TPU clusters, including pod slicing and topology planning Global Workload Orchestration: Bui

PythonAWSAzureGCP
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.1%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Lead at Baseten, you will lead the "engine room" of the company, architecting, securing, and optimizing the global GPU fleet that powers our customers' AI workloads. You’ll own the end-to-end journey of capacity management, from securing multi-million dollar GPU clusters to building the automation that ensures 99.9% uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering. You will act as the fleet orchestrator for the world's most advanced chips, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the next generation of hardware, like NVIDIA’s Blackwell (B200) architecture. EXAMPLE INITIATIVES The B200 Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's first Blackwell GPU clusters. Global Workload Orchestration: Building "Multi-cloud Capacity Management" systems to move customer workloads seamlessly across regions to optimize cost and latency. Precision GPU Triage: Developing automated Go-based operators to identify, cordon, and repair unhealthy H100 nodes in under an hour. The Supply Chain of Intelligence: Partnering with lead

PythonAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team This team builds and operates the systems that enable OpenAI researchers to run reliable, scalable, and efficient research workflows. The team sits close to research and works across infrastructure, systems, and automation to make sure researchers have the tools and environments they need to move quickly. The work spans software engineering, infrastructure, systems administration, cluster operations, and reliability engineering. As OpenAI’s infrastructure evolves from bespoke bare-metal systems toward more standard, scalable platforms, the team needs engineers who can understand how systems work end-to-end and build the right abstractions without reinventing the wheel. About the Role As a Software Engineer on this team, you will build and operate the infrastructure that supports frontier research and critical research-facing systems. You will work on systems that sit close to the metal, but the role is not limited to classic operations or sysadmin work. We are looking for someone who can reason about networking, bootstrapping, Kubernetes, scalability, automation, and reliability - while also writing software to make these systems better over time. This role is a strong fit for an independent, high-ownership engineer who enjoys reliability-heavy infrastructure work but still wants to build. You do not need to come in as a kernel expert or highly algorithmic optimization engineer, but you should be deeply curious about infrastructure, comfortable debugging complex systems, and excited to support researchers doing novel work. We expect you to: Build and operate reliable infrastructure for research workloads and research-facing services. Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services. Improve cluster bootstrapping, provisioning, automation, and deployment workflows. Debug issues across networking, compute, storage, orchestration, and service reliability layers.

AWSKubernetesCI/CDGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea

AWSAzureKubernetesCI/CD
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

PythonAWSLinuxRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops system and infrastructure solutions tailored to the demands of advanced AI workloads. We work across the full stack—from silicon to system integration—partnering closely with internal teams and external vendors to define and deliver next-generation AI infrastructure. Our team focuses on defining scalable, high-performance system architectures and reference designs that balance performance, cost, and operational efficiency across rapidly evolving technologies. About the Role We are seeking a 3P Architect to define and drive rack- and cluster-level reference designs in collaboration with external partners. This role is responsible for translating workload requirements and system-level goals into concrete architectures, aligning partners on critical design attributes, and ensuring vendor roadmaps meet our infrastructure needs. You will work closely with performance modeling and internal architecture teams to evaluate tradeoffs, while owning the end-to-end definition and execution of third-party system designs. This includes identifying gaps in current technologies, driving vendor development, and shaping future infrastructure capabilities. This role requires strong system intuition, cross-functional leadership, and the ability to operate effectively across internal teams and external ecosystems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Define rack- and cluster-level reference architectures for AI infrastructure deployments. Translate workload requirements into clear system design specifications and partner deliverables. Collaborate with performance modeling teams to evaluate architectural tradeoffs and system behaviors. Align internal stakeholders and external partners on critical system attributes (performance, cost, power, reliability, scalability). Identify gaps in current technology offerings and dr

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t

AWSKubernetesRestAI
🔔

Get new cluster hr head jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime