Jobiba hiring network

Cluster Hr Head Jobs

315 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cluster hr head jobs. Use filters to narrow by work mode, employment type, experience and date posted.

DU
DoorDash USA
📍 San Francisco• Full-time• From $102K/yr
14 days ago

About the Team The DoorDash Research Fellowship is a 3-month program (extendable to 6 months) looking for Summer and Fall 2026 cohorts, for researchers and engineers who want to work on the hardest applied ML and AI problems in local commerce. Fellows are given the resources, autonomy, and access to real-world operational data needed to pursue ambitious research directions — with the goal of producing work that influences both the field and how DoorDash operates at scale. This program is modeled on the best external research fellowships: fellows are treated as independent researchers, not as junior employees on a product team. You pick the problem (within a set of priority areas), you own the direction, and you publish or ship the outcome. You’re excited about this opportunity because you will receive… Dedicated compute allocation sized to the research agenda — GPU clusters for training and inference budgets for experimentation Full access to DoorDash's research infrastructure — our internal RL stack, training and evaluation pipelines, RL environments built on real operational systems, agent evaluation harnesses, and the tooling our own research teams use day-to-day. Fellows are first-class users, not sandboxed visitors. Access to DoorDash operational data — real-world datasets spanning logistics, merchant operations, consumer behavior, and marketplace dynamics, under appropriate data governance Research mentorship from senior researchers and engineering leaders at DoorDash, plus a named research sponsor for each fellow who meets with you weekly and is accountable for unblocking your work Speaker series featuring leading researchers and practitioners from academia and industry — faculty from top ML programs, research leads from frontier AI labs, and senior operators from across tech. Fellows get dedicated 1:1 time with speakers when possible. A cohort of fellows working alongside you — a small, tight-knit group of researchers tackling different problems but sharing

gitrestai
View job →
W
27 days ago

🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Join WRITER's security team as a staff detection and response engineer and help protect the AI infrastructure that's transforming how the world works. You'll build sophisticated detection systems that identify attacks targeting our AI platform, training data, and model deployments while creating automated response capabilities that scale with our explosive growth. This isn't just traditional security work – you're defending cutting-edge AI/AGI systems against adversaries who are evolving their tactics as fast as AI itself advances. This role combines hands-on security engineering with strategic thinking to stay ahead of novel threats that don't exist in textbooks yet. You'll be the operational arm of our security function, translating threat intelligence into real-time detections, coordinating incident response across multiple teams, and hunting for sophisticated attacks across GPU clusters and distributed training environments. If you're excited by the challen

REMOTEpythonaigo
View job →
W
Writer
📍 San Francisco• Full-time• Remote
27 days ago

🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Join WRITER's security team as a staff detection and response engineer and help protect the AI infrastructure that's transforming how the world works. You'll build sophisticated detection systems that identify attacks targeting our AI platform, training data, and model deployments while creating automated response capabilities that scale with our explosive growth. This isn't just traditional security work – you're defending cutting-edge AI/AGI systems against adversaries who are evolving their tactics as fast as AI itself advances. This role combines hands-on security engineering with strategic thinking to stay ahead of novel threats that don't exist in textbooks yet. You'll be the operational arm of our security function, translating threat intelligence into real-time detections, coordinating incident response across multiple teams, and hunting for sophisticated attacks across GPU clusters and distributed training environments. If you're excited by the challen

REMOTEpythonrestai
View job →
E
Everpure
📍 Bengaluru• Full-time
14 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Pure Solutions team as a Senior MLOps Solutions Engineer to architect and build high-scale, enterprise-grade AI/ML solutions. You will be instrumental in integrating Pure Storage platforms with the evolving open-source MLOps ecosystem (Kubeflow, MLflow, Ray) to operationalize the complete machine learning lifecycle. This role requires a creative technologist with deep Python expertise to drive innovation and enable our customers and partners to achieve production AI success. WHAT YOU'LL DO Design and Automate MLOps Pipelines: Lead the development of end-to-end MLOps workflows using CI/CD tools (Git/Jenkins) and orchestration platforms (MLflow/Kubeflow), specifically integrating Pure Storage's FlashBlade, FlashArray, and Portworx as the high-performance data plane for data ingestion, training, and inference. Build High-Performance AI/ML Reference Architectures: Create validated, repeatable deployment models using Infrastructure as Code (e.g., Ansible, Terraform) for AI/ML environments spanning bare metal, virtual machines, and GPU-accelerated Kubernetes clusters, ensuring optimal performance for distributed training. Optimize and Operationalize GPU Inference: Architect and implement solutions for high-throughput, low-latency model serving, utilizing technologies like NVIDIA Triton Inference Server and advanced optimization techniques (quantization, model sharding like DeepSpeed/Megatron-LM, and dynamic bat

pythonawskubernetes
View job →
B
Baseten
📍 San Francisco• Full-time• Remote
21 days ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Manager focused on TPUs at Baseten, you will lead the "engine room" for our non-NVIDIA accelerator fleet, architecting, securing, and optimizing the Google Cloud TPU (and broader emerging accelerator) capacity that powers our customers' AI workloads. You'll own the end-to-end journey of capacity management for this fleet, from securing large-scale TPU pod allocations to building the automation that ensures reliable uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering, with a specific focus on the TPU ecosystem. You will act as the fleet orchestrator for Google's TPU architecture, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics as we diversify beyond NVIDIA. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the latest generation of TPU hardware, like Google's Trillium (v6e) architecture, and partnering closely with the Model Performance (MP) team to ensure workloads are tuned for TPU-specific execution. EXAMPLE INITIATIVES The TPU Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's TPU clusters, including pod slicing and topology planning Global Workload Orchestration: Bui

REMOTEpythonawsazure
View job →
B
Baseten
📍 San Francisco• Full-time• Remote
23 days ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Lead at Baseten, you will lead the "engine room" of the company, architecting, securing, and optimizing the global GPU fleet that powers our customers' AI workloads. You’ll own the end-to-end journey of capacity management, from securing multi-million dollar GPU clusters to building the automation that ensures 99.9% uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering. You will act as the fleet orchestrator for the world's most advanced chips, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the next generation of hardware, like NVIDIA’s Blackwell (B200) architecture. EXAMPLE INITIATIVES The B200 Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's first Blackwell GPU clusters. Global Workload Orchestration: Building "Multi-cloud Capacity Management" systems to move customer workloads seamlessly across regions to optimize cost and latency. Precision GPU Triage: Developing automated Go-based operators to identify, cordon, and repair unhealthy H100 nodes in under an hour. The Supply Chain of Intelligence: Partnering with lead

REMOTEpythonawsazure
View job →
P
Pagerduty
📍 Atlanta• Full-time• From $98K/yr
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office, you'll help build and operate the foundational infrastructure that powers PagerDuty's real-time digital operations platform. Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You'll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases. Key Responsibilities ● Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems. ● Contribute to the reliability and scalability of PagerDuty's core platform by hardening existing systems and supporting the rollout of new infrastructure

pythonawsazure
View job →
P
Pagerduty
📍 Atlanta• Full-time• $113K – $171.6K/yr
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. As a Site Reliability Engineer II on the Core Infrastructure team in our Atlanta office, you'll help build and operate the foundational infrastructure that powers PagerDuty's real-time digital operations platform. Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You'll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases. Key Responsibilities ● Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems. ● Contribute to the reliability and scalability of PagerDuty's core platform by hardening existing systems and supporting the rollout of new infrastructur

pythonawsazure
View job →

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As the Engineering Manager, GitLab Delivery - Operate , you’ll guide a globally distributed team focused on making it easier for customers to deploy, upgrade, and run GitLab reliably in their own infrastructure. You’ll help shape the systems and tooling that support environments ranging from single-node virtual machines to large Kubernetes clusters, with a focus on reliability , operational simplicity , upgrade velocity , and zero-downtime capabilities across GitLab.com , GitLab Dedicated , and self-managed deployments. In this role, you’ll partner closely with a Product Manager and work across Infrastruc

kubernetesgitrest
View job →
M
10 days ago

About the Team Being part of Meesho's Fulfilment and Experience (F&E) team as Cluster Head LM, will zip you to the cockpit of our ever-burgeoning rocketship, where you get to directly shape the experience of the country's next billion e-commerce users. We are an eclectic mix of 100+ professionals with diverse skill sets ranging from running operations/support, supply chain know-how, analytics and the holy grail, first principles problem-solving. At Meesho, we’re trying to do what's never been done before - taking e-commerce to the masses. This leaves us with no choice but to completely reimagine logistics from the ground up, to cater to our customers' price and delivery expectations. That means a host of "zero-to-one" projects (takers, anyone?) to build a supply chain to change how folks think about e-commerce in India and globally. We are firm believers in fun at work. With monthly F&E happy hour sessions, informal team outings, and internal virtual water cooler chat sessions, there’s never a dull moment with us About the Role As Cluster Head LM, you’ll own the onboarding and training of partners. You’ll also drive key operational metrics by regularly visiting their facilities in different cities in your area. You’ll take complete ownership of processes allotted to you and work with various stakeholders to achieve team goals. You’ll continuously work towards identifying gaps and providing recommendations for improving our processes. What you will do Own the onboarding and training of new partners for Last Mile operations in your cluster. Identify and onboard new partners onto the network on an ongoing basis Track and own the performance of different partners in your cluster Visit facilities to conduct audits and solve operational gaps Ensure compliance with operational processes Own and drive critical operational metrics end to end and achieve performance targets Continuously work towards ide

supply chainlogistics
View job →
M
10 days ago

About the Team Being part of Meesho's Fulfilment and Experience (F&E) team as Cluster Head LM, will zip you to the cockpit of our ever-burgeoning rocketship, where you get to directly shape the experience of the country's next billion e-commerce users. We are an eclectic mix of 100+ professionals with diverse skill sets ranging from running operations/support, supply chain know-how, analytics and the holy grail, first principles problem-solving. At Meesho, we’re trying to do what's never been done before - taking e-commerce to the masses. This leaves us with no choice but to completely reimagine logistics from the ground up, to cater to our customers' price and delivery expectations. That means a host of "zero-to-one" projects (takers, anyone?) to build a supply chain to change how folks think about e-commerce in India and globally. We are firm believers in fun at work. With monthly F&E happy hour sessions, informal team outings, and internal virtual water cooler chat sessions, there’s never a dull moment with us About the Role As Cluster Head LM, you’ll own the onboarding and training of partners. You’ll also drive key operational metrics by regularly visiting their facilities in different cities in your area. You’ll take complete ownership of processes allotted to you and work with various stakeholders to achieve team goals. You’ll continuously work towards identifying gaps and providing recommendations for improving our processes. What you will do Own the onboarding and training of new partners for Last Mile operations in your cluster. Identify and onboard new partners onto the network on an ongoing basis Track and own the performance of different partners in your cluster Visit facilities to conduct audits and solve operational gaps Ensure compliance with operational processes Own and drive critical operational metrics end to end and achieve performance targets Continuously work towards identi

supply chainlogisticsrecruitment
View job →
PE
1mo ago

About the Team Being part of Meesho's Fulfilment and Experience (F&E) team as Cluster Head LM, will zip you to the cockpit of our ever-burgeoning rocketship, where you get to directly shape the experience of the country's next billion e-commerce users. We are an eclectic mix of 100+ professionals with diverse skill sets ranging from running operations/support, supply chain know-how, analytics and the holy grail, first principles problem-solving. At Meesho, we’re trying to do what's never been done before - taking e-commerce to the masses. This leaves us with no choice but to completely reimagine logistics from the ground up, to cater to our customers' price and delivery expectations. That means a host of "zero-to-one" projects (takers, anyone?) to build a supply chain to change how folks think about e-commerce in India and globally. We are firm believers in fun at work. With monthly F&E happy hour sessions, informal team outings, and internal virtual water cooler chat sessions, there’s never a dull moment with us About the Role As Cluster Head LM, you’ll own the onboarding and training of partners. You’ll also drive key operational metrics by regularly visiting their facilities in different cities in your area. You’ll take complete ownership of processes allotted to you and work with various stakeholders to achieve team goals. You’ll continuously work towards identifying gaps and providing recommendations for improving our processes.

aigosupply chain
View job →
AG
1mo ago

Regional Security Head is responsible for fostering a culture of security awareness and operational excellence across assigned geographic regions. This role entails overseeing all security functions at individual sites and clusters, ensuring compliance with security protocols and regulations. By driving strategic initiatives, mentoring security personnel, and enhancing stakeholder collaboration, the Regional Security Head aims to safeguard organizational assets, mitigate risks, and ensure business continuity. Source: Adani Group | Job ID: 55782

🔔

Get new cluster hr head jobs by email

Daily job updates · Unsubscribe anytime