Jobiba hiring network

Cluster Hr Head Jobs

315 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cluster hr head jobs. Use filters to narrow by work mode, employment type, experience and date posted.

M
Mongodb
📍 Ireland• Full-time
1mo ago

We are seeking a Staff engineer to design, build, and operate the internal and external Observability stack for the MongoDB platform. Tens of thousands of customers depend on our Observability stack to monitor their database clusters and to generate actionable alerts to safeguard critical workloads. This is an opportunity to join a team that is responsible for all Observability systems that support metrics, metric visualization, logs, traces, and alerts for MongoDB. We are looking for engineers with high standards, and experience in setting direction and technical leadership for large engineering teams in designing and operating complex distributed systems, with strict SLO on security, durability, availability and performance. As MongoDB Atlas and its supporting infrastructure continue to experience rapid growth, the demand for high-cardinality observability data for internal and external use cases means we need to continually innovate and scale our systems to the next level. For example, MongoDB Observability systems need to handle 10’s of billions of metrics time series, all whilst processing petabytes of logs, traces, and events. Our stack includes VictoriaMetrics, Splunk, Flink, WarpStream/Kafka, Java, Golang Fluentbit. In addition to owning critical components of our observability infrastructure, as a Staff engineer on the team, you’ll also work closely with other SWE, Product and SRE teams to promote and implement best practices in instrumenting and monitoring their services. This is a highly collaborative role, and you will get to own some of the most relied upon internal infrastructure at Mongo. Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to be a deeply technical leader on a collaborative team that applies low-level systems expertise to build the foundational infrastructure of a popular database, join us! Let’s build a faster, more reliable, and exceptionally observable database system together. W

javamongodbaws
View job →
M
Mongodb
📍 Ireland• Full-time
1mo ago

We are seeking a Staff engineer to design, build, and operate the internal and external Observability stack for the MongoDB platform. Tens of thousands of customers depend on our Observability stack to monitor their database clusters and to generate actionable alerts to safeguard critical workloads. The Collections team is a newly formed team within MongoDB's Observability & Adoption Org focused on making telemetry onboarding and collection significantly easier across MongoDB. We own key parts of the observability collection stack, including onboarding experience, telemetry collection agents across the data and control planes, and ingestion services for metrics, logs, and traces that support both internal and customer observability in MongoDB, driving insights, recommendations, and alerting. Our mission is to reduce friction for teams implementing and iterating on Observability while partnering closely with development teams to instrument their services using shared best practices, helping define the conventions our telemetry should follow, and building collection and ingestion systems that are stable, performant, secure, well-documented, and self-service. We also work closely with the Data Pipeline and Storage & Query teams to help ensure MongoDB has a stable and performant observability stack end to end. This is an opportunity to join a team shaping how observability works across MongoDB and to have outsized impact on both the developer experience and the reliability of the platform underneath it. As MongoDB Atlas and its supporting infrastructure continue to experience rapid growth, the demand for high-cardinality observability data for internal and external use cases means we need to continually innovate and scale our systems to the next level. For example, MongoDB Observability systems need to handle 10’s of billions of metrics time series, all whilst processing petabytes of logs, traces, and events. Our stack includes VictoriaMetrics, Grafana, Sp

javamongodbaws
View job →
M
1mo ago

The MongoDB Customer Observability Team is a diverse group of contributors working together to help our users manage MongoDB at global scale. The team is responsible for MongoDB Atlas: our database-as-a-service offering and fastest-growing product, which allows users to deploy globally distributed MongoDB clusters in just minutes. We're seeking a Senior Software Engineer to join our team to tackle exciting challenges within the Observability space. You'll contribute to developing tools and platforms that help our users understand the health and performance of their MongoDB deployments. This includes collecting metrics, monitoring slow queries, and offering actionable insights such as index and schema suggestions that improve the speed, efficiency, and overall reliability of their databases. This role provides a unique opportunity to drive engineering excellence across both dimensions of observability, leveraging technologies being developed within the Customer Observability group and contributing directly to the success of our customers and our product teams. If you're passionate about large-scale systems, digging deep into telemetry data, and building tools that make a real impact both internally and externally, we’d love to have you on board! We are looking to speak to candidates who are based in Dublin for our hybrid working model. We're looking for someone who Has at least 5 years of experience as a backend or full stack engineer Enjoys collaboration and being part of a team Is approachable, curious, and intellectually honest Is a backend engineer with a willingness to take on frontend tasks or a full-stack developer with a bias towards backend Has written backend systems in a compiled language (Java, C#, Go, etc.) Has experience with the design and architecture of a modern, scalable web application Enjoys chasing down difficult problems in a distributed environment and on an database diagnostic/operation level Always strives to expand their knowledg

typescriptjavareact
View job →
M
Mongodb
📍 Atlanta• Full-time
1mo ago

The MongoDB Database Experience team is a diverse group of contributors working together to help our users manage MongoDB at global scale. This includes building tools for MongoDB Atlas: our database as a service offering and fastest growing product which allows users to deploy fault-tolerant, globally distributed MongoDB clusters in just minutes. We're seeking a Senior Software Engineer to join our Developer Tools Team within the Database Experience department. The team is responsible for tools which allow users to interact with their data and to understand the health and performance of their MongoDB deployments. The Developer Tools Team works on a variety of APIs and customer-facing UIs empowering MongoDB users to interact with their stored data, whether that’s on desktop via MongoDB Compass or in the browser via Atlas Data Explorer. This role is fully remote for a candidate based in the United States with US citizenship. We're looking for someone who Is a full stack engineer with a willingness to take on both frontend tasks and backend tasks Has experience with the design and architecture of a modern, scalable, high availability web application Has experience working with databases Enjoys collaboration and being part of a team Is approachable, curious, and intellectually honest Would enjoy chasing down difficult problems in a distributed environment Always strives to expand their knowledge Nice to Haves Knowledge of database internals and tuning mechanisms, particularly indexing Previous work in TypeScript, React, and Node.js, particularly maintaining an ecosystem of dependencies Experience working with websocket based client to server communication Familiarity with developing and supporting microservice based architectures using Kubernetes About MongoDB MongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined the data platform for the AI era, enabling builders to create, transform, and d

typescriptreactnode.js
View job →
S
Squarespace
📍 Dublin• Full-time• €93K – €143K/yr
1mo ago

Squarespace provides innovative solutions to empower our customers to focus on building their brand and growing their businesses on our platform. The Databases team manages all of the backend infrastructure that Squarespace runs on – MongoDB, CockroachDB, and Kafka clusters, to name a few examples. We are an accomplished, diverse group of people who develop the services that guarantee reliable and scalable infrastructure for both our cross-functional partners in product engineering, as well as our end users on the Squarespace platform. We believe that infrastructure excellence doesn't stop at just building for today; it needs to have a solid foundation of scalability, reliability, and a robust developer experience for the future. This is a hybrid role working from our Dublin office 3 days per week. You will report to the Databases Senior Engineering Manager. You’ll Get To… Nurture high-performing software engineers by guiding navigation when there is ambiguity. Distill the scope of the team and help hire a balanced group of engineers that will excel as a unit. Grow the career development of direct reports through regular 1:1s with direct, actionable feedback. Celebrate wins that motivate the team’s positive culture and robust dynamic. Evaluate consistently to improve team efficiency and effectiveness when required. Evolve a deep understanding of local systems to identify appropriate architectural decisions. Thread with Product, Design & Engineering to champion, define and execute an optimal roadmap. Bond across Engineering, Product, Design, Marketing, Data Science and Business Operations. Who We’re Looking For 3+ years of recent experience managing a Product Engineering team of four or more engineers. 7+ years of industry experience deploying apps across large codebases with many contributors. Ability to fluently translate, document and present technical concepts to non-technical stakeholders. Strong technical foundations to navigate the inherent tra

mongodbaigo
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de

awsrestai
View job →
M
9 days ago

Enterprise Advanced is a distributed team across Europe and India that builds the software running MongoDB on any infrastructure, at global scale — from on-prem data centers to private cloud. You'll work primarily on Ops Manager and Automation, the systems that let customers deploy fault-tolerant, globally distributed MongoDB clusters in minutes. Our software manages some of the largest self-managed MongoDB deployments in the world, with production clusters running hundreds of shards and nodes under a single deployment. The main focus of this team is to adapt our software to manage MongoDB clusters which are deployed in data centers or private cloud platforms. You will work on the core functionality for all of our products, mainly on the Ops Manager , and Automation products. Our team's end users are some of the largest businesses in the world, deploying massive clusters and processing huge amounts of data. This role is based in our Gurgaon office, and can work in a hybrid fashion. This role will report to the Senior Engineering Manager also based in our Gurgaon office. What you’ll do Design, implement, test, and release features for Ops Manager Own end-to-end delivery of complex projects, from design through incremental shipping Troubleshoot and resolve issues surfaced in customer deployments running at scale Apply engineering judgment and MongoDB's core values across planning, design, and code review A great fit for this role will be You enjoy distributed-systems problems; consistency, fault tolerance, and scale are the daily reality, not edge cases People who like ambiguity and are comfortable defining their own approach with guidance, not step-by-step instruction You're flexible! You're willing to take on a wide variety of responsibilities, learning as you go You're a self-starter! You're comfortable organizing your own time, acting on feedback and prioritizing with guidance from senior members of your team Requirements 4+ years experience with a language

javascriptpythonjava
View job →

About the Role Together AI is building the AI Native Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art GPU cloud infrastructure. The Together Cloud team builds the [Together GPU Clusters](https://www.together.ai/gpu-clusters) flagship IaaS product that provides high-performance, AI-ready GPU clusters through a self-serve cloud console, along with the virtualized infrastructure layer powering Together's inference, RL, and fine-tuning products. As a Staff Software Engineer focusing on AI Compute in the Together Cloud org, you will set technical direction for and build major components of the next generation AI cloud platform – a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware: GB300s/VRs, BlueField DPUs, InfiniBand and dual/quad-plane RoCEv2 fabrics. That virtualized computing platform powers our own SaaS products – inference, RL, and fine-tuning – and serves external cloud customers through self-serve offerings such as on-demand/reserved Kubernetes/Slurm clusters, across dozens of data centers and hundreds of thousands of GPUs. This is an architect-and-build role. Fully automated bootstrapping of GPU data centers, high-performance virtualization of GPU compute and DC networking without compromising isolation or portability, and fault-tolerant decentralized control planes — you'll set the architecture for these across our global and in-DC services, and be a key owner of the hardest parts, in the code as well as the design. Your designs will span the IaaS layer of a greenfield Vera Rubin data center up to the global management plane that schedules capacity across all of them. At this level the job is as much leverage as code: the standards you set and the engineers you grow decide how fast the rest of Together Cloud ships. Responsibilities Own the GPU and network virtualization stack: the hypervisor, kernel, and SDN work that keeps

awsazuregcp
View job →

DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You'll Do: Deploy, configure, and maintain Kubernetes clusters for our microservices architecture. Utilize Git and Helm for version control and deployment management. Implement and manage monitoring solutions using Prometheus and Grafana. Work on continuous integration and continuous deployment (CI/CD) pipelines. Containerize applications using Docker and manage orchestration. Manage and optimize AWS services, including but not limited to EC2, S3, RDS, and AWS CDN. Maintain and optimize MySQL databases, Airflow, and Redis instances. Write automation scripts in Bash or Python for system administration tasks. Perform Linux administration tasks and troubleshoot system issues. Utilize Ansible and Terraform for configuration management and infrastructure as code. Demonstrate knowledge of networking and load-balancing principles. Collaborate with development teams to ensure applications meet reliability and performance standards. Who you are: Bachelor’s degree in engineering (CS / IT) or equivalent degree from a well-known Institute / University. 2+ years of experience in a Site Reliability Engineer role or similar. Proven experience with Kubernetes, Git, Helm, Prometheus, Grafana, CI/CD, Docker, and microservices architecture. Strong knowledge of AWS services, MySQL, Airflow, Redis, AWS CDN. Proficient in scripting languages such as Bash or Python. Hands-on experience with Linux administration. Familiarity with Ansible and Terraform fo

pythonmysqlredis
View job →
N
12 days ago

NVIDIA DGX Cloud is an AI Factory designed to power the next generation of AI and industrial-scale breakthroughs. As a Principal Engineer for Security Architecture, within our Security Engineering organization, you will own a core security domain of the AI factory: the architecture, the paved road that delivers it, and much of the code underneath. You will hold the security design bar across DGX Cloud from inside the teams doing the building, and this is a founding seat on a new team. Security Engineering is a new organization at DGX Cloud, accountable for the security outcome of the platform, and Security Architecture is the function inside it that holds the design bar. Security here is fleet horizontal and stack vertical, so your work will cross every DGX Cloud engineering organization: you will embed with the teams building GPU clusters, control planes, and services, join their designs as a participant rather than an approver, and leave behind systems in which an entire class of risk is no longer possible. There is no architecture review board here and no approval queue. You are a senior IC with deep security domain knowledge, and the security bar holds because you helped set it and then helped ship it. What You Will Be Doing: Own a Security Domain End to End: Take architectural ownership of a core domain of DGX Cloud security, from the design through the system running in production. That could be tenant and GPU workload isolation, workload identity, infrastructure and network, supply-chain provenance, hardened baselines and patching, or deploy-time policy and admission control. Embed with the Teams Building It: Join the design early, write the code, and help land it. The posture is not "you did this wrong." It is "here are the considerations we need to meet, I will help, let's go to work." Build Paved Roads, Not

REMOTEkuberneteslinuxartificial intelligence
View job →
N
12 days ago

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated

pythonkuberneteslinux
View job →

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity AI is a top strategic priority for New Relic, and the Bengaluru design team is at the center of it. The team works across three closely related product clusters: autonomous incident response (SRE Agent, Autopilot, and Intelligent RCA), the intelligence and platform layer that powers them (Ground Truth, Agentic Platform, and New Relic AI), and AIOps for event correlation and incident management. These products share a common design challenge: users need to trust systems that act autonomously, and building that trust through good design is genuinely hard work. This is an on-site role in Bengaluru. Your designers are there, and many of your engineering and product partners are too. Being present — in standups, reviews, and the quick conversations before a decision gets made — is part of how you'll build the relationships that make design effective. You'll also collaborate with design, product, and engineering partners in the US and Spain, so operating across time zones and communicating well in writing are part of the job. You'll manage a small team of designers and own the quality of the work across these products. The design questions here don't have established answers — how do you make an autonomous system legible? How do you build user trust in AI-generated root cause analysis? How do you design a handoff from machine decision to human judgment? If you're already working in AI product design, or actively building toward it, and you want to lead a team

airecruitment
View job →
O
Okta
📍 Bellevue• Full-time• From $194K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities—like reducing costs, and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimizing costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimize AWS cloud infrastructure, including EKS,ECS, S3, VPCs, RDS, IAM, and more. Implement b

pythonawsdocker
View job →
P
1mo ago

Role Summary This role serves as the single point of accountability between USMAPPS and the Specialty Care Business Unit, owning access strategy, payer marketing, and brand-level contracting and pricing strategy across the full Specialty Care portfolio. The role drives and owns outcomes, integrating all USMAPPS capabilities into one coherent access plan that directly supports the Specialty Care BU President and franchise leads. The VP is expected to operate at full strategic weight, lead a team aligned to Specialty Care brands and franchise clusters, and be the single person the Specialty Care BU holds accountable for market access results. The Vice President US Market Access Lead – Specialty Care reports directly to the Senior Vice President of US Market Access & Pfizer Patient Services (USMAPPS) with a dotted line to the Specialty Care BU President. This position requires close partnership with the Specialty Care BU President, franchise leads, Strategic Contracting & Analytics, Strategic Account Management, the Patient Services, and USMAPPS leadership. The role sits on the Specialty Care BU leadership team and on the USMAPPS Leadership Team. Role Responsibilities 1. Specialty Care Access Strategy Ownership </spa

airecruitment
View job →
P
Pfizer
📍 New York• $300.1K – $500.1K/yr
1mo ago

Role Summary This role serves as the single point of accountability between USMAPPS and the Primary Care Business Unit, owning access strategy, payer marketing, and brand-level contracting and pricing strategy across the full Primary Care portfolio. The role drives and owns outcomes, integrating all USMAPPS capabilities into one coherent access plan that directly supports the Primary Care BU President and franchise leads. The VP is expected to operate at full strategic weight, lead a team aligned to Primary Care brands and franchise clusters, and be the single person accountable the Primary Care BU holds accountable for market access results. The Vice President, US Market Access Lead – Primary Care reports directly to the Senior Vice President of US Market Access & Pfizer Patient Services (USMAPPS) with a dotted line to the Primary Care BU President . This position requires close partnership with the Primary Care BU President, franchise leads, Strategic Contracting & Analytics, Strategic Account Management, the Patient Services, and USMAPPS leadership. The role sits on the Primary Care BU leadership team and on the USMAPPS Leadership Team. Role Responsibilities 1. Primary Care Access Strategy Ownership </

airecruitment
View job →
🔔

Get new cluster hr head jobs by email

Daily job updates · Unsubscribe anytime