Join the MongoDB Networking & Observability team and help build the core of a distributed database! Our team focuses on creating and enhancing components which facilitate communication between distributed processes and make these processes, and their communication, easily observable. Networking Observability’s responsibilities include improving MongoDB networking, improving the efficiency of resource utilization, and building low-overhead observability features. Our team includes engineers located in New York City and fully remote engineers. We operate close to the bottom of the stack, and have a lot of influence over the availability, performance, and robustness of our open source database. Recently, we’ve improved connection handling, explored new networking architectures, and integrated OpenTelemetry to make issues easier to diagnose and connect MongoDB to modern observability tools. We are planning to further improve our networking’s stack performance, availability and scalability as well as further enhance our observability stack using open observability frameworks. Are you excited to help the MongoDB engineering team build a better database? We are! Join us today, and we can build a faster, more reliable, exceptionally observable, database system together. This role can be based out of our New York City office or remotely within the United States and Canada. Candidate Profile 3+ years of experience building distributed systems Passionate about delivering and deploying a product with cross-team stakeholders Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Hands-on experience with building production-level code. Experience in C++ is required Interest in furthering their knowledge of networking, observability and how computer architecture and internals impact the availability of SaaS Solid verbal and written communication skills and highly motivated to collaborate with colleagues Po
Jobiba hiring network
Cluster Hr Head Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster hr head jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Join the MongoDB Networking & Observability team and help build the core of a distributed database! Our team focuses on creating and enhancing components which facilitate communication between distributed processes and make these processes, and their communication, easily observable. Networking Observability’s responsibilities include improving MongoDB networking, improving the efficiency of resource utilization, and building low-overhead observability features. Our team includes engineers located in New York City and fully remote engineers. We operate close to the bottom of the stack, and have a lot of influence over the availability, performance, and robustness of our open source database. Recently, we’ve improved connection handling, explored new networking architectures, and integrated OpenTelemetry to make issues easier to diagnose and connect MongoDB to modern observability tools. We are planning to further improve our networking’s stack performance, availability and scalability as well as further enhance our observability stack using open observability frameworks. Are you excited to help the MongoDB engineering team build a better database? We are! Join us today, and we can build a faster, more reliable, exceptionally observable, database system together. This role will be based remotely in Canada. Candidate Profile 3+ years of experience building distributed systems Passionate about delivering and deploying a product with cross-team stakeholders Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Hands-on experience with building production-level code. Experience in C++ is required Interest in furthering their knowledge of networking, observability and how computer architecture and internals impact the availability of SaaS Solid verbal and written communication skills and highly motivated to collaborate with colleagues Position Expectations Understand and improve the current funct
The MongoDB Atlas Clusters Team is seeking a Senior Staff Product Manager to provide division-wide, product leadership for MongoDB Atlas which is our best-in-class, fully managed, multi-cloud database service that simplifies the complexities of deploying and managing highly available database deployments across multiple regions and cloud providers. As the most senior product management contributor, you will shape the multi-year strategy, product architecture, and market trajectory for the entire Atlas Clusters business unit. You will be the primary product authority on the most complex trade-offs, anticipating systemic risks and cutting through structural friction that blocks multiple teams and initiatives. This role requires collaboration with distinguished engineers to rigorously debate strategy and pressure-test assumptions and advise executive leadership on options that balance long-term monetization, strategy, and business outcomes. And, where relevant, owns the pricing and monetization strategy for the product line. We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. Responsibilities Product Vision and Roadmap: Spearhead strategic unity across multiple product areas over a multi-year horizon to drive a coherent product experience Business Leadership: Originate and de-risk the division’s most consequential strategic bets, owning rigorous business cases and advising leadership on multi-year economic outcomes Customer Advocacy: Engage with strategically significant customers at the senior level (e.g., advisory boards and executive briefings) to surface cross-portfolio opportunities and inform the roadmap PM Excellence: Elevate the global standard of product management by mentoring Staff and Senior PMs and pioneering new data-driven playbooks across the organization. Create frameworks and standards for customer and market research that others use for conducting high-quality customer and market research Cross-Fu
Key Responsibilities : Primary responsibilities :- Installation and configuration of MySQL instances on single or multiple ports. Hands-on experience of working with MysQL 5.7 and MySQL 8. Clear understanding of MysQL Replication process flows , threads , setting up multi node clusters and basic troubleshooting. Understanding of at least one of the backup and recovery methods for MySQL . Strong fundamentals of SQL and able to understand and tune complex SQL queries when needed. Strong fundamentals on the linux system side and monitoring tools like top , iostats , sar etc. At Least couple of years of production hands on experience on medium to big sized MySQL databases. Setting up and maintaining users and privileges management system and troubleshooting relevant access issues. Some exposure to external tools like Percona , ProxySQL , HAP etc. Understand the transaction flows and ACID compliance. Basic understanding of networking concepts . Performing on-call support and should be able to provide the first level support . Excellent verbal and written communication skills. Strong shell scripting skills . Good to have Python . Secondary responsibilities. :- Able to configure and setup NOSQL databases like Mongodb and Cassandra. Ability to learn new technologies along with a team and a positive outlook to understand problems from the business point of view. Qualifications: Proficiency in database management systems such as , MySQL or NoSQL databases. SQL programming and database design skills. Knowledge of database performance tuning and optimization techniques. Familiarity with database security best practices. Scripting and automation skills (Good to have- Python). Good problem-solving and analytical skills. Excellent communication and teamwork skills.
Key Responsibilities : Primary responsibilities :- Installation and configuration of MySQL instances on single or multiple ports. Hands-on experience of working with MysQL 5.7 and MySQL 8. Clear understanding of MysQL Replication process flows , threads , setting up multi node clusters and basic troubleshooting. Understanding of at least one of the backup and recovery methods for MySQL . Strong fundamentals of SQL and able to understand and tune complex SQL queries when needed. Strong fundamentals on the linux system side and monitoring tools like top , iostats , sar etc. At Least couple of years of production hands on experience on medium to big sized MySQL databases. Setting up and maintaining users and privileges management system and troubleshooting relevant access issues. Some exposure to external tools like Percona , ProxySQL , HAP etc. Understand the transaction flows and ACID compliance. Basic understanding of networking concepts . Performing on-call support and should be able to provide the first level support . Excellent verbal and written communication skills. Strong shell scripting skills . Good to have Python . Secondary responsibilities. :- Able to configure and setup NOSQL databases like Mongodb and Cassandra. Ability to learn new technologies along with a team and a positive outlook to understand problems from the business point of view. Qualifications: Bachelor's degree in Computer Science, Information Technology, or a related field (or equivalent experience). Proficiency in database management systems such as , MySQL or NoSQL databases. SQL programming and database design skills. Knowledge of database performance tuning and optimization techniques. Familiarity with database security best practices. Scripting and automation skills (Good to have- Python). Good problem-solving and analytical skills. Excellent communication and teamwork ski
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is building the world’s fastest, most efficient AI compute clusters. TT-Fabric is the high-performance nervous system of this platform: the low-level networking layer that lets thousands of RISC-V and AI processors snap together into a single, massively parallel distributed supercomputer. If you love squeezing nanoseconds out of hot paths, designing protocols that move data at absurd scale, and turning messy hardware constraints into elegant distributed systems, this is an opportunity to shape the fabric that future AI models will run on This role is hybrid based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who We Are Strong systems engineer with deep C or C++ experience and comfort working in low-level or bare-metal environments. Passionate about hardware-software interaction, performance tuning, and eliminating inefficiencies at the protocol level. Curious about networking, synchronization, and communication across large clusters. Comfortable reasoning from first principles and challenging industry conventions. Motivated by building infrastructure that directly impacts large-scale
About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE You will join the Portworx team in Everpure, which is responsible for delivering the highest quality Portworx Enterprise products. You will be contributing clean & robust code, be customer oriented and put quality first.A WHAT YOU'LL DO Designing and developing cloud native microservices and integrating new features to Portworx products Bringing a focus on design, development, unit/functional testing, code reviews, documentation, continuous integration and continuous deployment Debug product and performance issues in large scale clusters using AI tooling Collaborating with peers and stake-holders to take solutions from initial design to production Take full ownership of design and development activity by adapting to customer feedback and handling issues found in unit testing, system testing and customer deployments Experimenting with new technologies in order to push the state-of-the-art and innovate new solutions. We are primarily an in-office environment and therefore, you will be expected to work from the Bangalore office in compliance with Everpure's policies, unless you are on PTO, or work travel, or other approved leave. WHAT YOU BRING BS in Computer Science 8+ years of experience in Designing, Development and Testing of Enterprise products (Golang preferred). Good understanding of Microservice Architectures and Cloud Native platforms Designing and owning micro services to operate and scale in a di
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE You will join the Portworx team in Everpure, which is responsible for delivering the highest quality Portworx Enterprise products. You will be contributing clean & robust code, be customer oriented and put quality first.A WHAT YOU'LL DO Designing and developing cloud native microservices and integrating new features to Portworx products Bringing a focus on design, development, unit/functional testing, code reviews, documentation, continuous integration and continuous deployment Debug product and performance issues in large scale clusters using AI tooling Collaborating with peers and stake-holders to take solutions from initial design to production Take full ownership of design and development activity by adapting to customer feedback and handling issues found in unit testing, system testing and customer deployments Experimenting with new technologies in order to push the state-of-the-art and innovate new solutions. We are primarily an in-office environment and therefore, you will be expected to work from the Bangalore office in compliance with Everpure's policies, unless you are on PTO, or work travel, or other approved leave. WHAT YOU BRING BS in Computer Science 8+ years of experience in Designing, Development and Testing of Enterprise products (Golang preferred). Good understanding of Microservice Architectures and Cloud Native platforms Designing and owning micro services to operate and
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE You will join the Portworx team in Everpure, which is responsible for delivering the highest quality Portworx Enterprise products. You will be contributing clean & robust code, be customer oriented and put quality first.A WHAT YOU'LL DO Designing and developing cloud native microservices and integrating new features to Portworx products Bringing a focus on design, development, unit/functional testing, code reviews, documentation, continuous integration and continuous deployment Debug product and performance issues in large scale clusters using AI tooling Collaborating with peers and stake-holders to take solutions from initial design to production Take full ownership of design and development activity by adapting to customer feedback and handling issues found in unit testing, system testing and customer deployments Experimenting with new technologies in order to push the state-of-the-art and innovate new solutions. We are primarily an in-office environment and therefore, you will be expected to work from the Bangalore office in compliance with Everpure's policies, unless you are on PTO, or work travel, or other approved leave. WHAT YOU BRING BS in Computer Science 7+ years of experience in Designing, Development and Testing of Enterprise products (Golang preferred). Good understanding of Microservice Architectures and Cloud Native platforms Designing and owning micro services to operate and scale in a di
Roles and Responsibilities Installation and configuration of NoSQL instances on single or multiple ports. ? Hands on experience of production on medium to big sized NoSQL databases Setting up and maintaining users and privileges management systems and Troubleshooting relevant access issues. Understand the transaction flowsand ACID compliance. Performing on-call support and should be able to provide the first level support . Configure and setup NOSQL databases like mongodb and Cassandra. Automation of repetitive tasks. Qualifications & Experience 3-6 years of Hands-on experience of working with NoSQL DBA . Some exposure to external tools like Percona , ProxySQL , HAP etc. Understanding of networking concepts . verbal and written communication skills. Experience in tools like shell , python . perl etc for automation. fundamentals on the linux system side and monitoring tools like top , iostats , sar etc. Clear understanding of NoSQL Replication process flows , threads , setting up multi node clusters and basic troubleshooting. Understanding of at least one of the backup and recovery methods for MySQL, fundamentals of SQL. Understand and tune complex SQL queries when needed.
About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking an experienced Principal Hardware Diagnostics Engineer to design and develop diagnostics software used to monitor hardware health and diagnose system-level issues across Graphcore’s AI infrastructure platforms. This role focuses on building diagnostics agents, tools, and analytics frameworks that enable engineers and automation systems to identify, isolate, and resolve hardware issues across blade-level servers and rack-scale clusters. The Team Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. The Systems Engineering and Platform Validation team ensures Graphcore’s AI compute platforms are reliable, diagnosable, and operationally robust at scale. The team co
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe
About the Team The Storage organization builds and operates the online stateful systems and abstractions that DoorDash Engineering depends on: reliable, efficient, secure, and easy to use. Within Storage, the Distributed Caching team owns every caching offering at DoorDash end to end, including ElastiCache (Redis/Valkey), Boulder (our KVRocks-based key-value store for high-QPS feature serving), Entity Cache (read Bill Shen’s engineering blog post, “ High-Performance Proxy Cache for DoorDash Services ”), and the Distributed Lock Service, plus the smart clients (asgard-redis, valkey-go) that sit in front of them. These systems back critical product surfaces across DoorDash, Wolt, and Deliveroo: the team runs roughly 400 ElastiCache clusters serving hundreds of millions of GET requests per second in aggregate, and Boulder, our offline-to-online feature store, serves billions of feature lookups per second at peak. About the Role The team owns provisioning of clusters and the smart clients that sit in front of them, baking in sensible defaults so that other engineering teams get a turnkey caching solution instead of having to run their own. You'll help drive Boulder's evolution to scale further, improve cost efficiency, enhance performance, and support real-time updates; re-platform the Distributed Lock Service onto a strongly consistent backend; and build the self-serve tooling and recommendation engine that let customers describe a workload (QPS, TTL, payload size, latency profile) and get the right backend without talking to a human. You'll go deep on cache invalidation, replication, sharding, compaction, and failover, while shipping the guardrails, automation, and observability that keep this scale operable by a small team. You must be located in San Francisco, Seattle, or the New York Metro Area for this hybrid position. You will report to the Engineering Manager on the Distributed Caching team within the Storage organization. You’re excited about this opportunity b
Get new cluster hr head jobs by email
Daily job updates · Unsubscribe anytime