Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold builders and sharp problem-solvers who are wired to deliver great outcomes. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. The DevX team’s mission is to build and operate the core developer infrastructure at Robinhood. Our team owns and scales the systems that thousands of engineers rely on daily, partnering with software developers across the company to make development fast, reliable, and cost-efficient! As a Senior Software Developer, you will focus on building and scaling our developer infrastructure. You will manage and optimize monorepo builds with Bazel and own the remote build execution cluster that powers them, and you will scale CI pipelines to be fast, safe, and reliable across thousands of daily builds. You will also improve the developer onboarding experience and provide remote development environments to accelerate development workflows. In this role, you will ensure developers can code, test, and deploy with minimal friction. This is an incredible opportunity to make a massive difference for the entire engineering organization! This role is based in our Toronto, ON office, with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. What you’ll do Manage and
Jobiba hiring network
Cluster Hr Head Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster hr head jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Staff Machine Learning Engineer, Identity Verification As a Staff Machine Learning Engineer on the Identity Verification team within the Platform group, you'll own the ML systems that determine whether a person, document, and capture session are legitimate. Every signup, account recovery, and high-risk action at Coinbase depends on these models. You'll lead the technical strategy for IDV ML end-to-end, from architecture through production enforcement, protecting the integrity of millions of accounts. What you'll do: Own the full IDV ML stack, including document authenticity models, 1:1 and 1:N face-match, liveness detection, presentation-attack detection, and deepfake/injection detection from feature pipeline through threshold tuning and production enforcement. Build identity-graph systems using GNNs that cluster accounts sharing biometric, device, and document signals to detect synthetic-identity rings and coordinated fraud at onboarding. Develop behavioral and device-intelligence models for capture-session anomaly detection, bot-vs-human classification, and device-fingerprint-based risk scoring at real-time latency. Drive vendor ML strategy by benchmarking external models against a Coinbase-owned evaluation set, designing dynamic routing logic across providers and geographies, and building the in-house evaluation layer that catches regressions before they reac
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Software Engineer on the Compute Platform team within the Platform group, you'll own the primary compute orchestration infrastructure that every service at Coinbase runs on. Built largely on CNCF technologies including Kubernetes and Istio, this platform underpins the scalability, reliability, and efficiency of our entire product suite. You'll design and ship tooling, automation, and net-new capabilities that make it easy for hundreds of engineers to deploy and operate critical services, while partnering closely with Security, Reliability, and Observability teams to raise the bar across the stack. What you'll do: Own the design, build, and operation of Kubernetes cluster management tooling and automation that keeps our compute platform reliable and self-healing at scale. Build developer-facing tooling and workflows that improve how engineers across Coinbase interact with Kubernetes, with a heavy emphasis on integrating AI-driven processes and support. Deliver net-new compute capabilities for service owners, such as one-off jobs, cron scheduling, deployment strategies, EFS support, and automated right-sizing. Drive operational excellence by automating toil, reducing on-call burden, and continuously improving platform observability and incident response. Partner with Security, Reliability, and Observability teams to ensure the compute platform meets C
About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. Our team builds the software, tooling, and operational systems that help manage this fleet at scale. We work across production engineering, distributed systems, capacity management, and operational automation to improve reliability, reduce manual work, and make better use of available compute. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will develop the systems that help manage the GPU fleet powering ChatGPT, including tooling for fleet health, capacity planning, operational automation, and incident response. You will work closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and compute utilization. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference. Develop systems for capacity planning, fleet health monitoring, and resource utilization. Automate operational workflows, including incident detection, diagnosis, and response. Identify and address bottlenecks affecting fleet reliability, scalability, and performance. Partner with infrastructure, research, and product engineering teams to improve the compute platform. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems. Have a background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering. Have built software that automates operational workflows and reduces manual work. Have worked with distributed infrastructure, cluster orchestration, or large-scale int
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system
About the Team OpenAI’s Industrial Compute team is responsible for building and scaling large-scale compute capacity across first-party data centers, strategic partners, and industrial infrastructure environments. We focus on converting power, land, hardware, and operational execution into reliable compute capacity that can support frontier AI training and inference workloads. This team operates at the intersection of infrastructure delivery, hardware systems, utilities, supply chain, and capacity strategy—ensuring OpenAI can scale compute faster than traditional models allow. About the Role We are seeking a Tokens-as-a-Service (TaaS) Lead to drive the end-to-end conversion of industrial-scale infrastructure investments into usable token capacity for OpenAI workloads. In this role, you will own execution across complex compute programs where raw infrastructure capacity must be transformed into operational GPU throughput. You will coordinate across data center delivery, power, networking, hardware deployment, workload enablement, finance, and external partners to ensure capacity becomes productive tokens as quickly and efficiently as possible. This role is ideal for someone who can bridge physical infrastructure delivery with compute utilization outcomes. Success requires strong systems thinking, elite program leadership, and the ability to drive accountability across internal teams and strategic partners. In this role, you will Lead Tokens-as-a-Service programs across industrial compute environments, including first-party and partner-owned capacity. Convert delivered power, space, and hardware capacity into production-ready token throughput. Build integrated execution plans spanning construction, power energization, rack deployment, networking, cluster readiness, and workload onboarding. Partner with infrastructure engineering, hardware, networking, finance, supply chain, and operations teams. Drive external providers, EPCs, OEMs, utilities, and strategic partners t
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer, Snowflake Natsec Running Snowflake in public sectors in different countries and regions, even in different industry verticals, requires us to build a compliant, secure, and auditable infrastructure. Many key design decisions are deeply rooted in the Snowflake product architecture. As a Senior Software Engineer, you will be responsible for leading several key areas and collaborating with various engineering groups in addition to the Public Sector team. To be successful in the area, you will need to have (and continue to build) a broad and in-depth knowledge base on cloud infrastructure, privacy, and governance, compliance controls, data security and data residency in various aspects of Snowflake. AS A SENIOR SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Solve real business needs at large scale by applying your software engineering and analytical problem solving skills. Design, implement and maintain scalable distributed systems for our cloud automation platform that include cloud control plane, Kubernetes container platform and traffic and networking. Work directly with customers to quickly understand their critical problems and design and implement solutions Deploy and maintain availability of cloud compute servers and Kubernetes cluster that power the
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer, Snowflake Natsec Running Snowflake in public sectors in different countries and regions, even in different industry verticals, requires us to build a compliant, secure, and auditable infrastructure. Many key design decisions are deeply rooted in the Snowflake product architecture. As a Senior Software Engineer, you will be responsible for leading several key areas and collaborating with various engineering groups in addition to the Public Sector team. To be successful in the area, you will need to have (and continue to build) a broad and in-depth knowledge base on cloud infrastructure, privacy, and governance, compliance controls, data security and data residency in various aspects of Snowflake. AS A SENIOR SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Solve real business needs at large scale by applying your software engineering and analytical problem solving skills. Design, implement and maintain scalable distributed systems for our cloud automation platform that include cloud control plane, Kubernetes container platform and traffic and networking. Work directly with customers to quickly understand their critical problems and design and implement solutions Deploy and maintain availability of cloud compute servers and Kubernetes cluster that power the
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - Container Platform (Menlo Park) About the Role We build the foundational container platform that runs Snowflake's production, AI/ML, and CI workloads across AWS, Azure, and GCP, including a rapidly growing AI/ML footprint. Hundreds of large Kubernetes clusters under management and growing. The work is to make that fleet reliable, automated, and invisible to the thousands of engineers building on top of it. This is a staff-level role on a senior, high-performing platform team. You'll own hard problems end to end, drive technical direction across teams, and build the automation and platform abstractions that make operating at this scale sustainable. There is significant unsolved work ahead: improving the developer experience for thousands of internal engineers and continuing to scale the platform to meet Snowflake's growth. What You'll Do Own the design and delivery of large, complex platform initiatives spanning cluster lifecycle management, multi-cloud automation, and internal developer tooling. Identify and drive cross-team technical improvements across the platform, from architecture through adoption. Make and defend architectural trade-offs grounded in reliability, scalability, and operational reality. Act as a technical anchor for the team, dev
Data Engineering – Technical Lead About Us: Paytm is India’s leading digital payments and financial services company, which is focused on driving consumers and merchants to its platform by offering them a variety of payment use cases. To merchants, Paytm offers acquiring devices like Soundbox, EDC, QR and Payment Gateway where payment aggregation is done through PPI and also other banks’ financial instruments. To further enhance merchants’ business, Paytm offers merchants commerce services through advertising and Paytm Mini app store. Operating on this platform leverage, the company then offers credit services such as merchant loans, personal loans and BNPL, sourced by its financial partners. About the Role: This position requires someone to work on complex technical projects and closely work with peers in an innovative and fast-paced environment. For this role, we require someone with a strong product design sense & specialized in Hadoop and Spark technologies. Requirements: 4 to 8 years of experience in Big Data technologies. The position Grow our analytics capabilities with faster, more reliable tools, handling petabytes of data every day. Brainstorm and create new platforms that can help in our quest to make available to cluster users in all shapes and forms, with low latency and horizontal scalability. Make changes to our diagnosing any problems across the entire technical stack. Design and develop a real-time events pipeline for Data ingestion for real-time dash- boarding.Develop complex and efficient functions to transform raw data sources into powerful, reliable components of our data lake. Design & implement new components and various emerging technologies in Hadoop Eco- System, and successful execution of various projects. Be a brand ambassador for Paytm – Stay Hungry, Stay Humble, Stay Relevant! Skills that will help you succeed in this role: Fluent with Strong hands-on experience with Hadoop, MapReduce, Hive, Spark, PySpark etc.Excellent progr
As a Sr. Staff Technical Program Manager, you will partner with key Engineering, Product, Product Design, Marketing, and Analytics stakeholders to conduct data-driven experiments and deliver features for MongoDB. As a seasoned program leader, you will be responsible for one of our most mission-critical programs this fiscal year, and own executive level communication related to the program. We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model, or remote within Ireland. The right candidate for this role will be: Experienced with 15+ years of working in an engineering organization leading complex cross-functional technical programs Experienced with 10 years of Software development background, with Cloud storage and compute products Experience with Service Oriented Architecture and Cloud Infrastructure Able to leverage their knowledge and experience to influence technical discussions, summarizing outcomes and next steps Skilled at communicating across a diverse set of engineers and stakeholders Hyper-organized and capable of coordinating across multiple independent work streams and organizations A role model for effective execution practices, driven by an attuned sense of priority and urgency Able to leverage their experience in program delivery to influence improvements to our tools, operations, and architecture Trained in working with project tracking software (e.g. Jira, Rally, MS Project) Familiar with MongoDB or a comparable technology Interested in business automation work such as scripting in Python, Google Apps Script and Slack Position Expectations: Leverage technical acumen and analytical skills to drive engineering programs forward and maximize business value delivery Recognize patterns in a sea of information and take action accordingly Design, maintain, and improve the processes and tools that power program delivery Build strategic partnerships with Product and Engineering stakeholders Expand knowledge into new
MongoDB is seeking a Senior Software Engineer to join the Atlas Clusters Organization. The organization is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. This includes developing software to interface with the three major cloud providers (AWS, Azure, and GCP) in order to bring security, durability, availability, and performance to all deployments of MongoDB. We are forming a new Atlas Clusters team in the Dublin area. We are looking to speak to candidates who are based in Dublin for our hybrid working model. What you’ll do Build and design new features for MongoDB Atlas Contribute to and lead complex technical projects Work with stakeholders throughout MongoDB to build our roadmap and product offerings Work with customers and support engineers to fix issues and become part of our on-call rotation Collaborate with team members to develop our codebase, best practices, and design principles Foster an inclusive and respectful work environment according to MongoDB's Core Values We’re looking for someone who Has at least 6 years of professional software development experience Is skilled at writing large-scale, distributed backend systems in a compiled language (Go, Java, C#, etc.). Has experience with at least one major cloud provider technology (AWS, Azure, GCP) Has led the launch of a new module and maintained it in production Is eager to solve tough problems Has excellent communication skills Is curious, collaborative, and motivated Success Measures In 3 months, you'll have shipped code into production and collaborated with the team to solve tough problems In 6 months, you'll have contributed to a large project and joined our on-call rotation In 12 months, you'll have designed new features, led development work, and become a go-to expert on parts of the system About MongoDB MongoDB is built for
MongoDB is seeking an Engineering Manager to join the Atlas Clusters Organization. The organization is responsible for building MongoDB Atlas, our database-as-a-service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. This includes developing software to interface with the three major cloud providers (AWS, Azure, and GCP) in order to bring security, durability, availability, and performance to all deployments of MongoDB. The Atlas Clusters Fleet Signal Management team is a product engineering team that builds machinery to rollout hardware and software to the Atlas data plane, detect regressions at scale, and automate remediative actions. The core mission is to establish platform stability by owning the rollout and monitoring systems designed to swiftly detect and mitigate potential issues. We are forming a new Atlas Clusters Fleet Signal Management team in the Dublin area. The Engineering Manager who fills this position will be pivotal in growing that team. We are looking to speak to candidates who are based in Dublin for our hybrid working model. What you’ll do Lead a team of motivated individual contributors who are eager to learn and grow Contribute to the code, design, and architecture of the systems your team develops Work with stakeholders throughout MongoDB to build our roadmap and product offerings Work with customers and support engineers to fix issues and become part of our on-call rotation Collaborate with team members to develop our codebase, best practices, and design principles Foster an inclusive and respectful work environment according to MongoDB's Core V
Get new cluster hr head jobs by email
Daily job updates · Unsubscribe anytime