Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Database Reliability Engineer (DBRE) Experience Level: Mid–Senior (4+ years PostgreSQL experience) About the Role We are looking for a highly skilled Database Reliability Engineer (DBRE) with deep expertise in PostgreSQL at scale and solid experience with MySQL. In this role, you will design, operationalize, and optimize the data persistence layer that powers our large-scale, mission-critical systems. You will work closely with SRE, Platform, and Engineering teams to ensure performance, reliability, automation, and operational excellence across our database environment. This is a hands-on engineering role focused on building resilient data infrastructure, not just administering it. Responsibilities: Architecture, Reliability & Performance Design, implement, and operate highly available PostgreSQL clusters (physical replication, logical replication, sharding/partitioning, failover automation). Optimize query performance, indexing strategies, schema design, and storage engines. Perform capacity planning, growth forecasting, and workload modeling. Own high-availability strategies including automatic failover, multi-AZ/multi-region setups, and disaster recovery. Automation & Tooling Develop automation for any and all tasks including but not limited to: provisioning, configuration, backups, failovers, vacuum tuning, and schema management using tools such as Terraform, Ansible, Kubernetes Operators, or custom tooling. Build monitoring, alerti
Jobiba hiring network
Cluster Hr Head Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster hr head jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . About Pinterest Millions of people across the world come to Pinterest to find new ideas every day. It’s where they get inspiration, dream about new possibilities and plan for what matters most. Our mission is to help those people find their inspiration and create a life they love. As a Pinterest employee, you’ll be challenged to take on work that upholds this mission and pushes Pinterest forward. You’ll grow as a person and leader in your field, all the while helping users make their lives better in the positive corner of the internet. The MySQL Infra team manages one of Pinterest's most vital backend systems, offering huge scope and impact by serving hundreds of business-critical use cases across the entire company. Pinterest’s MySQL online database system operates at a massive scale, managing 100+ production clusters with 6 PB of data. This ro
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Storage Platform team builds and operates the platform that powers database access across Robinhood. We own relational (Postgres/Aurora), key-value (DynamoDB), and caching systems, along with the SDKs, control plane automation, and data plane services that enable safe and reliable access at scale. Our mission is to standardize and strengthen how services connect to storage, improve reliability and performance, and reduce operational overhead through automation. We manage thousands of databases and hundreds of caching clusters supporting millions of users and critical brokerage workloads. Availability is our highest priority — our systems are designed to meet strict uptime targets, including no downtime during market hours. As a Staff Software Engineer , you will design and evolve the core infrastructure that underpins Robinhood’s storage systems. You’ll lead complex distributed systems initiatives such as horizontal sharding, proxy-based query routing, connection pooling, and cross-shard transactions. You’ll work on improving database reliability, performance, and cost efficiency across multi-region deployments. This role has a direct impact on system availability, laten
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Storage Platform team builds and operates the platform that powers database access across Robinhood. We own relational (Postgres/Aurora), key-value (DynamoDB), and caching systems, along with the SDKs, control plane automation, and data plane services that enable safe and reliable access at scale. Our mission is to standardize and strengthen how services connect to storage, improve reliability and performance, and reduce operational overhead through automation. We manage thousands of databases and hundreds of caching clusters supporting millions of users and critical brokerage workloads. Availability is our highest priority — our systems are designed to meet strict uptime targets, including no downtime during market hours. As a Senior Software Engineer , you will build and improve core infrastructure used by many engineering teams, with a focus on reliability, performance, and operational excellence. You’ll deliver key components of data plane and control plane systems (for example: connection pooling, query routing, automation workflows, and observability) and help evolve patterns for safe, consistent database access. You’ll work closely with peers to design pragmatic s
About the Team The Agent Infrastructure team at OpenAI is responsible for building systems that enable training and deployment of highly useful AI agents, both internally and for the world. We work hand-in-hand with researchers to design and scale the environment in which agentic models are trained – providing a workspace for AI models to execute code, debug issues, and develop software just as human SWEs do. Our training environment for agentic models operates at an extremely high scale and has the flexibility to emulate any environment in which an agent might work. At the same time, our team builds and maintains OpenAI’s core platform for the deployment and execution of agents in production. Our systems power products such as Codex, Operator, tool use in ChatGPT, and future agentic products. Some of the most challenging technical problems in scaling the capabilities and utility of agents and agentic models lie in the infrastructure layer – and our team is focused on building the research and production systems that enable OpenAI to train the most capable models in the world, and maximize the utility of our agentic products for users around the world. About the Role As a Software Engineer on the Agent Infrastructure team, you will have the opportunity to work closely with both research and product at OpenAI - building and scaling systems to train highly capable agentic models, and building the platform and integrations to launch new agents to hundreds of millions of users worldwide. Your work will consist of both building new capabilities - standing up the infrastructure and integrations needed to train more complex agentic models - and rapidly scaling these new capabilities to some of the largest compute clusters in the world. At the same time, you’ll be instrumental to the launch of agentic products at OpenAI - building, maintaining, and scaling the production platform on which all agents run. We’re looking for people with deep experience building AI infrastructu
About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Senior Software Engineer, ML Systems & Training Infrastructure, you will be a deeply hands-on engineering force multiplier for the robotics team. You will help keep the training framework and surrounding infrastructure healthy, review and improve code quickly, debug failures across ML systems and infrastructure, and unblock researchers and engineers when the path from idea to working training job gets rough. We’re looking for people who love writing, reading, reviewing, and fixing code; who can get productive quickly in unfamiliar systems; and who bring strong practical judgment without a lot of ego or process overhead. This role will be based in San Francisco, CA and be expected in office 5 days per week and offer relocation assistance to new employees. In this role, you will: Review, improve, and clean up code across training frameworks and adjacent infrastructure. Identify risky or low-quality changes before they land, and raise the code quality bar without slowing the team down. Debug issues across ML training systems, GPUs, clusters, networking, and related infrastructure. Help researchers and engineers unblock broken training jobs, flaky workflows, and brittle internal tooling. Improve the reliability, maintainability, and usability of the robotics team’s training framework. Move quickly on practical engineering problems that directly affect team velocity. You might thrive in this role if you: Have strong software engineering fundamentals and excellent code review judgment. Have experience with ML systems, training fr
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet Hardware team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automation to reduce manual work
About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training Performance Engineer, you’ll drive efficiency improvements across our distributed training stack. You’ll analyze large-scale training runs, identify utilization gaps, and design optimizations that push the boundaries of throughput and uptime. This role blends deep systems understanding with practical performance engineering — analyzing GPU kernel performance, collective communication throughput, investigating I/O bottlenecks, and sharding our models so we can train them at massive scale. You’ll help ensure that our clusters are running at peak performance, enabling OpenAI to train larger, more capable models with the same compute budget. This role is based in San Francisco, CA. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you will: Profil
About the Team: The Database Systems team specializes in high-performance distributed databases. Our team built Rockset, the real-time search, analytics, and vector database that powers all vector search and retrieval augmented generation (RAG) at OpenAI. In addition to retrieval, as an online database, Rockset powers core functionality across all of OpenAI's product lines and many critical internal use cases. About the Role : We are looking for engineers passionate about distributed systems, close-to-the-metal performance optimization (our core engine is written in C++), and building scalable database infrastructure from the ground up. As an engineer on the Database Systems team, you'll contribute to the core database engine, driving improvements across ingestion, query execution, indexing, and storage. You'll partner with teams across OpenAI to unlock new product capabilities and help scale online database reliability and throughput as usage grows by orders of magnitude. In this role you will: Design, build, and operate high-performance distributed systems Identify and resolve performance bottlenecks to scale infrastructure to the next order of magnitude Define long-term technical direction and guide system evolution Collaborate with product, engineering, and research teams to deliver scalable and reliable infrastructure Dig deep into complex production issues across the stack Contribute to incident response, postmortems, and best practices for system reliability You might thrive in this role if you: Have significant experience building, scaling, and optimizing distributed systems at scale Are curious about database internals, storage engines, or low-latency query systems Enjoy debugging challenging performance issues in complex, high-throughput systems Have experience operating production clusters at scale (e.g., Kubernetes or other orchestration systems) Think rigorously about scalability, correctness, and reliability Thrive in fast-paced environments with high
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role We’re seeking an exceptional Staff - Principal level offensive security domain expert to build agents that continuously identify and coordinate remediation of vulnerabilities across OpenAI’s infrastructure and applications. You will be the technical owner of this effort, combining deep offensive security judgment with agent engineering to build a production system that can operate safely and reliably at scale. As OpenAI increasingly uses automation throughout the company, we believe our security testing must become increasingly automated as well. Advances in model capabilities create an opportunity to test more of our attack surface than would be possible through human effort alone and a need to ensure that we remain ahead of those same capabilities as they become available to attackers. In this role, you’ll build a portfolio of specialized agents that develop a deep understanding of OpenAI’s infrastructure, applications, processes, and security boundaries. These agents will combine internal context with feedback from running systems to explore our cloud environments, Kubernetes clusters, web applications, endpoints, external attack surface, and other high-value targets. The goal is for agents to not only discover vulnerabilities, but also to validate exploitability, document impact, drive remediation, and verify fixes. Success will be measured through outcomes like vulnerabilities fixed, attack surface covered, and performance on evals
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity As a Senior Backend Engineer on the Cloud Platform team, you will play a key role in building the core systems and services that power Postman’s internal platform. You’ll help create new backend services that manage how we deploy, scale, and operate our infrastructure and product services, leveraging Java, Spring Boot, and Hibernate (JPA) on top of cloud-native technologies like Kubernetes, ArgoCD, Istio, and Terraform. This role is highly impactful: the systems you build will be used across Postman engineering, enabling faster delivery, better scalability, and a stronger developer experience. You’ll also have the opportunity to contribute to open source, shaping tools that extend beyond Postman’s boundaries. What You’ll Do Design and develop backend services in Java and Spring Boot to support Postman’s internal Cloud Platform. Architect new services that manage service deployment, lifecycle, and scaling across Kubernetes clusters. Implement GitOps workflows (ArgoCD) to support continuous delivery. Integrate with cloud-native tooling such as Istio, Helm, and Terraform. Apply strong soft
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake is a high-growth, cloud-native data platform company committed to empowering enterprises to achieve their full potential. With a culture built on impact, innovation, and collaboration, we offer an environment where you can build large-scale systems, move fast, and take your technology career to the next level. We are seeking an outstanding Staff Software Engineer with a passion for large scale databases and distributed systems to help us take the FDB platform to the next level. A massive new market opportunity is being created at the intersection of Cloud and Data, and the Snowflake Data Cloud is leading the way, all powered by the database engine we are building from the ground up. Key to Snowflake’s Database Engine is our large scale distributed transactional Key-Value store - called FDB - which powers all of Snowflake’s products and services and is rapidly evolving to meet Snowflake’s future needs. FDB runs on multiple cloud providers including Amazon Web Services, Microsoft Azure and Google Cloud. The elastic infrastructure FDB runs on is being built from the ground up and is envisioned to be a cloud agnostic, fully automated manageability platform that provides: Autoscaling and auto-balancing of clusters based on utilization, traffic and workloads Auto-provis
About us Paytm is India's leading mobile payments and financial services distribution company. Pioneer of the mobile QR payments revolution in India, Paytm builds technologies that help small businesses with payments and commerce. Paytm’s mission is to serve half a billion Indians and bring them to the mainstream economy with the help of technology. Role Overview We are seeking a Database STL (Individual Contributor) with deep expertise in MySQL and strong working knowledge of MongoDB, PostgreSQL, and Cassandra. This role combines hands-on database administration and optimization with strategic ownership of database reliability, automation, and cloud adoption. The candidate will lead by example—driving technical excellence, influencing best practices, and partnering cross-functionally with DevOps, SRE, and product engineering teams to deliver highly available, secure, and scalable database platforms. Key Responsibilities 1. End-to-End Ownership of MySQL databases in production & staging—availability, performance, and reliability. 2. Architect, manage, and support MongoDB, PostgreSQL, and Cassandra clusters for scale and resilience. 3. Define and enforce backup, recovery, HA, and DR strategies across all critical database platforms. 4. Drive database performance engineering—tuning queries, optimizing schemas, indexing, and partitioning for high-volume workloads. 5. Own replication, clustering, and failover architectures ensuring business continuity. 6. Champion automation & AI-driven operations—design self-healing scripts, predictive scaling, and proactive monitoring solutions. Collaborate with Cloud/DevOps teams on AWS database services (RDS, Aurora, DynamoDB, EC2, S3) to optimize cost, security, and performance. 7. Establish monitoring dashboards & alerting mechanisms for slow queries, replication lag, deadlocks, and capacity planning. Ensure compliance & security standards—encryption, auditing, and regulatory requirements. 8. Lea
Position: Engineering Manager - Database Job Location: Noida Role Overview We are seeking a Database Engineering Manager (Individual Contributor) with deep expertise in MySQL and strong working knowledge of MongoDB, PostgreSQL, and Cassandra. This role combines hands-on database administration and optimization with strategic ownership of database reliability, automation, and cloud adoption. The candidate will lead by example—driving technical excellence, influencing best practices, and partnering cross-functionally with DevOps, SRE, and product engineering teams to deliver highly available, secure, and scalable database platforms. Key Responsibilities 1. End-to-End Ownership of MySQL databases in production & staging—availability, performance, and reliability. 2. Architect, manage, and support MongoDB, PostgreSQL, and Cassandra clusters for scale and resilience. 3. Define and enforce backup, recovery, HA, and DR strategies across all critical database platforms. 4. Drive database performance engineering—tuning queries, optimizing schemas, indexing, and partitioning for high-volume workloads. 5. Own replication, clustering, and failover architectures ensuring business continuity. 6. Champion automation & AI-driven operations—design self-healing scripts, predictive scaling, and proactive monitoring solutions. Collaborate with Cloud/DevOps teams on AWS database services (RDS, Aurora, DynamoDB, EC2, S3) to optimize cost, security, and performance. 7. Establish monitoring dashboards & alerting mechanisms for slow queries, replication lag, deadlocks, and capacity planning. Ensure compliance & security standards—encryption, auditing, and regulatory requirements. 8. Lead incident management & on-call rotations, ensuring rapid response and minimal MTTR. 9. Act as a strategic technical partner, contributing to database roadmaps, automation strategy, and adoption of AI-driven DBA practices. Required Skills & Experience 1. 6–10 years of p
Get new cluster hr head jobs by email
Daily job updates · Unsubscribe anytime