Come join and lead the Server Ingress Security team, where we are rearchitecting MongoDB Server’s ingress networking to make MongoDB clusters even more secure. This new team is building the Atlas Network Protection layer, a set of performant, security-critical services that harden MongoDB's pre-authentication attack surface and provides the ability to respond rapidly to emergent threats. We are looking for a talented Lead Engineer to join the team and be founding members, where you will play a crucial role in our multi-year roadmap. Our team champions a strong culture of inclusivity, diversity, and collaboration, and lives MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. If you want to lead a fast-growing team that applies security and systems engineering fundamentals to protect a popular database at scale, join us! We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. Candidate Profile 3+ years of experience managing a team of software engineers, including hiring, performance and growth management, compensation planning, and mentoring You have 8+ years of experience building production-quality systems software with large backend/compiled codebases, ideally in Rust. Bonus points for experience with performance profiling, network protocols, TLS, and connection management You have strong technical judgment that you use to effectively guide engineering decisions in security-sensitive or networking-adjacent domains You put the customer first and don't hesitate to cross team boundaries in search of the right solution Solid experience in designing, writing, testing, maintaining, and operating mission-critical software systems Bonus points Professional or advanced academic expertise in the domains of security or networking You enjoy coaching, career development, and creating growth opportunities to help your
Jobiba hiring network
Cluster Lead Facilities Services Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster lead facilities services jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Job Requisition ID # 26WD101306 Position overview Autodesk Flow is the connected platform behind how film and television get made — from the moment footage is captured on set, through review and approval, to final delivery. This is a dedicated ABM role for Flow, working hand in hand with the Flow sales organization. You will be their marketing counterpart: building the account clusters, programs and sales-facing materials that turn business priorities into engaged accounts and qualified pipeline. You will shape how the account-based motion works here — how accounts are scored and clustered, how programs are built around sales priorities, and how we measure what they produced. You will work in close partnership with a Field Marketing Manager, the wider Media & Entertainment marketing organization, and our central content team. Location :This position can be remote or hybrid. Responsibilities Partner with Sales: Act as the dedicated ABM partner to the Flow sales organization, building programs around their account priorities and revenue targets Run a recurring planning cadence with Sales — account co-planning, pipeline reviews and quarterly business reviews — reviewing engagement, buying signals and next best actions Own lead routing and funnel optimization for Flow, making sure the handoff from marketing program to seller follow-up is fast, clean and measured Partner across Marketing, Industry Strategy and Technical Sales to coordinate account engagement Build and prioritize account clusters: Build account clusters around shared buying t
MongoDB is seeking a Staff Software Engineer to join the Atlas Clusters Organization. The organization is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. This includes developing software to interface with the three major cloud providers (AWS, Azure, and GCP) in order to bring security, durability, availability, and performance to all deployments of MongoDB. This engineer will also work on our Atlas Data Federation & Archiving product. Atlas Data Federation & Archiving allows customers to move data from hot to cold storage and run federated queries over that data. We are forming a new Atlas Clusters team in the Dublin area. We are looking to speak to candidates who are based in Dublin and would like a hybrid or in-office working model. What you’ll do Build and design new features for MongoDB Atlas and Atlas Data Federation & Archiving Contribute to and lead complex technical projects Work with stakeholders throughout MongoDB to build our roadmap and product offerings Work with customers and support engineers to fix issues and become part of our on-call rotation Collaborate with team members to develop our codebase, best practices, and design principles Foster an inclusive and respectful work environment according to MongoDB's Core Values We’re looking for someone who Has at least 10+ years of professional software development experience Is skilled at writing large-scale, distributed backend systems in a compiled language (Go, Java, C#, etc) Has experience with at least one major cloud provider technology (AWS, Azure, GCP) Has led the launch of a new module and maintained it in production Is eager to solve tough problems Has excellent communication skills Is curious, collaborative, and motivated Success Measures In 3 months, you'll have shipped code into production and c
The MongoDB Atlas team is a diverse group of contributors working together to help our users manage MongoDB at global scale. We are responsible for MongoDB Atlas: our database as a service offering and fastest growing product which allows users to deploy fault-tolerant, globally distributed MongoDB clusters in just minutes. We're seeking a Senior Engineer to join the Atlas Identity and Access Management (IAM) team. IAM is a platform and a product team. We serve internal engineers by providing them a secure and durable suite of services, and we serve external customers by providing them user facing features and products. We are the owners of Atlas’ authentication (OAuth, SSO, Federated Identity) and authorization (RBAC, ABAC) systems, along with many others. The IAM team’s mission is to enable customers to securely build their applications with Atlas through our best in class user experience. We are looking to speak to candidates who are based in New York City, NY for our hybrid working model. Role Responsibilities Design, architect, build, and deliver core pieces of IAM Lead projects from specification to delivery Mentor and grow other team members Improve our codebase, best practices, and design principles Define your top priorities and focuses, communicate them, and execute against them Lead and contribute to complex technical projects and initiatives Candidate Profile 5+ years experience of software engineering, primarily focused on backend systems Proficient in a modern compiled programming language (Java, Go, C#, C++, etc.) Willingness to learn JavaScript and/or TypeScript along with modern frontend technologies (React, Redux, etc.); prior experience a plus Excellent communication skills, both written and verbal Desire to collaborate with colleagues and mentor fellow engineers Is curious, collaborative, empathetic, and intellectually honest Has a passion for problem solving and learning new things in the domains of computer science and software engineering Expe
The Security Libraries team owns the customer-side integrations behind Datadog’s run-time security products — App & API Protection , Workload Protection , and Code Security — shipping and maintaining security capabilities across seven open-source language libraries ( .NET , Java , Go , Node.js , Python , Ruby , PHP ) and a set of HTTP proxy integrations (Envoy, NGINX, and HAProxy), running inside thousands of production clusters worldwide. Recent work spans exploit prevention (RASP), WAF detections, API Security, code security (IAST and SCA), and AI-assisted onboarding. As Engineering Manager, you’ll lead part of this polyglot team, setting the technical bar and team culture while driving the pace at which new detections and AI-assisted capabilities reach customers. This is a hands-on role: you’ll balance people leadership, product and roadmap ownership, and the operational health of code that runs in production at massive scale, with room to grow into more of Datadog’s security portfolio over time. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead, grow, and develop a team of roughly 4-8 library engineers — coaching, giving direct feedback, and empowering senior ICs as technical leaders Own team delivery and productivity: planning, milestones, reviews, and the on-call rotation Set product direction with Product Management and balance priorities across App & API Protection, Workload Protection, and Code Security Stay technically close to the work — apply strong technical judgment, contribute code where it matters most, and keep quality and architecture high Build strong relationships and drive alignment across the language teams and product, backend, and frontend partners Shape team identity and culture, and own accountability when problems oc
The Atlas API Experience (APIx) teams are part of MongoDB Atlas Data Services, a diverse group of individuals who develop the capabilities to run MongoDB globally (see MongoDB Atlas ). Our software and services allow users to deploy fault-tolerant, scalable, globally distributed MongoDB clusters in minutes. APIx’s mission is to create delightful experiences that bring developers along on the journey from eager beginner to sophisticated MongoDB expert! As part of our team in APIx, you will be responsible for improving & extending our API platform & downstream tooling, such as the Atlas CLI . Our mission is to help Atlas customers automate their workloads easily through APIs & support internal teams contributing to the API. We are looking for passionate, intrinsically-motivated software engineers who lead by example, raise the bar for those around them, and want to make a broad impact. No prior experience with MongoDB technologies is required! During interviews, you will meet most of our team and have the opportunity to ask questions about working at MongoDB. We pride ourselves on our team's culture and on being an inclusive and collaborative group that Builds Together . This role can be based in our Dublin office or remotely within Ireland. What will you do? Propose, design, implement and support product features for the MongoDB Atlas API Platform & MongoDB Atlas DevTools, such as AtlasCLI Build tooling that enables MongoDB users and developers to succeed, using programming languages such as Java, Go, Python & Javascript/Typescript Design and develop software integration components utilizing MongoDB’s technology (i.e. database, mobile, search, etc.) in larger contexts and integrate into partner frameworks or solutions Mentor and provide technical guidance to other engineers through code and design reviews, pairing, and knowledge sharing Lead sophisticated projects end to end, breaking large efforts into incrementally shippable deliverables Investi
The MongoDB Atlas team is a diverse group of contributors working together to help our users manage MongoDB at a global scale. We are responsible for MongoDB Atlas: our database-as-a-service offering and fastest-growing product, which allows users to deploy fault-tolerant, globally distributed MongoDB clusters in just minutes. We're seeking a Software Engineer to join the Atlas Identity and Access Management (IAM) team. The IAM team owns authentication and authorization for all of MongoDB Atlas, spanning from the underlying platform to user-facing features and products. IAM is both a platform and a product team, serving internal engineers as well as external customers. We enjoy the challenge of keeping MongoDB Atlas secure while providing a best-in-class user experience for users ranging from startup to enterprise. We are looking to speak to candidates who are based in New York City, NY for our hybrid working model. Role Responsibilities Lead medium-sized, full-stack projects from technical design to delivery Collaborate with colleagues at all stages of the project lifecycle (ideation, requirements gathering, design, execution, and delivery) Own our leadership principles , and exemplify them in your work Continue to learn and be curious in growing your career! Candidate Profile 2+ years of professional experience building full-stack web applications Proficient in a modern compiled programming language (Java, Go, C#, C++, etc.) Willingness to learn JavaScript and/or TypeScript along with modern frontend technologies (React, Redux, etc.); prior experience a plus Excellent communication skills, both written and verbal Is collaborative, empathetic, and intellectually honest Success Measures In 1 month, you’ll have shipped code into production In 3 months, you’ll have collaborated and delivered on a project with other engineers on the team In 6 months, you’ll have led the technical design, execution, and delivery of a project About MongoDB MongoDB is built for change, em
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity AI is a top strategic priority for New Relic, and the Bengaluru design team is at the center of it. The team works across three closely related product clusters: autonomous incident response (SRE Agent, Autopilot, and Intelligent RCA), the intelligence and platform layer that powers them (Ground Truth, Agentic Platform, and New Relic AI), and AIOps for event correlation and incident management. These products share a common design challenge: users need to trust systems that act autonomously, and building that trust through good design is genuinely hard work. This is an on-site role in Bengaluru. Your designers are there, and many of your engineering and product partners are too. Being present — in standups, reviews, and the quick conversations before a decision gets made — is part of how you'll build the relationships that make design effective. You'll also collaborate with design, product, and engineering partners in the US and Spain, so operating across time zones and communicating well in writing are part of the job. You'll manage a small team of designers and own the quality of the work across these products. The design questions here don't have established answers — how do you make an autonomous system legible? How do you build user trust in AI-generated root cause analysis? How do you design a handoff from machine decision to human judgment? If you're already working in AI product design, or actively building toward it, and you want to lead a team
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Manager focused on TPUs at Baseten, you will lead the "engine room" for our non-NVIDIA accelerator fleet, architecting, securing, and optimizing the Google Cloud TPU (and broader emerging accelerator) capacity that powers our customers' AI workloads. You'll own the end-to-end journey of capacity management for this fleet, from securing large-scale TPU pod allocations to building the automation that ensures reliable uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering, with a specific focus on the TPU ecosystem. You will act as the fleet orchestrator for Google's TPU architecture, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics as we diversify beyond NVIDIA. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the latest generation of TPU hardware, like Google's Trillium (v6e) architecture, and partnering closely with the Model Performance (MP) team to ensure workloads are tuned for TPU-specific execution. EXAMPLE INITIATIVES The TPU Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's TPU clusters, including pod slicing and topology planning Global Workload Orchestration: Bui
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Lead at Baseten, you will lead the "engine room" of the company, architecting, securing, and optimizing the global GPU fleet that powers our customers' AI workloads. You’ll own the end-to-end journey of capacity management, from securing multi-million dollar GPU clusters to building the automation that ensures 99.9% uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering. You will act as the fleet orchestrator for the world's most advanced chips, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the next generation of hardware, like NVIDIA’s Blackwell (B200) architecture. EXAMPLE INITIATIVES The B200 Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's first Blackwell GPU clusters. Global Workload Orchestration: Building "Multi-cloud Capacity Management" systems to move customer workloads seamlessly across regions to optimize cost and latency. Precision GPU Triage: Developing automated Go-based operators to identify, cordon, and repair unhealthy H100 nodes in under an hour. The Supply Chain of Intelligence: Partnering with lead
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Principal Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. As a principal engineer, you will set technical direction and drive execution on high-impact infrastructure security programs, partnering across various orgs at OpenAI to deliver durable controls that raise the security bar at OpenAI scale. In this role, you will: Own end-to-end security outcomes for one or more critical infrastructure areas, including multi-quarter strategy, roadmap, and delivery. Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Lead cross-functional programs to deploy security enhancements and control changes across broad-scale infrastructure, balancing security guarantees with reliability and velocity. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches
Position: Engineering Manager - Database Job Location: Noida Role Overview We are seeking a Database Engineering Manager (Individual Contributor) with deep expertise in MySQL and strong working knowledge of MongoDB, PostgreSQL, and Cassandra. This role combines hands-on database administration and optimization with strategic ownership of database reliability, automation, and cloud adoption. The candidate will lead by example—driving technical excellence, influencing best practices, and partnering cross-functionally with DevOps, SRE, and product engineering teams to deliver highly available, secure, and scalable database platforms. Key Responsibilities 1. End-to-End Ownership of MySQL databases in production & staging—availability, performance, and reliability. 2. Architect, manage, and support MongoDB, PostgreSQL, and Cassandra clusters for scale and resilience. 3. Define and enforce backup, recovery, HA, and DR strategies across all critical database platforms. 4. Drive database performance engineering—tuning queries, optimizing schemas, indexing, and partitioning for high-volume workloads. 5. Own replication, clustering, and failover architectures ensuring business continuity. 6. Champion automation & AI-driven operations—design self-healing scripts, predictive scaling, and proactive monitoring solutions. Collaborate with Cloud/DevOps teams on AWS database services (RDS, Aurora, DynamoDB, EC2, S3) to optimize cost, security, and performance. 7. Establish monitoring dashboards & alerting mechanisms for slow queries, replication lag, deadlocks, and capacity planning. Ensure compliance & security standards—encryption, auditing, and regulatory requirements. 8. Lead incident management & on-call rotations, ensuring rapid response and minimal MTTR. 9. Act as a strategic technical partner, contributing to database roadmaps, automation strategy, and adoption of AI-driven DBA practices. Required Skills & Experience 1. 6–10 years of p
We are seeking a highly skilled and hard-working Senior Test Developer / test engineer to join our multifaceted Enterprise Software QA team. This role offers an outstanding opportunity to leave your mark on the design, construction, optimization and testing of large-scale infrastructure for various foundational NVIDIA unified cloud services and data center offerings. If you are a dedicated engineer with strong expertise in cloud infrastructure and distributed systems and want to apply your skills with AI tools, this role could fit you perfectly. You will thrive in an exciting, innovative environment. What you'll be doing: Work with development teams on test plans for all layers of SW stack for cloud infrastructure, execution, reviews, failure analysis and assessing overall quality and risk. Work with customer PMs on software issues including technical feedback from OEMs and CSPs. Develop key benchmarks to track execution and deploy process improvements to improve efficiency Leverage AI skills to expedite the test scope, test plan, execution and automation workflows. Lead NVIDIA Cloud and Data Center bring up activities which will involve validation, reporting, working with engineering to debug issues, providing design input at times, adding coverage in different areas. Design, develop and maintain CI/CD pipelines for continuous testing in cloud environments when needed. Perform performance, scalability, and reliability testing of cloud services. Implement and maintain test environments in cloud platforms such as AWS, Azure, or Google Cloud. Supervise the infrastructure to alert on significant events, ensuring the highest level of system performance and reliability. Work with various different partner teams to ensure availability of clusters to test on and take the lead in resolve all issues. Working with tea
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Pure Solutions team as a Senior MLOps Solutions Engineer to architect and build high-scale, enterprise-grade AI/ML solutions. You will be instrumental in integrating Pure Storage platforms with the evolving open-source MLOps ecosystem (Kubeflow, MLflow, Ray) to operationalize the complete machine learning lifecycle. This role requires a creative technologist with deep Python expertise to drive innovation and enable our customers and partners to achieve production AI success. WHAT YOU'LL DO Design and Automate MLOps Pipelines: Lead the development of end-to-end MLOps workflows using CI/CD tools (Git/Jenkins) and orchestration platforms (MLflow/Kubeflow), specifically integrating Pure Storage's FlashBlade, FlashArray, and Portworx as the high-performance data plane for data ingestion, training, and inference. Build High-Performance AI/ML Reference Architectures: Create validated, repeatable deployment models using Infrastructure as Code (e.g., Ansible, Terraform) for AI/ML environments spanning bare metal, virtual machines, and GPU-accelerated Kubernetes clusters, ensuring optimal performance for distributed training. Optimize and Operationalize GPU Inference: Architect and implement solutions for high-throughput, low-latency model serving, utilizing technologies like NVIDIA Triton Inference Server and advanced optimization techniques (quantization, model sharding like DeepSpeed/Megatron-LM, and dynamic bat
Get new cluster lead facilities services jobs by email
Daily job updates · Unsubscribe anytime