The Green PVC Complex Head will lead the operational and maintenance activities for the Green PVC project at Mundra, Gujarat, ensuring seamless pre-commissioning, commissioning, and stabilization of the facility. This role is pivotal in optimizing operations, maximizing efficiency, and driving value engineering for a petrochemical complex with a capacity of 1 MMT per annum in its initial phase. The position supports Adani Petchem’s strategic goal of establishing a world-class petrochemical cluster, contributing to India’s ambition of becoming a global petrochemical hub. Source: Adani Group | Job ID: 54238
Jobiba hiring network
Cluster Lead Facilities Services Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster lead facilities services jobs. Use filters to narrow by work mode, employment type, experience and date posted.
NVIDIA is looking for a hands-on Solutions Architect Manager to lead a team of GPU, networking & software solution architects and engineers. Do you want to build and lead a group that designs, debugs, and deploys new AI hardware and software technologies into production in customer data centers? As part of the NVIDIA SA organization, you will drive people and technical leadership for end-to-end solutions deployments at some of NVIDIA's most strategic technology customers, while directly contributing to designs and deep-dive debugging and shaping our product roadmap with customer feedback. What you will be doing: Recruit & manage a team of solutions architects, system/network and software engineers focused on large-scale GPU and AI networking deployments. Set priorities, allocate resources, mentor, and ensure high-quality customer delivery across multiple concurrent projects - while remaining directly involved in key technical reviews, design decisions, and critical debug efforts. Provide deep subject-matter expertise in advanced GPU and network systems and serve as the senior technical point of contact for strategic customers. Personally lead and guide complex compute/network configuration and performance debugging, working side-by-side with your team to deliver performant, reliable clusters. Guide your team as they lead network / compute / software architecture discussions, and support server, network, and cluster bring-up, including on-site data center work where needed. Systematically collect and synthesize customer-specific requirements across your portfolio. Partner with GPU/Network Systems Engineering, Product Management, and Sales to influence roadmap priorities and packaging of reference designs and solutions. Demonstrate SME in advanced GPU & network systems and be a trusted technical advisor to NVIDIA's strategic customers. Bring customer-sp
About the role (Remote) We're hiring a Senior Product Designer to join our global UX team. You'll own design within your product cluster — working directly with the engineers and PMs in those areas to define problems, shape solutions, and ship work that holds up. This is an end-to-end role. You lead discovery, define the interaction model, produce specs, and stay close through implementation. You're also expected to contribute to the systems and standards the whole team relies on. The Senior Product Designer will report to our Product Design Manager, and will work with our EU, UK and US based designers. You're part of a distributed team, so async clarity and proactive communication are as important as craft. What you’ll do Lead the design of complex AI-powered interactions in your product cluster, accounting for trust, transparency, fallback states, and human-in-the-loop considerations Embrace the use of modern AI-based tools like Claude, Figma Make, Cursor, Vercel, Loveable or equivalent not as theatre, but as a meaningful multiplier to your workflow. Partner with our UX research team to lead interviews, usability tests, and competitive analysis and translate findings into clear product direction and guidance Develop presentations, wireframes, mockups, and prototypes that effectively communicate interaction and design intent across web and mobile surfaces. Partner with PMs to shape requirements and with engineers to protect design intent through delivery Contribute to Showpad's design systems — adding components, documenting patterns, raising consistency across product areas Communicate design rationale through strategic storytelling: you can walk a PM, an engineer, and a senior stakeholder through the same decision and each one gets what they need Establish scalable AI-assisted workflows for your own work — research synthesis, rapid prototyping, concept generation — and help others on the team use them effectively Deliver UX solutions that measurably improve
Atlas Growth is a cloud engineering group whose mission is to guide customers through their app development journey—from cluster configuration, data modeling and load testing, to running a production workload at scale. We use an in-house experiments platform which helps us validate our features quickly, releasing only the work that positively impacts our customers. Our engineers participate in cross-functional “squads” with product, design, analytics, and research focusing on a single metric (e.g. retention). Our engineering team is part of a larger Atlas Core Engineering org, building foundational elements of MongoDB’s developer data platform. Atlas Growth 2 builds customer-facing features in Atlas and sits alongside other Growth engineering teams. Recent projects include an AI Chatbot for cluster creation, a recommendation system that offers tips for better database performance, and a pricing page designed to optimize conversion rates. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Role Overview Atlas Growth seeks a mid-level software engineer (Software Engineer 3). SE3s are solid contributors to projects they work on and often lead projects of their own. They act in accordance with MongoDB’s core values and leadership principles, and are actively working toward a Senior role. Candidate Profile 3+ years of software engineering experience, with fluency in TypeScript/JavaScript, and experience with a modern framework (e.g. React) Proficiency in Java, Go, C++/C, or a similar compiled language is a plus Experience writing database queries, either document-based or relational Experience writing and reviewing technical specs, and leading small projects Interest in A/B testing or product design Expectations Contribute readable and well-tested code to ongoing projects Collaborate closely with product and design partners to implement and iterate on new customer facing features Write scope and technical spec docs for new projects
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! About the Role We are seeking a seasoned Manager, Software Engineering with 12+ years of experience to lead our Database Engineering and Cloud Infrastructure team. In this role, you will lead a team of high-performing engineers responsible for architecting, scaling, and optimizing multi-cloud relational and in-memory database platforms. You will bridge technical execution, engineering leadership, and strategic infrastructure planning across AWS and Azure environments. Key Responsibilities Technical Leadership & Architecture Lead the architectural design and operations of enterprise-grade, multi-cloud relational databases across AWS (RDS PostgreSQL, MySQL, Aurora) and Azure (Database for PostgreSQL/MySQL, Azure SQL Managed Instance). Drive high-availability architecture strategies, including Multi-AZ deployments, auto-failover groups, read replica scaling, and cross-region disaster recovery (DR). Oversee zero-downtime operations, including major-version engine upgrades, schema migrations, and blue/green deployment strategies. In-Memory Infrastructure & Open-Source Strategy Manage scale operations for in-memory datastores (AWS ElastiCache, Azure Cache for Redis), focusing on cluster mode operations, eviction policies, and persistence tuning. Spearhead open-source caching initiatives and migration pathways from Redis to Valkey (e.g., AWS ElastiCache for Valkey) using zero-downtime tools like RedisShake to ensure open-source license compliance and optimize cloud spend. Aut
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale is seeking a Staff Software Engineer to lead the technical vision for our Infrastructure team. As a Staff Engineer, you will be responsible for the architectural evolution of our control plane and data plane, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. You will act as a force multiplier, setting the standards for Kubernetes-based cloud-native infrastructure while mentoring engineers and driving cross-functional alignment across the Ray open-source community and our proprietary product teams. Key Responsibilities Architectural Leadership: Define and drive the multi-year technical roadmap for services that orchestrate Ray clusters across diverse cloud and on-premises environments. Systemic Optimization: Lead the design and optimization of high-performance control plane components specifically tailored for large-scale, heterogeneous AI/ML workloads. Platform Reliability: Establish the organization-wide standards for the reliability, scalability, and observability of Anyscale-managed infrastructure. Strategic Integration: Direct the long-term strategy for accelerator integration (GPUs, TPUs) and container management to ens
About the Role: We are looking for a Senior DevOps Engineer to join our DevOps team at K Health. You will own and evolve the infrastructure underpinning a healthcare AI platform serving patients and enterprise health system partners. This is a high-ownership role: you will architect and operate cloud environments across K Health and its enterprise partners, lead complex infrastructure migrations, drive disaster recovery programs, and help build the next generation of AI-powered operations tooling. You will also mentor junior engineers and collaborate closely with product and engineering teams across the company. This is a hybrid role based in New York City (4 days/week in office) and includes participation in a daytime on-call rotation. What you will do: Own the design, implementation, and evolution of our GKE-based Kubernetes infrastructure across K Health and enterprise partner environments. Build and maintain our Terraform modular infrastructure library, including reusable modules with automated testing, across GCP, Cloudflare, and AWS. Architect, build, and maintain GitLab CI/CD shared pipeline templates used by all engineering teams (build, test, security scanning, deployment). Own and maintain self-hosted infrastructure software running in-cluster, including GitLab, ArgoCD, Langfuse, DependencyTrack, NGINX Ingress, and others. Implement and support security and compliance controls across infrastructure and the software supply chain - secrets management, pipeline secret detection, container scanning, SOC2 and HIPAA. Drive disaster recovery readiness: design failover scenarios, author runbooks, and lead periodic DR tests. Lead development of AI-powered operations tooling and agentic infrastructure. Monitor, troubleshoot, and improve production system reliability; respond to incidents during on-call shifts. Mentor junior DevOps engineers and establish team-wide engineering standards. What we are looking for: 5+ years of experience in DevOps, platform engineering,
About the Team OpenAI's Industrial Compute organization builds and operates the infrastructure required to train and serve frontier AI models. The Capacity Planning team connects rapidly changing research and product demand with the compute, networking, storage, power, data center, hardware, and operational resources required to make that demand executable. About the Role We are seeking a Technical Program Manager to build and lead capacity planning across OpenAI's large-scale AI infrastructure. You will translate uncertain workload demand into clear infrastructure requirements, allocation decisions, supply commitments, activation priorities, and long-range capacity strategies. This role sits at the intersection of research, engineering, infrastructure, finance, sourcing, deployment, and operations. You will create the planning models, operating cadences, governance mechanisms, and source-of-truth systems that allow teams to understand what capacity is required, what is available, what is at risk, and what decisions must be made. This is not a finance-only forecasting or reporting role. Success requires technical fluency across the infrastructure stack, strong analytical judgment, and the ability to move consequential decisions forward when requirements, timelines, and supply conditions change quickly. Key Responsibilities Own capacity-planning processes across near-term workload allocation, quarterly execution, and longer-range infrastructure horizons. Translate research, training, inference, and product demand into compute, accelerator, cluster, networking, storage, rack, power, and site requirements. Develop scenarios that make assumptions, confidence levels, constraints, sensitivities, and decision points explicit. Reconcile requested demand against contracted, delivered, installed, activated, and workload-usable capacity. Partner with research and engineering teams to understand workload priorities, technical dependencies, utilization patterns, and changing req
About the Team The Core Services organization builds and runs the mission-critical online services that product teams rely on in production. We own foundational distributed systems and platform capabilities that enable reliable execution, high-performance services, and large-scale file/data needs across our products. This team is distinct from developer infrastructure and data infrastructure—our focus is production service foundations and core runtime services. About the Role We’re hiring an Engineering Manager, Core Services to help lead teams responsible for highly reliable, high-scale distributed systems that sit on the critical path for OpenAI products. Your team will own foundational production systems that OpenAI’s product engineering teams build on. You’ll collaborate closely with product and infrastructure partners to ship reliable services quickly, and help scale systems and teams as OpenAI grows. You’ll partner closely with senior engineering leaders to scale the org, mature operations, and drive major platform initiatives. This role requires strong technical ability. You’ll be responsible for: Managing and growing a high-performing team of infrastructure engineers. Leading teams building and operating large, critical production platforms, including cluster reliability, scaling, and rollout safety. Building and operating mission-critical distributed systems with strong operational rigor (SLOs, incident response, capacity planning, reliability). Setting technical direction for platform foundations such as workflow/orchestration capabilities, large-scale file/blob/storage services, and core service foundations. Partnering with a broad set of stakeholders, including product engineering, adjacent infrastructure teams, and (where relevant) finance/cost partners. Coaching, mentoring, and developing engineers and emerging leaders. You might thrive in this role if you: Have significant experience leading teams that run mission-critical infrastructure in production
At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale's need to detect and respond to security events across its production and corporate environments is growing as the company scales. We're looking for a Senior Detection and Response Engineer to own detection engineering and to lead incident response when it counts, coordinating the response and driving it to resolution. This is a high-ownership role with real room to shape how detection and response works at Anyscale. You will own the detection pipeline, the response runbooks, and incident response, reporting to the Head of Security and partnering with engineering. This role is based in India. In your first year, success looks like strong detection coverage across our cloud, endpoint, and runtime telemetry, a working correlation and alerting pipeline, and incident response runbooks that have been exercised in practice. What You'll Do Own and build detection coverage across cloud, endpoint, and runtime telemetry. Own a centralized correlation and alerting capability that turns telemetry into actionable detections. Own incident response: runbooks, escalation paths, and coordination during an incident, across corporate and production environments. Drive detection of anomalous activity across the environments
The worldwide data management software market is massive. At MongoDB we are transforming industries and empowering developers to build amazing apps that people use every day. We are the leading modern data platform and the first database provider to IPO in over 20 years. Join our team and be at the forefront of innovation and creativity. MongoDB is seeking a Software Engineer 3 to join the Atlas Clusters Platform team. The team is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. The Atlas Clusters Platform team develops the foundational orchestration platform behind MongoDB Atlas. Our systems drive cluster planning and execution, evolve the Atlas control plane toward service-oriented architecture, and provide critical infrastructure that help Atlas run safely and efficiently across cloud environments. We are looking to speak to candidates who are based in New York City for our hybrid working model. What you’ll do Build and design new features for MongoDB Atlas Contribute to and lead complex technical projects Work closely with product and design teams, considering the user’s perspective while building technical solutions Work with customers and support engineers to fix issues Collaborate with team members to develop our codebase, best practices, and design principles Learn from and mentor other team members We’re looking for someone who Has at least 3 years of professional software development experience Is skilled at writing large-scale, distributed backend systems in a compiled language (Java, C#, Go, etc.) Is comfortable working across the stack of a modern web application (e.g. React, TypeScript, Kubernetes) Has experience with at least one major cloud provider technology (AWS, Azure, GCP) Has led the launch of a new feature and maintained it in production Is eager to solve tough problems Has excellent
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Anyscale is looking for an experienced and hands-on engineering leader to lead our Customer Engineering team. This is a critical leadership role within our Go-To-Market organization, responsible for delivering exceptional technical support while helping ensure customer experiences directly influence the evolution of our platform. You will lead a highly technical team responsible for supporting customers running production AI workloads on Anyscale. Your team will resolve complex technical issues, manage customer escalations, and partner closely with Product and Engineering to ensure customer feedback is translated into meaningful product improvements. Success in this role requires balancing operational excellence with strong technical leadership. Beyond resolving individual customer issues, you will help the team identify recurring patterns, improve support workflows, expand customer self-service, and leverage automation, diagnostics, and engineering best practices to improve both the customer experience and the product over time. As opportunities arise, your team may also contribute tooling, documentation, automation, or occasional product fixes that help eliminate recurring sources of customer fri
At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: As a Forward Deployed Engineer at Anyscale, you will partner directly with our most strategic customers, including Spanish-speaking customers across Latin America and other regions, to ensure they achieve meaningful business outcomes with Ray and the Anyscale platform. Embedded within customer teams, you’ll act as a trusted advisor, aligning technical solutions with customer priorities, accelerating time-to-value, and driving adoption at scale. You’ll work across customer organizations — from technical leadership to individual contributors — to scope and deliver impactful solutions. By connecting insights from the field back to our product and engineering teams, you’ll help shape Anyscale’s roadmap and ensure we remain focused on solving our customers’ most critical challenges. In this role, you will: Work onsite with key customers to lead proof-of-value engagements, deployments, and enterprise adoption Translate business objectives into technical solutions that demonstrate clear ROI and strategic impact Build and deliver high-impact demos, reference architectures, and enablement programs tailored to customer needs Act as a trusted advisor across all levels of the organization, ensuring confidence in Anyscale and
Anyscale Platform Engineering Leader About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for an experienced Engineering leader to lead our Infrastructure, SRE and Enterprise Governance Engineering teams. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud using Ray - the popular open source platform used by companies like Netflix, Uber, Instacart and others - seamless. In this position, you will guide the vision, technical direction of the team, and recruit, enable a high-performing engineering team that delivers critical values to developers and Anyscale customers by solving complex distributed systems challenges. You will oversee and drive the strategy and execution of components which includes cluster launcher, cloud providers (AWS/GCP/Azure/etc.), Kubernetes support, cluster autoscaling, control plane, data plane, reliability, billing stack, production database and related components. You will closely work with our customers and our field engineering team to solve their problems, understand their challenges and make sure they are successful. We'd love to hear from you if you have: Solid engineering management experience leading produ
The opportunity Datadog’s Infrastructure products help engineers understand and operate the systems their applications depend on. Our customers work in complex environments like Kubernetes and serverless, where infrastructure changes constantly, information is dense, and decisions about reliability, performance, and cost are closely connected. We’re looking for a Staff Product Designer to join Modern Compute, with an initial focus on Containers Autoscaling. Autoscaling helps engineering teams make better decisions about how their applications and infrastructure use resources. Designing these experiences requires making deeply technical systems understandable, helping customers act with confidence, and fitting into the tools and workflows they already use. The team is rethinking how workload and cluster autoscaling come together as a more coherent product experience. This includes how customers get started, understand recommendations, evaluate value, and safely apply changes across their environments. The work also connects to other parts of Datadog, including observability, Cloud Cost Management, permissions, and AI-assisted workflows. As a Staff Product Designer, you will help define that direction and lead the work from early problem framing through shipped product. You will partner closely with product and engineering, bring a high level of interaction and visual craft to complex workflows, and help raise the quality of design across Modern Compute. At Datadog, we place value in our office culture, the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to help our Datadogs find a work-life rhythm that works for them. What you’ll do Lead end-to-end product design for Modern Compute, initially focused on our Autoscaling product. Help define the product direction for an area that is still evolving, from early framing and exploration through detailed design and delivery. Design clear, trustwort
Get new cluster lead facilities services jobs by email
Daily job updates · Unsubscribe anytime