ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. Product at Baseten Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the people who defines it. You'll work directly with our founders and with some of the best systems and AI engineers and you'll set the standard for what product looks like here. PMs at Baseten don't sit above engineers - you earn ownership by being technical, finding the truth in front of customers, building great cross-functional relationships, and just shipping great product experiences. The role Once a model is deployed, keeping it fast, reliable, and economical at scale is where production inference is won or lost. You'll own the surface that makes that happen: how deployments autoscale, how traffic is routed, how the system fails over, and how workloads scale across clusters and regions. You'll own these as products end to end - both how they work under the hood and how customers configure and observe them - and you'll help set and define the roadmap that infrastructure and product teams alike can build towards. This space is largely still evolving - think Cloud Infrastructure in mid-2000s. Your job is to make it 10x easier to reliably scale and serve AI models in production and set the market standard. Impact and outcomes you'll drive You will own how workloads scale and where they land — autosca
Jobiba hiring network
Cluster Lead Facilities Services Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster lead facilities services jobs. Use filters to narrow by work mode, employment type, experience and date posted.
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Join WRITER's security team as a staff detection and response engineer and help protect the AI infrastructure that's transforming how the world works. You'll build sophisticated detection systems that identify attacks targeting our AI platform, training data, and model deployments while creating automated response capabilities that scale with our explosive growth. This isn't just traditional security work – you're defending cutting-edge AI/AGI systems against adversaries who are evolving their tactics as fast as AI itself advances. This role combines hands-on security engineering with strategic thinking to stay ahead of novel threats that don't exist in textbooks yet. You'll be the operational arm of our security function, translating threat intelligence into real-time detections, coordinating incident response across multiple teams, and hunting for sophisticated attacks across GPU clusters and distributed training environments. If you're excited by the challen
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Join WRITER's security team as a staff detection and response engineer and help protect the AI infrastructure that's transforming how the world works. You'll build sophisticated detection systems that identify attacks targeting our AI platform, training data, and model deployments while creating automated response capabilities that scale with our explosive growth. This isn't just traditional security work – you're defending cutting-edge AI/AGI systems against adversaries who are evolving their tactics as fast as AI itself advances. This role combines hands-on security engineering with strategic thinking to stay ahead of novel threats that don't exist in textbooks yet. You'll be the operational arm of our security function, translating threat intelligence into real-time detections, coordinating incident response across multiple teams, and hunting for sophisticated attacks across GPU clusters and distributed training environments. If you're excited by the challen
Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. We're looking for an engineer to own the deployment and operational infrastructure of Multigres, our distributed Postgres platform. You'll be responsible for building and maintaining the Multigres Operator, ensuring reliable cloud deployments, and creating the tooling that powers our Kubernetes-based infrastructure. What You’ll Be Responsible for: Build and maintain the Multigres Operator - Maintain our Go-based Kubernetes operator that orchestrates distributed Postgres deployments Architect cloud deployment infrastructure - Design and implement robust deployment patterns for EKS and other Kubernetes platforms Manage storage and networking layers - Work with CSI drivers, persistent volumes, and cross-cloud networking to ensure data reliability and connectivity Develop deployment tooling - Create internal tools and automation for provisioning, scaling, and managing Multigres clusters Ensure operational excellence - Build monitoring, alerting, and diagnostic capabilities into the deployment layer Collaborate across teams - Work with database engineers, SRE, and product teams to deliver seamless deployment experiences You Might Be a Good Fit If You have: Strong systems programming skills - Proficiency in Go and experience building production-grade operators or controllers Deep Kubernetes expertise - Hands-on experience with Kubernetes internals, custom resources, and cloud-managed Kubernetes services (EKS, GKE, AKS) Database operations knowledge - Understanding of database deployment patterns, backup/restore, replication, and high availability Distributed systems experience - Familiarity with consensus protocols, failure scenarios, and designing for resilience Cloud infrastructure background - Experience with cl
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. As an engineer within Palantir's Infrastructure teams, you'll have the opportunity to grow more quickly than you ever imagined as you contribute high-quality code directly to: The shared infrastructure underpinning Palantir Foundry, Palantir Gotham and Palantir Apollo — platforms deployed at the most important institutions across the public and private sectors Rubix and Mission Manager, our new internal-infrastructure business line, used by advanced civil and defence agencies worldwide to power their infrastructure in highly sensitive environments The substrate on which Palantir deploys Foundry and Gotham, powering workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters This means driving investments that improve the velocity and quality of our engineering. Infrastructure at Palantir spans our Foundations, Production Infrastructure and Foundry teams. Teams within Palantir's Foundations organisation are made up of a small number of engineers, each focused on one of four major categories of our infrastructure: Backend Infrastructure Developer Infrastructure Frontend Infrastructure Storage Infrastructure Production Infrastructure organisation, made up of small teams of engineers working on: Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters Apollo: secure, fleet-wide deployment and change-management for complex microservice suites Signals: our full suite of observability and alerting tools Foundry itself is also a developer
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. As a Software Engineer Intern, you’ll have an opportunity to grow more quickly than you ever envisioned as you contribute high-quality code directly to: • Rubix and Apollo, platforms deployed at the most important institutions across the public and private sectors. • Shaping Mission Manager, our new internal-infrastructure business line, used by advanced civil and defense agencies worldwide to power their infrastructure in highly sensitive environments • Building the core capabilities used by advanced civil and defense agencies worldwide to power their infrastructure • Providing the substrate on which Palantir deploys its other platforms, Foundry and Gotham, which power workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters. You’ll join our Production Infrastructure organization, made up of small teams of engineers working on: • Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full suite of observability and alerting tools Core Responsibilities As a Software Engineer Intern at Palantir, you’ll own every phase of the product lifecycle—from generating ideas and designing prototypes to executing features and shipping releases—while being paired with a dedicated mentor who champions your growth. You’ll work hand-in-hand with both technical and non-technical colleagues to uncover real customer problems and
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. In this role, you’ll have an opportunity to grow more quickly than you ever envisioned as you contribute high-quality code directly to: • Rubix and Apollo, platforms deployed at the most important institutions across the public and private sectors • Shaping Mission Manager, our new internal-infrastructure business line, used by advanced civil and defense agencies worldwide to power their infrastructure in highly sensitive environments • Building the core capabilities used by advanced civil and defense agencies worldwide to power their infrastructure • Providing the substrate on which Palantir deploys its other platforms, Foundry and Gotham, which power workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters You’ll join our Production Infrastructure organization, made up of small teams of engineers working on: • Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full suite of observability and alerting tools Core Responsibilities As a Software Engineer at Palantir, you’ll own every phase of the product lifecycle—from generating ideas and designing prototypes to executing features and shipping releases—while being paired with a dedicated mentor who champions your growth. You’ll work hand-in-hand with both technical and non-technical colleagues to uncover real customer problems and deliver solutions that ad
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. As a Software Engineer Intern, you’ll have an opportunity to grow more quickly than you ever envisioned as you contribute high-quality code directly to: • Rubix and Apollo, platforms deployed at the most important institutions across the public and private sectors. • Shaping Mission Manager, our new internal-infrastructure business line, used by advanced civil and defense agencies worldwide to power their infrastructure in highly sensitive environments • Building the core capabilities used by advanced civil and defense agencies worldwide to power their infrastructure • Providing the substrate on which Palantir deploys its other platforms, Foundry and Gotham, which power workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters. You’ll join our Production Infrastructure organization, made up of small teams of engineers working on: • Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full suite of observability and alerting tools Core Responsibilities As a Software Engineer Intern at Palantir, you’ll own every phase of the product lifecycle—from generating ideas and designing prototypes to executing features and shipping releases—while being paired with a dedicated mentor who champions your growth. You’ll work hand-in-hand with both technical and non-technical colleagues to uncover real customer problems and
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. In this role, you’ll have an opportunity to grow more quickly than you ever envisioned as you contribute high-quality code directly to: • Rubix and Apollo, platforms deployed at the most important institutions across the public and private sectors • Shaping Mission Manager, our new internal-infrastructure business line, used by advanced civil and defense agencies worldwide to power their infrastructure in highly sensitive environments • Building the core capabilities used by advanced civil and defense agencies worldwide to power their infrastructure • Providing the substrate on which Palantir deploys its other platforms, Foundry and Gotham, which power workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters You’ll join our Production Infrastructure organization, made up of small teams of engineers working on: • Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full suite of observability and alerting tools Core Responsibilities As a Software Engineer at Palantir, you’ll own every phase of the product lifecycle—from generating ideas and designing prototypes to executing features and shipping releases—while being paired with a dedicated mentor who champions your growth. You’ll work hand-in-hand with both technical and non-technical colleagues to uncover real customer problems and deliver solutions that ad
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. As an engineer within Palantir's Infrastructure teams, you'll have the opportunity to grow more quickly than you ever imagined as you contribute high-quality code directly to: The shared infrastructure underpinning Palantir Foundry, Palantir Gotham and Palantir Apollo — platforms deployed at the most important institutions across the public and private sectors Rubix and Mission Manager, our new internal-infrastructure business line, used by advanced civil and defence agencies worldwide to power their infrastructure in highly sensitive environments The substrate on which Palantir deploys Foundry and Gotham, powering workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters This means driving investments that improve the velocity and quality of our engineering. Infrastructure at Palantir spans our Foundations, Production Infrastructure and Foundry teams. Teams within Palantir's Foundations organisation are made up of a small number of engineers, each focused on one of four major categories of our infrastructure: Backend Infrastructure Developer Infrastructure Frontend Infrastructure Storage Infrastructure Production Infrastructure organisation, made up of small teams of engineers working on: Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters Apollo: secure, fleet-wide deployment and change-management for complex microservice suites Signals: our full suite of observability and alerting tools Foundry itself is also a developer
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. In this role, you’ll have an opportunity to grow more quickly than you ever envisioned as you contribute high-quality code directly to: • Rubix and Apollo, platforms deployed at the most important institutions across the public and private sectors • Shaping Mission Manager, our new internal-infrastructure business line, used by advanced civil and defense agencies worldwide to power their infrastructure in highly sensitive environments • Building the core capabilities used by advanced civil and defense agencies worldwide to power their infrastructure • Providing the substrate on which Palantir deploys its other platforms, Foundry and Gotham, which power workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters You’ll join our Production Infrastructure organization, made up of small teams of engineers working on: • Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full suite of observability and alerting tools Core Responsibilities As a Software Engineer at Palantir, you’ll own every phase of the product lifecycle—from generating ideas and designing prototypes to executing features and shipping releases—while being paired with a dedicated mentor who champions your growth. You’ll work hand-in-hand with both technical and non-technical colleagues to uncover real customer problems and deliver solutions that ad
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. As a Software Engineer Intern, you’ll have an opportunity to grow more quickly than you ever envisioned as you contribute high-quality code directly to: • Rubix and Apollo, platforms deployed at the most important institutions across the public and private sectors. • Shaping Mission Manager, our new internal-infrastructure business line, used by advanced civil and defense agencies worldwide to power their infrastructure in highly sensitive environments • Building the core capabilities used by advanced civil and defense agencies worldwide to power their infrastructure • Providing the substrate on which Palantir deploys its other platforms, Foundry and Gotham, which power workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters. You’ll join our Production Infrastructure organization, made up of small teams of engineers working on: • Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full suite of observability and alerting tools Core Responsibilities As a Software Engineer Intern at Palantir, you’ll own every phase of the product lifecycle—from generating ideas and designing prototypes to executing features and shipping releases—while being paired with a dedicated mentor who champions your growth. You’ll work hand-in-hand with both technical and non-technical colleagues to uncover real customer problems and
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team… GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the industry, powering the object, block, and file storage platforms that underpin hosting, applications, internal infrastructure, and next-generation AI/HPC workloads. If you're passionate about distributed systems, large-scale storage architecture, and solving complex reliability challenges, you'll work on infrastructure that few engineers ever experience. At GoDaddy, Ceph isn't a side project — it's a critical platform. Our environment spans 80+ production clusters, 20,000+ OSDs, and approximately 300 PB of raw storage capacity, supporting tens of billions of objects across multiple continents. The scale demands deep technical expertise in storage architecture, automation, observability, and performance engineering. As a Senior Site Reliability Engineer, you'll be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You'll tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide. What You'll Get to Do… Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage wor
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Location: Austin, TX Job Title: Senior Systems Engineer, Cloudflare Tunnel Role Summary: As a Senior Systems Engineer on the Cloudflare Tunnel team, you will drive the technical vision and architect for future scale, ensuring our product securely connects any machine to the Cloudflare network. You will be responsible for the strategic design of systems across our high-performance global edge network and microservice clusters, providing c
Get new cluster lead facilities services jobs by email
Daily job updates · Unsubscribe anytime