At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: As a Forward Deployed Engineer at Anyscale, you will partner directly with our most strategic customers, including Spanish-speaking customers across Latin America and other regions, to ensure they achieve meaningful business outcomes with Ray and the Anyscale platform. Embedded within customer teams, you’ll act as a trusted advisor, aligning technical solutions with customer priorities, accelerating time-to-value, and driving adoption at scale. You’ll work across customer organizations — from technical leadership to individual contributors — to scope and deliver impactful solutions. By connecting insights from the field back to our product and engineering teams, you’ll help shape Anyscale’s roadmap and ensure we remain focused on solving our customers’ most critical challenges. In this role, you will: Work onsite with key customers to lead proof-of-value engagements, deployments, and enterprise adoption Translate business objectives into technical solutions that demonstrate clear ROI and strategic impact Build and deliver high-impact demos, reference architectures, and enablement programs tailored to customer needs Act as a trusted advisor across all levels of the organization, ensuring confidence in Anyscale and
Jobiba hiring network
Cluster Head Last Mile Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster head last mile jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Anyscale Platform Engineering Leader About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for an experienced Engineering leader to lead our Infrastructure, SRE and Enterprise Governance Engineering teams. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud using Ray - the popular open source platform used by companies like Netflix, Uber, Instacart and others - seamless. In this position, you will guide the vision, technical direction of the team, and recruit, enable a high-performing engineering team that delivers critical values to developers and Anyscale customers by solving complex distributed systems challenges. You will oversee and drive the strategy and execution of components which includes cluster launcher, cloud providers (AWS/GCP/Azure/etc.), Kubernetes support, cluster autoscaling, control plane, data plane, reliability, billing stack, production database and related components. You will closely work with our customers and our field engineering team to solve their problems, understand their challenges and make sure they are successful. We'd love to hear from you if you have: Solid engineering management experience leading produ
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale is seeking a Staff Software Engineer to lead the technical vision for our Infrastructure team. As a Staff Engineer, you will be responsible for the architectural evolution of our control plane and data plane, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. You will act as a force multiplier, setting the standards for Kubernetes-based cloud-native infrastructure while mentoring engineers and driving cross-functional alignment across the Ray open-source community and our proprietary product teams. Key Responsibilities Architectural Leadership: Define and drive the multi-year technical roadmap for services that orchestrate Ray clusters across diverse cloud and on-premises environments. Systemic Optimization: Lead the design and optimization of high-performance control plane components specifically tailored for large-scale, heterogeneous AI/ML workloads. Platform Reliability: Establish the organization-wide standards for the reliability, scalability, and observability of Anyscale-managed infrastructure. Strategic Integration: Direct the long-term strategy for accelerator integration (GPUs, TPUs) and container management to ens
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role If you joined us today, you get to author the strategy by which each of Palantir’s software platforms - Foundry, Gotham, Apollo - achieves full containerization across an intimidating diversity of infrastructure types. Tomorrow, you get to do the same for Palantir’s ever expanding customer community. You will drive those goals by building elegant, robust APIs powered by K8s controllers which bridge the gap between a raw Kubernetes cluster and a fully-featured, infrastructure agnostic runtime that can scale to the operational needs of 100s of specialized microservices. Joining you on that journey is a highly motivated team with a diverse group of backgrounds and skillsets, brimming with ambition.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role If you joined us today, you get to author the strategy by which each of Palantir’s software platforms - Foundry, Gotham, Apollo - achieves full containerization across an intimidating diversity of infrastructure types. Tomorrow, you get to do the same for Palantir’s ever expanding customer community. You will drive those goals by building elegant, robust APIs powered by K8s controllers which bridge the gap between a raw Kubernetes cluster and a fully-featured, infrastructure agnostic runtime that can scale to the operational needs of 100s of specialized microservices. Joining you on that journey is a highly motivated team with a diverse group of backgrounds and skillsets, brimming with ambition.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Apollo is Palantir’s autonomous software management and deployment platform. It enables seamless, continuous delivery of mission-critical software (Foundry, Gotham, AIP) across a vast range of environments: on-prem, public cloud, disconnected (air-gapped) networks, and highly regulated settings (including IL-5 and FedRAMP). As a Software Engineer on the Apollo team, you’ll build and operate a large-scale distributed system to allow the remote operation and maintenance of Kubernetes clusters. Our mission is to extract the entire state of a cluster into a portable, high-performance artifact within minutes, enabling full and almost instant cluster reconstruction from the ground up—all while pushing the limits of speed, reliability, and scale. You’ll design and implement backup and restore solutions for Kubernetes, leveraging proprietary compression infrastructure tailored to Palantir’s unique deployment models. You’ll also build and optimize our container artifact store, which is based on the OCI (Open Container Initiative) distribution spec—the industry standard for storing and distributing container images and artifacts. You’ll own the backbone of every environment Apollo supports, from hyperscalers to Army trucks. If you’re excited by challenges at the intersection of container technologies like OCI and docker, storage, and distributed systems, you’ll find opportunities here to dive deep into storage formats and low-level optimizations, where milliseconds matter. As we increasingly automate cluster creation and management on diverse hardware, you’ll play a key role in scaling Palantir’s presence at the edge and solving tough distributed systems proble
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Apollo is Palantir’s autonomous software management and deployment platform. It enables seamless, continuous delivery of mission-critical software (Foundry, Gotham, AIP) across a vast range of environments: on-prem, public cloud, disconnected (air-gapped) networks, and highly regulated settings (including IL-5 and FedRAMP). As a Software Engineer on the Apollo team, you’ll build and operate a large-scale distributed system to allow the remote operation and maintenance of Kubernetes clusters. Our mission is to extract the entire state of a cluster into a portable, high-performance artifact within minutes, enabling full and almost instant cluster reconstruction from the ground up—all while pushing the limits of speed, reliability, and scale. You’ll design and implement backup and restore solutions for Kubernetes, leveraging proprietary compression infrastructure tailored to Palantir’s unique deployment models. You’ll also build and optimize our container artifact store, which is based on the OCI (Open Container Initiative) distribution spec—the industry standard for storing and distributing container images and artifacts. You’ll own the backbone of every environment Apollo supports, from hyperscalers to Army trucks. If you’re excited by challenges at the intersection of container technologies like OCI and docker, storage, and distributed systems, you’ll find opportunities here to dive deep into storage formats and low-level optimizations, where milliseconds matter. As we increasingly automate cluster creation and management on diverse hardware, you’ll play a key role in scaling Palantir’s presence at the edge and solving tough distributed systems proble
The opportunity Datadog’s Infrastructure products help engineers understand and operate the systems their applications depend on. Our customers work in complex environments like Kubernetes and serverless, where infrastructure changes constantly, information is dense, and decisions about reliability, performance, and cost are closely connected. We’re looking for a Staff Product Designer to join Modern Compute, with an initial focus on Containers Autoscaling. Autoscaling helps engineering teams make better decisions about how their applications and infrastructure use resources. Designing these experiences requires making deeply technical systems understandable, helping customers act with confidence, and fitting into the tools and workflows they already use. The team is rethinking how workload and cluster autoscaling come together as a more coherent product experience. This includes how customers get started, understand recommendations, evaluate value, and safely apply changes across their environments. The work also connects to other parts of Datadog, including observability, Cloud Cost Management, permissions, and AI-assisted workflows. As a Staff Product Designer, you will help define that direction and lead the work from early problem framing through shipped product. You will partner closely with product and engineering, bring a high level of interaction and visual craft to complex workflows, and help raise the quality of design across Modern Compute. At Datadog, we place value in our office culture, the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to help our Datadogs find a work-life rhythm that works for them. What you’ll do Lead end-to-end product design for Modern Compute, initially focused on our Autoscaling product. Help define the product direction for an area that is still evolving, from early framing and exploration through detailed design and delivery. Design clear, trustwort
The worldwide data management software market is massive (IDC forecasts it to be $138 billion by 2026). At MongoDB, we are transforming industries and empowering developers to build amazing apps that people use every day. We are the leading modern data platform and the first database provider to IPO in over 20 years. Join our team and be at the center of innovation and creativity. MongoDB is seeking a Sr. Staff Software Engineer to join the Atlas Core Data Services organization. The organization is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product, along with the API Platform and Developer Tools. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. The Atlas Core Data Services organization builds the software that manages the Atlas cluster infrastructure hosted on the three major cloud providers (AWS, Azure, and GCP), as well as the software that manages the MongoDB database hosted on that infrastructure. We are constantly challenged to design features that ensure Atlas clusters are secure, available, durable, and performant while running large-scale, critical workloads. The Sr. Staff Engineer in this role will drive innovation across the organization and the company, setting technical standards and direction that enable future growth and velocity. We are looking for engineers with the experience and high standards needed to lead at that scale. Our organization champions a strong culture of inclusivity, diversity, and collaboration. If you want to be a deeply technical leader on a collaborative team that applies systems expertise to build the foundational infrastructure of a popular database, join us. Let's build a faster, more reliable, and highly scalable database platform together. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Responsibilities Define standards and vision for the mission-critical Atlas SaaS data
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Fleet Management team provides the core runtime environment that empowers our developers to build and ship products to delight our customers. We manage the end-to-end lifecycle of our Kubernetes fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As our infrastructure scales to support new use cases and products, we are spearheading a migration from Terraform-based Infrastructure as Code (IaC) to an Operator-driven lifecycle management model. This role can be based out of our Austin, Boston, Los Angeles, New York City, Raleigh, or San Francisco offices, remotely in the United States region, or our European office in Dublin. Responsibilities Contribute to developing and maintaining a scalable and secure runtime environment on top of Kubernetes that supports product needs across MongoDB Provide internal support for our Kubernetes ecosystem, partnering with engineering teams to help them solve domain-specific problems Participate in a 24/7 on-call rotation to resolve critical issues Prioritize blameless post-mortems and dedicate engineering time to systemic fixes, ensuring you aren’t paged for the same issue twice You may be a good fit if you Have 6+ years of experience in software development and operating distributed systems Are proficient in Go, Python, or a similar language, with a strong commitment to code quality and testing practices (writing unit, integration, and E2E tests) Have deep experience using and extending containerization technologies, preferably Kubernetes Have a solid understanding
MongoDB is seeking an Engineering Manager to join the Atlas Growth Team. The team is responsible for improving and creating new features for MongoDB Atlas, our developer data platform that accounts for 65% of the company’s revenue. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. The Atlas Growth team focuses on improving the customer experience via new product experiences and refinements to existing user flows. The team works closely with cross-functional partners, as well as other MongoDB engineering teams to bring new visions to life. We are constantly challenged to design features such as new onboarding flows, monetization improvements, and advanced cluster management tools for our large B2B customer base. We are looking to speak to candidates who are based in Dublin for our hybrid working model. What You’ll Do Manage a team of engineers to design, build and test new features for MongoDB Atlas Contribute to and lead complex technical projects Work with cross-functional stakeholders to design the team's roadmap, defining delivery dates that balance technical feasibility with the pace of the market Work closely with product, design and analytics teams, considering the user’s perspective while building technical solutions Collaborate with team members to develop our codebase, best practices, and design principles Learn from and mentor an impassioned array of team members We’re Looking for Someone Who Has at least 5 years of professional software development experience Has at least 2 years of people management experience Is skilled at writing large-scale, distributed backend systems in a compiled language (Java, C#, Go, etc.) Is comfortable working across the stack of a modern web application (e.g. React, TypeScript, React Testing Library) Has experience with at least one major cloud provider technology (AWS, Azure, GCP) Has a deep understanding of product analytics Has experience with A/B testing and
About the Team This team builds and operates the systems that enable OpenAI researchers to run reliable, scalable, and efficient research workflows. The team sits close to research and works across infrastructure, systems, and automation to make sure researchers have the tools and environments they need to move quickly. The work spans software engineering, infrastructure, systems administration, cluster operations, and reliability engineering. As OpenAI’s infrastructure evolves from bespoke bare-metal systems toward more standard, scalable platforms, the team needs engineers who can understand how systems work end-to-end and build the right abstractions without reinventing the wheel. About the Role As a Software Engineer on this team, you will build and operate the infrastructure that supports frontier research and critical research-facing systems. You will work on systems that sit close to the metal, but the role is not limited to classic operations or sysadmin work. We are looking for someone who can reason about networking, bootstrapping, Kubernetes, scalability, automation, and reliability - while also writing software to make these systems better over time. This role is a strong fit for an independent, high-ownership engineer who enjoys reliability-heavy infrastructure work but still wants to build. You do not need to come in as a kernel expert or highly algorithmic optimization engineer, but you should be deeply curious about infrastructure, comfortable debugging complex systems, and excited to support researchers doing novel work. We expect you to: Build and operate reliable infrastructure for research workloads and research-facing services. Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services. Improve cluster bootstrapping, provisioning, automation, and deployment workflows. Debug issues across networking, compute, storage, orchestration, and service reliability layers.
This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions tailored to the demands of advanced AI workloads. We work across the full stack—from silicon to system integration—partnering closely with internal teams and external vendors to define and deliver next-generation AI infrastructure. Our team focuses on defining scalable, high-performance system architectures and reference designs that balance performance, cost, and operational efficiency across rapidly evolving technologies. About the Role We are seeking a 3P Architect to define and drive rack- and cluster-level reference designs in collaboration with external partners. This role is responsible for translating workload requirements and system-level goals into concrete architectures, aligning partners on critical design attributes, and ensuring vendor roadmaps meet our infrastructure needs. You will work closely with performance modeling and internal architecture teams to evaluate tradeoffs, while owning the end-to-end definition and execution of third-party system designs. This includes identifying gaps in current technologies, driving vendor development, and shaping future infrastructure capabilities. This role requires strong system intuition, cross-functional leadership, and the ability to operate effectively across internal teams and external ecosystems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Define rack- and cluster-level reference architectures for AI infrastructure deployments. Translate workload requirements into clear system design specifications and partner deliverables. Collaborate with performance modeling teams to evaluate architectural tradeoffs and system behaviors. Align internal stakeholders and external partners on critical system attributes (performance, cost, power, reliability, scalability). Identify gaps in current technology offerings and dr
Get new cluster head last mile jobs by email
Daily job updates · Unsubscribe anytime