Jobiba hiring network

Cluster Head Last Mile Jobs

315 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cluster head last mile jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About the Role Together AI is building the AI Native Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art GPU cloud infrastructure. The Together Cloud team builds the [Together GPU Clusters](https://www.together.ai/gpu-clusters) flagship IaaS product that provides high-performance, AI-ready GPU clusters through a self-serve cloud console, along with the virtualized infrastructure layer powering Together's inference, RL, and fine-tuning products. As a Staff Software Engineer focusing on AI Compute in the Together Cloud org, you will set technical direction for and build major components of the next generation AI cloud platform – a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware: GB300s/VRs, BlueField DPUs, InfiniBand and dual/quad-plane RoCEv2 fabrics. That virtualized computing platform powers our own SaaS products – inference, RL, and fine-tuning – and serves external cloud customers through self-serve offerings such as on-demand/reserved Kubernetes/Slurm clusters, across dozens of data centers and hundreds of thousands of GPUs. This is an architect-and-build role. Fully automated bootstrapping of GPU data centers, high-performance virtualization of GPU compute and DC networking without compromising isolation or portability, and fault-tolerant decentralized control planes — you'll set the architecture for these across our global and in-DC services, and be a key owner of the hardest parts, in the code as well as the design. Your designs will span the IaaS layer of a greenfield Vera Rubin data center up to the global management plane that schedules capacity across all of them. At this level the job is as much leverage as code: the standards you set and the engineers you grow decide how fast the rest of Together Cloud ships. Responsibilities Own the GPU and network virtualization stack: the hypervisor, kernel, and SDN work that keeps

awsazuregcp
View job →

DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You'll Do: Deploy, configure, and maintain Kubernetes clusters for our microservices architecture. Utilize Git and Helm for version control and deployment management. Implement and manage monitoring solutions using Prometheus and Grafana. Work on continuous integration and continuous deployment (CI/CD) pipelines. Containerize applications using Docker and manage orchestration. Manage and optimize AWS services, including but not limited to EC2, S3, RDS, and AWS CDN. Maintain and optimize MySQL databases, Airflow, and Redis instances. Write automation scripts in Bash or Python for system administration tasks. Perform Linux administration tasks and troubleshoot system issues. Utilize Ansible and Terraform for configuration management and infrastructure as code. Demonstrate knowledge of networking and load-balancing principles. Collaborate with development teams to ensure applications meet reliability and performance standards. Who you are: Bachelor’s degree in engineering (CS / IT) or equivalent degree from a well-known Institute / University. 2+ years of experience in a Site Reliability Engineer role or similar. Proven experience with Kubernetes, Git, Helm, Prometheus, Grafana, CI/CD, Docker, and microservices architecture. Strong knowledge of AWS services, MySQL, Airflow, Redis, AWS CDN. Proficient in scripting languages such as Bash or Python. Hands-on experience with Linux administration. Familiarity with Ansible and Terraform fo

pythonmysqlredis
View job →
N
12 days ago

NVIDIA DGX Cloud is an AI Factory designed to power the next generation of AI and industrial-scale breakthroughs. As a Principal Engineer for Security Architecture, within our Security Engineering organization, you will own a core security domain of the AI factory: the architecture, the paved road that delivers it, and much of the code underneath. You will hold the security design bar across DGX Cloud from inside the teams doing the building, and this is a founding seat on a new team. Security Engineering is a new organization at DGX Cloud, accountable for the security outcome of the platform, and Security Architecture is the function inside it that holds the design bar. Security here is fleet horizontal and stack vertical, so your work will cross every DGX Cloud engineering organization: you will embed with the teams building GPU clusters, control planes, and services, join their designs as a participant rather than an approver, and leave behind systems in which an entire class of risk is no longer possible. There is no architecture review board here and no approval queue. You are a senior IC with deep security domain knowledge, and the security bar holds because you helped set it and then helped ship it. What You Will Be Doing: Own a Security Domain End to End: Take architectural ownership of a core domain of DGX Cloud security, from the design through the system running in production. That could be tenant and GPU workload isolation, workload identity, infrastructure and network, supply-chain provenance, hardened baselines and patching, or deploy-time policy and admission control. Embed with the Teams Building It: Join the design early, write the code, and help land it. The posture is not "you did this wrong." It is "here are the considerations we need to meet, I will help, let's go to work." Build Paved Roads, Not

REMOTEkuberneteslinuxartificial intelligence
View job →
N
12 days ago

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated

pythonkuberneteslinux
View job →

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity AI is a top strategic priority for New Relic, and the Bengaluru design team is at the center of it. The team works across three closely related product clusters: autonomous incident response (SRE Agent, Autopilot, and Intelligent RCA), the intelligence and platform layer that powers them (Ground Truth, Agentic Platform, and New Relic AI), and AIOps for event correlation and incident management. These products share a common design challenge: users need to trust systems that act autonomously, and building that trust through good design is genuinely hard work. This is an on-site role in Bengaluru. Your designers are there, and many of your engineering and product partners are too. Being present — in standups, reviews, and the quick conversations before a decision gets made — is part of how you'll build the relationships that make design effective. You'll also collaborate with design, product, and engineering partners in the US and Spain, so operating across time zones and communicating well in writing are part of the job. You'll manage a small team of designers and own the quality of the work across these products. The design questions here don't have established answers — how do you make an autonomous system legible? How do you build user trust in AI-generated root cause analysis? How do you design a handoff from machine decision to human judgment? If you're already working in AI product design, or actively building toward it, and you want to lead a team

airecruitment
View job →
H
Hyreo
📍 Bengaluru• Full-time
15 days ago

Key Responsibilities : Primary responsibilities :- Installation and configuration of MySQL instances on single or multiple ports. Hands-on experience of working with MysQL 5.7 and MySQL 8. Clear understanding of MysQL Replication process flows , threads , setting up multi node clusters and basic troubleshooting. Understanding of at least one of the backup and recovery methods for MySQL . Strong fundamentals of SQL and able to understand and tune complex SQL queries when needed. Strong fundamentals on the linux system side and monitoring tools like top , iostats , sar etc. At Least couple of years of production hands on experience on medium to big sized MySQL databases. Setting up and maintaining users and privileges management system and troubleshooting relevant access issues. Some exposure to external tools like Percona , ProxySQL , HAP etc. Understand the transaction flows and ACID compliance. Basic understanding of networking concepts . Performing on-call support and should be able to provide the first level support . Excellent verbal and written communication skills. Strong shell scripting skills . Good to have Python . Secondary responsibilities. :- Able to configure and setup NOSQL databases like Mongodb and Cassandra. Ability to learn new technologies along with a team and a positive outlook to understand problems from the business point of view. Qualifications: Proficiency in database management systems such as , MySQL or NoSQL databases. SQL programming and database design skills. Knowledge of database performance tuning and optimization techniques. Familiarity with database security best practices. Scripting and automation skills (Good to have- Python). Good problem-solving and analytical skills. Excellent communication and teamwork skills.

pythonsqlmysql
View job →
H
Hyreo
📍 Bengaluru• Full-time
15 days ago

Key Responsibilities : Primary responsibilities :- Installation and configuration of MySQL instances on single or multiple ports. Hands-on experience of working with MysQL 5.7 and MySQL 8. Clear understanding of MysQL Replication process flows , threads , setting up multi node clusters and basic troubleshooting. Understanding of at least one of the backup and recovery methods for MySQL . Strong fundamentals of SQL and able to understand and tune complex SQL queries when needed. Strong fundamentals on the linux system side and monitoring tools like top , iostats , sar etc. At Least couple of years of production hands on experience on medium to big sized MySQL databases. Setting up and maintaining users and privileges management system and troubleshooting relevant access issues. Some exposure to external tools like Percona , ProxySQL , HAP etc. Understand the transaction flows and ACID compliance. Basic understanding of networking concepts . Performing on-call support and should be able to provide the first level support . Excellent verbal and written communication skills. Strong shell scripting skills . Good to have Python . Secondary responsibilities. :- Able to configure and setup NOSQL databases like Mongodb and Cassandra. Ability to learn new technologies along with a team and a positive outlook to understand problems from the business point of view. Qualifications: Bachelor's degree in Computer Science, Information Technology, or a related field (or equivalent experience). Proficiency in database management systems such as , MySQL or NoSQL databases. SQL programming and database design skills. Knowledge of database performance tuning and optimization techniques. Familiarity with database security best practices. Scripting and automation skills (Good to have- Python). Good problem-solving and analytical skills. Excellent communication and teamwork ski

pythonsqlmysql
View job →
TA
16 days ago

About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw

pythonkuberneteslinux
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
16 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is building the world’s fastest, most efficient AI compute clusters. TT-Fabric is the high-performance nervous system of this platform: the low-level networking layer that lets thousands of RISC-V and AI processors snap together into a single, massively parallel distributed supercomputer. If you love squeezing nanoseconds out of hot paths, designing protocols that move data at absurd scale, and turning messy hardware constraints into elegant distributed systems, this is an opportunity to shape the fabric that future AI models will run on This role is hybrid based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who We Are Strong systems engineer with deep C or C++ experience and comfort working in low-level or bare-metal environments. Passionate about hardware-software interaction, performance tuning, and eliminating inefficiencies at the protocol level. Curious about networking, synchronization, and communication across large clusters. Comfortable reasoning from first principles and challenging industry conventions. Motivated by building infrastructure that directly impacts large-scale

awsaic++
View job →
DU
16 days ago

About the Team The Storage organization builds and operates the online stateful systems and abstractions that DoorDash Engineering depends on: reliable, efficient, secure, and easy to use. Within Storage, the Distributed Caching team owns every caching offering at DoorDash end to end, including ElastiCache (Redis/Valkey), Boulder (our KVRocks-based key-value store for high-QPS feature serving), Entity Cache (read Bill Shen’s engineering blog post, “ High-Performance Proxy Cache for DoorDash Services ”), and the Distributed Lock Service, plus the smart clients (asgard-redis, valkey-go) that sit in front of them. These systems back critical product surfaces across DoorDash, Wolt, and Deliveroo: the team runs roughly 400 ElastiCache clusters serving hundreds of millions of GET requests per second in aggregate, and Boulder, our offline-to-online feature store, serves billions of feature lookups per second at peak. About the Role The team owns provisioning of clusters and the smart clients that sit in front of them, baking in sensible defaults so that other engineering teams get a turnkey caching solution instead of having to run their own. You'll help drive Boulder's evolution to scale further, improve cost efficiency, enhance performance, and support real-time updates; re-platform the Distributed Lock Service onto a strongly consistent backend; and build the self-serve tooling and recommendation engine that let customers describe a workload (QPS, TTL, payload size, latency profile) and get the right backend without talking to a human. You'll go deep on cache invalidation, replication, sharding, compaction, and failover, while shipping the guardrails, automation, and observability that keep this scale operable by a small team. You must be located in San Francisco, Seattle, or the New York Metro Area for this hybrid position. You will report to the Engineering Manager on the Distributed Caching team within the Storage organization. You’re excited about this opportunity b

javaredisaws
View job →
DU
DoorDash USA
📍 San Francisco• Full-time• From $102K/yr
16 days ago

About the Team The DoorDash Research Fellowship is a 3-month program (extendable to 6 months) looking for Summer and Fall 2026 cohorts, for researchers and engineers who want to work on the hardest applied ML and AI problems in local commerce. Fellows are given the resources, autonomy, and access to real-world operational data needed to pursue ambitious research directions — with the goal of producing work that influences both the field and how DoorDash operates at scale. This program is modeled on the best external research fellowships: fellows are treated as independent researchers, not as junior employees on a product team. You pick the problem (within a set of priority areas), you own the direction, and you publish or ship the outcome. You’re excited about this opportunity because you will receive… Dedicated compute allocation sized to the research agenda — GPU clusters for training and inference budgets for experimentation Full access to DoorDash's research infrastructure — our internal RL stack, training and evaluation pipelines, RL environments built on real operational systems, agent evaluation harnesses, and the tooling our own research teams use day-to-day. Fellows are first-class users, not sandboxed visitors. Access to DoorDash operational data — real-world datasets spanning logistics, merchant operations, consumer behavior, and marketplace dynamics, under appropriate data governance Research mentorship from senior researchers and engineering leaders at DoorDash, plus a named research sponsor for each fellow who meets with you weekly and is accountable for unblocking your work Speaker series featuring leading researchers and practitioners from academia and industry — faculty from top ML programs, research leads from frontier AI labs, and senior operators from across tech. Fellows get dedicated 1:1 time with speakers when possible. A cohort of fellows working alongside you — a small, tight-knit group of researchers tackling different problems but sharing

gitrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
16 days ago

About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,

awskuberneteslinux
View job →
O
Okta
📍 Bellevue• Full-time• From $194K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities—like reducing costs, and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimizing costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimize AWS cloud infrastructure, including EKS,ECS, S3, VPCs, RDS, IAM, and more. Implement b

pythonawsdocker
View job →
P
1mo ago

Role Summary This role serves as the single point of accountability between USMAPPS and the Specialty Care Business Unit, owning access strategy, payer marketing, and brand-level contracting and pricing strategy across the full Specialty Care portfolio. The role drives and owns outcomes, integrating all USMAPPS capabilities into one coherent access plan that directly supports the Specialty Care BU President and franchise leads. The VP is expected to operate at full strategic weight, lead a team aligned to Specialty Care brands and franchise clusters, and be the single person the Specialty Care BU holds accountable for market access results. The Vice President US Market Access Lead – Specialty Care reports directly to the Senior Vice President of US Market Access & Pfizer Patient Services (USMAPPS) with a dotted line to the Specialty Care BU President. This position requires close partnership with the Specialty Care BU President, franchise leads, Strategic Contracting & Analytics, Strategic Account Management, the Patient Services, and USMAPPS leadership. The role sits on the Specialty Care BU leadership team and on the USMAPPS Leadership Team. Role Responsibilities 1. Specialty Care Access Strategy Ownership </spa

airecruitment
View job →
P
Pfizer
📍 New York• $300.1K – $500.1K/yr
1mo ago

Role Summary This role serves as the single point of accountability between USMAPPS and the Primary Care Business Unit, owning access strategy, payer marketing, and brand-level contracting and pricing strategy across the full Primary Care portfolio. The role drives and owns outcomes, integrating all USMAPPS capabilities into one coherent access plan that directly supports the Primary Care BU President and franchise leads. The VP is expected to operate at full strategic weight, lead a team aligned to Primary Care brands and franchise clusters, and be the single person accountable the Primary Care BU holds accountable for market access results. The Vice President, US Market Access Lead – Primary Care reports directly to the Senior Vice President of US Market Access & Pfizer Patient Services (USMAPPS) with a dotted line to the Primary Care BU President . This position requires close partnership with the Primary Care BU President, franchise leads, Strategic Contracting & Analytics, Strategic Account Management, the Patient Services, and USMAPPS leadership. The role sits on the Primary Care BU leadership team and on the USMAPPS Leadership Team. Role Responsibilities 1. Primary Care Access Strategy Ownership </

airecruitment
View job →
🔔

Get new cluster head last mile jobs by email

Daily job updates · Unsubscribe anytime