About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t
Jobiba hiring network
Cluster Head Last Mile Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster head last mile jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical. The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly. About the Role We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity. In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity. This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions. This role is based in San Francisco, CA, with travel as needed. Key Responsibilities Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint Drive readiness to convert contracted compute capacity into schedulable production clusters Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives Build integrated schedules spanning procurement, logistics, installation, st
NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world. NVIDIA GH200 superchip provides performance and productivity required for strong scaling for HPC and generative AI workload. Scale out is inherent to design of this massive superchip. We are looking for expert engineers to come and help design rack level solutions for next generation scaling AI supercomputing platforms. We are looking for a strong technical architect to own end to end manageability architecture for these products in data centers. You will work with various component leads internally and externally, drive customer use cases, align architecture with customer requirements and release best products to market. Join us at the forefront of technological advancement. What you’ll be doing: Drive server management for large clusters and data centers deploying GPUs and Grace solution from Nvidia. Work with data center architects and cloud customers to narrow down on requirements for implementation to ensure speed of light product development. Work with internal teams to make sure requirements are designed and implemented in right way with each firmware and software module Collaborate with other leads to design & build data center health management workflow. Drive reliability and optimization in firmware architecture from a data center view point. Work closely with cluster bring up team and resolve is
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer (Java) Job Description Summary Who is Mastercard? Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital economy that benefits everyone, everywhere by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. With connections across more than 210 countries and territories, we are building a sustainable world that unlocks priceless possibilities for all. Overview • The Decision Management program enables intelligent decision-based products through streaming analytics with the ability to govern these decisions and manage their outcomes with business agility. • This program leverages business rules & AI engines, a streaming big data cluster, an in-memory data grids, APIs, & UIs to deliver real time decisions at global scale • This person will be responsible for mentoring the team as well as stay hands on. We are looking for a Senior Software Developer to join our DMP team in Vancouver office.<
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Department: Treasury Reporting To: Manager - Treasury Job Location: Mumbai Education Qualification Required: B Com/M Com Number of Years of Experience Required: 2-4Yrs Summary: The Treasury Analyst is responsible for the day-to-day management of treasury operations, ensuring efficient cash flow, accurate financial reporting, and compliance with internal controls and regulations. Key Responsibilities: Daily: Cash Management: Prepare and share daily cash position reports with OpCo and regional teams. Report cluster and cash balances to GL and Treasury Manager. Manage fund flow and cash flow, including loan disbursements, payments, rollovers, and follow-ups. Share bank debits with the Payments team and credits with the Collection team. Share bank statements (collections) with relevant teams. Accounting and Reporting: Account for and approve treasury-related entries and share with the GL team for posting. Prepare Bank Reconciliation Statements (BRSs). Follow up with Collection and Payment teams to resolve open items in BRS. Arrange entries for Conversion of EEFC balances to INR as needed. Tran
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As a Datacenter Liquid Cooling Architect, you will define, design, and architect next-generation liquid cooling infrastructure for Tenstorrent’s large-scale AI training and inference clusters. You will partner with systems engineering, mechanical engineering, software, and cross-functional design teams to develop chassis-, rack-, and cluster-scale cooling solutions, including CDU integration, telemetry and control, leak detection, and resilient operating strategies. This role will help shape reliable AI datacenter architectures and deployments for both internal and external customers. This role is on-site, based out of Toronto, Canada, Austin, Texas or Santa Clara, California. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A datacenter and system thermal design professional with 10+ years of experience architecting cooling infrastructure for complex computing environments. An experienced liquid cooling architect who can design chassis- and rack-scale solutions for large AI training and inference clusters. A systems thinker who understands how mechanical, electrical, software, facility, and systems engineering decisions come toge
NVIDIA is looking for a hands-on Solutions Architect Manager to lead a team of GPU, networking & software solution architects and engineers. Do you want to build and lead a group that designs, debugs, and deploys new AI hardware and software technologies into production in customer data centers? As part of the NVIDIA SA organization, you will drive people and technical leadership for end-to-end solutions deployments at some of NVIDIA's most strategic technology customers, while directly contributing to designs and deep-dive debugging and shaping our product roadmap with customer feedback. What you will be doing: Recruit & manage a team of solutions architects, system/network and software engineers focused on large-scale GPU and AI networking deployments. Set priorities, allocate resources, mentor, and ensure high-quality customer delivery across multiple concurrent projects - while remaining directly involved in key technical reviews, design decisions, and critical debug efforts. Provide deep subject-matter expertise in advanced GPU and network systems and serve as the senior technical point of contact for strategic customers. Personally lead and guide complex compute/network configuration and performance debugging, working side-by-side with your team to deliver performant, reliable clusters. Guide your team as they lead network / compute / software architecture discussions, and support server, network, and cluster bring-up, including on-site data center work where needed. Systematically collect and synthesize customer-specific requirements across your portfolio. Partner with GPU/Network Systems Engineering, Product Management, and Sales to influence roadmap priorities and packaging of reference designs and solutions. Demonstrate SME in advanced GPU & network systems and be a trusted technical advisor to NVIDIA's strategic customers. Bring customer-sp
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! About the Role We are seeking a seasoned Manager, Software Engineering with 12+ years of experience to lead our Database Engineering and Cloud Infrastructure team. In this role, you will lead a team of high-performing engineers responsible for architecting, scaling, and optimizing multi-cloud relational and in-memory database platforms. You will bridge technical execution, engineering leadership, and strategic infrastructure planning across AWS and Azure environments. Key Responsibilities Technical Leadership & Architecture Lead the architectural design and operations of enterprise-grade, multi-cloud relational databases across AWS (RDS PostgreSQL, MySQL, Aurora) and Azure (Database for PostgreSQL/MySQL, Azure SQL Managed Instance). Drive high-availability architecture strategies, including Multi-AZ deployments, auto-failover groups, read replica scaling, and cross-region disaster recovery (DR). Oversee zero-downtime operations, including major-version engine upgrades, schema migrations, and blue/green deployment strategies. In-Memory Infrastructure & Open-Source Strategy Manage scale operations for in-memory datastores (AWS ElastiCache, Azure Cache for Redis), focusing on cluster mode operations, eviction policies, and persistence tuning. Spearhead open-source caching initiatives and migration pathways from Redis to Valkey (e.g., AWS ElastiCache for Valkey) using zero-downtime tools like RedisShake to ensure open-source license compliance and optimize cloud spend. Aut
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The Sandbox service team at SpaceXAI builds and maintains a secure, scalable system that gives our models safe, controlled access to computational environments. This infrastructure powers critical workloads across training and product, enabling models to run code, build software, interact with tools, and even control applications with user interfaces. We provision containers and virtual machines on large-scale clusters, granting models interactive control over these remote environments. Our work spans the full stack: from orchestrating massive jobs and resource scheduling at the cluster level, to fine-tuning filesystem performance on nodes. The Sandbox service enables Grok to safely run and test code in real-time for user queries, and supports reinforcement learning in training, where models interactively explore tools ranging from compilers to productivity apps. BASIC QUALIFICATIONS: Expert knowledge of Rust, C++ or Go Familiarity with Python Deep experience with either Linux or Windows systems (familiarity with both is a strong plus) Experience with virtualisation and containerisation technologies (e.g., cgroups, KVM, gVisor, QEMU) Solid knowledge of the networking stack COMPENSATION AND BENEFITS: £107,000 -
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE The CPBU Manageability team is responsible for Middleware of FlashArray and FlashBlade products. We are building a scale-out all-flash file and object store, designed for the modern world. To really understand how our customers work with data, we are deeply immersed in AI, modern backup, log analytics with Splunk and Elastic, data pipeline with Kafka, cluster computing with Spark, and many more use cases. You will love it on the CPBU Manageability team if you: want to understand how modern applications work with data and how we can make it better. are ready to dive into a complex problem and be the one who will drive it to a resolution. enjoy working with distributed systems, algorithms, operating systems, Linux kernel, database internals, hypervisors, containers, compilers and hardware… or at least some of those. want to work with other great engineers and develop or refine skills that will serve your entire career. enjoy working in a collaborative team environment in an open office. If this describes you, let's talk! You can take a part in changing how the world works with data. WHAT YOU'LL DO Own and deliver innovation end-to-end, from concept to shipped product Design, develop and maintain customer-facing and internal-facing API and command line interface for end to end configuration and management of FlashArray and FlashBlade products using Java, Python and beyond Experimenting with new technologies and archi
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the high-growth FlashBlade team as a foundational engineering leader building our next-generation, scale-out all-flash file and object storage platform. In this role, you will architect, design, and deliver high-performance distributed systems optimized for modern, data-intensive workloads like AI, log analytics, and cluster computing. You will own critical technological roadmaps end-to-end, directly influencing our core architecture while mentoring top-tier talent. Partnering closely with Product Management, System Validation, and Customer Support, your mission is to drive scalable innovation that redefines enterprise data storage and delivers six-nines reliability to our global customers. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management,
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the high-growth FlashBlade team as a foundational engineering leader building our next-generation, scale-out all-flash file and object storage platform. In this role, you will architect, design, and deliver high-performance distributed systems optimized for modern, data-intensive workloads like AI, log analytics, and cluster computing. You will own critical technological roadmaps end-to-end, directly influencing our core architecture while mentoring top-tier talent. Partnering closely with Product Management, System Validation, and Customer Support, your mission is to drive scalable innovation that redefines enterprise data storage and delivers six-nines reliability to our global customers. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management,
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the high-growth FlashBlade team as a foundational engineering leader building our next-generation, scale-out all-flash file and object storage platform. In this role, you will architect, design, and deliver high-performance distributed systems optimized for modern, data-intensive workloads like AI, log analytics, and cluster computing. You will own critical technological roadmaps end-to-end, directly influencing our core architecture while mentoring top-tier talent. Partnering closely with Product Management, System Validation, and Customer Support, your mission is to drive scalable innovation that redefines enterprise data storage and delivers six-nines reliability to our global customers. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management,
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the high-growth FlashBlade team as a foundational engineering leader building our next-generation, scale-out all-flash file and object storage platform. In this role, you will architect, design, and deliver high-performance distributed systems optimized for modern, data-intensive workloads like AI, log analytics, and cluster computing. You will own critical technological roadmaps end-to-end, directly influencing our core architecture while mentoring top-tier talent. Partnering closely with Product Management, System Validation, and Customer Support, your mission is to drive scalable innovation that redefines enterprise data storage and delivers six-nines reliability to our global customers. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management,
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the high-growth FlashBlade team as a foundational engineering leader building our next-generation, scale-out all-flash file and object storage platform. In this role, you will architect, design, and deliver high-performance distributed systems optimized for modern, data-intensive workloads like AI, log analytics, and cluster computing. You will own critical technological roadmaps end-to-end, directly influencing our core architecture while mentoring top-tier talent. Partnering closely with Product Management, System Validation, and Customer Support, your mission is to drive scalable innovation that redefines enterprise data storage and delivers six-nines reliability to our global customers. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management,
Get new cluster head last mile jobs by email
Daily job updates · Unsubscribe anytime