Jobiba hiring network

Cluster Head Last Mile Telangana Jobs

315 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cluster head last mile telangana jobs. Use filters to narrow by work mode, employment type, experience and date posted.

E
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... This is an opportunity to be one of the seed members of a growing product team in the FlashBlade BU, one of the fastest growing at Pure. With the FlashBlade product, we are building a scale-out all-flash file and object store, designed for the modern world. To really understand how our customers work with data, we are deeply immersed in AI, modern backup, log analytics with Splunk and Elastic, data pipeline with Kafka, cluster computing with Spark, and many more use cases. You will love it on the FlashBlade team if you: Want to understand how modern applications - like AI or Splunk - work with data and how we can make it better. Enjoy working with distributed systems, algorithms, operating systems, Linux kernel, database internals, hypervisors, containers, compilers and hardware... or at least some of those. Are ready to dive into a complex problem and be the one who will drive it to a resolution. Want to work with other great engineers and develop or refine skills that will serve your entire career. If you, like us, say “bring it on” to exciting challenges that change the world, we have endless opportunities where you can make your mark. WHAT YOU’LL NEED TO BRING TO THIS ROLE... Design, collaborate and implement creative new algorithms and technologies for high-performance, highly reliable systems (think six 9’s). Own and deliver innovation end-to-end, from concept to shipped prod

pythonjavaaws
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Powering the METCA AI & Sovereign Cloud Revolution As Everpure , we are transcending traditional storage to deliver the Enterprise Data Cloud - a unified architecture engineered to fuel the world’s most ambitious AI, Deep Learning, and HPC projects. The METCA region is a global epicenter for AI infrastructure investment, characterised by massive capital flowing into Tier-1 sovereign GPU clouds (CSPs, Neo-scalers) and enterprise AI factories. We are seeking a Battle-Trained, High-Conviction Hunting Systems Engineer (SE) to serve as our technical tip of the spear. This is not a passive, box-pushing relationship management role. You will partner aggressively with an Enterprise Account Executive to target, break into, and land the largest AI infrastructure projects in the market, displacing legacy architectures and securing net-new footprints. WHAT YOU'LL DO Execute High-Impact Hunting: Partner closely with Account Executives to actively map out and break into net-new enterprise accounts, sovereign GPU clouds, and high-performance computing clusters. Architect the AI Factory: Design high-performance, multi-tenant data pipelines. Move beyond basic storage architecture to design full-stack environments, optimising how data nodes interact within massive GPU fabrics. Drive Technical Consensus: Lead deep-dive architectural workshops with customer GPU cluster architects while simultaneously translating complex en

awskubernetesrest
View job →
E
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... This is an opportunity to be one of the seed members of a growing product team in the FlashBlade BU, one of the fastest growing at Pure. With the FlashBlade product, we are building a scale-out all-flash file and object store, designed for the modern world. To really understand how our customers work with data, we are deeply immersed in AI, modern backup, log analytics with Splunk and Elastic, data pipeline with Kafka, cluster computing with Spark, and many more use cases. You will love it on the FlashBlade team if you: Want to understand how modern applications - like AI or Splunk - work with data and how we can make it better. Enjoy working with distributed systems, algorithms, operating systems, Linux kernel, database internals, hypervisors, containers, compilers and hardware... or at least some of those. Are ready to dive into a complex problem and be the one who will drive it to a resolution. Want to work with other great engineers and develop or refine skills that will serve your entire career. If you, like us, say “bring it on” to exciting challenges that change the world, we have endless opportunities where you can make your mark. WHAT YOU’LL NEED TO BRING TO THIS ROLE... Design, collaborate and implement creative new algorithms and technologies for high-performance, highly reliable systems (think six 9’s). Own and deliver innovation end-to-end, from concept to shipped prod

pythonjavaaws
View job →
E
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... This is an opportunity to be one of the seed members of a growing product team in the FlashBlade BU, one of the fastest growing at Pure. With the FlashBlade product, we are building a scale-out all-flash file and object store, designed for the modern world. To really understand how our customers work with data, we are deeply immersed in AI, modern backup, log analytics with Splunk and Elastic, data pipeline with Kafka, cluster computing with Spark, and many more use cases. You will love it on the FlashBlade team if you: Want to understand how modern applications - like AI or Splunk - work with data and how we can make it better. Enjoy working with distributed systems, algorithms, operating systems, Linux kernel, database internals, hypervisors, containers, compilers and hardware... or at least some of those. Are ready to dive into a complex problem and be the one who will drive it to a resolution. Want to work with other great engineers and develop or refine skills that will serve your entire career. If you, like us, say “bring it on” to exciting challenges that change the world, we have endless opportunities where you can make your mark. WHAT YOU’LL NEED TO BRING TO THIS ROLE... Design, collaborate and implement creative new algorithms and technologies for high-performance, highly reliable systems (think six 9’s). Own and deliver innovation end-to-end, from concept to shipped prod

pythonjavaaws
View job →
E
Everpure
📍 Bengaluru• Full-time
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the high-growth FlashBlade team as a foundational engineering leader building our next-generation, scale-out all-flash file and object storage platform. In this role, you will architect, design, and deliver high-performance distributed systems optimized for modern, data-intensive workloads like AI, log analytics, and cluster computing. You will own critical technological roadmaps end-to-end, directly influencing our core architecture while mentoring top-tier talent. Partnering closely with Product Management, System Validation, and Customer Support, your mission is to drive scalable innovation that redefines enterprise data storage and delivers six-nines reliability to our global customers. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management,

pythonjavaaws
View job →
E
Everpure
📍 Bengaluru• Full-time
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the high-growth FlashBlade team as a foundational engineering leader building our next-generation, scale-out all-flash file and object storage platform. In this role, you will architect, design, and deliver high-performance distributed systems optimized for modern, data-intensive workloads like AI, log analytics, and cluster computing. You will own critical technological roadmaps end-to-end, directly influencing our core architecture while mentoring top-tier talent. Partnering closely with Product Management, System Validation, and Customer Support, your mission is to drive scalable innovation that redefines enterprise data storage and delivers six-nines reliability to our global customers. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management,

pythonjavaaws
View job →
P
Point72
📍 Bengaluru• Full-time
16 days ago

JOB TITLE Cloud Compute Engineer A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. WHAT YOU'LL DO Design, build, and operate Kubernetes clusters on Amazon EKS, including cluster lifecycle management, networking, autoscaling, and workload scheduling. Manage and optimize EC2-based compute infrastructure, including instance selection, placement strategies, capacity planning, and utilization analysis. Operate and improve ECS-based services where applicable, ensuring consistency across our container runtime environments. Develop and maintain Infrastructure as Code (IaC) using Terraform to provision and manage compute resources at scale. Collaborate with development and platform teams to define compute patterns, containerization standards, and deployment best practices. Monitor compute environments for availability, performance, and cost, driving continuous optimization across the fleet. Contribute to architectural decisions around workload placement, multi-tenancy, OS image and container li

pythonawsazure
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
16 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As our TT-Distributed Software Engineer, you will develop and optimize distributed software systems that power the most efficient and highest-performing AI and HPC clusters. In this role, you'll work on distributed programming across multiple nodes, utilizing systems programming, inter-node communication, and Tenstorrent’s scalable architectures to advance the state-of-the-art distributed inference and training infrastructure. This role is hybrid, based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Strong C or C++ engineer with solid foundations in systems programming, operating systems, and distributed systems principles. Enthusiastic about distributed computing, including IPC, socket programming, and cluster resource coordination. Comfortable reasoning about scalability, fault tolerance, and performance across multi-node environments. Curious and first-principles thinker who challenges conventional approaches to distributed system design. Motivated to grow into a deep technical expert in large-scale distributed AI infrastructure. What We Need Architect, implement, and optim

awsaic++
View job →
KH
K Health
📍 Tel Aviv• Full-time
16 days ago

About the Role: We are looking for a Senior DevOps Engineer to join our DevOps team at K Health. You will own and evolve the infrastructure underpinning a healthcare AI platform serving patients and enterprise health system partners. This is a high-ownership role: you will architect and operate cloud environments across K Health and its enterprise partners, lead complex infrastructure migrations, drive disaster recovery programs, and help build the next generation of AI-powered operations tooling. You will also mentor junior engineers and collaborate closely with product and engineering teams across the company. This is a hybrid role based in New York City (4 days/week in office) and includes participation in a daytime on-call rotation. What you will do: Own the design, implementation, and evolution of our GKE-based Kubernetes infrastructure across K Health and enterprise partner environments. Build and maintain our Terraform modular infrastructure library, including reusable modules with automated testing, across GCP, Cloudflare, and AWS. Architect, build, and maintain GitLab CI/CD shared pipeline templates used by all engineering teams (build, test, security scanning, deployment). Own and maintain self-hosted infrastructure software running in-cluster, including GitLab, ArgoCD, Langfuse, DependencyTrack, NGINX Ingress, and others. Implement and support security and compliance controls across infrastructure and the software supply chain - secrets management, pipeline secret detection, container scanning, SOC2 and HIPAA. Drive disaster recovery readiness: design failover scenarios, author runbooks, and lead periodic DR tests. Lead development of AI-powered operations tooling and agentic infrastructure. Monitor, troubleshoot, and improve production system reliability; respond to incidents during on-call shifts. Mentor junior DevOps engineers and establish team-wide engineering standards. What we are looking for: 5+ years of experience in DevOps, platform engineering,

pythonsqlpostgresql
View job →

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Developer Infrastructure org is the engine behind Robinhood's entire engineering organization — a collection of tightly integrated teams whose collective mission is to make every engineer at Robinhood faster, more reliable, and exponentially more productive. The org spans four major teams: DevX (Developer Experience), TestX (Test Infrastructure), Backend Platform, and Mobile Platform. DevX owns Robinhood's monorepo and Bazel-based build infrastructure — the critical layer between a developer writing code and that code being ready to ship — along with the company's full CI/CD pipeline and remote build execution cluster. TestX owns the infrastructure behind Robinhood's entire test experience: the integration test environments, and personal development environments that serve as miniature simulations of the full Robinhood system, giving engineers a safe, isolated space to test their code end-to-end before it ever touches production. Backend Platform and Mobile Platform own the core language runtimes, libraries, IDEs, and developer toolchains across Python, Go, TypeScript, Swift, and Android. Together, these teams share a single north star: leveraging AI and agentic systems

typescriptpythonaws
View job →

About the Team OpenAI's Industrial Compute organization builds and operates the infrastructure required to train and serve frontier AI models. The Capacity Planning team connects rapidly changing research and product demand with the compute, networking, storage, power, data center, hardware, and operational resources required to make that demand executable. About the Role We are seeking a Technical Program Manager to build and lead capacity planning across OpenAI's large-scale AI infrastructure. You will translate uncertain workload demand into clear infrastructure requirements, allocation decisions, supply commitments, activation priorities, and long-range capacity strategies. This role sits at the intersection of research, engineering, infrastructure, finance, sourcing, deployment, and operations. You will create the planning models, operating cadences, governance mechanisms, and source-of-truth systems that allow teams to understand what capacity is required, what is available, what is at risk, and what decisions must be made. This is not a finance-only forecasting or reporting role. Success requires technical fluency across the infrastructure stack, strong analytical judgment, and the ability to move consequential decisions forward when requirements, timelines, and supply conditions change quickly. Key Responsibilities Own capacity-planning processes across near-term workload allocation, quarterly execution, and longer-range infrastructure horizons. Translate research, training, inference, and product demand into compute, accelerator, cluster, networking, storage, rack, power, and site requirements. Develop scenarios that make assumptions, confidence levels, constraints, sensitivities, and decision points explicit. Reconcile requested demand against contracted, delivered, installed, activated, and workload-usable capacity. Partner with research and engineering teams to understand workload priorities, technical dependencies, utilization patterns, and changing req

pythonsqlaws
View job →
B
Baseten
📍 San Francisco• Full-time• Remote
29 days ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments. We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment. This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment.. RESPONSIBILITIES Core Responsibilities: Drive suppliers to keep the maximum amount of the GPU fleet online and healthy. Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs. Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier. SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short. Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability. Scope and

REMOTEmachine learningaigo
View job →
B
Baseten
📍 San Francisco• Full-time• Remote
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're hiring a Product Data Scientist to establish how product decisions at Baseten are made with data. You'll work directly with Product and Engineering, alongside GTM to determine measurement, strategy, experimentation and implementation. This is a foundational, hands-on role. You'll define what success looks like across a technical, usage-based platform and turn ambiguous questions into analyses, forecasts, and experiments that shape product strategy. You'll work from clickstream and product events through inference telemetry and observability data, helping Baseten make faster decisions about reliability, performance, adoption and developer experience. RESPONSIBILITIES Partner directly with Product and Engineering: frame the questions that matter, define success criteria, and turn analysis into roadmap, launch, and prioritization decisions. Define how product success is measured: establish metrics across activation, adoption, retention, expansion, reliability and user experience. Support experimentation and launches: design measurement plans, analyze A/B experiments and controlled rollouts, and translate results into product decisions. Diagnose reliability and scaling behavior: join customer signals with request, replica, deployment, and cluster telemetry to find patterns in release bottlenecks, unhealthy replicas, and models without traffic. Define the enterprise customer journey and measure feature adoption

REMOTEpythonsqlmachine learning
View job →
M
Mongodb
📍 Gurugram• Full-time
1mo ago

MongoDB Technical Services Engineers use their exceptional problem solving and customer service skills, along with their deep technical experience, to advise customers and to solve their complex MongoDB problems. Technical Service Engineers are experts in the entire MongoDB ecosystem - database server, drivers, cloud, and infrastructure. This also includes services such as Atlas (database as a service), or Cloud Manager (which helps customers with automation, backup and monitoring of their MongoDB systems). Our engineers combine their MongoDB expertise with passion, initiative, teamwork, and a great sense of humor to help our customers be successful with MongoDB. The position will be based in our Gurugram office. The standard work week will be Monday to Friday, every week. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Cool things you’ll do You'll be working alongside our largest customers, solving their complex challenges - resolving questions on architecture, performance, recovery, security, and everything in between. You'll be an expert resource on best practices in running MongoDB at scale, whatever that scale may be. You'll be an advocate for customers' needs - interfacing with our product management and development teams on their behalf. And you'll contribute to internal projects, including software development of support tools for performance, benchmarking, and diagnostics. What you need You should have 5+ years of proven experience, we consider all candidates with an eye for those who are self-taught, insatiably curious, and multi-faceted. The ideal candidates should have strong technical experience in more than one of the following areas Systems engineering experience, including Linux performance, memory management, I/O tuning, configuration, security, networking, clusters, and troubleshooting Understand core Kubernetes concepts, including containers, namespaces, custom resources and multi-cluster dep

javascriptpythonjava
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Core Services organization builds and runs the mission-critical online services that product teams rely on in production. We own foundational distributed systems and platform capabilities that enable reliable execution, high-performance services, and large-scale file/data needs across our products. This team is distinct from developer infrastructure and data infrastructure—our focus is production service foundations and core runtime services. About the Role We’re hiring an Engineering Manager, Core Services to help lead teams responsible for highly reliable, high-scale distributed systems that sit on the critical path for OpenAI products. Your team will own foundational production systems that OpenAI’s product engineering teams build on. You’ll collaborate closely with product and infrastructure partners to ship reliable services quickly, and help scale systems and teams as OpenAI grows. You’ll partner closely with senior engineering leaders to scale the org, mature operations, and drive major platform initiatives. This role requires strong technical ability. You’ll be responsible for: Managing and growing a high-performing team of infrastructure engineers. Leading teams building and operating large, critical production platforms, including cluster reliability, scaling, and rollout safety. Building and operating mission-critical distributed systems with strong operational rigor (SLOs, incident response, capacity planning, reliability). Setting technical direction for platform foundations such as workflow/orchestration capabilities, large-scale file/blob/storage services, and core service foundations. Partnering with a broad set of stakeholders, including product engineering, adjacent infrastructure teams, and (where relevant) finance/cost partners. Coaching, mentoring, and developing engineers and emerging leaders. You might thrive in this role if you: Have significant experience leading teams that run mission-critical infrastructure in production

awsrestai
View job →
🔔

Get new cluster head last mile telangana jobs by email

Daily job updates · Unsubscribe anytime