About Anyscale At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role The Customer Engineer will play a crucial role in the customers’ post-sale journey - helping them to onboard, adopt and grow on Anyscale, troubleshooting and resolving open customer tickets and driving consumption. Anyscale is an ever evolving platform and hence will require close co-ordination with our engineering teams to debug complex issues. This is an exciting role for those who are technically curious and passionate about ML/AI, LLM, vLLM and the role of AI in next generation applications. It’s an opportunity to make a significant impact in a collaborative, fast-paced environment while building a new segment in this space. In this role, you’ll be able to Resolve customer issues and help in their successful adoption of Anyscale platform Be a technical advisor, and internal champion for our key customers Own customer issues end-to-end, from troubleshooting, triaging, escalations and eventual resolution Participate in our follow-the-sun customer support model to ensure continuity in resolving high priority tickets Keep track of open customer bugs and feature requests to influence prioritization and provide timely customer updates upon resolution Contribute towards improvement of internal tools an
Jobiba hiring network
Cluster Head Last Mile Telangana Jobs
315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster head last mile telangana jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Business Systems drives efficiency across Datadog through business process analysis, systems automation and integrations, AI agent and MCP development, and vendor/software review. The team is increasingly embedded in cross-functional initiatives across People, Finance, GTM , Legal, Recruiting, and Technical Solutions — translating ambiguous business problems into scoped, buildable solutions and owning delivery end-to-end. This is not a generalist BSA role. Each Senior BSA will own a cluster of business functions end-to-end, acting as an internal product manager for their domain rather than processing inbound requests reactively. There are two openings, each covering a different domain: People, Recruiting, Legal or Finance, GTM, Procurement As the role evolves alongside AI, Senior BSAs leverage agentic tools for data aggregation and context gathering while remaining the critical human-in-the-loop layer — owning business context and product outlook, validating use cases, managing stakeholder relationships, and making the judgment calls agents can't. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Discovery and scoping: Translate ambiguous asks from business stakeholders into well-defined requirements that Business Systems Engineers can build against. Many stakeholders don't know what they want or the cross-functional impact of what they're asking for — you surface both before engineering begins. Cross-functional visibility: Identify dependencies, downstream impacts, and integration considerations that requesting teams miss. Proactive opportunity identification: Develop deep domain knowledge to identify automation and AI opportunities before they become inbound requests, shifting the team from reactive in
At NVIDIA, we push the boundaries of computing innovation. Our ASIC Verification Engineers focus on developing the world’s top SoCs and GPUs. Joining us as a Senior ASIC Verification Engineer - GPU means working on modern technology powering consumer graphics and AI applications. This position is ideal for those passionate about technology and eager to impact computing’s future. What you'll be doing: As a key member of our ASIC Verification team, you will verify the design and implementation of the industry's leading GPUs. You will be responsible for verifying the ASIC build, architecture, golden models, and micro-architecture using advanced verification methodologies such as UVM or equivalent. Understand the design and implementation of your unit/cluster/chip, define the verification scope, develop the verification infrastructure, and verify the correctness of the design. Collaborate with architects, designers, and pre- and post-silicon verification teams to accomplish your task. What we need to see: Bachelor's Degree in EE, CS, or CE or equivalent experience. 5+ years of relevant experience. Experience in verification using random stimulus along with functional coverage and assertion-based verification methodologies. Experience with design and verification tools (VCS or equivalent simulation tools, debug tools like Verdi, Indago, GDB). Expertise in System Verilog or similar HVL. Strong debugging and analytical skills. Perl and C/C++ programming language experience desirable. Strong communication skills and the ability & desire to work as a great teammate are huge pluses. Experience in crafting test bench environments for unit and system level verification. #LI-Hybrid Your base salary will be det
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Manager, Business Strategy COVEC Manager Business Strategy COVEC Overview Reporting directly to the Vice President, Business Strategy, South LAC, with a dual reporting line to the COVEC Cluster Lead, this role leads Mastercard’s strategy agenda across Colombia, Ecuador, Venezuela, Guyana, and Suriname, while supporting broader South Division priorities. As a member of both the South LAC Strategy team and the COVEC Leadership Team, the role partners across the organization to shape strategic priorities, identify new growth opportunities, evaluate investments and business initiatives, support key decision-making processes, and drive execution against critical business objectives. The position works closely with leadership teams across markets, products, services, operations, finance, legal, public affairs, people, and other functions to ensure alignment, mobilize resources, and deliver measurable business impact. The role serves as a trusted advisor to senior leadership, helping translate market opportunities, industry trends, and business challenges into actionable strategies that accelerate growth, strengthen Mastercard’s competitive position, and support the delivery of cluster and divisional objectives. Key Responsibilities • Lead the development and execution of Mastercard’s strat
We build and operate the compute infrastructure our researchers run on, supporting large-scale processing of historical market data and model training on our own hardware across multiple data centers. Our environment includes bare-metal Linux, virtualization, storage, and GPU clusters, where performance, reliability, and predictable system behavior are critical. Our Infrastructure team covers monitoring and automation, distributed storage, hardware and OS provisioning, GPU clusters and workload scheduling, high-speed networking, L2/L3 Linux support, and security engineering. Engineers here own their tasks end to end, so there's room to go deeper in your area and pick up the parts you haven't touched yet. We’re looking for a Linux Infrastructure Engineer who can work hands-on with server and cluster environments, from deployment and configuration to performance tuning, troubleshooting, and ongoing improvement What You’ll Be Doing: Deploying, configuring, and maintaining Linux-based bare-metal servers across our data centers Building and operating clustered environments, including virtualization, storage, GPU compute, and database clusters Troubleshooting complex Linux, hardware, networking, and cluster-level issues Performance tuning for throughput, latency, stability, and resource utilization Monitoring infrastructure health and performance, identifying bottlenecks, and preventing recurring issues Supporting the full server lifecycle: provisioning, setup, upgrades, and maintenance Improving reliability and predictability during failures, maintenance, and scaling Automating provisioning, configuration, and operational tasks, primarily using Ansible and scripting What We Look For In You: Strong hands-on Linux administration and troubleshooting experience Production experience with on-premise, bare-metal infrastructure Good understanding of Linux performance and bottleneck analysis Experience with: infrastructure monitoring and troubleshooting production issues,
About the Role: We're hiring Senior and Staff Data Platform Engineers to join the Data Infrastructure teams in Toronto. Together these teams own the infrastructure that processes billions of events per day: Spark-on-Kubernetes, Flink and Kinesis pipelines, a multi-petabyte Delta Lake, a large-scale MemoryDB feature store, Databricks multi-environment operations, and the catalog and lifecycle systems that govern it. The team is small and senior. Each engineer owns major platform components: you design it, build it, and support it in production. This is a hybrid-role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: Spark-on-Kubernetes — EKS-based compute platform for Spark workloads: cluster configuration, Pod Identity IAM, job environment setup, Kustomize overlays, and shadow canary validation Event ingestion — Rust services and Flink jobs processing billions of events per day over Kinesis; throughput, reliability, on-call response, and AI-assisted operational tooling to reduce toil Platform infrastructure — Terraform modules for environment provisioning, cross-account AWS IAM, ARC runner infrastructure, and CI/CD for data platform changes Feature store and ML compute — Flink-based real-time feature pipelines feeding a large-scale MemoryDB cluster; GPU capacity governance and Databricks multi-environment operations for ML training workloads Workflow orchestration and CDC — Airflow-based DAG deployment, change data capture pipeline operations, and data quality monitoring Your Background: 3+ years building and operating production data platform infrastructure at the cluster or platform level, across Spark, Flink, Kinesis, Kubernetes, or equivalent Deep experience in at least one of: Spark-on-K8s cluster operations, Rust-based data or systems engineering, Kubernetes platform engineering and IaC, or data catalog and governance tooling Production AWS experience or equivalent: EKS, S3, Kinesis, and mu
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for a Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the Infra team, we build the scalable, secure, and robust backbone that enables this vision. Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads. We are seeking a talented engineers with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale’s cloud platform. You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers. A snapshot of projects you may work on Design, build, and scale services that orches
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we're building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Anyscale is looking for a Software Engineer to join the ML Developer Experience (MLDevX) team. MLDevX owns the experience layer of the Anyscale platform: the interfaces through which users and coding agents discover, configure, run, observe, debug, and productionize AI workloads. Every user journey crosses this layer through the CLI, SDKs, APIs, UI, Workspaces, MCP, or the workflows and integrations built on top of them. Together, these form the user’s primary interface into Anyscale, turning distributed computing from a systems problem back into a coding problem. We build the common contracts, tools, control-plane services, and architecture that power these surfaces. You will work across the stack from developer tooling to cloud infrastructure and the Ray runtime. Manage long-running operations and make failures across jobs, tasks, actors, nodes, and GPUs easier to diagnose. The systems you build must scale with the platform, remain predictable through failures, and be intuitive for developers, programmable for applications, and operable by coding agents. This is a high impact individual-contributor role with end-to-end ownership. You will work directly with users and field teams to identify high
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: As a Site Reliability Engineer, you will play a crucial role in ensuring the smooth operation of all user-facing services and other Anyscale production systems. Anyscale values diversity and inclusion, and we encourage applications from individuals of all backgrounds. This includes processes for provisioning, negotiating prices, managing costs, seeing opportunities for teams to reduce wastage by finding applications across the company. You will apply sound engineering principles, operational discipline, and mature automation to our environments and the Anyscale codebase as we scale. As part of this role, you will: Develop a unified perspective on how cloud components are utilized across the company, taking into account diverse needs and requirements. Ensure that deployment methodologies align with the company's reliability goals. Build systems that promote understanding of production environments, facilitating quick identification of issues through robust observability infrastructure for metrics, logging, and tracing. Create monitoring and alerting systems at different levels, enabling teams to easily contribute and enhance the overall monitoring capabilities. Establish testing infrastructure to s
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI’s AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production. In this role, you will: Design and implement the LLM inference runtime for frontier models running on custom silicon. Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization. Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads. Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack. Enable new
NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel
The worldwide data management software market is massive. At MongoDB we are transforming industries and empowering developers to build amazing apps that people use every day. We are the leading modern data platform and the first database provider to IPO in over 20 years. Join our team and be at the forefront of innovation and creativity. MongoDB is seeking a Software Engineer 3 to join the Atlas Clusters Platform team. The team is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. The Atlas Clusters Platform team develops the foundational orchestration platform behind MongoDB Atlas. Our systems drive cluster planning and execution, evolve the Atlas control plane toward service-oriented architecture, and provide critical infrastructure that help Atlas run safely and efficiently across cloud environments. We are looking to speak to candidates who are based in New York City for our hybrid working model. What you’ll do Build and design new features for MongoDB Atlas Contribute to and lead complex technical projects Work closely with product and design teams, considering the user’s perspective while building technical solutions Work with customers and support engineers to fix issues Collaborate with team members to develop our codebase, best practices, and design principles Learn from and mentor other team members We’re looking for someone who Has at least 3 years of professional software development experience Is skilled at writing large-scale, distributed backend systems in a compiled language (Java, C#, Go, etc.) Is comfortable working across the stack of a modern web application (e.g. React, TypeScript, Kubernetes) Has experience with at least one major cloud provider technology (AWS, Azure, GCP) Has led the launch of a new feature and maintained it in production Is eager to solve tough problems Has excellent
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About Ray Data Team: Ray Data is Python-native data processing engine that is a one stop shop for all AI data processing needs. Ray Data provides performant, first-class integration with cutting edge AI frameworks using both multi-modal and structured data. The Ray Data team currently develops and maintains Ray Data . We are a team of engineers passionate about building a Data processing engine which is a one-stop shop for all of your ML/AI needs. We are looking for exceptional engineers to build, optimize, and scale Ray for modern and increasingly complex AI workloads. As part of this role, you will: Improve the performance of Ray Data and multi-modal batch inference use cases. Ensure efficient scaling across different stages of the Data pipeline in a heterogeneous environment. Building data loading solutions for production training workloads. Focus on stability and fault tolerance at high scale Working with customers and new age AI native companies in scaling their AI workloads. We'd love to hear from you if have: At least 3-4 years of relevant work experience Solid background in building scalable and fault-tolerant distributed systems Experience with data processing, database internals. Passionate about large
The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 55616
The Shift Engineer - Automation is responsible for ensuring the efficient and continuous operation of the MSEL module line through stabilization initiatives, effective team building, and optimized spare management. This role focuses on maintaining high equipment availability for production while fostering a culture of safety and sustainability within the organization. By leveraging SAP-PM tools and implementing best practices, the Shift Engineer contributes to the overall operational excellence and productivity of the cluster. Source: Adani Group | Job ID: 55614
Get new cluster head last mile telangana jobs by email
Daily job updates · Unsubscribe anytime