About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine
Jobs in India
Deployment Lead in India
231 active opportunities · Updated October 2026
Showing
15 jobs
Explore current deployment lead jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.
About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
Role: Senior AI Engineer Location: Hyderabad, India (Hybrid) Department: Product Development About the Role GHX is building a cutting-edge LLM-powered document understanding platform focused on classification, structured data extraction, and intelligent orchestration at scale. This is a high-impact AI engineering role where you will own the full lifecycle—from problem framing to production deployment . Initially, you will focus on prompt engineering and evaluation systems , building the quality foundation for AI performance. Over time, the role expands into agent orchestration, system architecture, and migration of rule-based systems to LLM-driven pipelines . A strong foundation in software engineering (5+ years) is essential. This role demands engineering rigor across both traditional system design and AI system behavior . Core Responsibilities 1. Prompt Engineering Design prompts for diverse document classification and extraction tasks Treat prompts as formal specifications (precise, structured, and edge-case-aware) Develop few-shot, chain-of-thought, and structured output templates Manage prompt lifecycle: versioning, testing, and rollback 2. LLM Output Evaluation Create and maintain ground truth datasets Build automated evaluation pipelines (precision, recall, field-level accuracy) Identify and resolve conceptually incorrect outputs despite surface correctness 3. AI Agent Orchestration Design multi-agent workflows for document processing Implement tool-use patterns and integrate MCP servers Optimize orchestration for scale and efficiency 4. Software Engineering Develop production-grade APIs and backend services Apply Clean Architecture / DDD principles Write maintainable, testable Python code Contribute to CI/CD, deployment, and observability systems 5. Stakeholder Collaboration Act as a bridge between business stakeholders and AI systems Translate product requirements into technical architectures Communicate system behavior, limitations, and quality
Description: Graviton Research Capital LLP, Gurgaon is looking to hire Software Engineers for our Core Technology team which has some of the best programmers in India working on cutting edge technologies to build a super fast and robust trading infrastructure handling millions of dollars worth of trading transactions every day. As a Senior Software Engineer with Graviton your responsibilities will include: Designing and implementing a high-frequency automated trading system, that trades on multiple exchanges Building live reporting and administration tools for the trading system Performance optimization and improving the overall latency of systems, through algorithm research and using cutting edge tools and techniques End-to-end ownership of modules, including designing, development, deployment and support Growing the team through involvement in the regular hiring process and occasional campus recruitments Requirements : The ideal requirements for our candidates are: A degree in Computer Science 3-5 yrs Experience with C/C++ and object-oriented programming Experience in HFT industry Expertise in algorithms and data structures Excellent problem solving skills Strong communication skills A working knowledge of Linux systems Any of the following is a plus: A good understanding of TCP/IP and Ethernet Knowledge of any other programming language e.g. Java, Scala, Python, bash, Lisp, etc. Familiarity with parallel programming models and parallel algorithms Experience with big data environments e.g. Hadoop, Spark etc. Benefits: Our open and collaborative work culture gives you the freedom to innovate and experiment. Our cubicle free offices, non-hierarchical work culture and insistence to hire the very best creates a melting pot for great ideas and technological innovations. Everyone on the team is approachable, there is nothing better than working with friends! Our perks have you covered. Competitive compensation Annual international team outing Fully covered commuti
SDLC Engineer A Career with Point72’s Technology Team As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. What You’ll Do Provide expert technical support for our enterprise Software Development Life Cycle (SDLC) platforms including Jenkins, GitHub, Bitbucket, and AWS-based CI/CD pipelines Support and optimize our artifact management solutions (Artifactory) and static code analysis tools (SonarQube) Partner directly with business teams and development clients to address SDLC platform challenges and deliver effective solutions Implement and maintain security controls across our development toolchain and infrastructure Develop automation solutions to eliminate toil and enhance developer productivity Troubleshoot complex build, deployment, and integration issues across our development environments Contribute to continuous improvement of our AWS-based development infrastructure Maintain documentation and knowledge base for supported platforms and tools What’s required 5+ years of experience in software engineering, DevOps, or SRE roles Strong technical expertise in AWS services and cloud-native architectures Experience with container technologies (Docker, ECS, EKS) and container orchestration Deep understanding of Git workflows, branching strategies, and version control best practices Strong hands-on programming/scripting skills (Python, Go, or similar), with ability to debug and build automation for CI/CD pipelines Experience with infrastructure as code (Terraform, CloudFormation) in AWS environments Hands-on experience with CICD solutio
Here's a summary of the role: Do you love building scalable cloud platforms and solving complex engineering problems with modern technologies? As a Senior Software Engineer at Diligent, you'll design and deliver high-performing , serverless applications that power our global SaaS platform. You'll work extensively with TypeScript, Node.js, AWS, and event-driven microservices, owning services from design to deployment and production monitoring. This is an opportunity to influence technical decisions, mentor engineers, and explore how AI can transform software development and engineering productivity. If you're passionate about cloud-native architectures, distributed systems, and building software that scales to millions of users, we'd love to meet you. Here's a breakdown of what you'll do (not all of it, just the important stuff): Design and build scalable backend services and event-driven microservices using TypeScript and AWS. Develop secure APIs and integrations that power reporting, analytics, and dashboard experiences. Build and maintain serverless solutions using AWS services such as Lambda, EventBridge , SQS, and DynamoDB. Drive engineering excellence through testing, observability, automation, and production readiness practices. Contribute to infrastructure-as-code and CI/CD pipelines using AWS CDK and modern DevOps practices. Mentor engineers, participate in architecture discussions, and champion the use of AI tools to improve development efficiency. These are the essentials you'll need to get an interview: 6-8 years of professional software engineering experience. Strong experience with TypeScript, Node.js, and modern backend development patterns. Hands-on experience building cloud-native applications on AWS. Strong understanding of serverless architectures and event-driven microserv
Opportunity Overview: We’re looking for a senior-level automation engineer who will help raise the bar on release quality, environment reliability, and change safety across Cohere’s platform. You’ll partner closely with Product, Engineering, Platform, and SRE to build scalable automation, guardrails, and validation systems that reduce production risk while increasing delivery velocity. This is not a “test scripts only” role. You’ll shape automation strategy, embed quality into the SDLC, and help define how changes move safely from dev → staging → UAT → prod in a fast-moving healthcare platform. You’ll help define how quality scales as Cohere grows. This role has real influence over release safety, platform reliability, and how engineering teams ship software in a regulated, high-impact domain. You won’t just test features — you’ll shape how Cohere delivers them safely to production. What you’ll do: Own and evolve Cohere’s end-to-end test automation strategy across UI, API, config changes, and critical workflows Design and maintain scalable E2E automation frameworks for multi-tenant, payer-specific workflows Build automated validation for deployment guardrails, release readiness, and production change safety Partner with Platform/DevOps to integrate automation into CI/CD pipelines and deployment workflows Create automated coverage for high-risk paths (authorization flows, partner integrations, file pipelines, feature flags, config changes) Drive test reliability, flake reduction, and actionable failure signals Define and enforce quality gates for prod releases, blue/green and canary deployments, and config changes Collaborate with Product and Engineering to ensure business outcomes are testable, measurable, and observable Improve test data management and environment stability to enable reliable automation at scale Mentor engineers on testability, automation best practices, and quality-first development Partner with SRE and Security to ensure production readines
Opportunity Overview: We’re looking for a senior-level automation engineer who will help raise the bar on release quality, environment reliability, and change safety across Cohere’s platform. You’ll partner closely with Product, Engineering, Platform, and SRE to build scalable automation, guardrails, and validation systems that reduce production risk while increasing delivery velocity. This is not a “test scripts only” role. You’ll shape automation strategy, embed quality into the SDLC, and help define how changes move safely from dev → staging → UAT → prod in a fast-moving healthcare platform. You’ll help define how quality scales as Cohere grows. This role has real influence over release safety, platform reliability, and how engineering teams ship software in a regulated, high-impact domain. You won’t just test features — you’ll shape how Cohere delivers them safely to production. What you’ll do: Own and evolve Cohere’s end-to-end test automation strategy across UI, API, config changes, and critical workflows Design and maintain scalable E2E automation frameworks for multi-tenant, payer-specific workflows Build automated validation for deployment guardrails, release readiness, and production change safety Partner with Platform/DevOps to integrate automation into CI/CD pipelines and deployment workflows Create automated coverage for high-risk paths (authorization flows, partner integrations, file pipelines, feature flags, config changes) Drive test reliability, flake reduction, and actionable failure signals Define and enforce quality gates for prod releases, blue/green and canary deployments, and config changes Collaborate with Product and Engineering to ensure business outcomes are testable, measurable, and observable Improve test data management and environment stability to enable reliable automation at scale Mentor engineers on testability, automation best practices, and quality-first development Partner with SRE and Security to ensure production readines
Opportunity Overview: We are seeking a Senior Data Engineer to contribute to the design and delivery of our cloud-native healthcare data platform. You will implement scalable data solutions built on AWS, Apache Iceberg, Lake Formation, Glue Catalog, Athena, dbt, and modern orchestration frameworks. This role combines strong hands-on engineering with collaboration across platform, analytics, and business teams. What You'll Do Data Engineering Delivery Deliver complex data engineering projects in collaboration with cross-functional teams Drive technical execution from design through production deployment Implement scalable data patterns and reusable frameworks Design and implement batch and near-real-time pipelines Build reusable ingestion, transformation, validation, and publishing frameworks Support modernization of legacy workloads Contribute to Apache Iceberg implementation and optimization Apply standards for schema evolution, partitioning, compaction, and metadata management Ensure efficient storage and query performance Implement data quality frameworks and validation layers Support observability and monitoring practices Contribute to operational excellence and reliability improvements Participate in architecture and design discussions Conduct and participate in code reviews Mentor junior engineers and share best practices ISMS roles and responsibilities Good knowledge of Information security Oversee specific business processes within the ISMS. Responsible to manage the ISMS documentation, conduct risk assessments, and implement risk treatment plans. Risk Owners are responsible for identifying, assessing, and managing risks within their areas of responsibility. They are also responsible for implementing risk treatment plans. Conduct the BCP and other test related to information security continuity along with CISO Responsible for monitoring and reporting on the performance of the ISMS. Responsible for implementation of security policies and procedures and report
Opportunity Overview: We’re looking for a senior-level automation engineer who will help raise the bar on release quality, environment reliability, and change safety across Cohere’s platform. You’ll partner closely with Product, Engineering, Platform, and SRE to build scalable automation, guardrails, and validation systems that reduce production risk while increasing delivery velocity. This is not a “test scripts only” role. You’ll shape automation strategy, embed quality into the SDLC, and help define how changes move safely from dev → staging → UAT → prod in a fast-moving healthcare platform. You’ll help define how quality scales as Cohere grows. This role has real influence over release safety, platform reliability, and how engineering teams ship software in a regulated, high-impact domain. You won’t just test features — you’ll shape how Cohere delivers them safely to production. What you’ll do: Own and evolve Cohere’s end-to-end test automation strategy across UI, API, config changes, and critical workflows Design and maintain scalable E2E automation frameworks for multi-tenant, payer-specific workflows Build automated validation for deployment guardrails, release readiness, and production change safety Partner with Platform/DevOps to integrate automation into CI/CD pipelines and deployment workflows Create automated coverage for high-risk paths (authorization flows, partner integrations, file pipelines, feature flags, config changes) Drive test reliability, flake reduction, and actionable failure signals Define and enforce quality gates for prod releases, blue/green and canary deployments, and config changes Collaborate with Product and Engineering to ensure business outcomes are testable, measurable, and observable Improve test data management and environment stability to enable reliable automation at scale Mentor engineers on testability, automation best practices, and quality-first development Partner with SRE and Security to ensure production readines
Job Title: Full Stack Software Development Engineer II (SDE II) Team: Product Engineering Location: Bangalore, India Employment Type: Full-time About the Role As a Full Stack SDE II in Sigmoid's Product Engineering team, you will take ownership of building scalable, reliable, and user-centric web applications. You will collaborate closely with product managers, designers, and system architects to drive features from design to deployment, ensuring robust frontend performance, clean backend services, and efficient database architectures. Key Responsibilities Full Stack Development Design and build responsive, high-performance web interfaces using React, TypeScript, Tailwind CSS, and Shadcn UI, backed by reliable API services using Node.js or Python (FastAPI). Database & API Design Model, query, and optimize relational data structures in PostgreSQL. Design, build, and maintain clean RESTful APIs and microservices. Code Quality & Best Practices Apply SOLID principles, basic design patterns, and clean code principles to write maintainable, scalable, and self-documenting code. Testing & Quality Assurance Implement comprehensive test suites using Jest, React Testing Library, or backend testing frameworks (e.g., PyTest) to maintain high code coverage and prevent regressions. Agile Execution Actively participate in Agile ceremonies (sprints, stand-ups, retrospectives, and planning) to deliver quality code incrementally. Collaboration & Mentorship Partner with cross-functional teams and mentor junior engineers through constructive code reviews and technical guidance. Requirements & Qualifications Technical Skills Frontend 3–5 years of hands-on experience with React, TypeScript, Tailwind CSS, and modern UI component libraries like Shadcn UI. State & Data Fetching Experience with modern client-side state management and data-fetching libraries (TanStack Query, Zustand, or Redux Toolkit). Backend Strong expertise in server-side development using Node.js or
Assistant Hospitality Manager Job Description Assist the Hospitality Manager in managing day-to-day catering and hospitality operations at the site. Supervise cafeteria food service housekeeping and front-of-house operations to ensure smooth service delivery. Ensure compliance with SOPs food safety standards FSSAI requirements HACCP guidelines and company policies. Monitor food quality portion control presentation hygiene and service standards. Coordinate with chefs service staff housekeeping stores vendors and client representatives for smooth operations. Manage and supervise staff attendance deployment grooming discipline and performance. Conduct regular kitchen and cafeteria inspections and ensure corrective actions are implemented. Monitor food temperatures cleaning schedules pest control sanitation and hygiene records. Assist in menu planning menu implementation and customer feedback management. Handle client complaints and customer feedback and coordinate with the operations team for timely resolution. Monitor stock levels FIFO/FEFO practices food wastage and inventory control. Coordinate with procurement and vendors to ensure timely availability of food consumables and operational supplies. Prepare and maintain daily/monthly operational reports checklists audit records and MIS. Support internal client and statutory audits including preparation and closure of observations. Assist with cost control manpower planning and budget management. Conduct staff briefings and training on food safety hygiene service standards and workplace procedures. Ensure proper execution of special events VIP services meetings and catering requirements. Address: Karjat Facility: UNIVERSAL BUSINESS SCHOOL - MUMBAI Qualification: Graduate Experience: 4 - 5 years Source: Sodexo India | Job Code: IJP574532
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! As a Senior Salesforce Developer at New Relic, you will be responsible for building and maintaining our Salesforce org. You will work closely with our Sales Revenue & Finance teams to automate processes and build custom applications on the Salesforce platform. In this role, you will be responsible for the full development life cycle, from Design and Solution to deployment and maintenance. The ideal candidate will have 5+ years of experience as a Salesforce developer and be proficient in Apex, Lightning Web Components and Visualforce. They will also have experience with integrations and be able to work with other team members or independently. Duties & Responsibilities Design, develop, test, deploy, and maintain high-quality Salesforce.com solutions Configure Salesforce.com to meet business requirements Write Apex classes, triggers, Visualforce pages, and Lightning Web Components Integrate Salesforce.com with other systems (as needed) Perform data migrations Stay up to date on the latest Salesforce.com features Handle multiple projects simultaneously Work closely with business analysts, project managers, and other developers Adhere to coding standards and best practices Required Skills and Qualifications 5+ years of experience in Salesforce development, configuration, and customization Advanced Apex programming skills, including batch jobs, triggers, web services, and unit testing Experience with Lightning Components, Visualforce, and Force.com integration techno
About the Role Adobe is seeking a Machine Learning Engineer to join the Adobe Genuine Engineering team. This group protects Adobe's ecosystem from fraud, abuse, and misuse using intelligent systems worldwide. In this position, you will build and develop machine learning models from scratch, including custom transformer-based frameworks, to identify fraudulent actions, stop account sharing, and protect the experience of hundreds of millions of users. You will manage the entire model lifecycle: raw behavioral data and feature engineering, architecture development, large-scale GPU training, deployment, and monitoring. The team is actively building in-house behavioral foundation models that learn identity-preserving representations from long sequences of user activity. This is a role for an engineer who wants to own deep learning systems end-to-end — not consume pre-built ones. Key Responsibilities Build and train deep learning models from scratch, including custom transformer and attention-based architectures for long behavioral event sequences. Own the full training stack: event tokenization, temporal and positional embeddings, self-supervised pretraining (e.g., masked modeling, contrastive learning), and downstream fine-tuning. Train large models efficiently on GPU infrastructure using mixed-precision training, gradient accumulation/checkpointing, efficient attention, and distributed strategies (DDP, FSDP, or equivalent). Build and optimize feature pipelines on Databricks and Spark, transforming raw behavioral events into high-quality model inputs. Translate prototypes into production ML systems — scalable, reliable, and observable — and drive inference performance through architectural and serving-side optimization. Contribute to MLOps practices: experiment tracking, model versioning, CI/CD, automated retraining, and
About the Role Adobe is seeking a Machine Learning Engineer to join the Adobe Genuine Engineering team. This group protects Adobe's ecosystem from fraud, abuse, and misuse using intelligent systems worldwide. In this position, you will build and develop machine learning models from scratch, including custom transformer-based frameworks, to identify fraudulent actions, stop account sharing, and protect the experience of hundreds of millions of users. You will manage the entire model lifecycle: raw behavioral data and feature engineering, architecture development, large-scale GPU training, deployment, and monitoring. The team is actively building in-house behavioral foundation models that learn identity-preserving representations from long sequences of user activity. This is a role for an engineer who wants to own deep learning systems end-to-end — not consume pre-built ones. Key Responsibilities Build and train deep learning models from scratch, including custom transformer and attention-based architectures for long behavioral event sequences. Own the full training stack: event tokenization, temporal and positional embeddings, self-supervised pretraining (e.g., masked modeling, contrastive learning), and downstream fine-tuning. Train large models efficiently on GPU infrastructure using mixed-precision training, gradient accumulation/checkpointing, efficient attention, and distributed strategies (DDP, FSDP, or equivalent). Build and optimize feature pipelines on Databricks and Spark, transforming raw behavioral events into high-quality model inputs. Translate prototypes into production ML systems — scalable, reliable, and observable — and drive inference performance through architectural and serving-side optimization. Contribute to MLOps practices: experiment tracking, model versioning, CI/CD, automated retraining, and
Other cities to consider
More places hiring for this role
Get new deployment lead jobs in India by email
Daily job updates · Unsubscribe anytime