About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
Jobs in India
Ai Infrastructure Engineer in India
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current ai infrastructure engineer jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.
About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. What You’ll Do: Architect the Future of AI Infrastructure: You will design, build, and own the end-to-end platform that supports the entire lifecycle of our ML models—from massive-scale distributed training to ultra-low-latency, highly-available inference. Optimize and Serve Cutting-Edge Models: You'll implement and scale sophisticated inference stacks for LLMs using frameworks like vLLM, TensorRT-LLM, or SGLang . You’ll solve complex challenges in throughput, latency, token streaming, and automated scaling to deliver a seamless user experience. Empower AI Innovation: You will act as a strategic partner to our AI Research and Data Science teams. You’ll create a seamless developer experience that accelerates their ability to experiment, fine-tune, and deploy groundbreaking models with velocity and confidence. Automate Everything: You'll develop robust CI/CD/CT (Continuous Training) pipelines using tools like Argo Workflows, ArgoCD, and GitHub Actions to automate model validation, deployment, and lifecycle management, ensuring our systems are both agile and rock-solid. What are we looking for Experience: 5+ years in infrastructure or software engineering, with at least 2+ years laser-focused on MLOps or ML infrastructu
About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Infrastructure Engineer Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role The Hyderabad Infra team builds and maintains Notion's internal async task runner and configuration management platform. The async task runner plays a critical role in ensuring our millions of users have a fast, reliable, and secure experience. The configuration management platform enables safe, explicit configuration management for Notion's product and infrastructure engineers. As part of the Hyderabad Infra Team, you’ll have a unique opportunity to shape how Notion manages and scales its async task runner and configuration management platform, enabling innovation across the company. What You'll Achieve You will contribute to the evolution and maintenance of our async task runner to meet the needs of over 100 million global users and support the rapid growth of our product and business. With guidance from senior team members, you'll help ensure our systems remain reliable, efficient, and scalable. You'll evaluate and integrate new technologies to keep us ahead of emerging challenges. Your work will empower our engineering team to build features confidently while you grow your skills in distributed systems and infrastr
From ₹20L/yr
XO Health believes healthcare is fixable. Become part of the community changing the face of the industry. XO Health is the first health plan designed by and for self-insured employers that delivers a more unified health experience for everyone – from those who receive care, to those who deliver it, to those who pay for it. We are growing a multi-disciplinary team of diverse and digitally empowered employees ready to rebuild trust in healthcare through comprehensive and unified transformation. CyberSecurity & Infrastructure Engineer - India (Remote) About the Role : The Cybersecurity & Infrastructure Engineer is responsible for designing, implementing, monitoring, and securing the organization's hybrid cloud and infrastructure environments. This role serves as a technical leader for cybersecurity operations, cloud security, compliance initiatives, infrastructure engineering, and incident response. The position combines hands-on infrastructure administration with cybersecurity engineering responsibilities across Microsoft Azure, Microsoft Sentinel, Microsoft 365, Entra ID, AWS, networking, endpoints, and security platforms. Engineering plays a key role in maintaining compliance with SOC 2 Type II controls, improving cyber resilience, supporting audits, and advancing the organization's security maturity. Responsibilities: SOC2 Audit Support organization's annual SOC 2 Type II audit program, including control design, evidence collection, remediation management, auditor coordination, and continuous compliance monitoring. Partner with business and technology stakeholders to ensure security, availability, confidentiality, and change management controls are effectively implemented and operating throughout the audit period. Drive successful completion of SOC 2 Type II examinations with minimal findings by maintaining an audit-ready environment, strengthening internal controls, and promoting a culture of security and compliance. Develop and maintain policies, pr
About Bolna Bolna is Voice AI infrastructure built for India. We help businesses deploy intelligent voice agents that can call, converse, and convert in any language, at scale. From collections to customer support to sales, our agents handle millions of conversations so humans don't have to. We're a YC F25 company, backed by General Catalyst, with 1,050+ paying customers and growing fast. Our team of ~25 is based in Bengaluru. The Role You're the person who makes our voice agents actually sound good and actually work for real customers. As an AI Solutions Engineer, you'll sit at the intersection of our product and our customers. Your primary job is to design, write, and iterate on the prompts and tools that power Bolna's voice agents, making them smarter, more natural, and more effective for each use case. You'll work closely with customers to understand their goals, build agent flows, test conversations, and fix what breaks. No heavy coding required-if you can vibe-code a basic script or write a solid system prompt, you're qualified. What You'll Do Write, test, and iterate on system prompts for voice agents across industries like D2C, fintech, healthcare, and logistics Listen to real call recordings, identify where agents fail, and fix them Build conversation flows and call pathways for new customer deployments Help onboard new customers-understand their use case, set up their agent, and get it live Maintain a growing library of prompts, templates, and best practices across Bolna's verticals Red-team agents-try to break them, find edge cases, and make them bulletproof Work with vernacular inputs: test agents in Hindi, Hinglish, and regional languages Feed insights back to product and engineering-you'll see what customers need before anyone else does What We're Looking For Must-have: You've spent serious time prompting ChatGPT, Claude, or similar LLMs-not just casually, but to actually build or solve something You're obsessive about language-you notice when a senten
Staff SRE for Cloud Network Infrastructure Team (Managing edge in multi-cloud(AWS & GCP), mTLS, Global Routing, AI Automation)
OktaSecure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities, such as reducing costs and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, global traffic routing, and resilient cloud infrastructure problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new network topologies, multi-cloud ecosystems, and agentic engineering models. Position Overview: The Staff Software Engineer - Infrastructure will play a foundational role in architecting, evolving, and securing Okta's global core network fabric and multi-cloud platform layers. This position focuses on building highly resilient, edge management in AWS and GCP cloud, executing enterprise-wide cloud migrations (AWS to GCP), enforcing absolute Zero-Trust primitives (mTLS / TLS 1.3), and integrating cutting-edge AI automation layers (Agentic SRE) into our day-to-day operations. As a Staff Engineer on this team, you will act as a key technical anchor in India, working closely with global counterparts to maintain Okta's high-availability SLAs while protecting the platform against active global DDoS attacks. Key Responsibilities: Global Ingress & Routerless Evolution (CFSaaS): Design and engineer Next-Gen traffic
About Us Blueshift is the Intelligent Customer Engagement Platform (CEP), headquartered in San Francisco, that empowers leading B2C brands to drive truly personalized, 1:1 marketing across every channel. Founded by repeat entrepreneurs who previously built Mertado (acquired by Groupon) and were part of the early team at Kosmix (acquired by Walmart), Blueshift leverages AI, including Predictive, Generative, and Agentic AI, to automate customer engagement for clients like ClearScore, LendingTree, Udacity, and U.S. News. Backed by top-tier VCs including Nexus Venture Partners, Storm Ventures, and SoftBank Venture Asia, the company has raised a total of $65 million in venture funding and is consistently recognized as a market leader and a Deloitte Technology Fast 500 award recipient. Blueshift is actively scaling its development center in Pune, India. As part of our team, you will drive innovation in cutting-edge technologies including machine learning, artificial intelligence, big data, and large-scale distributed data systems. This is an exciting career path for motivated individuals looking to build complex, impactful solutions that define the future of customer engagement. AI Solutions Engineer II As a Software Engineer in the AI Solutions team , you occupy a unique techno-functional position. You are not a researcher; you are an implementation specialist and problem-solver . You bridge the gap between our core AI infrastructure and real-world customer impact. You aren't just writing code; you are applying data engineering, analysis, and AI knowledge to help global brands realize the full potential of AI-driven marketing. Responsibilities End-to-End Solution Delivery: Lead the full lifecycle of custom AI projects—from initial customer design and technical architecture to testing and production implementation. Production Stewardship: Take ownership of the "last mile" of delivery. This includes triaging technical tickets, analyzing logs (Datadog/Kibana), and deb
About the Role Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. You’ll work across two areas: Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack. Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve. We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure. This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation , solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure. responsible for delivering the software but also for operating and supporting it in production. Why this Role You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective. You’ll have the opportunity to build foundational systems from the ground up, work on infrastructur
Role: Senior Cloud Engineer Location: Hyderabad, India (Hybrid) Department: Product Development About GHX: GHX (Global Healthcare Exchange) is a leading healthcare technology company on a mission to simplify the business of healthcare and improve patient outcomes. Founded in 2000, GHX has built the GHX Global Network — the world’s largest cloud-based supply chain community connecting healthcare providers, suppliers, distributors, and partners to automate key processes, reduce costs, and increase operational efficiency. Its solutions span electronic trading, procurement automation, inventory and contract management, business intelligence, and data synchronization, helping healthcare organizations improve productivity and focus more on patient care. Over the years, GHX has enabled significant cost savings for the industry and continues to innovate with intelligent automation and AI-driven capabilities. Website: https://www.ghx.com/ LinkedIn: https://www.linkedin.com/company/ghx/ About Role: The Senior Cloud Engineer leads the design, implementation, and operations of the organization’s cloud infrastructure. This role is responsible for complex projects, high-level architectural planning, and making strategic technology decisions that support scalability, performance, security, and cost optimization. Acting as a technical leader, the Senior Cloud Engineer provides mentorship to junior and mid-level engineers while driving innovation, resiliency, and compliance in cloud environments. Key Responsibilities: Design and Implementation Lead the design and implementation of cloud-native architectures and hybrid cloud solutions. Oversee cloud engineering projects, ensuring solutions are secure, resilient, and cost-effective. Contribute to high-level architectural planning and strategic decision-making around technology adoption. Review and approve design proposals, Infrastructure as Code (IaC) tem
₹2K – ₹2K/yr
Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is growing quickly. You will play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features, while also balancing scalability, reusability, and performance. Role Overview: We're looking for a Staff Platform Engineer to serve as the technical backbone of our Engineering organization. You'll own the technical strategy, and delivery of our platform — spanning architecture, DevOps, SRE, security, Dev-ex. This is a hands-on staff level role: you'll set technical direction, drive cross-team alignment, and be the senior escalation point for platform challenges. What you’ll do: Drive platform reliability, scalability, security, and cost efficiency across all environments. Technical Leadership: Provide technical leadership for platform components, Influence the technical strategy and architecture of our cloud platform, from CI/CD pipelines to observability and incident response. Design and implement platform components and reusable integration patterns that minimize custom development efforts, reduce the time spent on repetitive tasks, and ensure that integrations scale across multiple healthcare systems Partner closely with Architecture, DevOps, SRE, and Security teams to deliver cohesive platform solutions Cross-Functional Collaboration: Work closely with product teams, and solutions architects to understand integration needs and ensure the platform meets current and future business requirements. Serve as a senior escalation point for infrastructure and platform incidents Establish frameworks for: AI governance and compliance. Observability of systems. Traceability of decisions and outputs. Ensure enterprise readiness with security, auditability, and reliability in production environments. Security & Compliance : Ensure all p
A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO As Database Support Engineer, you’ll support various critical database platforms across Development, QA, UAT, and Production environments. The role partners closely with application teams, application support, and database engineers and operates within a Follow‑the‑Sun model to ensure availability, performance, and reliability of database services. Key responsibilities include: • Provide operational support for enterprise database platforms in both on-prem private cloud and public cloud • Monitor database health, capacity, performance, and availability, and respond to alerts, diagnose issues, and perform timely remediation • Perform routine maintenance activities (patching, upgrades, housekeeping etc) • Troubleshoot database‑related incidents and collaborate on root cause analysis • Work closely with application owners, application support teams, and DB Engineers • Provide guidance on database best practices and operational standards • Participate in cross‑team problem resolution and continuous improvement initiatives • Contribute to design, implementation and testing of automation and self service capabilities of DB platforms • Drive continuous improvement, identifying opportunities to reduce toil and increase platform efficiency. • Participate in a Follow‑the‑Sun operating model, including shift‑based coverage and handoffs WHAT’S REQUIRED • Bachelor’s degr
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What you will be doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuous
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity As a Senior Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with 5–8 years of relevant experience who can leverage open-source and custom tooling to orchestrate and maintain large-scale Kubernetes environments. You will play a "Captain" role—leading critical deliverables and mentoring junior engineers while maintaining the reliability of our global fleet. What You'll Do Architectural Leadership: Drive the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management. Platform Orchestration: Lead complex, large-scale infrastructure shifts. Operational Excellence: Take ownership of incident response, author comprehensive retrospectives, and implement systemic hardening to prevent recurrence using advanced overcommit strategies. This Role Requires Experience: 5–8 years in a DevOps, Site Reliability, or Infrastructure Engineering role. Kubernetes Mastery: Deep internals knowledge of Kubernetes and hands-on experience writing custom operators. Tooling Proficiency: Strong experience building production-grade tools and services, specifically for infrastructure automation. Operations-Heavy Mindset: A proven track record of Day 1/Day 2 operations for a large-scale Kubernetes fleet, handling high-severity incidents, and improving SLA compliance through auto
Other cities to consider
More places hiring for this role
Get new ai infrastructure engineer jobs in India by email
Daily job updates · Unsubscribe anytime