Jobs in India

Senior Infrastructure Engineer in India

1,091 active opportunities · Updated October 2026

Explore current senior infrastructure engineer jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.

E(
📍 Bengaluru, Karnataka, India· Full-time
✓ Quality checkedExact matchCompany trend -94.4%
Quick readExact title match for your search

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Infrastructure Engineer Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning

PythonSQLAWSAzure
G
📍 Pune, Maharashtra, India
✓ High-confidence listingCompany trend 0%
Quick readStrong listing-quality and freshness signals

Location Details: Pune, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a hybrid position. You’ll divide your time between working remotely from your home and an office, so you should live within commuting distance. Hybrid teams may work in-office as much as a few times a week or as little as once a month or quarter, as decided by leadership. The hiring manager can share more about what hybrid work might look like for this team. Join our Team Our team builds and operates the foundational infrastructure platforms that power GoDaddy's engineering organization. We own critical services including secrets management, software distribution, host security controls, and live patching for thousands of Linux systems running on OpenStack. This role sits at the intersection of Linux engineering, platform engineering, reliability engineering, and security. You will help define how core infrastructure services are designed, operated, automated, and scaled across the enterprise! What you'll get to do... Design, build, and operate highly available, scalable, and secure infrastructure platforms supporting large-scale Linux environments, with a focus on reliability, resiliency, and operational efficiency Lead the architecture, implementation, and operation of infrastructure services, including OpenStack, enterprise secrets management, package management, software promotion pipelines, and platform lifecycle management Develop and maintain automation solutions using infrastructure-as-code, Ansible, Python, Go, and self-service capabilities to improve efficiency and reduce operational overhead Build and improve observability and reliability practices through monitoring, logging, alerting, dashboards, managing incidents, analyzing underlying causes, disaster recovery, and service health reporting

PythonGitLinuxAI
TA
📍 India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine

PythonKubernetesCI/CDAI
G
📍 Hyderabad, TELANGANA, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Role: Senior Cloud Engineer Location: Hyderabad, India (Hybrid) Department: Product Development About GHX: GHX (Global Healthcare Exchange) is a leading healthcare technology company on a mission to simplify the business of healthcare and improve patient outcomes. Founded in 2000, GHX has built the GHX Global Network — the world’s largest cloud-based supply chain community connecting healthcare providers, suppliers, distributors, and partners to automate key processes, reduce costs, and increase operational efficiency. Its solutions span electronic trading, procurement automation, inventory and contract management, business intelligence, and data synchronization, helping healthcare organizations improve productivity and focus more on patient care. Over the years, GHX has enabled significant cost savings for the industry and continues to innovate with intelligent automation and AI-driven capabilities. Website: https://www.ghx.com/ LinkedIn: https://www.linkedin.com/company/ghx/ About Role: The Senior Cloud Engineer leads the design, implementation, and operations of the organization’s cloud infrastructure. This role is responsible for complex projects, high-level architectural planning, and making strategic technology decisions that support scalability, performance, security, and cost optimization. Acting as a technical leader, the Senior Cloud Engineer provides mentorship to junior and mid-level engineers while driving innovation, resiliency, and compliance in cloud environments. Key Responsibilities: Design and Implementation Lead the design and implementation of cloud-native architectures and hybrid cloud solutions. Oversee cloud engineering projects, ensuring solutions are secure, resilient, and cost-effective. Contribute to high-level architectural planning and strategic decision-making around technology adoption. Review and approve design proposals, Infrastructure as Code (IaC) tem

AWSAzureGCPKubernetes
N
📍 Bengaluru, Bengaluru, India
✓ Quality checkedCompany trend -100%

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What you will be doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuous

PythonDockerKubernetesAI
NR
📍 Bengaluru, India· Full-time
✓ Quality checkedCompany trend -36.4%

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity As a Senior Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with 5–8 years of relevant experience who can leverage open-source and custom tooling to orchestrate and maintain large-scale Kubernetes environments. You will play a "Captain" role—leading critical deliverables and mentoring junior engineers while maintaining the reliability of our global fleet. What You'll Do Architectural Leadership: Drive the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management. Platform Orchestration: Lead complex, large-scale infrastructure shifts. Operational Excellence: Take ownership of incident response, author comprehensive retrospectives, and implement systemic hardening to prevent recurrence using advanced overcommit strategies. This Role Requires Experience: 5–8 years in a DevOps, Site Reliability, or Infrastructure Engineering role. Kubernetes Mastery: Deep internals knowledge of Kubernetes and hands-on experience writing custom operators. Tooling Proficiency: Strong experience building production-grade tools and services, specifically for infrastructure automation. Operations-Heavy Mindset: A proven track record of Day 1/Day 2 operations for a large-scale Kubernetes fleet, handling high-severity incidents, and improving SLA compliance through auto

AWSAzureKubernetesCI/CD
N
📍 Hyderabad, India, India· Full-time
✓ Quality checkedCompany trend -100%

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role The Hyderabad Infra team builds and maintains Notion's internal async task runner and configuration management platform. The async task runner plays a critical role in ensuring our millions of users have a fast, reliable, and secure experience. The configuration management platform enables safe, explicit configuration management for Notion's product and infrastructure engineers. As part of the Hyderabad Infra Team, you’ll have a unique opportunity to shape how Notion manages and scales its async task runner and configuration management platform, enabling innovation across the company. What You'll Achieve You will contribute to the evolution and maintenance of our async task runner to meet the needs of over 100 million global users and support the rapid growth of our product and business. With guidance from senior team members, you'll help ensure our systems remain reliable, efficient, and scalable. You'll evaluate and integrate new technologies to keep us ahead of emerging challenges. Your work will empower our engineering team to build features confidently while you grow your skills in distributed systems and infrastr

CH
📍 Hyderabad, TELANGANA, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is growing quickly. You will play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features, while also balancing scalability, reusability, and performance. Role Overview: We're looking for a Staff Platform Engineer to serve as the technical backbone of our Engineering organization. You'll own the technical strategy, and delivery of our platform — spanning architecture, DevOps, SRE, security, Dev-ex. This is a hands-on staff level role: you'll set technical direction, drive cross-team alignment, and be the senior escalation point for platform challenges. What you’ll do: Drive platform reliability, scalability, security, and cost efficiency across all environments. Technical Leadership: Provide technical leadership for platform components, Influence the technical strategy and architecture of our cloud platform, from CI/CD pipelines to observability and incident response. Design and implement platform components and reusable integration patterns that minimize custom development efforts, reduce the time spent on repetitive tasks, and ensure that integrations scale across multiple healthcare systems Partner closely with Architecture, DevOps, SRE, and Security teams to deliver cohesive platform solutions Cross-Functional Collaboration: Work closely with product teams, and solutions architects to understand integration needs and ensure the platform meets current and future business requirements. Serve as a senior escalation point for infrastructure and platform incidents Establish frameworks for: AI governance and compliance. Observability of systems. Traceability of decisions and outputs. Ensure enterprise readiness with security, auditability, and reliability in production environments. Security & Compliance : Ensure all p

AWSCI/CDGitRest
P
📍 Bengaluru, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

JOB TITLE Mobility Engineer, End User Technology A CAreer with point72’s Technology TEAM As Point72 continues to reimagine the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting and discovering new ways to provide exceptional end user experience while embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our team while satisfying your own intellectual curiosity. Our Technology Infrastructure team engineers and operates the foundational technology platforms that power all of the firm’s applications and businesses. Our disciplines span a broad array of technologies from datacenter infrastructure to large scale cloud services, with the shared goal of providing the most reliable, performant, modern technology platforms to improve time-to-market for our business. We also deliver end-user technology solutions to support the evolving collaboration and productivity needs of our global teams. Our team focuses on innovation and challenging the current state of our infrastructure technology in a fast-paced, dynamic, and collaborative working environment. What you’ll do We are seeking a Mobility Engineer with deep expertise in Microsoft Intune to own and evolve Point72’s modern endpoint and mobility platform. This role is heavily focused on mobility and modern device management, while also supporting Windows, macOS, mobile devices, and Windows 365 Cloud PCs.The ideal candidate will be responsible for the end-to-end lifecycle management of corporate and BYOD endpoints, ensuring a secure, seamless, and high-performing end-user experience across the firm. Lead the end‑to‑end management of mobile and endpoint devices using Microsoft Intune and other enterprise MDM tools. Design, deploy, and support Intune for iOS, Android, Windows, macOS, and Win

AzureAgileAIGo
D
📍 Bengaluru, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. What You’ll Do: Architect the Future of AI Infrastructure: You will design, build, and own the end-to-end platform that supports the entire lifecycle of our ML models—from massive-scale distributed training to ultra-low-latency, highly-available inference. Optimize and Serve Cutting-Edge Models: You'll implement and scale sophisticated inference stacks for LLMs using frameworks like vLLM, TensorRT-LLM, or SGLang . You’ll solve complex challenges in throughput, latency, token streaming, and automated scaling to deliver a seamless user experience. Empower AI Innovation: You will act as a strategic partner to our AI Research and Data Science teams. You’ll create a seamless developer experience that accelerates their ability to experiment, fine-tune, and deploy groundbreaking models with velocity and confidence. Automate Everything: You'll develop robust CI/CD/CT (Continuous Training) pipelines using tools like Argo Workflows, ArgoCD, and GitHub Actions to automate model validation, deployment, and lifecycle management, ensuring our systems are both agile and rock-solid. What are we looking for Experience: 5+ years in infrastructure or software engineering, with at least 2+ years laser-focused on MLOps or ML infrastructu

PythonKubernetesCI/CDGit
E
📍 Bengaluru, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As a hands-on senior leader for our India SecOps team, you will shape and safeguard Everpure’s security posture at the intersection of detection engineering, threat hunting, attack surface management, and incident response. Positioned as a strategic cornerstone in Bangalore, you will empower an elite engineering team, optimize critical SecOps pipelines, and partner cross-functionally across global engineering and infrastructure groups. By driving execution excellence and high team morale, you ensure our enterprise platform and global telemetry remain resilient against evolving threats. WHAT YOU'LL DO Scale & Lead SecOps Operations: Architect, mentor, and grow the India SecOps team to foster an environment of high morale, technical excellence, and rapid execution across detection engineering and incident response. Proactively Manage & Remediate Attack Surface: Own end-to-end Attack Surface Management (ASM) across cloud environments, SaaS applications, endpoints, and secrets management to measurably minimize enterprise exposure and mitigate risk. Optimize Telemetry & Incident Response: Mature SIEM and SOAR automation pipelines to drastically reduce mean time to detect, contain, and respond (MTTD/MTTC/MTTR) while continuously elevating alert fidelity and signal confidence. Drive Cross-Functional Alignment & RCA Postmortems: Lead continuous validation through purple-teaming and incident postmortems alo

AWSAzureCI/CDRest
G
📍 Mumbai, Maharashtra, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Overview: The role is responsible for managing and supporting the operations of Global Network Engineering in a hybrid/cloud environment and provides senior-level expertise to the network engineering team. We're looking for an expert in on-premises networking, cloud infrastructure networking (Azure/AWS), and telecommunications (VOIP) in global environments, with knowledge and focus on zero-trust networking methodologies. This position requires off-hours on-call availability. This is a remote position with a 3 PM - 12 AM Shift. Candidates may need to visit the Mumbai office as and when needed in the general shift. What You'll Do: Manage physical network management and support covering Guidepoint offices, including, but not limited to, firewalls, routers, switches, etc. Responsible for managing the Enterprise WiFi Operations ZScaler Network management Monitor global traffic latency, issues, and application performance Update and design next-generation network models Network detection and response Intrusion detection and prevention technologies Azure Cloud Traffic integrating with On-prem traffic Azure Web Security CASB models and methodologies What You Have: 8+ years of experience with core networking in an enterprise environment, ideally global network design and security methodologies. 5+ years of working experience with advanced Azure cloud networking technologies. Must have substantial experience with managing WiFi Operations. CCNP Certified or equivalent experience 3+ years of hands-on experience, preferably with Sonic-wall/Netgear. Solid engineering knowledge of routing, switching (VLANs), Azure networking, firewalls, and traffic management. Update network security based on VNETs and Subnets, with traffic monitoring, etc. Knowledgeable in Azure Application Gateways, Front Door, and other security tools. Nice to have: Must know telecommunication technologies (VOIP, SIP Trunking, Microsoft Teams) Network security experience (vulnerability manage

AWSAzureAIRust
O
📍 Bengaluru, India· Full-time
✓ Quality checkedCompany trend -71.5%

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are seeking an experienced and technically influential Senior Software Development Engineer to join our Cloud Tooling and Pipelines team. This pivotal team is responsible for the design, development, and maintenance of our core Continuous Delivery (CD) platform (leveraging Spinnaker and custom tooling), Infrastructure as Code (IaC) execution engines (primarily Terraform), and a suite of supporting microservices. These systems are critical for enabling and managing our extensive resource footprint across AWS ECS and EKS. As a Senior Software Development Engineer, you will be a key contributor, driving the implementation of scalable, reliable, and secure software solutions that automate infrastructure provisioning and application deployments. Your deep expertise in software engineering principles and cloud-native development will be essential in building and enhancing our critical tooling for infrastructure provisioning, vulnerability management, and IaC deployments. You will also play a vital role in mentoring other engineers and influencing the team's technical roadmap. If you have a strong passion for building robust software systems that empower operational efficiency at scale, we encourage you to apply. Key Responsibilities Design and Develop Core Platform Components: Lead the design and development of scalable and reliable microservices and tools that form the backbone of Okta's Continuous Delivery (CD) platform (including components for Spinna

PythonJavaSQLMySQL
O
📍 Bengaluru, India· Full-time
✓ Quality checkedCompany trend -71.5%

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Datastores Engineer, Platform Infrastructure The Auth0 platform secures more than 100 million logins each day for customers all around the world - and we're growing fast! The Platform Infrastructure team enables Auth0 engineers to move faster by giving them tools to easily deploy and manage their services on AWS and Azure. This is a role with a huge impact. You will get to work with engineers throughout the organization and what you build will be a foundational piece of the infrastructure that allows Auth0 to scale for years to come. We are looking for Engineer who are passionate about distributed systems, availability, and delivering customer value to join our Platform Infrastructure Datastores team. Because we build and support the overall Auth0 platform, the ideal candidate is someone who is passionate about infrastructure, operations, databases and not intimidated by cross-organization coordination and collaboration. You will: Develop our large, distributed and highly-available infrastructure. Implement platform tools that allow feature teams to deploy and manage the datastores for their services. Research new technologies to accelerate new environment creation. Carry cross team initiatives from end to end: code reviews, design reviews, operational robustness, security hygiene, etc. Participate in the team's on-call rotation. You might be a good fit if you: Have 5-8 years of software development experience. Are proficient in or have a desire

SQLPostgreSQLMongoDBRedis
N
📍 Bengaluru, Bengaluru, India
✓ Quality checkedCompany trend -100%

We are seeking a Senior Software Engineer with strong infrastructure expertise to design, build, and operate the next generation of our enterprise Observability, Automation, and AI-driven Reliability Platform. This role will build highly scalable distributed systems and platform services spanning Storage, Compute, Network, VMware, OpenShift, and bare-metal infrastructure. The engineer will help transform infrastructure operations from reactive monitoring and manual remediation to proactive, predictive, and AI-driven autonomous operations. What You Will Be Doing: Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at large scale. Develop reusable platform services, APIs, automation frameworks, and control planes that enable self-service, reduce operational toil, and automate infrastructure operations across multiple engineering teams. Build scalable telemetry and event-processing systems spanning metrics, logs, traces, events, topology, and alerts, with the performance and efficiency to process billions of infrastructure signals. Build intelligent and AI-native reliability capabilities, including agentic workflows for anomaly detection, forecasting, root-cause analysis, automated debugging, and closed-loop remediation. Drive technical architecture and engineering direction across Storage, Compute, Network, and Platform domains, solving complex and ambiguous problems that span multiple teams. Engineer for production at scale, with strong focus on software quality, scalability, security, performance, observability, maintainability, and operational readiness. Provide technical leadership and mentorship, influence engineerin

PythonKubernetesAITerraform
🔔

Get new senior infrastructure engineer jobs in India by email

Daily job updates · Unsubscribe anytime