About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine
Jobs in India
Senior Infrastructure Software Engineer in India
1,107 active opportunities · Updated October 2026
Showing
15 jobs
Explore current senior infrastructure software engineer jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.
About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. What You’ll Do: Architect the Future of AI Infrastructure: You will design, build, and own the end-to-end platform that supports the entire lifecycle of our ML models—from massive-scale distributed training to ultra-low-latency, highly-available inference. Optimize and Serve Cutting-Edge Models: You'll implement and scale sophisticated inference stacks for LLMs using frameworks like vLLM, TensorRT-LLM, or SGLang . You’ll solve complex challenges in throughput, latency, token streaming, and automated scaling to deliver a seamless user experience. Empower AI Innovation: You will act as a strategic partner to our AI Research and Data Science teams. You’ll create a seamless developer experience that accelerates their ability to experiment, fine-tune, and deploy groundbreaking models with velocity and confidence. Automate Everything: You'll develop robust CI/CD/CT (Continuous Training) pipelines using tools like Argo Workflows, ArgoCD, and GitHub Actions to automate model validation, deployment, and lifecycle management, ensuring our systems are both agile and rock-solid. What are we looking for Experience: 5+ years in infrastructure or software engineering, with at least 2+ years laser-focused on MLOps or ML infrastructu
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. We are seeking a Senior Engineer to help build next-generation Database Observability products. In this high-impact role, you will own the entire lifecycle of critical data flows from developing lightweight database agents and high-throughput ingestion pipelines to building intelligent DB Recommendation engines and autonomous DB AI Agents. You will build cutting-edge capability that moves observability from passive monitoring to automated, AI-driven database diagnostic and remediation workflows. If you are energized by hard distributed systems problems, extreme scalability, and harnessing agentic AI to solve complex developer challenges, this is a career-defining opportunity to shape the future of intelligent database observability at New Relic.We look forward to talking with you! What you'll do . Architect, develop, and scale high-throughput, low-late
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. The Infrastructure product organization develops New Relic infrastructure instrumentation agents, next generation data processing and management services, vulnerability management, and security testing capabilities for on-prem and cloud customers. We work with data at a scale using a diverse tech stack (Go, Java, JavaScript, React GraphQL, Kubernetes, many public cloud web services, and more). As a senior backend engineer, you will help us build and extend next generation solutions such as a control plane for customers to manage their data pipelines at scale. New Relic is looking for engineers who are interested in building a brand-new observability experience. This high-impact engineering position is a phenomenal opportunity to own and build a set of next generation services and capabilities for the company. We are searching for a motivated engineer who is ready for a career-defining role in their next opportunity. We look forward to talking with you! What you'll do ● Design, Build, maintain, and scale back-end services and their support tools. ● Participate in architectural definitions with a high degr
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. Database Observability is a critical pillar of New Relic's platform strategy. We are looking for an experienced Engineering Manager (M3) to lead a senior, high-performing team building next-generation Database Observability products. Your team will own the entire lifecycle of critical telemetry data flows from lightweight database agents and high-throughput ingestion pipelines to intelligent DB recommendation engines and autonomous DB AI Agents. You will lead a team that includes senior and Lead-level engineers with deep domain expertise in distributed systems and AI. Your primary value will come from setting strategic technical direction, enabling their best work, and fostering a high-accountability culture while partnering closely with Product and Design to deliver features that directly drive New Relic's Database Observability. What you'll do Manage a full-stack engineering team (6–8 engineers) spanning backend systems, database telemetry, agent engineering, and UI workflows. Own end-to-end delivery sprint planning, roadmap execution, system quality, and operational excellence for critical database ingesti
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. Service Levels is one of New Relic's most commercially critical products — it powers SLO compliance and reliability measurement for thousands of customers. We're looking for an experienced Engineering Manager to lead a senior, high-performing team building the next generation of service level management at scale. You'll lead a team that includes lead-level engineers with deep domain expertise, and your value will come from enabling their best work — not directing it. You'll own delivery, quality, and team health while partnering closely with product and design to ship features that directly impact New Relic's commercial momentum. What you'll do Lead a full-stack engineering team of 6-8 across backend (Java/Spring Boot) and frontend (React/TypeScript) Own end-to-end delivery — sprint planning, execution, quality, and release Set clear expectations, manage performance equitably, and develop engineers at every level Partner with Product Manager and XD to translate requirements into technically sound, deliverable plans Drive architectural discussions and hold the team to engineering excellence standards Identify
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are seeking an experienced and technically influential Senior Software Development Engineer to join our Cloud Tooling and Pipelines team. This pivotal team is responsible for the design, development, and maintenance of our core Continuous Delivery (CD) platform (leveraging Spinnaker and custom tooling), Infrastructure as Code (IaC) execution engines (primarily Terraform), and a suite of supporting microservices. These systems are critical for enabling and managing our extensive resource footprint across AWS ECS and EKS. As a Senior Software Development Engineer, you will be a key contributor, driving the implementation of scalable, reliable, and secure software solutions that automate infrastructure provisioning and application deployments. Your deep expertise in software engineering principles and cloud-native development will be essential in building and enhancing our critical tooling for infrastructure provisioning, vulnerability management, and IaC deployments. You will also play a vital role in mentoring other engineers and influencing the team's technical roadmap. If you have a strong passion for building robust software systems that empower operational efficiency at scale, we encourage you to apply. Key Responsibilities Design and Develop Core Platform Components: Lead the design and development of scalable and reliable microservices and tools that form the backbone of Okta's Continuous Delivery (CD) platform (including components for Spinna
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Datastores Engineer, Platform Infrastructure The Auth0 platform secures more than 100 million logins each day for customers all around the world - and we're growing fast! The Platform Infrastructure team enables Auth0 engineers to move faster by giving them tools to easily deploy and manage their services on AWS and Azure. This is a role with a huge impact. You will get to work with engineers throughout the organization and what you build will be a foundational piece of the infrastructure that allows Auth0 to scale for years to come. We are looking for Engineer who are passionate about distributed systems, availability, and delivering customer value to join our Platform Infrastructure Datastores team. Because we build and support the overall Auth0 platform, the ideal candidate is someone who is passionate about infrastructure, operations, databases and not intimidated by cross-organization coordination and collaboration. You will: Develop our large, distributed and highly-available infrastructure. Implement platform tools that allow feature teams to deploy and manage the datastores for their services. Research new technologies to accelerate new environment creation. Carry cross team initiatives from end to end: code reviews, design reviews, operational robustness, security hygiene, etc. Participate in the team's on-call rotation. You might be a good fit if you: Have 5-8 years of software development experience. Are proficient in or have a desire
We are seeking a Senior Software Engineer with strong infrastructure expertise to design, build, and operate the next generation of our enterprise Observability, Automation, and AI-driven Reliability Platform. This role will build highly scalable distributed systems and platform services spanning Storage, Compute, Network, VMware, OpenShift, and bare-metal infrastructure. The engineer will help transform infrastructure operations from reactive monitoring and manual remediation to proactive, predictive, and AI-driven autonomous operations. What You Will Be Doing: Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at large scale. Develop reusable platform services, APIs, automation frameworks, and control planes that enable self-service, reduce operational toil, and automate infrastructure operations across multiple engineering teams. Build scalable telemetry and event-processing systems spanning metrics, logs, traces, events, topology, and alerts, with the performance and efficiency to process billions of infrastructure signals. Build intelligent and AI-native reliability capabilities, including agentic workflows for anomaly detection, forecasting, root-cause analysis, automated debugging, and closed-loop remediation. Drive technical architecture and engineering direction across Storage, Compute, Network, and Platform domains, solving complex and ambiguous problems that span multiple teams. Engineer for production at scale, with strong focus on software quality, scalability, security, performance, observability, maintainability, and operational readiness. Provide technical leadership and mentorship, influence engineerin
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. About the Role We are looking for Staff System Software Engineer in Test to join our team. In this role, you will be responsible for design, development, automation and reporting of Integration and system tests spanning across firmware and device drivers. This role requires you to have significant technical breadth and deep understanding of low-level system software specifically in server class systems. You will be part of a new team responsible for integration of different system software deliverables and development of system tests spanning all the components. You will contribute to shaping the test strategy , guide best practices and solve complex problems while maintaining a strong hands-on focus. You will partner with development and other QA teams to deliver high quality scalable and reliable solutions. About the Team Integration and system test team is responsible for verification and validation of integrated components across Board management controller (BMC), Firmware and Linux device driver. The team is also responsible for management and maintenance of common tools and pipel
Location Details: Pune, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a hybrid position. You’ll divide your time between working remotely from your home and an office, so you should live within commuting distance. Hybrid teams may work in-office as much as a few times a week or as little as once a month or quarter, as decided by leadership. The hiring manager can share more about what hybrid work might look like for this team. Join our Team Our team builds and operates the foundational infrastructure platforms that power GoDaddy's engineering organization. We own critical services including secrets management, software distribution, host security controls, and live patching for thousands of Linux systems running on OpenStack. This role sits at the intersection of Linux engineering, platform engineering, reliability engineering, and security. You will help define how core infrastructure services are designed, operated, automated, and scaled across the enterprise! What you'll get to do... Design, build, and operate highly available, scalable, and secure infrastructure platforms supporting large-scale Linux environments, with a focus on reliability, resiliency, and operational efficiency Lead the architecture, implementation, and operation of infrastructure services, including OpenStack, enterprise secrets management, package management, software promotion pipelines, and platform lifecycle management Develop and maintain automation solutions using infrastructure-as-code, Ansible, Python, Go, and self-service capabilities to improve efficiency and reduce operational overhead Build and improve observability and reliability practices through monitoring, logging, alerting, dashboards, managing incidents, analyzing underlying causes, disaster recovery, and service health reporting
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know Okta Okta is The World’s Identity Company. We free everyone to safely use any technology anywhere, on any device or app. Our Workforce and Customer Identity Clouds enable secure yet flexible access, authentication, and automation that transforms how people move through the digital world, putting Identity at the heart of business security and growth. At Okta, we celebrate a variety of perspectives and experiences. We are not looking for someone who checks every single box we’re looking for lifelong learners and people who can make us better with their unique experiences. Join our team! We’re building a world where Identity belongs to you. About Technology Data and Intelligence at Okta At Okta, the Technology Data and Intelligence (TDI) team drives internal efficiency through secure, scalable, and innovative systems. TDI partners with teams across the company to build and support the infrastructure, automation, and enterprise applications that keep operations running smoothly. Focused on enabling productivity and aligning technology with business goals, TDI plays a vital role in both day-to-day operations and long-term strategic growth. The Senior Software Engineer Opportunity We are looking for a Senior Software Engineer to join our growing team in TDI and help scale our internal business solutions with a sharp focus on security, reliability, scalability, and intelligent automation. You will be responsible for designing and developing customizati
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Job Overview: We are looking for a Senior Engineer to join the FGA DevEx team and help evolve our end to end developer experience across both OSS and SaaS. This team owns the SDKs in Go, JavaScript, .NET, Python, Java and other languages, along with CLI workflows, IDE integrations, GitHub automation, developer documentation, and release strategy. All development is done in the open as open source, and we actively welcome and review community contributions. Our guiding principle is One developer experience, many deployment models. As a Senior Engineer, you will take ownership of significant portions of the SDK and tooling ecosystem, ensure high quality implementations across languages, and contribute to a consistent and reliable developer experience. Responsibilities: Maintain and enhance existing SDKs for FGA in Go, JavaScript, .NET, Python, and Java, leveraging our SDK generator framework. Customize and refine SDK templates and wrappers to ensure consistency across languages and support configuration overrides such as store ID, authorization model ID, headers, and parallelization limits. Implement and improve core SDK features including client credentials authentication flows, robust error mapping, retry logic with jitter, and rate limiting safeguards. Implement advanced capabilities such as BatchCheck, ListRelations, and non transactional write operations with appropriate parallelization and performance considerations. Contribute to the SDK generator tool
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Title: Senior Software Engineer, Lifecycle Management Company Description: Okta is the leading independent provider of enterprise identity. The Okta Identity Cloud enables organizations to securely connect the right people to the right technologies at the right time. With over 6,500 pre-built integrations to applications and infrastructure providers, Okta customers can easily and securely use the best technologies for their business. Over 7,950 organizations, including 20th Century Fox, JetBlue, Nordstrom, Slack, Teach for America and Twilio, trust Okta to help protect the identities of their workforces and customers Position Description: We are looking for an experienced Senior Software Engineer to work on our Onboarding and Lifecycle Management (LCM) Platform team with focus on enhancing and managing services for importing, syncing and provisioning identities and access policies i.e., users, groups, roles, entitlements, etc. These features allow customers the flexibility to link and enhance their business processes with Okta’s identity management product. This role is to build, design solutions, and maintain our platform for scale. The ideal candidate will be naturally curious and has experience building software systems to manage and deploy reliable and performant infrastructure and product code at scale on a cloud infrastructure. Job Duties and Responsibilities: Work with senior engineering team in major development projects, design and implementation C
Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. Location- Chennai Team: Engineering Enablement Group As a Senior Software Engineer in our Engineering Enablement Group, you will lead the re-design and evolution of our Mobile Branding framework — the system that enables customers to create custom-branded versions of the Appian mobile application for both iOS and Android. You will drive the architectural modernization of the end-to-end branding pipeline, from the customer-facing Forum application and provisioning tools to the backend build service running on Mac EC2 runners in AWS. By leveraging modern microservices, CI/CD automation, and cloud-native infrastructure, you will transform the current system into a more reliable, scalable, and maintainable platform that reduces manual intervention and accelerates customer delivery. We are looking for a technical leader who can bridge the gap between complex Ruby/Bash-based tooling, Appian process models, and AWS infrastructure to deliver a seamless mobile branding experience. Primary Qualifications: 6-9 Strong working experience with Android and iOS frameworks and mobile application development workflows. Familiarity with mobile build systems (Fastlane, Xcode, Gradle) and code-signing workflows. Experience with proficiency in Python, with experience in Ruby, Bash, or Go being a plus. Advanced experience with AWS infrastructure (S3, Lambda, EC2) and CI/CD pipeline design. Strong end-to-end knowledge of pipeline creation, deployment automation, and infrastructure-as-code (Terraform). Familiarity with monitoring, observability, and performanc
Other cities to consider
More places hiring for this role
Get new senior infrastructure software engineer jobs in India by email
Daily job updates · Unsubscribe anytime