NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What you will be doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuous
Jobs in India
Incident Commander in India
164 active opportunities · Updated October 2026
Showing
15 jobs
Explore current incident commander jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services. What you’ll be doing: Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale. Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation. Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data. Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems. Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation. What we need to see: BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10+ years operating production infrastructure or platform services. Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges. <l
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput, and 99.999 availability. We're looking for a technical leader to help us to continue to scale the service with great people and reliable, cost-effective and efficient infrastructure, processes and tooling. As the Director of Site Reliability Engineering you will oversee the SRE organization focused on Okta platform, Databases, Edge networking, K8s platform, CI/CD, Observability, FinOps, and automation platform & tooling. Job Duties and Responsibilities: Build and lead a high-caliber India-based SRE organization supporting Okta’s production fleet. Partner with global engineering, product, and infrastructure leaders to deliver resilient, scalable, and secure services. Define and execute the India SRE strategy in alignment with global reliability goals. Lead post-incident reviews, drive root-cause analysis, and ensure long-term corrective actions. Participate in incident management, on-call rotations, and blameless RCAs. Implement automation and observability to reduce manual toil and improve operational efficiency. Drive adoption of modern infrastructure practices: infrastructure as code (Terraform), container orchestration (Kubernetes), and AI within Infrastructure org. H
NVIDIA pioneers computer graphics, gaming, AI, and accelerated computing. We are looking for a Technical Platform Operations Lead to join our team and play an important role in scaling Sales AI applications and platforms. This position offers the opportunity to shape how these solutions operate after launch and help ensure they remain reliable, secure, well governed, widely adopted, and continuously improved. You will collaborate with Sales, Product, Engineering, Data, Security, and IT teams to strengthen platform health, improve the user experience, and increase business impact. What you’ll be doing: Lead end-to-end post-launch operations for Sales AI applications, including availability, performance, support readiness, releases, upgrades, and lifecycle planning. Develop effective processes for incident response, problem management, changes, and issue resolution. Coordinate timely recovery and lasting improvements. Analyze service-level indicators and objectives, adoption metrics, dashboards, alerts, and user feedback to identify risks, performance degradation, and usage gaps. Collaborate with partner teams to translate operational signals and user needs into prioritized improvements and roadmap inputs. Improve adoption and business value through usage analytics, enablement, feedback loops, and user experience enhancements. Establish governance practices for security, access controls, compliance, documentation, and platform support. Develop automation, observability, and self-service capabilities that simplify operations and reduce repetitive work and recurring incidents. Prepare new AI capabilities and releases for production with runbooks, monitoring, rollback plans, support models, and partner enablement. What we need to see: 8+ years of experience in technical operations, pl
Position : Senior ServiceNow Developer Exp Level: 7+ years of experience Location: Hyderabad / Visakhapatnam Shift Timings: 2:00PM to 11:00PM IST Skills: #ServiceNow, #Rest & Soap API's, #Javascript Key Responsibilities: Design, develop, and implement complex ServiceNow solutions using Flow Designer, Business Rules, Client Scripts, UI Policies, Script Includes, Integrations, and custom applications. Lead technical design for enhancements across modules such as ITSM, CMDB, Asset or others as required. Develop and maintain #ServiceNow catalog items, workflows, record producers, and custom applications following platform best practices. Integrations Build and support integrations with external systems using REST/SOAP APIs, MID Server, IntegrationHub, web services, and authentication methods like OAuth and SAML. Troubleshoot integration failures and optimize performance. Required Qualifications: 5+ years of hands-on ServiceNow development experience. Strong understanding of JavaScript, AngularJS, Glide API, and general web technologies. Expertise in multiple ServiceNow modules like ITSM, ITOM, HRSD, CSM, SecOps. (ITSM is required; others a plus). Experience with Service Portal, Workspace, and UI Builder. Roll : service Location : Visakhapatnam and Hyderabad Employment Type : Full-Time Experience : 5+ Years Skiils : SAP FI,SAP ABAP Job Summary: Provide support and enhancements for SAP FI and ABAP applications. Analyze and implement simple change requests (e.g., adding/changing a field, field validations, report updates, screen modifications). Perform basic ABAP development, testing, and transport management. Support incident resolution and user queries. Coordinate with onsite teams and business stakeholders. Preferred Skills: SAP FI functional knowledge (GL, AP, AR, Asset Accounting). Basic ABAP development and debugging. Experience with small enhancements and support activities. Knowledge of SAP interfaces/integrations. Preferred: Experience with S
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. We are seeking a highly motivated Senior Analyst to own and evolve a best-in-class tooling ecosystem that powers Roblox's Safety, Privacy, Trust & Safety, and Support Operations teams. In this role, you will sit at the intersection of operations and technology — driving systematic improvements, scaling tooling infrastructure, and ensuring agents across global vendor sites have the tools they need to operate efficiently and safely at Roblox's scale. You will join a growing team of administrators, incident managers, and product support specialists in India, providing comprehensive tooling and admin support to our global operations teams. You will help build a premier suite of agent tooling by managing systematic changes, identifying improvements, and delivering exceptional service to our internal customers. You will Administer and configure applications — including Zendesk, Absorb, JIRA, and other third-party and internally developed tools — to support 24/7 operations across multiple vendor sites. Design and implement queue and routing systems that reduce agent decision fatigue, enable quality support interactions, and support seamless multi-site operations. Analyze and translate b
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: BizTech fosters culture and connection at Airbnb by providing reliable corporate tools, innovative products, and technical support for all teams. We drive technical breakthroughs and strategies that redefine what it means to belong anywhere, delivering greater value for the business and our people. The Global Operations team at BizTech manages production services across Airbnb’s corporate environment, delivering reliable operations through Observability, Incident Management, Core Operations, and AI-enabled automation. We partner across BizTech to scale service quality, efficiency, and resilience. The Difference You Will Make: As an Operations Engineer, you'll apply AI at the forefront of BizTech's operational health: using LLM-powered triage and intelligent automation to resolve tickets, speed up incident response, and build self-healing observability that catches problems before they escalate. AI fluency is core to this role, not an add-on. You'll prototype agentic workflows, embed AI into runbooks and diagnostics, and continuously look for repetitive work automation can take over. Success looks like a shrinking backlog of recurring ticket categories through AI-assisted automation, faster MTTR powered by intelligent alerting and root-cause suggestions, and dashboards/reportin
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role The Observability, Monitoring, and Integrations team manages observability, monitoring, and detection for systems that support customer purchasing and use of GitLab. As a Staff Backend Engineer on the Fulfillment Workflow Monitoring (Catch All) team, you'll set the technical direction for the telemetry, detection, and reconciliation tooling that identifies billing, data, and event anomalies across CustomersDot, Salesforce, and Zuora before they can affect revenue or the customer experience. This is a greenfield team. You'll help build it from the ground up, shaping its operating rhythm, incident response
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity As a Senior Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with 5–8 years of relevant experience who can leverage open-source and custom tooling to orchestrate and maintain large-scale Kubernetes environments. You will play a "Captain" role—leading critical deliverables and mentoring junior engineers while maintaining the reliability of our global fleet. What You'll Do Architectural Leadership: Drive the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management. Platform Orchestration: Lead complex, large-scale infrastructure shifts. Operational Excellence: Take ownership of incident response, author comprehensive retrospectives, and implement systemic hardening to prevent recurrence using advanced overcommit strategies. This Role Requires Experience: 5–8 years in a DevOps, Site Reliability, or Infrastructure Engineering role. Kubernetes Mastery: Deep internals knowledge of Kubernetes and hands-on experience writing custom operators. Tooling Proficiency: Strong experience building production-grade tools and services, specifically for infrastructure automation. Operations-Heavy Mindset: A proven track record of Day 1/Day 2 operations for a large-scale Kubernetes fleet, handling high-severity incidents, and improving SLA compliance through auto
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. We are seeking a highly motivated Senior Analyst to own and evolve a best-in-class tooling ecosystem that powers Roblox's Safety, Privacy, Trust & Safety, and Support Operations teams. In this role, you will sit at the intersection of operations and technology — driving systematic improvements, scaling tooling infrastructure, and ensuring agents across global vendor sites have the tools they need to operate efficiently and safely at Roblox's scale. You will join a growing team of administrators, incident managers, and product support specialists in India, providing comprehensive tooling and admin support to our global operations teams. You will help build a premier suite of agent tooling by managing systematic changes, identifying improvements, and delivering exc
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. We are seeking a highly motivated Analyst to support and continuously improve the tooling ecosystem that powers Roblox's Safety, Privacy, Trust & Safety, and Support Operations teams. In this role, you will work at the intersection of operations and technology, helping scale tooling, improve operational workflows, and ensure agents across global vendor sites have the tools they need to operate efficiently. You will join a growing team of administrators, incident managers, and product support specialists in India, providing comprehensive tooling and admin support to our global operations teams. You will contribute to building and maintaining a robust suite of agent tooling by supporting systematic changes, identifying improvement opportunities, and delivering exceptional service to internal stakeholders. Work Schedule : This role is currently aligned to a 2:00 PM to 11:00 PM IST schedule. As part of supporting a 24×7 operational environment, working hours may change based on business requirements, and you should be flexible to work across rotational shifts, including weekends and holidays, as needed. YOU WILL Administer and configure applications – including Zendesk, Absorb, Jira,
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t
Other cities to consider
More places hiring for this role
Get new incident commander jobs in India by email
Daily job updates · Unsubscribe anytime