JOB TITLE IT Operations Engineer, Application Support A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO • Provide technical support for software applications and investigate, diagnose, and resolve application issues • Automate start-of-day and end-of-day checks for key applications. • Log and track incidents across applications in the production environment. • Implement monitoring and automation initiatives and develop custom solutions using Python, Shell, and/or Powershell scripts. • Create tactical support tools and scripts to improve the incident investigation process and enable • transparency into potential business impacts. • Prioritize and categorize incidents based on severity and impact. • Collaborate with the development team to improve applications based on user feedback. • Create and maintain documentation for responding to common errors and application incidents. • Assist with software applications deployment and configuration . • Provide training and assistance to users to ensure effective use of applications and systems. • Develop knowledge base resources to empower users to independently resolve common problems. WHAT’S REQUIRED • Bachelor's degree in computer science, information technology, or a related field. • Experience supporting middle- and back-office applications created in .Net, Java, C# etc. Ability to debug apps using of code, logs, alerts etc. • Literacy in complex SQL procedures/queries. • Ability to diagnose and troubleshoot technical issues.
Jobs in India
Incident Commander in Bengaluru
45 active opportunities · Updated October 2026
Showing
15 jobs
Explore current incident commander jobs in Bengaluru. Filter by work mode, employment type, experience, department, date posted and distance.
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . Role: Staff NOC Analyst (5 - 8 years) Location: Bangalore (24/7 Shift Environment) Role Summary We are looking for a Cloud Operations & Staff NOC Analyst who will act as the first line of operational defense for enterprise infrastructure, cloud platforms, and applications. This role requires strong real-time monitoring, incident response, and troubleshooting capabilities, along with a proactive mindset toward improving operational processes and reducing alert noise. Key Responsibilities Monitoring & Incident Management Monitor infrastructure, applications, and cloud platforms using tools such as New Relic, Datadog, Prometheus/Grafana, AWS CloudWatch, or GCP Monitoring Perform real-time alert triage, validation, and troubleshooting to restore services quickly Act as the first responder for incidents , ensuring minimal downtime and impact Identify false positives and reduce alert noise through analysis and tuning Incident Handling & Escalation Own and manage high-priority incidents (P1/P2), including: Driving incident bridges Coordinating with cross-functional
Role: DataOps Solution Architect — Azure Data Platform (Fabric & Databricks) Owns end-to-end architecture of the Azure data platform (Fabric, Databricks, Azure services) and sets the governance, security, and DataOps standards engineering teams implement. Design authority and primary architecture advisor to the client, not the day-to-day builder. Experience Required 8-10+ years overall in data engineering / cloud architecture, including 3+ years as a Solution Architect or in a comparable design-authority role. Microsoft Fabric: 2+ years hands-on. Azure Databricks: 4+ years hands-on. Proven track record designing and delivering enterprise-scale Azure data platforms end-to-end. Azure is mandatory ; experience across multiple client engagements or a consulting/SI background is a plus. Core Technical Expertise Azure Platform Fabric & Databricks architecture. ADF. ADLS. Entra ID. Key Vault. Monitor/Log Analytics. Microsoft Fabric Workspace & environment strategy. Lakehouse/Warehouse. OneLake. Pipelines/Notebooks. REST APIs/CLI. Governance. Azure Databricks Workspace strategy. Unity Catalog. RBAC. Jobs/Workflows. Delta Lake. Asset Bundles. ML platform integration. DevOps & IaC Azure DevOps/GitHub CI/CD. Terraform & Bicep. Branching/release strategy. Reusable IaC modules. Security & Networking Entra ID. RBAC architecture. Private Endpoints/DNS. Key Vault integration. DataOps Standards CI/CD & test gates for pipelines. Data-quality frameworks. Observability/alerting architecture. SLAs/SLOs. Incident-management practices. FinOps guardrails. Key Responsibilities Own end-to-end architecture; translate requirements into scalable, secure, cost-effective designs. Design enterprise-scale Fabric/Databricks architecture including workspace, networking, security, and governance. Set standards for CI/CD, IaC, DataOps, and environment management; ensure enterprise security and regulatory compliance. Evaluate technology options; identify architectural risks a
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. As the Director of Engineering at Smartsheet India, you will build capabilities to empower the world's largest companies to transform their approach to work. You will guide teams that own the grid ecosystem - defining how data linking, synchronization, and grid infrastructure evolve as a cohesive platform. You will ensure architectural decisions are coherent and avoid fragmentation. You will be willing to challenge technical choices. Platform Reliability & Operational Excellence: You will be accountable for the availability and performance of foundational services that other teams depend on. Drive a high bar for on-call health, incident response, and SLA/SLO definition across all the services. You will manage cross-pillar/cross-domain dependencies, negotiate API contracts, and prevent the grid ecosystem from becoming a delivery bottleneck. You will balance the needs of user-facing product features with infrastructural stability and operational health. You will ensure career growth paths are clear for engineers across that spectrum, and develop a strong sense of customer centricity and pillar identity for your teams. You will be comfortable accepting responsibility for impact, service availability, and the effectiveness of your teams. You will be comfortable being at the forefront of AI adoption for delivery and operations, leaning in and helping the team leverage AI for optimum delivery in their ways of working. You are passionate about continuous improvement and have built learning organizations that keep up w
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What you will be doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuous
NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services. What you’ll be doing: Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale. Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation. Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data. Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems. Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation. What we need to see: BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10+ years operating production infrastructure or platform services. Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges. <l
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput, and 99.999 availability. We're looking for a technical leader to help us to continue to scale the service with great people and reliable, cost-effective and efficient infrastructure, processes and tooling. As the Director of Site Reliability Engineering you will oversee the SRE organization focused on Okta platform, Databases, Edge networking, K8s platform, CI/CD, Observability, FinOps, and automation platform & tooling. Job Duties and Responsibilities: Build and lead a high-caliber India-based SRE organization supporting Okta’s production fleet. Partner with global engineering, product, and infrastructure leaders to deliver resilient, scalable, and secure services. Define and execute the India SRE strategy in alignment with global reliability goals. Lead post-incident reviews, drive root-cause analysis, and ensure long-term corrective actions. Participate in incident management, on-call rotations, and blameless RCAs. Implement automation and observability to reduce manual toil and improve operational efficiency. Drive adoption of modern infrastructure practices: infrastructure as code (Terraform), container orchestration (Kubernetes), and AI within Infrastructure org. H
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role The Observability, Monitoring, and Integrations team manages observability, monitoring, and detection for systems that support customer purchasing and use of GitLab. As a Staff Backend Engineer on the Fulfillment Workflow Monitoring (Catch All) team, you'll set the technical direction for the telemetry, detection, and reconciliation tooling that identifies billing, data, and event anomalies across CustomersDot, Salesforce, and Zuora before they can affect revenue or the customer experience. This is a greenfield team. You'll help build it from the ground up, shaping its operating rhythm, incident response
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity As a Senior Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with 5–8 years of relevant experience who can leverage open-source and custom tooling to orchestrate and maintain large-scale Kubernetes environments. You will play a "Captain" role—leading critical deliverables and mentoring junior engineers while maintaining the reliability of our global fleet. What You'll Do Architectural Leadership: Drive the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management. Platform Orchestration: Lead complex, large-scale infrastructure shifts. Operational Excellence: Take ownership of incident response, author comprehensive retrospectives, and implement systemic hardening to prevent recurrence using advanced overcommit strategies. This Role Requires Experience: 5–8 years in a DevOps, Site Reliability, or Infrastructure Engineering role. Kubernetes Mastery: Deep internals knowledge of Kubernetes and hands-on experience writing custom operators. Tooling Proficiency: Strong experience building production-grade tools and services, specifically for infrastructure automation. Operations-Heavy Mindset: A proven track record of Day 1/Day 2 operations for a large-scale Kubernetes fleet, handling high-severity incidents, and improving SLA compliance through auto
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies — from the world's largest enterprises to the most ambitious startups — use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the Organization The Core Infrastructure organization operates the foundational systems that power Stripe globally — including databases (MongoDB, PostgreSQL), high availability and disaster recovery (HADR), AWS cloud infrastructure, Linux servers, container orchestration, mesh networking, service discovery, and network edge infrastructure. Within Core Infra, the Regional Enablement Platform (REP) team helps Stripe launch and operate new regions without learning about broken dependencies from users. REP builds the regionalization, validation, deploy-safety, and operator tooling needed to answer practical launch-readiness questions: can critical payment paths run from the new region, which services still depend on a remote control plane, what breaks under packet loss or failover, and what must be fixed before deploys, launches, traffic shifts, or failovers proceed. The team uses traffic replay, synthetics, failover drills, dependency analysis, CI/CD gates, and incident data to turn those findings into platform fixes, service-owner asks, and reusable readiness checks across networking, HADR, and service teams. This role is based in Bangalore and serves as a senior technical anchor for Core Infrastructure in India, with direct cross-region influence across AMER, EU, and APAC. What you'll do As a Staff Engineer on REP, you will play a key leadership role in enabling Stripe's infrastructure to power all of our products, globally and at scale. You will
Other cities to consider
More places hiring for this role
Get new incident commander jobs in Bengaluru, India by email
Daily job updates · Unsubscribe anytime