SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . Role: Staff NOC Analyst (5 - 8 years) Location: Bangalore (24/7 Shift Environment) Role Summary We are looking for a Cloud Operations & Staff NOC Analyst who will act as the first line of operational defense for enterprise infrastructure, cloud platforms, and applications. This role requires strong real-time monitoring, incident response, and troubleshooting capabilities, along with a proactive mindset toward improving operational processes and reducing alert noise. Key Responsibilities Monitoring & Incident Management Monitor infrastructure, applications, and cloud platforms using tools such as New Relic, Datadog, Prometheus/Grafana, AWS CloudWatch, or GCP Monitoring Perform real-time alert triage, validation, and troubleshooting to restore services quickly Act as the first responder for incidents , ensuring minimal downtime and impact Identify false positives and reduce alert noise through analysis and tuning Incident Handling & Escalation Own and manage high-priority incidents (P1/P2), including: Driving incident bridges Coordinating with cross-functional
Jobs in India
Incident Response Analyst in Bengaluru
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current incident response analyst jobs in Bengaluru. Filter by work mode, employment type, experience, department, date posted and distance.
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity AI is a top strategic priority for New Relic, and the Bengaluru design team is at the center of it. The team works across three closely related product clusters: autonomous incident response (SRE Agent, Autopilot, and Intelligent RCA), the intelligence and platform layer that powers them (Ground Truth, Agentic Platform, and New Relic AI), and AIOps for event correlation and incident management. These products share a common design challenge: users need to trust systems that act autonomously, and building that trust through good design is genuinely hard work. This is an on-site role in Bengaluru. Your designers are there, and many of your engineering and product partners are too. Being present — in standups, reviews, and the quick conversations before a decision gets made — is part of how you'll build the relationships that make design effective. You'll also collaborate with design, product, and engineering partners in the US and Spain, so operating across time zones and communicating well in writing are part of the job. You'll manage a small team of designers and own the quality of the work across these products. The design questions here don't have established answers — how do you make an autonomous system legible? How do you build user trust in AI-generated root cause analysis? How do you design a handoff from machine decision to human judgment? If you're already working in AI product design, or actively building toward it, and you want to lead a team
Who are we? FalconX is a pioneering team of operators, investors, and builders committed to revolutionizing institutional access to the crypto markets. Operating at the intersection of traditional finance and cutting-edge technology, FalconX addresses the industry's foremost challenges: Navigating the digital asset market can be complex and fragmented, with limited products and services that support trading strategies, structures, and liquidity found in conventional financial markets. As a comprehensive solution for all digital asset strategies from start to scale, FalconX operates as the connective tissue empowering clients with seamless navigation through the ever- evolving cryptocurrency landscape. Responsibilities Be part of a trading systems engineering team, dedicated to building out the core trading platforms. Work closely with cross functional teams to improve the system reliability, scalability and security. Engage in and improve the quality supporting the platform. Build and manage systems, infrastructure and applications through automation. Provide operational support to internal teams working on the platform. Work on improvements to bring in high efficiency, reduce latency, deploy systems faster. Practice sustainable incident response and blameless postmortems. Together with your engineering team, you will share an on-call rotation and be an escalation contact for service incidents. Implement and maintain rigorous security best practices across all infrastructure, with a focus on minimizing attack surface and ensuring data integrity. Monitor system health and performance with a keen eye for identifying and resolving issues before they affect trading activity. Manage user queries and service requests (often requiring in depth analysis of the technical and/or business logic of our systems). Proactive approach to problem analysis and resolution of production incidents. Manage Issue tracking and prioritisation of day to day production incidents. Manage platf
A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO Monitor, triage, and assess incidents across global markets and technology platforms, validating severity, business impact, and escalation requirements. Act as command-and-control lead for major incidents, driving structured response, rapid restoration, and informed decision-making. Coordinate recovery efforts across L1/L2/L3 support teams, engineering, infrastructure, and third-party vendors to minimize business disruption. Lead internal and external stakeholder communications, including senior leadership, during incidents and service disruptions. Trigger and govern disaster recovery (DR) failover in line with predefined criteria, controls, and governance standards. Ensure accurate, timely, and audit-ready incident documentation, including detailed timelines, evidence, and impact analysis. Facilitate post-incident reviews (RCA/post-mortems), ensure root causes are clearly identified, and track corrective and preventative actions to closure. Govern end-to-end change management, including change risk assessment, CAB facilitation, approval workflows, maintenance windows, and change-related incident management. Drive continuous improvement in incident, change, and event management processes through automation, standardization, and improved tooling and alert quality. Partner with engineering and operations teams to design scalable, resilient processes, promote best practices, and enhance operational matur
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As a hands-on senior leader for our India SecOps team, you will shape and safeguard Everpure’s security posture at the intersection of detection engineering, threat hunting, attack surface management, and incident response. Positioned as a strategic cornerstone in Bangalore, you will empower an elite engineering team, optimize critical SecOps pipelines, and partner cross-functionally across global engineering and infrastructure groups. By driving execution excellence and high team morale, you ensure our enterprise platform and global telemetry remain resilient against evolving threats. WHAT YOU'LL DO Scale & Lead SecOps Operations: Architect, mentor, and grow the India SecOps team to foster an environment of high morale, technical excellence, and rapid execution across detection engineering and incident response. Proactively Manage & Remediate Attack Surface: Own end-to-end Attack Surface Management (ASM) across cloud environments, SaaS applications, endpoints, and secrets management to measurably minimize enterprise exposure and mitigate risk. Optimize Telemetry & Incident Response: Mature SIEM and SOAR automation pipelines to drastically reduce mean time to detect, contain, and respond (MTTD/MTTC/MTTR) while continuously elevating alert fidelity and signal confidence. Drive Cross-Functional Alignment & RCA Postmortems: Lead continuous validation through purple-teaming and incident postmortems alo
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Exa team and lead the charge in redefining enterprise storage by unifying block, file, and object protocols across hybrid-cloud environments. You will combine deep technical expertise in distributed systems with hands-on people leadership to guide architectural decisions and mentor high-impact engineers. This is a unique opportunity to build new engineering teams from the ground up and drive industry-leading innovation alongside Product and Architecture partners. Your work will directly impact how customers consume, scale, and operate mission-critical storage infrastructure. WHAT YOU'LL DO Establish and scale a new engineering organization focused on critical Exa platform services, ensuring a foundation of long-term success, technical excellence, and a high-performing culture. Own the successful delivery of complex, high-scale engineering features for the Exa platform, ensuring world-class security, reliability, and availability across multi-array and hybrid-cloud deployments. Define the technical vision and execution roadmap in close partnership with Product Management and Architecture, translating customer needs into a measurable business impact for Pure Storage. Drive a culture of operational rigor, owning the refinement of engineering processes around observability, CI/CD, and incident response, while actively mentoring the next generation of technical leads. WHAT YOU BRING Leadership and Scaling: Pro
DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. About the Role We're seeking an experienced DevOps/ Site Reliability Engineering (SRE) Engineer to join DataHub and drive the reliability, scalability, and operational excellence of our platform offerings. In this role, you'll work on technical initiatives across DataHub Cloud and our emerging enterprise deployment solution, which provides customers with enhanced control and flexibility for running DataHub in their preferred environments. Key Responsibilities Enterprise Platform Development: Partner with product and engineering teams to influence the development of advanced deployment capabilities. Collaborate with cross-functional teams to help build systems for seamless installation, upgrade, and rollback processes across various environments. Influence the design and help implement comprehensive monitoring and health check systems for distributed deployments. Partner with engineering teams to help develop self-healing and automated remediation capabilities. Platform Reliability and Operations: Establish and maintain SLAs/SLOs for both cloud and enterprise offerings. Lead incident response and post-mortem processes to drive continuous improvement. Optimise system performance, capacity planning, and cost efficiency. Work closely with product, engineerin
Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. Why is this a critical role at OneTrust? What is the challenge / type of challenges someone in this role will have the opportunity to take on? What impact does someone in this role have the opportunity to make within the company, for our customers and on a larger scale in the evolution of privacy and trust? An awareness of current issues affecting the industry and its technologies Create innovative, scalable, fault-tolerant software solutions for our clients and customer base Expand existing software to meet the changing needs of our key demographics What does this person do each day/each week? Describe a true to life day in the life for someone in this role. What do they do and how do they do it? Goal is to paint a real, genuine picture so candidates can see themselves in the role. Support production customers by monitoring and maintaining our cloud application & cloud infrastructure hosting it Build scripts for operational automation and incident response Handle processes surrounding cloud application deployment for our agile release Work with the monitoring, tuning, maintenanc
AI/ML – Investment Services A Career with Point72's AI/ML – Investment Services Team The AI/ML – Investment Services team at Point72 spearheads the development of cutting-edge AI solutions that seek to transform our business processes and enhance enterprise intelligence. The team aims to bridge the gap between business challenges and technological innovation, collaborating with stakeholders across the firm and leveraging expertise in generative AI, data engineering, and machine learning. WHAT YOU'LL DO Build and scale core backend services and platforms that power generative AI applications and data infrastructure used across the firm’s investment workflows Design and implement high-throughput, low-latency data pipelines to ingest, normalize, and serve both structured and unstructured data Develop robust APIs and microservices to support model inference, feature serving, and downstream applications Integrate generative AI tools and model-serving workflows into production, including embedding stores, retrieval components, and fine-tuning pipelines Optimize system performance, cost, and reliability through profiling, capacity planning, and architectural improvements Implement automated testing, continuous delivery pipelines, monitoring, and incident response practices to maintain production health Partner with data scientists, AI engineers, product owners, and operations to translate models and prototypes into scalable, production-grade solutions Mentor engineers, lead code reviews, and establish engineering best practices for maintainability, security, and observability Own end-to-end delivery, operational runbooks, and metrics-driven measurement of feature impact and system reliability WHAT'S REQUIRED Bachelor’s degree in computer science, software engineering, or a related technical field Minimum 5+ years of professional experience building backend systems and production services Demonstrated experience designing and operating large-scale data engineering pipelines
JOB TITLE Observability Engineer A CAREER WITH POINT72'S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. WHAT YOU’LL DO Design observability capabilities that give engineering teams clear insight into application health, platform performance, and issues affecting users Build scalable collection pipelines for metrics, logs, and traces across cloud-based and on-premises environments Develop actionable alerting standards that reduce noise, shorten incident response, and highlight the most important signals Partner with application and infrastructure teams to define service health indicators and improve operational readiness before production launches Automate monitoring configuration, dashboard deployment, and reliability checks to support consistent observability across the technology environment Analyze production incidents to identify telemetry gaps and improve detection, diagnosis, and recovery Create dashboards and reporting views that help teams understand trends, capacity risks, and reliability outcomes Establish practical observability
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. As the Director of Engineering at Smartsheet India, you will build capabilities to empower the world's largest companies to transform their approach to work. You will guide teams that own the grid ecosystem - defining how data linking, synchronization, and grid infrastructure evolve as a cohesive platform. You will ensure architectural decisions are coherent and avoid fragmentation. You will be willing to challenge technical choices. Platform Reliability & Operational Excellence: You will be accountable for the availability and performance of foundational services that other teams depend on. Drive a high bar for on-call health, incident response, and SLA/SLO definition across all the services. You will manage cross-pillar/cross-domain dependencies, negotiate API contracts, and prevent the grid ecosystem from becoming a delivery bottleneck. You will balance the needs of user-facing product features with infrastructural stability and operational health. You will ensure career growth paths are clear for engineers across that spectrum, and develop a strong sense of customer centricity and pillar identity for your teams. You will be comfortable accepting responsibility for impact, service availability, and the effectiveness of your teams. You will be comfortable being at the forefront of AI adoption for delivery and operations, leaning in and helping the team leverage AI for optimum delivery in their ways of working. You are passionate about continuous improvement and have built learning organizations that keep up w
Toradex is a global company strongly focused on engineering & technology. We’re powered by a diverse & uniquely gifted workforce. We pursue the best people to propel our innovative vision of embedded computing and IoT. If you’re interested in being a driving force at an agile technology company, engineering clever computing solutions & helping other companies bring their products to life, we should talk. Description We are looking for a DevOps Engineer to strengthen our cloud operations and engineering practices, with a focus on reliable website delivery, secure AWS foundations, and fast but controlled delivery of new services. The position combines AWS operations, infrastructure as code, CI/CD, automation, and pragmatic software engineering. The person should be confident working with services for edge delivery, compute, storage, databases, DNS, security, and observability without relying on manual console changes as the default operating model. The role also supports on-premises to cloud migration, global service optimization, and practical responses to increasing AI-driven traffic. We value candidates who can use modern AI-assisted development effectively to spin up proof-of-concept projects quickly, while still applying disciplined Git, review, security, and deployment practices. About you You enjoy building stable, secure, and maintainable infrastructure that supports business-critical services. You can work independently and take ownership of cloud environments, deployments, and operational improvements. You are comfortable balancing speed, reliability, cost, and security when making technical decisions. You communicate clearly with technical and non-technical stakeholders and explain trade-offs in a practical way. You document your work well and create clear runbooks and support material for future maintenance. You are methodical when troubleshooting incidents and stay calm when systems are under pressure. You are curious about modern traffic patt
About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As the Escalation Manager, Weekend Global Coverage, you will serve as the crucial command lead safeguarding customer trust during high-stakes, critical events across AMER, APJ, and EMEA . Operating within our global follow-the-sun coverage model, you are the decisive voice when a Sev1 or critical customer event occurs, rapidly uniting engineering, support, and executive leadership to stabilize complex enterprise scenarios. This high-impact role sets the standard for weekend operational excellence, turning chaotic technical crises into clear, swift paths to resolution and ensuring seamless cross-regional transitions . WHAT YOU'LL DO Command Critical Escalations: Serve as the accountable incident commander for major customer escalations (CIEs/Sev1s) during the weekend US coverage window, establishing clear decision rights, working cadences, and swift mitigation strategies . Unify Cross-Functional Teams: Mobilize and align Support, Escalations Engineering, Service Delivery, and Account teams to ensure technical responders have immediate context and unblocked paths to restore customer environments . Lead Executive & Customer Communications: Translate fast-moving, complex technical diagnostic details into precise, highly coherent updates for executive leadership and global enterprise customers, keeping stakeholders continuously aligned . Drive Global Follow-the-Sun Handovers: Execute complete, high-quality shift ha
Other cities to consider
More places hiring for this role
Get new incident response analyst jobs in Bengaluru, India by email
Daily job updates · Unsubscribe anytime