At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Customer Experience Engineering Team builds the internal and external technologies that scale Snowflake’s global support and sales organizations. We empower our technical experts by providing the advanced tools they need to resolve complex issues and drive customer success. Our team specializes in software engineering, data-driven decisions, ML, and LLM-based solutions . We build production-grade systems to automate manual processes and augment the capabilities of our technical staff. Our current focus includes: LLMs : Developing and deploying LLM and agent-based architectures for streamlining troubleshooting Scalable Evaluations : Implementing large-scale evaluations to ensure the quality and reliability of our internal and external tools Process Automation : Designing intelligent workflows that eliminate bottlenecks and allow our experts to focus on the most technical aspects of the Snowflake platform Incident discovery: using embeddings, LLMs, clustering, and agents to detect potential widespread issues more quickly Now, the team is growing, and we are looking for a Software Engineer to join us. In this role, you will work closely with the state of the art LLM models, fine-tune them, develop agents, apply various clusterings, summarizations, embeddings, and so on. Ev
Jobiba hiring network
Senior Incident Response Analyst React Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior incident response analyst react jobs. Use filters to narrow by work mode, employment type, experience and date posted.
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: The Technology Risk & Controls Analyst operates as part of the second line Technology Risk & Controls function, working under the direction of the Technology Risk & Controls Manager, Tech Ops. The role supports the delivery of independent assurance over Technology Operations control environments, including infrastructure, cloud platforms, identity and access management, service management and core operational processes. The role is focused on executing high-quality control testing, audit support and remediation assurance activities in line with WPP standards and methodologies. The Analyst works closely with Technology Operations teams, Financial Risk & Control colleagues and auditors to evidence control performance, support audit readiness (including SOX 404) and contribute to a consistent, well-documented and sustainable control environment. What you'll be doing: • Perform control design and operating effectiveness testing across Technology Operations processes, including change management, access management, incident/problem management, backup and resil
Position : Senior ServiceNow Developer Exp Level: 7+ years of experience Location: Hyderabad / Visakhapatnam Shift Timings: 2:00PM to 11:00PM IST Skills: #ServiceNow, #Rest & Soap API's, #Javascript Key Responsibilities: Design, develop, and implement complex ServiceNow solutions using Flow Designer, Business Rules, Client Scripts, UI Policies, Script Includes, Integrations, and custom applications. Lead technical design for enhancements across modules such as ITSM, CMDB, Asset or others as required. Develop and maintain #ServiceNow catalog items, workflows, record producers, and custom applications following platform best practices. Integrations Build and support integrations with external systems using REST/SOAP APIs, MID Server, IntegrationHub, web services, and authentication methods like OAuth and SAML. Troubleshoot integration failures and optimize performance. Required Qualifications: 5+ years of hands-on ServiceNow development experience. Strong understanding of JavaScript, AngularJS, Glide API, and general web technologies. Expertise in multiple ServiceNow modules like ITSM, ITOM, HRSD, CSM, SecOps. (ITSM is required; others a plus). Experience with Service Portal, Workspace, and UI Builder. Roll : service Location : Visakhapatnam and Hyderabad Employment Type : Full-Time Experience : 5+ Years Skiils : SAP FI,SAP ABAP Job Summary: Provide support and enhancements for SAP FI and ABAP applications. Analyze and implement simple change requests (e.g., adding/changing a field, field validations, report updates, screen modifications). Perform basic ABAP development, testing, and transport management. Support incident resolution and user queries. Coordinate with onsite teams and business stakeholders. Preferred Skills: SAP FI functional knowledge (GL, AP, AR, Asset Accounting). Basic ABAP development and debugging. Experience with small enhancements and support activities. Knowledge of SAP interfaces/integrations. Preferred: Experience with S
PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. About the role PagerDuty’s Operations Cloud runs on a platform that ingests billions of signals and turns them into real-time action for thousands of customers. We’re looking for an early-career AI/ML Engineer who is excited to grow at the intersection of two disciplines: large-scale distributed systems and machine learning. In this role you will help build and ship AI systems that run in production at PagerDuty’s scale — powering Incident Management AI Agents, event intelligence, and the LLM-powered capabilities embedded across our platform. You’ll work alongside senior engineers on real production problems, learning how AI features go from a prototype to something that serves reliably at scale. We are looking for a candidate who is genuinely excited about building with modern AI — LLMs, agents, and retrieval — eager to learn how resilient, high-throughput systems are built, and motivated to grow into an engineer who is strong in both. What you’ll do Contribute to AI-powered features — LLM agents, retrieval, and event intelligence — that operate on high-volume, real-time data, with support and guidanc
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As the Logistics Security Program Manager for EMEA, this role is an essential part of Micron’s Global Security team. The person leads the management and continuous refinement of the company’s important logistics security efforts across the EMEA area. This position serves as the regional authority, advancing risk-focused security methods that protect high-value shipments, improve supply chain durability, and lower transportation security risks in intricate multimodal logistics networks. As a senior individual contributor, the Logistics Security Program Manager takes charge of regional program initiatives on their own. They apply solid judgment to shifting threat conditions and collaborate with colleagues across functions and external partners to produce security results. The position involves balancing security, operational efficiency, and business continuity while transforming regional risks into scalable, practical controls that advance cargo visibility, shipment protection, and incident readiness. Responsibilities: Act as the EMEA logistics security authority, guiding the creation and implementation of risk-focused security programs for valuable and sensitive shipments involving carriers, freight forwarders, and logistics providers. Develop, apply, and manage shipment security controls, including tracking, telematics, geofencing, chain of custody, tamper detection, monitoring, critical issue handling, recovery processes, and carrier compliance requirements. Conduct carrier, route, l
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Safety is fundamental to the Lyft experience and to the trust riders, drivers, and our broader communities place in our platform. Lyft’s Safety and Customer Care team works through product, technology, data, policy, and operations to prevent harm before it happens, respond compassionately when it does, and continuously learn from every incident in efforts to make Lyft the safest way to get around. We are looking for an experienced safety leader with deep Trust & Safety experience to help turn Lyft’s Trust & Safety strategy into impact at scale, with the final shape of this role informed by the strengths of the leader we bring on. This is a highly visible leadership role within Lyft’s Safety & Customer Care organization. You will report to the Senior Director and General Manager of Trust & Safety and sit on the Trust & Safety leadership team. You will additionally partner deeply with Product, Engineering, Data Science, Research & Design, Legal, Compliance, Risk, Communications, Finance, and other teams across Lyft. The ideal candidate is an exceptional operator and people leader with Trust & Safety expertise, strong strategic judgment, and a track record of leading complex, high-stakes Trust & Safety organizations at scale. You can move fluidly between setting strategy with senior executives, developing leaders, overseeing large global operations and budgets, and diving into an emerging safety issue when the situation demands it. Responsibilities Lead Global Safety Compliance & Risk Oversee Lyft’s global safety compliance, transparency, and risk functions, supporting strong functional leaders and individual contributors on your team and strengthening capabilities as Lyft scales globally. Set direction across compliance readiness, safety transparency repo
Job description • You will perform the day-to-day data centre/computer operations (such as running the batch programs, printing of reports, systems and data backup, tapes/cartridges preparation and storage, etc). • You will be required to pro-actively monitor the data centre systems' uptime and performance (such as Server, Midrange and/or Mainframe systems, Networking equipment, Telecommunication connections, Email and Security Systems, Data Centre environment, etc) to prevent any down time or low performance. • You will diagnose and log down hardware problems into incident ticketing system, coordinate the problem resolution with the respective vendor or second level support personnel. • You will liaise with the vendors and other technical teams on any data centre related maintenance and installation activities. • You will also assist to provide helpdesk and basic IT troubleshooting support. • You will also assist in other IT related projects implementations. Job Requirements: • "A" Level Holders, ITE Graduates, Diploma and Degree Holders are welcome to apply. Those with at least 2 years of working experience will be considered for senior roles. • Experience in operating Windows Server, UNIX and/or AS/400 systems. • Good in work prioritization and able to maintain the service levels. • A detailed and systematic person, following the data centre operating procedures. • Willing to perform 12-hour Shift Work.
We’re looking for a Director, Customer Support Escalation Center to help us develop and execute on a comprehensive, customer-first strategy and approach to our Global Hootsuite’s Customer Support Escalation Center. Reporting to the Vice President, Customer Support, you’ll plan, organize, monitor, track and report on the service levels delivered by the escalation center. In this role, you will be accountable for continuous service improvements to the escalation center operational framework and improve Hootsuite’s ability to effectively service our customers. In this role, you will partner with a wide range of stakeholders, including senior leaders to deliver an exceptional customer support experience. In line with Hootsuite's distributed workforce strategy, our flexible work arrangement allows for a hybrid model. This role is open to applicants located in Bucharest, Romania. In this role, you will report to the VP, Customer Support. Please note: this is a contract role until August 2027. WHAT YOU’LL DO: Develop strategies, enhance performance, and foster alignment of objectives within the Customer Support Escalation Center. This involves overseeing the global bug and incident management process, advanced technical and 24x7 on call Business Critical Support process.Establish and operationalize high levels of customer service standards. Develop and implement systems, processes and people to consistently deliver services that are tied to organizational strategies and KPI’s. Drive continuous improvement through anticipating growth, identifying gaps, and offering proactive guidance to minimize potential risks to our customers. Build robust cross-functional connections among the Customer Office, Revenue, Product, and Development leadership to tackle intricate escalated bugs and incidents, eliminate customer pain points, and suggest methods for alleviating customer experience issues and enhancing outcomes via insights from escalations and root
About THG Ingenuity THG Ingenuity is a fully integrated digital commerce ecosystem, designed to power brands without limits. Our global end-to-end tech platform is comprised of three products: THG Commerce, THG Studios, THG Fulfilment. Each represents a single, unified solution, overcoming challenges and taking brands direct-to-consumer. Our client portfolio includes globally recognised brands such as Coca-Cola, Nestle, Elemis, Homebase, and Proctor & Gamble. Database Platform Manager Company: THG Ingenuity Location: Manchester (Head Office) Reports to: Director of Data Role Overview We are looking for a strong technical leader to run our database platform. Our estate runs on Google Cloud Platform following a recent migration to self-managed services, and there is a real opportunity here to shape the next phase. How we consolidate and modernise the estate, where managed and cloud-native services earn their place, which new technologies are worth adopting, and how the platform scales with the business are all live questions. You will lead the thinking on them and work with stakeholders across the business to agree and deliver the roadmap. You will lead our DBA team, who own every database across the group regardless of the application running on it and who run the platform as a 24/7 service. You will work hand in hand with our Data Reliability Engineering team, whose Principal Engineer is your peer and whose focus is automation, fleet reliability and SLOs. This role is weighted toward hands-on technical depth. You should be as credible in a design review or an incident bridge as you are in a planning session with senior stakeholders. Key Responsibilities Technical direction The technical roadmap for the estate. We want someone genuinely interested in emerging database technologies who evaluates them on merit and can articulate the pros and c
We're looking for an Engineering Manager II to own and grow the Observability Pipelines engineering org at a pivotal moment in the product's lifecycle. Observability Pipelines is Datadog's on-premise, vendor-agnostic telemetry pipeline product, with a lot still to build as it grows and scales. It sits at the center of a fast-consolidating market, is central to Datadog's data pipeline optimization story for Logs and Metrics customers. This is a build-and-scale opportunity: you'll grow the management and technical leadership layers, co-own the roadmap with Product, and define how this org operates as it continues to expand. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You'll Do: Directly manage the OP org including EM1s across NYC and Paris, set technical direction, and be the connective tissue across a distributed team Build out the management and technical leadership layers as the org continues to grow - today ~20 ICs Partner directly with Product to co-own the roadmap and strategy, helping decide where OP’s engineering investment goes next Set and evolve the operating rhythm across the group: planning cadence, on-call and incident standards, and cross-team alignment Own key cross-org relationships with the SaaS Logs Pipelines team, the BYOC team, and the Vector open-source community Coach managers and senior engineers, and build the succession and growth plans that let the org scale beyond you Who You Are: Experienced managing managers across distributed teams, with a track record of raising the bar on how those teams operate, not just delivering through them Back
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies — from the world's largest enterprises to the most ambitious startups — use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the Organization The Core Infrastructure organization operates the foundational systems that power Stripe globally — including databases (MongoDB, PostgreSQL), high availability and disaster recovery (HADR), AWS cloud infrastructure, Linux servers, container orchestration, mesh networking, service discovery, and network edge infrastructure. Within Core Infra, the Regional Enablement Platform (REP) team helps Stripe launch and operate new regions without learning about broken dependencies from users. REP builds the regionalization, validation, deploy-safety, and operator tooling needed to answer practical launch-readiness questions: can critical payment paths run from the new region, which services still depend on a remote control plane, what breaks under packet loss or failover, and what must be fixed before deploys, launches, traffic shifts, or failovers proceed. The team uses traffic replay, synthetics, failover drills, dependency analysis, CI/CD gates, and incident data to turn those findings into platform fixes, service-owner asks, and reusable readiness checks across networking, HADR, and service teams. This role is based in Bangalore and serves as a senior technical anchor for Core Infrastructure in India, with direct cross-region influence across AMER, EU, and APAC. What you'll do As a Staff Engineer on REP, you will play a key leadership role in enabling Stripe's infrastructure to power all of our products, globally and at scale. You will
About the Team The Applied team brings OpenAI’s technology to the world through products used by hundreds of millions of people and by developers and businesses building on our APIs. We work across research, engineering, product, policy, safety, and operations to deploy frontier AI systems responsibly and safely. The Trust & Safety Data Engineering team builds the data foundations that help OpenAI understand, detect, investigate, and mitigate abuse and safety risks across our products. We partner with Integrity, Investigations, Safety Systems, Product Policy, Privacy, Data Science, Engineering, and Data Platform to create reliable, privacy-safe datasets and pipelines for fraud and abuse detection, enforcement workflows, safety measurement, ML feature generation, launch readiness, and transparency reporting. About the Role We are hiring a Technical Lead Manager to lead and grow the Trust & Safety Data Engineering team. This is a hands-on leadership role for someone who can set strategy, shape data architecture, align senior stakeholders, coach engineers, and drive execution on high-impact data systems. You will help turn fragmented launch and incident support into durable, reusable, privacy-safe data foundations that Trust & Safety teams can rely on. The systems your team builds will help OpenAI detect risk, investigate abuse, power operational workflows, develop and evaluate safety models, measure interventions, support product launches, and report accurately on platform integrity. In This Role, You Will Lead and grow a high-performing Trust & Safety Data Engineering team. Define the roadmap and technical strategy for Trust & Safety data systems. Build canonical, privacy-safe datasets and pipelines for abuse detection, fraud detection, risk signals, enforcement, scaled review, transparency reporting, and safety monitoring. Create reusable foundations for Trust & Safety model development, including features, labels, training data, backtesting,
About the Role Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. You’ll work across two areas: Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack. Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve. We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure. This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation , solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure. responsible for delivering the software but also for operating and supporting it in production. Why this Role You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective. You’ll have the opportunity to build foundational systems from the ground up, work on infrastructur
Firefighter is responsible for responding to aircraft and structural fire emergencies, medical emergencies, hazardous material incidents, and other emergencies within the airport premises. This role involves ensuring the safety and security of passengers, crew, and airport personnel by performing firefighting, rescue and emergency medical services. Source: Adani Group | Job ID: 48510
We are seeking a highly skilled and experienced Staff Network Site Reliability Engineer (SRE) to join our Enterprise Network Operations and SRE team. In this role, you will be pivotal in implementing our vision for a reliable and efficient network infrastructure. The ideal candidate is passionate about network operations and committed to enhancing the user experience. You'll have the opportunity to solve complex network challenges using hands-on debugging and by focusing on network automation, observability, documentation, and operational excellence. This is a critical position focused on ensuring user satisfaction and brilliance in network operations. What you'll be doing: Owning the operational aspect of the network infrastructure, ensuring its high availability and reliability, actively working on network incidents and service requests. Partnering with architecture and deployment teams to guarantee that new implementations are supportable and align with production standards. Advocating for and implementing automation to reduce toil and improve operational efficiency. Minimizing manual operational tasks to achieve and maintain Service Level Objectives (SLOs). Monitoring network performance, identifying areas for improvement, and collaborating with relevant teams to implement refinements. Proactively identifying and mitigating network risks to promote continuous improvement. Collaborating with domain experts across functions to resolve production issues swiftly and effectively, ensuring customer happiness. Conducting blameless postmortems and following through on Root Cause Analyses (RCAs). Discovering opportunities for operational improvements and teaming up with colleagues to devise solutions that enhance excellence and sustainability in network operations. Developing knowledge base articles for automa
Get new senior incident response analyst react jobs by email
Daily job updates · Unsubscribe anytime