Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Incident Commander About the job This position is needed to help own and strategically evolve Twilio's incident response function as one of the company's most senior Incident Commanders, working alongside other senior ICs on the team. You will be capable of facilitating incidents across the full range of severities, including our highest-severity incidents. You will orchestrate response across a wide range of business functions, not just engineering. You'll be one of the senior voices VPs and C-suite leaders turn to for a clear, accurate read on what's happening and what's next. You approach every incident with genuine curiosity — pushing past the first plausible explanation to find what's really going on — and bring that same curiosity to the retrospective once the dust settles. Beyond the incident room, you'll help own the 12-month incident response process and tooling strategy for the company: setting company-w
Jobiba hiring network
Incident Commander Jobs
589 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current incident commander jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . Who we are At Twilio, we’re shaping the future of communications
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Abuse Operations is the front-line incident response and remediation function handling active product abuse and fraud impacting Stripe and its merchants. This multi-disciplinary group, spanning Incident Managers, Investigators, Forward Deployed Security Engineers, and Data Scientists, neutralizes active attacks, gathers requirements for operational tooling, and leads incidents. The team works directly with impacted merchants to resolve incidents and policy abuse rapidly. Operating primarily across Eastern, Pacific and Western European time zones, these team members regularly coordinate with global stakeholders across the world. What you’ll do In this role, you will play a critical part in safeguarding our financial ecosystem by investigating high-risk accounts, identifying complex fraud patterns, performing post-incident analyses, and driving cross-functional improvements to scale fraud detection. Building on these core operational duties, you will leverage your fraud, abuse, or product trust experience to improve incident response capabilities across Stripe by managing the entire fraud and abuse incident response process, developing response plans, leading workstreams, and serving as incident commander to ensure timely resolution. Furthermore, you will conduct gamedays to pressure-test response processes, drive proactive improvements, and help automate response workflows using agentic approaches ensuring we neutralize threats with speed
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: The Senior Security Incident Responder is a lead technical authority for incident response execution, responsible for handling the most complex, high-impact, and business-critical security incidents across WPP. The role does not have line management responsibility; people management remains with the Security Incident Management Lead. What you'll be doing: KEY RESPONSIBILITIES Advanced Incident Detection, Analysis & Response - Lead investigations for high-severity and complex security incidents. - Perform deep technical analysis using SIEM, SOAR, EDR/XDR, identity, email, and cloud telemetry. - Execute and oversee containment, eradication, and recovery actions. - Act as technical incident commander when delegated. Escalation Handling & Stakeholder Coordination - Serve as the primary escalation point for complex incidents. - Coordinate with Legal, Privacy, Risk, Technology Operations, and agency teams. - Provide clear technical updates to senior stakeholders. Forensics, Evidence Handling & Assurance - Lead forensic evidence collection, preservation, and analysis. - Ensure d
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a founding member of the Security Operations team in EMEA, you will join us at an exciting time in Roblox’s SIRT & SOC program. You will design and build the detections, automation and tooling that let a lean team monitor and protect players, developers, employees and the platform at global scale and serve as security incident commander in the region. This is a highly autonomous role where you will be a primary decision-maker, core to our mission to maintain a highly capable 24/7/365 monitoring and response capability. While you will work in close collaboration with peers at our US West Coast Headquarters, the time difference requires an engineer who can operate independently, making critical decisions without immediate oversight. We favor engineering our way out of toil through automation, orchestration, detections-as-code, and risk-based prioritization, while retaining the deep technical skills required to conduct detailed, hands-on analysis and lead response end-to-end when the situation warrants. Work Environment: This role is based in London, UK. You will be working from a dedicated, private space located within a shared office environment, designed to enable collaboration
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As the Escalation Manager, Weekend Global Coverage, you will serve as the crucial command lead safeguarding customer trust during high-stakes, critical events across AMER, APJ, and EMEA . Operating within our global follow-the-sun coverage model, you are the decisive voice when a Sev1 or critical customer event occurs, rapidly uniting engineering, support, and executive leadership to stabilize complex enterprise scenarios. This high-impact role sets the standard for weekend operational excellence, turning chaotic technical crises into clear, swift paths to resolution and ensuring seamless cross-regional transitions . WHAT YOU'LL DO Command Critical Escalations: Serve as the accountable incident commander for major customer escalations (CIEs/Sev1s) during the weekend US coverage window, establishing clear decision rights, working cadences, and swift mitigation strategies . Unify Cross-Functional Teams: Mobilize and align Support, Escalations Engineering, Service Delivery, and Account teams to ensure technical responders have immediate context and unblocked paths to restore customer environments . Lead Executive & Customer Communications: Translate fast-moving, complex technical diagnostic details into precise, highly coherent updates for executive leadership and global enterprise customers, keeping stakeholders continuously aligned . Drive Global Follow-the-Sun Handovers: Execute complete, high-quality shift ha
We are fueled by a moral imperative to advance mankind, and it all begins with our people, our product, and our purpose. Passion isn’t something we turn on and off; it’s woven into everything we do. If you thrive in high-challenge environments, are inspired by exceptional teammates, and are driven to grow beyond what you thought possible, MX is where you belong. Come build the future with us. Join an award-winning company that isn’t just shaping the financial industry, but transforming it in ways that create meaningful, lasting impact for millions of people. At MX, reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions, and customers feel every second of downtime. We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand. We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise, and leadership gets coverage and health as a program metric. This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought. Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, and data on PostgreSQL an
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: We are looking for a Senior Software Engineer to join our Site Reliability Engineering team. As a Senior Software Engineer in Production SRE, you will be responsible for developing and maintaining the tools and systems that enable our engineering teams to operate our services reliably and at scale. You will work closely with our SREs and other engineering teams to ensure our services are properly instrumented and able to scale with our growing business. The Difference You Will Make: In this role, your expertise in developing and maintaining tools and systems will be instrumental in bolstering our services' reliability and improving how the company manages incidents broadly. By collaborating closely with other engineering teams you will help establish a culture of reliability throughout the organization by providing a comprehensive incident management platform that is being used for instrumentation, operability, and around incidents. Your ability to identify opportunities for improvement and drive their implementation will contribute significantly to our overall operational efficiency and growth, ensuring that our services remain resilient as our business continues to expand. Additionally, as an essential part of this role, you will serve as an active member of the Production SRE team, responding to and managing high severity incidents. Your vast technical experience and leadership skills will be invaluable as you step into the role of Incident Commander during these critical events. You will guide cross-functional teams during crisis situations and ensure timely resolution, minimizi
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As a Manager On Duty within Everpure’s Global Technical Support organization, you will serve as the critical mission commander orchestrating real-time resolution for enterprise-level operational incidents. Acting as the ultimate bridge between global customers, account management teams, and internal engineering resources, your goal is to safeguard business continuity and de-escalate high-priority customer emergencies. You will drive operational stability across shift handoffs, direct cross-functional resource allocation during active incidents, and champion systemic improvements across the customer experience landscape. WHAT YOU'LL DO Lead Live Crisis & Incident Escalations: Direct real-time triage, cross-functional resource allocation, and technical alignment between engineering, customer support, and sales leadership during critical incident escalations to accelerate resolution times for high-value accounts. Bridge Global Shift Hand-offs: Coordinate continuous 24/7 engagement models across international time zones with fellow Managers On Duty, ensuring seamless tracking, transfer, and execution of commitments for critical customer accounts. Drive Root Cause & Accountability Commitments: Oversee post-escalation deliverables—including technical root cause analyses (RCA) and executive follow-up meetings—to restore operational confidence and reinforce long-term partner trust. Transform Trend Insights into Op
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team & role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Robinhood Command Center (RCC) is a newly formed reliability team that serves as the front line for detecting, coordinating, and mitigating production incidents across Robinhood. As part of Robinhood’s broader reliability initiative, RCC works closely with product engineering, reliability, observability, infrastructure, and business teams to reduce customer impact and shorten incident duration. As a Senior Engineer, you will be part of the founding RCC team, helping define how Robinhood responds to and learns from incidents at scale. This is a highly visible role focused on incident leadership, operational excellence, and reliability tooling. You will not own product services or core infrastructure, but you will own the processes and tools that enable fast, high-quality incident response. This role is based in our Menlo Park, California office, with in-person attendance expected at least 3 days per week. What you'll do: Serve as a senior technical leader driving the long-term reliability and observability strategy across Robinhood’s infrastructure Partner closely across many different types of engineers to raise the bar for operational excellence and incident r
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team & role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Robinhood Command Center (RCC) is a newly formed reliability team that serves as the front line for detecting, coordinating, and mitigating production incidents across Robinhood. As part of Robinhood’s broader reliability initiative, RCC works closely with product engineering, reliability, observability, infrastructure, and business teams to reduce customer impact and shorten incident duration. As a Senior Engineer, you will be part of the founding RCC team, helping define how Robinhood responds to and learns from incidents at scale. This is a highly visible role focused on incident leadership, operational excellence, and reliability tooling. You will not own product services or core infrastructure, but you will own the processes and tools that enable fast, high-quality incident response. This role is based in our New York, New York office, with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. What you'll do: Serve as a senior tech
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. About the Team The Security Operations (SecOps) team at Robinhood proactively safeguards our platform and millions of customers. We monitor, detect, and respond to security threats in real time while staying ahead of risks through threat intelligence, Red Team operations, and research partnerships. We are building the next generation of security operations—leveraging AI-driven automation, Autonomic Security Operations (ASO), and innovative detection frameworks to set the standard across the cybersecurity industry! About the Role As a Staff Security Engineer (IC6) on the Detection & Response team, you will drive our incident response strategy, build robust detection engineering frameworks, and mentor engineers across the organization. In this high-impact role, you will command high-stress incident responses, eliminate operational noise by developing an AI-native detection platform, and help shape the broader AI-agentic ecosystem for SecOps. You will work closely with cross-functional partners in Proactive Security, Security Engineering, Insider Trust, Infrastructure, Legal, and Communications to protect Robinhood’s ecosystem. This role is based in our Bellevue, WA a
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As our Senior Manager of Global Moderation Escalations, you aren't just managing a localized queue—you are the global architect of our crisis and escalation response. Leading highly specialized teams across the US, India, and beyond, you will transform real-time incident management into a proactive, insight-driven safety engine. By bridging Product, Engineering, and Executive Leadership, you will steer the strategy that protects millions of users daily. If you are ready to command a global front-line team and shape the future of online safety, come build with us. You will report to the Senior Manager, Moderation Operations. Build and drive a world-class global escalations and insights program, evolving our escalation systems from v0 into a sophisticated, industry-leading defect reduction engine. Command and align distributed teams across the US, India, and other global regions, establishing a seamless "Follow-the-Sun" operational model. Transition the program from transactional response to proactive prevention by identifying root causes of moderation errors and partnering with cross-functional teams to implement systemic fixes. Recruit, hire, and lead a specialized team for escalation and i
Location Details: Remote - India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. The Remediation Support Analyst plays a key role in supporting website security customers by managing the incident response lifecycle. This includes identifying malware, containing and removing threats, restoring site functionality, and recommending prevention measures. The role focuses on website clean-ups and requires knowledge of PHP, CMS platforms, and how malicious code operates. Each customer interaction is an opportunity to resolve issues and build technical expertise. What you'll get to do... Perform website clean-up and troubleshooting via support ticket Follow processes to identify and remove malware while delivering strong customer service Analyze code to detect malicious activity Contribute findings to improve automation and processes Support platforms including WordPress, Joomla, Drupal, and Magento Collaborate across teams to share security insights Your experience should include... Ability to identify and decode complex, multi-step obfuscated malware Experience writing PHP snippets and Shell scripts and able to read and interpret PHP and Java script Experience with Windows web stack (IIS, SQL Server, ASP.NET) Ability to read and write regular expressions (Regex) Database management experience and using advanced command-line tools (e.g. process tracing, advanced search and replace) Experience with cPanel/WHM or similar hosting control panels Solid understanding of web security principles and malware threats Experience troubleshooting websites across CMS platforms Working knowledge of Linux/UNIX environments and understandi
A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO Monitor, triage, and assess incidents across global markets and technology platforms, validating severity, business impact, and escalation requirements. Act as command-and-control lead for major incidents, driving structured response, rapid restoration, and informed decision-making. Coordinate recovery efforts across L1/L2/L3 support teams, engineering, infrastructure, and third-party vendors to minimize business disruption. Lead internal and external stakeholder communications, including senior leadership, during incidents and service disruptions. Trigger and govern disaster recovery (DR) failover in line with predefined criteria, controls, and governance standards. Ensure accurate, timely, and audit-ready incident documentation, including detailed timelines, evidence, and impact analysis. Facilitate post-incident reviews (RCA/post-mortems), ensure root causes are clearly identified, and track corrective and preventative actions to closure. Govern end-to-end change management, including change risk assessment, CAB facilitation, approval workflows, maintenance windows, and change-related incident management. Drive continuous improvement in incident, change, and event management processes through automation, standardization, and improved tooling and alert quality. Partner with engineering and operations teams to design scalable, resilient processes, promote best practices, and enhance operational matur
Get new incident commander jobs by email
Daily job updates · Unsubscribe anytime