Jobs in Canada

Incident Commander in Canada

41 active opportunities · Updated October 2026

Explore current incident commander jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 Menlo Park, Canada· Full-time
✓ Quality checkedCompany trend -82.8%

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team & role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Robinhood Command Center (RCC) is a newly formed reliability team that serves as the front line for detecting, coordinating, and mitigating production incidents across Robinhood. As part of Robinhood’s broader reliability initiative, RCC works closely with product engineering, reliability, observability, infrastructure, and business teams to reduce customer impact and shorten incident duration. As a Senior Engineer, you will be part of the founding RCC team, helping define how Robinhood responds to and learns from incidents at scale. This is a highly visible role focused on incident leadership, operational excellence, and reliability tooling. You will not own product services or core infrastructure, but you will own the processes and tools that enable fast, high-quality incident response. This role is based in our Menlo Park, California office, with in-person attendance expected at least 3 days per week. What you'll do: Serve as a senior technical leader driving the long-term reliability and observability strategy across Robinhood’s infrastructure Partner closely across many different types of engineers to raise the bar for operational excellence and incident r

VueAWSAIGo
R
📍 Bellevue, WA; Menlo Park, Canada· Full-time
✓ Quality checkedCompany trend -82.8%

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. About the Team The Security Operations (SecOps) team at Robinhood proactively safeguards our platform and millions of customers. We monitor, detect, and respond to security threats in real time while staying ahead of risks through threat intelligence, Red Team operations, and research partnerships. We are building the next generation of security operations—leveraging AI-driven automation, Autonomic Security Operations (ASO), and innovative detection frameworks to set the standard across the cybersecurity industry! About the Role As a Staff Security Engineer (IC6) on the Detection & Response team, you will drive our incident response strategy, build robust detection engineering frameworks, and mentor engineers across the organization. In this high-impact role, you will command high-stress incident responses, eliminate operational noise by developing an AI-native detection platform, and help shape the broader AI-agentic ecosystem for SecOps. You will work closely with cross-functional partners in Proactive Security, Security Engineering, Insider Trust, Infrastructure, Legal, and Communications to protect Robinhood’s ecosystem. This role is based in our Bellevue, WA a

VueAWSKubernetesAI
R
📍 Denver, CO; Menlo Park, Canada· Full-time
✓ Quality checkedCompany trend -82.8%

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Security Operations (SecOps) team protects Robinhood and its customers by detecting, investigating, and responding to security threats across our production systems, endpoints, and cloud environments. We combine detection engineering, incident response, and threat intelligence to identify emerging risks, strengthen our defenses, and reduce the impact of attacks before they reach our customers. As we continue to evolve our security platform, we're embracing AI and automation to help our teams move faster, uncover threats more effectively, and scale our defenses against increasingly sophisticated adversaries. As the Manager of Response, Automation, Intelligence and Detection Engineering (RAID) within SecOps, you will lead teams across North America, shaping the strategy, execution, and long term evolution of our defensive security capabilities. You'll be responsible for building high performing teams, maturing our detection and incident response programs, and ensuring we stay ahead of an increasingly sophisticated threat landscape. Working closely with Security, Engineering, Infrastructure, and Trust & Safety, you'll translate emerging threats into scalable defenses wh

VueAWSAIGo
MR
📍 Toronto, Ontario, Canada· Full-time
✓ Quality checked

Toronto, ON Want to work at a multinational investment bank? Join mthree as a Production Support Analyst and fast-track your career by working with one of our leading global banking clients in Toronto. You’ll support critical Capital Markets trading systems in a fast-paced commodities and futures environment, working closely with front office users to ensure stability, performance, and continuous improvement of production platforms. What you’ll do: This role focuses on supporting large-scale production systems within Capital Markets, ensuring end-to-end stability of trading applications while driving operational excellence. Provide specialized application support for Capital Markets trading platforms (approx. 60%) Perform end-of-day monitoring, incident management, root cause analysis, and problem resolution Support commodities and futures trading activities including trade flow/STP, P&L valuation, and breaks Work directly with Front Office business users in a high-pressure trading environment Monitor and support market data feeds and core trading applications Analyze existing systems and liaise with Business and Development teams to drive system improvements Manage user access and entitlements (approx. 20%) Coordinate with internal teams and external vendors (approx. 20%) Participate in weekend work as required for projects, releases, and disaster recovery testing Provide 24x7 on-call support based on assigned shift for end-of-day monitoring Deliver to tight timelines in a fast-paced trading environment What we’re looking for: (must have) Experience working in the financial services industry, ideally within Capital Markets / Futures Strong knowledge of Commodities and Futures products Proven experience supporting large-scale production systems Hands-on experience with end-of-day processing, monitoring, incident management, and root cause analysis Understanding of trade lifecycle concepts (trade flow/STP, P&L valuation, breaks) Experience with tools such as

PythonSQLRestAI
F
📍 Toronto, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Forma.ai: Forma.ai is a Series B startup that's revolutionizing how sales compensation is designed, managed and optimized. We handle billions in annual managed commissions for market leaders like Edmentum, Stryker, and Autodesk. Our growth has been fuelled by our passion for fundamentally changing and shaping how companies use sales intelligence to drive business strategy. We’re welcoming equally driven individuals who are excited about creating something big! The Opportunity As a Staff Security Engineer, you will be a hands-on technical leader strengthening security across Forma's application, cloud infrastructure, development lifecycle, internal systems, and incident-response practices. Security today is shared across Engineering and DevOps. You'll work closely with both teams and have real room to shape how Forma approaches security as we grow. Depending on your interests and the needs of the business, the role could develop into a deeper individual-contributor position or help build a dedicated security team. You'll work directly with Engineering, DevOps, IT, Product, Legal, and Privacy to identify risks, design practical controls, automate security processes, and help teams ship secure and reliable software. What you'll do Cloud and infrastructure security Design and implement security controls across Forma's AWS environments, with a focus on IAM, least-privilege access, service identities, and account boundaries. Embed security requirements into Terraform and other Infrastructure as Code, and improve secrets, certificate, encryption-key, and credential management. Build automated checks for insecure configurations, excessive permissions, exposed resources, and configuration drift across Kubernetes, containers, serverless workloads, networking, and data services. Application, data, and AI security Run threat modelling and security architecture reviews for new products, services, APIs, data pipelines, and third-party integrations

PythonAWSKubernetesCI/CD
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $134.4K/yr

Quick readStrong listing-quality and freshness signals

At Scale, we believe that the next frontier of artificial intelligence is embodied. The Physical AI team is focused on building general AI that can reason and act in the physical world. By leveraging Scale’s massive, industry-leading data infrastructure, we are partnering with frontier labs to build Foundation Models for Physical AI that will redefine the future of automation. To support our rapid hardware-software iteration cycles and ensure a world-class R&D environment, we are looking for a Safety Coordinator / Lab Lead to anchor our physical testing operations. Role Overview As the Safety Coordinator / Lab Lead , you will play a mission-critical role in scaling our physical testing infrastructure safely and efficiently. This is a high-impact position where your highest-priority responsibility will be owning the end-to-end execution of safety audits and incident documentation . Operating at the intersection of cutting-edge AI foundation models and complex robotics hardware, you will ensure our researchers, engineers, and autonomous systems interact in a secure, compliant, and highly organized environment. Core Responsibilities Priority Focus: Safety Audits & Incident Documentation Rigorous Safety Audits: Design, schedule, and execute routine safety audits across all physical testing environments, robot cells, and hardware workspaces to ensure continuous compliance with internal benchmarks and industrial safety standards. Incident & Near-Miss Documentation: Own the end-to-end incident management pipeline. Act as the primary point of contact for documenting, archiving, and analyzing any lab incidents, mechanical anomalies, or near-misses. Root-Cause Analysis (RCA): Lead structured post-incident investigations to identify systematic risks, authoring comprehensive RCA reports and implementing Corrective and Preventive Actions (CAPA). Data-Driven Risk Mitigation: Treat safety data as a core operational asset—tracking safety metrics and audit trends to proa

AWSRestAIGo
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $1.1M/yr

Quick readStrong listing-quality and freshness signals

About the Team DoorDash's Protective Services function safeguards executives, their families, and residences through executive protection, residential security, threat management, and secure transportation across a three-brand global enterprise. The team operates 24/7 and is built on the principle that protection starts before an incident, not after. The people who do this work well combine operational discipline with personal composure, take ownership of outcomes rather than tasks, and represent the program in every interaction with the principal population. About the Role The Protective Services Agent provides close protection for DoorDash executives and their families across all operating environments, including corporate headquarters, domestic and international travel, company events, and residential settings. The role requires a trained executive protection professional who develops and implements security plans, conducts advances and risk assessments, coordinates with law enforcement and venue partners, and serves as the principal's primary security presence. Personnel must be capable of managing incidents and serving as first responders on scene. Operations regularly involve extended duty periods, irregular schedules, and short-notice deployments. Potential travel up to 25% of the time. You are excited about this opportunity because you will… Provide close protection across all principal environments. Serve as the primary security presence for executives and their families across headquarters, travel, events, and residential operations. Maintain continuous situational awareness, assess threats in real time, and execute protective protocols effectively and unobtrusively. Plan and execute security operations. Develop and implement comprehensive security plans for executive movements and events, incorporating site surveys, route planning, risk assessments, staffing requirements, screening procedures, and emergency response strategies. Conduct advances for all pr

AWSGitRestAI
O
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -67.1%

From C$110K/yr

Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Team The Auth0 Platform Tools team owns the incident management tooling, Slack-based tooling, StatusPage, and local development environments that Auth0 engineers rely on every day. That includes incident.io and the services we have built around it, Statuspage, custom Slack bot applications that automate our incident response and engineering operations workflows, the customer-facing web application behind status.auth0.com, Vivaldi, and Tilt - the tools engineers use to run Auth0 locally. We are seeking an engineer to help build new features across all of these tools. Our stack is primarily TypeScript and Node.js, with a React and Next.js front end, backed by Postgres and Redis, and deployed on Kubernetes on AWS. A significant portion of our incident and engineering operations automation is built on Tines, a no-code automation platform. Prior no-code experience is welcome, but we expect you to learn Tines here and become effective with it. Current initiatives include extending our incident tooling to meet FedRAMP requirements, taking full ownership of the status page, and improving how we communicate incident status to customers. There is real room to improve along the way, from test coverage to resilience to inherited technical debt. We build for two audiences: Auth0 engineers, who depend on our tooling every day, and Auth0's customers, who rely on the status page during incidents. We are looking for an engineer who cares about both and enjoys working wi

TypeScriptPythonReactNode.js
T
📍 San Francisco, Canada
✓ High-confidence listingCompany trend -95.1%
Quick readStrong listing-quality and freshness signals

About Us Twitch is the world’s biggest live streaming service, with global communities built around gaming, entertainment, music, sports, cooking, and more. It is where thousands of communities come together for whatever, every day. We’re about community, inside and out. You’ll find coworkers who are eager to team up, collaborate, and smash (or elegantly solve) problems together. We’re on a quest to empower live communities, so if this sounds good to you, see what we’re up to on LinkedIn and X , and discover the projects we’re solving on our Blog . Be sure to explore our Interviewing Guide to learn how to ace our interview process. About the Role Twitch Security Platform builds and operates the technology at the intersection of security, privacy, software engineering, and data engineering. As a Senior Engineering Manager, Security Platform, you will guide a diverse group of software engineers to build critical tools and automation solutions used across Twitch. You will report to the Director of Security, Identity, and Privacy Platforms. You will manage engineers responsible for operating our security data platform handling billions of weekly events, a multi-petabyte data lake, and numerous analysis tools that provide critical data across the organization to support important security programs. You will also lead engineers responsible for writing software to accomplish security at scale through automation, libraries, and services. With this team, you will lead the full software development lifecycle from product discovery, roadmap prioritization, and execution. Programs you support will include Fraud, Service-to-Service Auth, Detection and Incident Response, Vulnerability Management, Engineering Intelligence, and Privacy. You are experienced in software and data engineering, you bring awareness of the industry's cutting edge, and you're passionate about security. If this sounds accurate, come join us! You can work from San Francisc

HI
📍 San Francisco, Canada
✓ High-confidence listing

$149.9K – $270K/yr

Quick readStrong listing-quality and freshness signals

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Senior Platform Engineer at HP IQ, you will help build and evolve the infrastructure, tooling, and shared platform capabilities that enable our engineering teams to develop and operate reliable, secure, and scalable services across cloud and edge environments . You will work closely with application, services, AI/ML, and security teams to improve developer velocity, production readiness, reliability, and operational efficiency across a heterogeneous infrastructure footprint. What You Might Do Design, build, and maintain shared infrastructure and platform capabilities across cloud and edge environments. Build automation and self-service tooling that improves engineering velocity and operational consistency. Develop and maintain Infrastructure-as-Code, deployment workflows, and environment provisioning. Partner with engineering teams on production readiness, including reliability, security, observability, scalability, and recovery. Improve monitoring, alerting, incident response, and operational tooling across distributed environments. Automate repetitive operational t

PythonKubernetesAI
T-
📍 Toronto, Canada· Full-time
✓ High-confidence listing

From C$2M/yr

Quick readStrong listing-quality and freshness signals

About the Role: As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams to build world-class machine learning inference platforms. These platforms power essential services like personalized recommendations, search, and content understanding across Tubi. A core responsibility of this team is developing and maintaining low-latency ML model serving systems that support Deep Learning, LLM, and Search models. This involves building self-service infrastructure and critical components such as the inference engine, feature store, vector store, and experimentation engine. You will improve the way we deploy and operate our services and even contribute to open-source projects. This role grants the architectural freedom to explore new frameworks, lead critical cross-functional projects, and transform the capabilities of our ML and Product teams. Responsibilities: Design and build scalable, high throughput, and low latency distributed systems using Scala Build reusable components and services that serve various ML applications like Personalization, Search, Ads and Exploration Partner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art. Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessary Mentor other engineers on the team on system design, effective incident management, interviewing, leveraging LLMs for work, etc. Collaborate with ML, Product, and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi. Your Background: Experience designing and building scalable, distributed systems in any modern backend language (e.g., Scala, Java, Python, Go, C++); experience with Scala or JVM b

PythonJavaSQLRedis
TI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$136.3K – $273.9K/yr

Quick readStrong listing-quality and freshness signals

TextNow is on a mission to make communications affordable and accessible for everyone. As a full MVNO operating our own mobile core network over LTE and 5G NSA, we have the unique advantage of controlling our network infrastructure end-to-end. We operate the HSS, PGW, and other critical network functions, giving us the flexibility to innovate and deliver exceptional service to millions of users. About the Role Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. T extNow is looking for a new SecOps team member to secure, monitor , and enable automated response within our infrastructure. What You’ll Do Ensure Secure & Reliable Systems: Design, implement, and maintain security-focused infrastructure to protect TextNow’s services while ensuring reliability and scalability. Security Automation & Infrastructure as Code: Develop and enforce best practices using Terraform, Ansible, Crowdstrike , and AWS security tools , ensuring secure configurations, automated compliance checks, and infrastructure as code. Threat Detection & Incident Response: Participate in an on-call rotation to respond to security incidents, investigate vulnerabilities, and implement proactive measures to prevent future threats. Work closely with engineering teams to remediate security risks. Monitoring & Logging for Security: Improve observability by implementing security monitoring solutions, logging best practices, and alerting mechanisms to detect anomalies and suspicious activity. Access Control & Identity Management: Manage IAM roles, permissions, and policies to ensure least privilege access and enforce security controls across cloud and internal systems. Collaboration & Security Advocacy: Wo

AWSCI/CDGitAI
TI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

C$113.4K – C$162K/yr

Quick readStrong listing-quality and freshness signals

We believe communication belongs to everyone. We exist to democratize phone service. TextNow is evolving the way the world connects and that's because we're made up of people with curious minds who bring an optimistic, yet critical lens into the work we do. We're the largest provider of free phone service in the nation. And we're just getting started. Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between! This role is about impact at scale. You’ll shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded into everyday engineering practice. Using AI is not optional, it’s expected. From design and architecture to implementation, testing, debugging, documentation, and operational analysis, you will actively leverage AI tools to increase velocity, improve code quality, and make better technical decisions. We provide a robust suite of AI-powered development tools and workflows to support you, and we expect you to continuously evolve how you use them to raise the bar for efficiency, clarity, and product excellence across the organization. What You'll Do Ensure System Reliability: Design, build, and maintain scalable, resilient, and highly available systems to support TextNow’s infrastructure and services. Automation & Infrastructure as Code: Develop and maintain automation using Terraform, Ansible, and other tools to enable efficient deployment, scaling, and operations of cloud-based systems (AWS preferred). Incident Response & On-Call Support: Participate in an on-call rotation, troubleshoot issues, and drive incident resolution to minimize downtime and improve syste

AWSCI/CDGitAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $1.6M/yr

Quick readStrong listing-quality and freshness signals

About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus

PythonJavaSQLAWS
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $897.6K/yr

Quick readStrong listing-quality and freshness signals

About the Team As one of DoorDash's core operations teams, Customer Experience ensures that when issues arise across the platform, there is a reliable and effective support system in place. Our team designs, manages, and continuously improves DoorDash's global support network, with the goal of delivering a high-quality and consistent customer experience. This role sits within the Safety Customer Experience team, focused on the most critical and high-risk incidents on the platform. The team owns some of the most sensitive customer interactions at DoorDash, where thoughtful operational and product decisions directly improve customer trust and platform safety. About the Role You'll operate at the intersection of Product, Operations, and Customer Experience to improve how DoorDash prevents, identifies, and responds to the most critical safety incidents on the platform. You'll partner closely with Product, Policy, Analytics, and Operations to design scalable solutions that deliver accurate, timely support during customers' highest-stakes moments. This role requires a detail-oriented operator who can navigate complex problem spaces, design scalable processes, and execute with precision in a fast-paced environment. You’ll be expected to take ownership of ambiguous, high-stakes problems and translate them into structured, actionable solutions. You’re excited about this opportunity because you will… Drive Safety Strategy – Partner with cross-functional teams to identify and execute initiatives that improve how DoorDash prevents, identifies, and responds to safety incidents. Own End-to-End Experience – Design and optimize the full lifecycle of safety incident handling, including reporting, workflows, support execution, and tooling. Drive Data-Driven Decisions – Leverage data and case-level insights to identify root causes, measure performance, and prioritize opportunities to improve the safety customer experience. Influence Cross-Functionally – Collaborate with Product, Engin

AWSGitRestAI
🔔

Get new incident commander jobs in Canada by email

Daily job updates · Unsubscribe anytime