Jobs in Canada

Incident Commander in San Francisco

13 active opportunities · Updated October 2026

Explore current incident commander jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $134.4K/yr

Quick readStrong listing-quality and freshness signals

At Scale, we believe that the next frontier of artificial intelligence is embodied. The Physical AI team is focused on building general AI that can reason and act in the physical world. By leveraging Scale’s massive, industry-leading data infrastructure, we are partnering with frontier labs to build Foundation Models for Physical AI that will redefine the future of automation. To support our rapid hardware-software iteration cycles and ensure a world-class R&D environment, we are looking for a Safety Coordinator / Lab Lead to anchor our physical testing operations. Role Overview As the Safety Coordinator / Lab Lead , you will play a mission-critical role in scaling our physical testing infrastructure safely and efficiently. This is a high-impact position where your highest-priority responsibility will be owning the end-to-end execution of safety audits and incident documentation . Operating at the intersection of cutting-edge AI foundation models and complex robotics hardware, you will ensure our researchers, engineers, and autonomous systems interact in a secure, compliant, and highly organized environment. Core Responsibilities Priority Focus: Safety Audits & Incident Documentation Rigorous Safety Audits: Design, schedule, and execute routine safety audits across all physical testing environments, robot cells, and hardware workspaces to ensure continuous compliance with internal benchmarks and industrial safety standards. Incident & Near-Miss Documentation: Own the end-to-end incident management pipeline. Act as the primary point of contact for documenting, archiving, and analyzing any lab incidents, mechanical anomalies, or near-misses. Root-Cause Analysis (RCA): Lead structured post-incident investigations to identify systematic risks, authoring comprehensive RCA reports and implementing Corrective and Preventive Actions (CAPA). Data-Driven Risk Mitigation: Treat safety data as a core operational asset—tracking safety metrics and audit trends to proa

AWSRestAIGo
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $1.1M/yr

Quick readStrong listing-quality and freshness signals

About the Team DoorDash's Protective Services function safeguards executives, their families, and residences through executive protection, residential security, threat management, and secure transportation across a three-brand global enterprise. The team operates 24/7 and is built on the principle that protection starts before an incident, not after. The people who do this work well combine operational discipline with personal composure, take ownership of outcomes rather than tasks, and represent the program in every interaction with the principal population. About the Role The Protective Services Agent provides close protection for DoorDash executives and their families across all operating environments, including corporate headquarters, domestic and international travel, company events, and residential settings. The role requires a trained executive protection professional who develops and implements security plans, conducts advances and risk assessments, coordinates with law enforcement and venue partners, and serves as the principal's primary security presence. Personnel must be capable of managing incidents and serving as first responders on scene. Operations regularly involve extended duty periods, irregular schedules, and short-notice deployments. Potential travel up to 25% of the time. You are excited about this opportunity because you will… Provide close protection across all principal environments. Serve as the primary security presence for executives and their families across headquarters, travel, events, and residential operations. Maintain continuous situational awareness, assess threats in real time, and execute protective protocols effectively and unobtrusively. Plan and execute security operations. Develop and implement comprehensive security plans for executive movements and events, incorporating site surveys, route planning, risk assessments, staffing requirements, screening procedures, and emergency response strategies. Conduct advances for all pr

AWSGitRestAI
T
📍 San Francisco, Canada
✓ High-confidence listingCompany trend -95.1%
Quick readStrong listing-quality and freshness signals

About Us Twitch is the world’s biggest live streaming service, with global communities built around gaming, entertainment, music, sports, cooking, and more. It is where thousands of communities come together for whatever, every day. We’re about community, inside and out. You’ll find coworkers who are eager to team up, collaborate, and smash (or elegantly solve) problems together. We’re on a quest to empower live communities, so if this sounds good to you, see what we’re up to on LinkedIn and X , and discover the projects we’re solving on our Blog . Be sure to explore our Interviewing Guide to learn how to ace our interview process. About the Role Twitch Security Platform builds and operates the technology at the intersection of security, privacy, software engineering, and data engineering. As a Senior Engineering Manager, Security Platform, you will guide a diverse group of software engineers to build critical tools and automation solutions used across Twitch. You will report to the Director of Security, Identity, and Privacy Platforms. You will manage engineers responsible for operating our security data platform handling billions of weekly events, a multi-petabyte data lake, and numerous analysis tools that provide critical data across the organization to support important security programs. You will also lead engineers responsible for writing software to accomplish security at scale through automation, libraries, and services. With this team, you will lead the full software development lifecycle from product discovery, roadmap prioritization, and execution. Programs you support will include Fraud, Service-to-Service Auth, Detection and Incident Response, Vulnerability Management, Engineering Intelligence, and Privacy. You are experienced in software and data engineering, you bring awareness of the industry's cutting edge, and you're passionate about security. If this sounds accurate, come join us! You can work from San Francisc

HI
📍 San Francisco, Canada
✓ High-confidence listing

$149.9K – $270K/yr

Quick readStrong listing-quality and freshness signals

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Senior Platform Engineer at HP IQ, you will help build and evolve the infrastructure, tooling, and shared platform capabilities that enable our engineering teams to develop and operate reliable, secure, and scalable services across cloud and edge environments . You will work closely with application, services, AI/ML, and security teams to improve developer velocity, production readiness, reliability, and operational efficiency across a heterogeneous infrastructure footprint. What You Might Do Design, build, and maintain shared infrastructure and platform capabilities across cloud and edge environments. Build automation and self-service tooling that improves engineering velocity and operational consistency. Develop and maintain Infrastructure-as-Code, deployment workflows, and environment provisioning. Partner with engineering teams on production readiness, including reliability, security, observability, scalability, and recovery. Improve monitoring, alerting, incident response, and operational tooling across distributed environments. Automate repetitive operational t

PythonKubernetesAI
TI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$136.3K – $273.9K/yr

Quick readStrong listing-quality and freshness signals

TextNow is on a mission to make communications affordable and accessible for everyone. As a full MVNO operating our own mobile core network over LTE and 5G NSA, we have the unique advantage of controlling our network infrastructure end-to-end. We operate the HSS, PGW, and other critical network functions, giving us the flexibility to innovate and deliver exceptional service to millions of users. About the Role Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. T extNow is looking for a new SecOps team member to secure, monitor , and enable automated response within our infrastructure. What You’ll Do Ensure Secure & Reliable Systems: Design, implement, and maintain security-focused infrastructure to protect TextNow’s services while ensuring reliability and scalability. Security Automation & Infrastructure as Code: Develop and enforce best practices using Terraform, Ansible, Crowdstrike , and AWS security tools , ensuring secure configurations, automated compliance checks, and infrastructure as code. Threat Detection & Incident Response: Participate in an on-call rotation to respond to security incidents, investigate vulnerabilities, and implement proactive measures to prevent future threats. Work closely with engineering teams to remediate security risks. Monitoring & Logging for Security: Improve observability by implementing security monitoring solutions, logging best practices, and alerting mechanisms to detect anomalies and suspicious activity. Access Control & Identity Management: Manage IAM roles, permissions, and policies to ensure least privilege access and enforce security controls across cloud and internal systems. Collaboration & Security Advocacy: Wo

AWSCI/CDGitAI
TI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

C$113.4K – C$162K/yr

Quick readStrong listing-quality and freshness signals

We believe communication belongs to everyone. We exist to democratize phone service. TextNow is evolving the way the world connects and that's because we're made up of people with curious minds who bring an optimistic, yet critical lens into the work we do. We're the largest provider of free phone service in the nation. And we're just getting started. Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between! This role is about impact at scale. You’ll shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded into everyday engineering practice. Using AI is not optional, it’s expected. From design and architecture to implementation, testing, debugging, documentation, and operational analysis, you will actively leverage AI tools to increase velocity, improve code quality, and make better technical decisions. We provide a robust suite of AI-powered development tools and workflows to support you, and we expect you to continuously evolve how you use them to raise the bar for efficiency, clarity, and product excellence across the organization. What You'll Do Ensure System Reliability: Design, build, and maintain scalable, resilient, and highly available systems to support TextNow’s infrastructure and services. Automation & Infrastructure as Code: Develop and maintain automation using Terraform, Ansible, and other tools to enable efficient deployment, scaling, and operations of cloud-based systems (AWS preferred). Incident Response & On-Call Support: Participate in an on-call rotation, troubleshoot issues, and drive incident resolution to minimize downtime and improve syste

AWSCI/CDGitAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $1.6M/yr

Quick readStrong listing-quality and freshness signals

About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus

PythonJavaSQLAWS
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $897.6K/yr

Quick readStrong listing-quality and freshness signals

About the Team As one of DoorDash's core operations teams, Customer Experience ensures that when issues arise across the platform, there is a reliable and effective support system in place. Our team designs, manages, and continuously improves DoorDash's global support network, with the goal of delivering a high-quality and consistent customer experience. This role sits within the Safety Customer Experience team, focused on the most critical and high-risk incidents on the platform. The team owns some of the most sensitive customer interactions at DoorDash, where thoughtful operational and product decisions directly improve customer trust and platform safety. About the Role You'll operate at the intersection of Product, Operations, and Customer Experience to improve how DoorDash prevents, identifies, and responds to the most critical safety incidents on the platform. You'll partner closely with Product, Policy, Analytics, and Operations to design scalable solutions that deliver accurate, timely support during customers' highest-stakes moments. This role requires a detail-oriented operator who can navigate complex problem spaces, design scalable processes, and execute with precision in a fast-paced environment. You’ll be expected to take ownership of ambiguous, high-stakes problems and translate them into structured, actionable solutions. You’re excited about this opportunity because you will… Drive Safety Strategy – Partner with cross-functional teams to identify and execute initiatives that improve how DoorDash prevents, identifies, and responds to safety incidents. Own End-to-End Experience – Design and optimize the full lifecycle of safety incident handling, including reporting, workflows, support execution, and tooling. Drive Data-Driven Decisions – Leverage data and case-level insights to identify root causes, measure performance, and prioritize opportunities to improve the safety customer experience. Influence Cross-Functionally – Collaborate with Product, Engin

AWSGitRestAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $962.4K/yr

Quick readStrong listing-quality and freshness signals

About the Team The Global Safety and Security team advances DoorDash by protecting people, property, operations, brand, and reputation. Within that team, Protective Services safeguards executives and residences through executive protection, intelligence collection, threat monitoring, residential security, and secure transportation. We manage risk and deliver value through technology and a people-first approach, and we strive to always be ahead of the curve and there for our people, anytime, anywhere. About the Role We are hiring Protective Services Associates to join the Protective Services team, providing residential security coverage and on-site protective presence in support of executives. This is the entry point on our cross-trained career ladder. Associates build proficiency across residential security, monitoring and threat reporting, and close-protection fundamentals, with a structured path to Senior Associate and beyond as you grow. Responsibilities Residential Security Coverage: Provide dedicated, round-the-clock physical security presence at a residential property as part of a rotating shift team. Monitoring & Threat Reporting: Monitor camera systems and access points, document and report security-relevant activity, and hand off findings to the team's intelligence function. Access Control: Screen and coordinate visitors, deliveries, and vendors according to established protocols. Logistics Coordination: Plan and manage the logistics of residential security operations, including scheduling, shift handoffs, and coverage during events or executive travel. Coordinate with drivers and other Protective Services team members to align movements, timing, and coverage. Incident Response: Respond to security incidents or anomalies, following established escalation procedures. Professional Conduct: Serve as a trusted, hands-on point of contact for executives, anticipating needs, exercising sound judgment in real time, and delivering discreet, concierge-level servi

AWSGitRestAI
L
📍 San Francisco, Canada· Full-time
✓ High-confidence listingCompany trend -74.4%
Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. We are hiring a Machine Learning Engineer to join our ETA team. Our team builds and maintains Lyft's system responsible for estimating/predicting ETAs for every ride request on our platform. ETAs play a critical role in matching decisions, pricing estimates and overall user experience. Low latency, high reliability and high accuracy are paramount for our success. If you are a critical thinker with experience in machine learning workflows and writing reliable code, passionate about solving business problems using data and working in a dynamic, creative, and collaborative environment, we are searching for you. Our technology stack runs on AWS, Kubernetes, Go, Spark, Python and Apache Airflow. In this role, you will work with incredibly passionate and talented colleagues from software engineering, machine learning and data science on building rideshare experiences that delight millions of riders and drivers. Responsibilities: Perform data analysis and build proof-of-concept to explore and compare ML and non-ML solutions Be able to make effective tradeoffs between model accuracy, its productization complexity and runtime performance Develop statistical, machine learning, or optimization models Write production quality code that can scale well to serve millions of requests per day Participate in code reviews, design reviews, production on-call support and incident triaging process. Write well-crafted, well-tested, readable, maintainable code Experience: B.S., M.S., or Ph.D. in Computer Science or other quantitative fields or related work experience 3+ years of Machine Learning experience Nice-to-have: Experience with big data processing / distributed data pipelines and tools such as Apache Airflow and Spark Ability to work in distributed teams spread across time zones. (North America and

PythonAWSKubernetesMachine Learning
L
📍 San Francisco, Canada· Full-time
✓ High-confidence listingCompany trend -74.4%
Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Lyft Business Product Platform team builds the systems and experiences that power Lyft's B2B products — enabling companies, organizations, and their employees to seamlessly access Lyft's transportation network. We sit at the intersection of product and platform, owning both the customer-facing features and the underlying infrastructure that makes them reliable at scale. Our work directly impacts how businesses integrate with Lyft, how admins manage their programs, and how millions of riders get where they need to go. Responsibilities: Drive architecture and technical design for systems that are highly available, scalable, and built to last — not just for today's requirements but for where the product is heading Own features end-to-end: from shaping the technical spec and design through to production rollout and operational health Think critically about how AI capabilities can be incorporated into Lyft Business products to improve the experience for business admins and riders — and bring that perspective into roadmap and architecture conversations Make well-reasoned trade-off decisions and communicate them clearly to peers, leads, and cross-functional partners Write clean, well-tested, maintainable code and hold a high bar for the same in code reviews Partner across engineering, product, and design to align on direction and get buy-in on technical approaches Proactively engage in incident response, contributing both to resolution and to long-term reliability improvements Grow the team's technical culture through design reviews, tech talks, and mentorship Experience: 5+ years of software engineering experience, with a track record of designing and shipping production systems at scale Strong system design instincts — you can reason through distributed systems trade-offs, identify failure modes, and

SQLAWSAzureGCP
A
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$165K – $247K/yr

Quick readStrong listing-quality and freshness signals

Amplitude is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Winter 2026 Report, Amplitude is the best-in-class solution for product, data, and marketing teams. Learn more at amplitude.com . As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do. Amplitude’s Commitment to Diversity Equity & Inclusion (DEI): Amplitude believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive. About the Role Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Senior Platform Engineer, you'll own medium-to-high-complexity platform projects end-to-end and help shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll partner with Staff engineers and product teams to make Kubernetes effortless across the engineering org, building self-service automation and scalable AWS infrastructure that lets product teams ship faster, safer, and with less cognitive load. If you're excited about building the systems that other engineers will rely on every day, this role is for you. Key Resp

PythonAWSGCPKubernetes
🔔

Get new incident commander jobs in San Francisco, Canada by email

Daily job updates · Unsubscribe anytime