Jobiba hiring network

Incident Commander Jobs

589 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current incident commander jobs. Use filters to narrow by work mode, employment type, experience and date posted.

L
21 days ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft’s Safety Team is looking for a caring and empathetic individual to be part of the highest level of Safety response. This person will assist and support current Compassionate Care Sr. Associates and Specialists as the first response to customers involved in the most egregious incidents that occur on the Lyft platform. You must be able to provide a sense of comfort to customers involved in traumatic situations, as well as identify potential ways we can support them as they move forward and recover. You must be comfortable working with Safety and Customer Cares team members, acting as the subject matter expert for all things Safety-related. Communicating confidently, efficiently and effectively is key, turning complex incidents and customer needs into concise reports for Safety leadership. We’re looking for someone who has the confidence and creativity to continue elevating Lyft’s Safety Support as the standard in world-class customer care. Responsibilities: Be a strong, caring, and empathetic victim advocate for the Lyft community. Act as the highest level of escalation for our Safety Support teams. Focus on building trust and long-term relationships with customers during extremely sensitive situations to ensure they feel supported. Respond with empathy and sound judgment, working with the customer end to end. Identify potential risks, mediate, and diffuse escalated situations. Support other tiers of Safety Support with case reviews. Partner with the Lyft Operations Center on credible threats to team members or locations Partner with the Crisis Communications team on safety incidents involving media. Partner with the Safety Policy and Community Compliance team on escalated cases that require review. Learn and maintain advanced levels of training focused on supporting and being a victim advoc

airustexcel
View job →
D
Datadog
📍 New York• Full-time• From $156K/yr
26 days ago

As a Product Manager – IaC Detection, you will define, build, and launch capabilities that proactively detect infrastructure issues in code (e.g. Terraform, Helm) before they can be deployed into production and escalate into production incidents. The Infrastructure Monitoring team has pioneered shift-left detection in the industry with Bits Infrastructure Operations , and we’re looking for a Product Manager to expand this capability to a broader set of use cases Customers (and thus developers) are increasingly standardizing on IaC tools to deploy and maintain ever-growing infrastructure in the cloud. At the same time, SREs and Infra teams struggle with an increasing number of production incidents. By shifting-left and identifying high-impact infra changes before they are deployed, we help reduce production incidents, reduce waste, and free up SRE time to focus on value-added tasks. You will own the roadmap to expand IaC detection to a broader set of use cases, including cost detection, blast radius impact, as well as configuration changes on infrastructure powering applications like nginx, postgres and more. You’ll partner closely with Engineering, Design, and customers to build and iterate on the roadmap, build product market fit, drive customer adoption (including internal usage), and focus on coverage and correctness of the AI system. This is an opportunity to lead an initiative at the intersection of AI, infrastructure operations, and autonomous observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the product roadmap for IaC Detection, enabling customers to proactively detect and catch high-impact infrastructure and configuration changes before they are deployed into production and escalate into incidents. Define the end-to

gitairust
View job →
R
Replit
📍 New York• Full-time• Remote
29 days ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. As a Premium Support Engineer at Replit, you’ll be the front line for our highest-value customers — delivering fast, expert, and reliable technical support when it matters most. You’ll handle complex product issues, guide customers through critical incidents, and ensure every interaction meets the highest standard of quality and speed. Replit is at the forefront of AI-driven software development, and how we support customers is constantly evolving. You’ll play a critical role in shaping how Premium Support adapts to new products, new customer expectations, and AI-assisted workflows, operating effectively in ambiguity and driving clarity for your team. You’ll combine deep technical troubleshooting with calm, confident communication to keep builders moving — whether it’s an enterprise team deploying at scale or a top-tier developer relying on Replit to power their business. In this role you will: Provide swift, high-priority support to Premium customers, responding within strict SLAs. Diagnose, reproduce, and resolve complex technical issues across the Replit platform. Escalate and track high-impact issues with Product and Engineering, ensuring timely fixes and transparent communication. Lead customer-facing communications during outages or incidents. Identify recurring issues and collaborate internally to reduce time-to-resolution. Contribute to internal tooling, automation, and documentation that improves team efficiency. Partner with Engineering, Product, Sales and other internal teams to ensure Premium customers receive a consistent, high-quality experience. Help onboard and mentor other support engineers, raising the team’s overall bar for responsiveness and quality. Required skills and experience: 3+ years in techn

REMOTEjavascriptpythonjava
View job →
R
29 days ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. As a Premium Support Engineer at Replit, you’ll be the front line for our highest-value customers — delivering fast, expert, and reliable technical support when it matters most. You’ll handle complex product issues, guide customers through critical incidents, and ensure every interaction meets the highest standard of quality and speed. Replit is at the forefront of AI-driven software development, and how we support customers is constantly evolving. You’ll play a critical role in shaping how Premium Support adapts to new products, new customer expectations, and AI-assisted workflows, operating effectively in ambiguity and driving clarity for your team. You’ll combine deep technical troubleshooting with calm, confident communication to keep builders moving — whether it’s an enterprise team deploying at scale or a top-tier developer relying on Replit to power their business. In this role you will: Provide swift, high-priority support to Premium customers, responding within strict SLAs. Diagnose, reproduce, and resolve complex technical issues across the Replit platform. Escalate and track high-impact issues with Product and Engineering, ensuring timely fixes and transparent communication. Lead customer-facing communications during outages or incidents. Identify recurring issues and collaborate internally to reduce time-to-resolution. Contribute to internal tooling, automation, and documentation that improves team efficiency. Partner with Engineering, Product, Sales and other internal teams to ensure Premium customers receive a consistent, high-quality experience. Help onboard and mentor other support engineers, raising the team’s overall bar for responsiveness and quality. Required skills and experience: 3+ years in techn

REMOTEjavascriptpythonjava
View job →
O
1mo ago

About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight: Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful autonomy (for example future versions of auto-review ). About the Role We’re looking for strong executors with excellent judgment, comfort with ambiguity, and an understanding of frontier model research. You don’t need prior safety or alignment experience, we also welcome people that recently realized that alignment and safety is a critical area to contribute to. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Train and evaluate frontier models to reduce harmful or misaligned agent actions, forming clear hypotheses and executing independently through ambiguity. Mine incidents and build scalable measurement, data-processing, and evaluation systems that turn real failures into repeatable safety signals. Collaborate closely with post-training, capabilities, oversight, and pre-training partners to ship research-backed mitigations into large-scale training and agent systems. You might thrive in this role if you: Have demonstrated strength in research engineering, ML en

awsrestai
View job →

About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight : Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful agent autonomy (for example future versions of auto-review ). About the Role This role focuses on oversight and system-level mitigations that enable increasingly capable agents to operate safely and autonomously in real environments. We prioritize building oversight systems that are used in practice today, both internally and externally (see our recent work on action monitoring for codex and former code review ). We also study longer-term questions about how increasingly capable agentis systems can be supervised, constrained, and corrected. We’re looking for a safety&security minded researcher or engineer who can reason rigorously about security boundaries and agent behavior, then build and test practical mitigations. A background in AI control or security is welcome but not required. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and evaluate system-level controls for agent actions like agent-based review. Plan how they fit in a broader syste

awsrestai
View job →

We are seeking a highly skilled and experienced Staff Network Site Reliability Engineer (SRE) to join our Enterprise Network Operations and SRE team. In this role, you will be pivotal in implementing our vision for a reliable and efficient network infrastructure. The ideal candidate is passionate about network operations and committed to enhancing the user experience. You'll have the opportunity to solve complex network challenges using hands-on debugging and by focusing on network automation, observability, documentation, and operational excellence. This is a critical position focused on ensuring user satisfaction and brilliance in network operations. What you'll be doing: Owning the operational aspect of the network infrastructure, ensuring its high availability and reliability, actively working on network incidents and service requests. Partnering with architecture and deployment teams to guarantee that new implementations are supportable and align with production standards. Advocating for and implementing automation to reduce toil and improve operational efficiency. Minimizing manual operational tasks to achieve and maintain Service Level Objectives (SLOs). Monitoring network performance, identifying areas for improvement, and collaborating with relevant teams to implement refinements. Proactively identifying and mitigating network risks to promote continuous improvement. Collaborating with domain experts across functions to resolve production issues swiftly and effectively, ensuring customer happiness. Conducting blameless postmortems and following through on Root Cause Analyses (RCAs). Discovering opportunities for operational improvements and teaming up with colleagues to devise solutions that enhance excellence and sustainability in network operations. Developing knowledge base articles for automa

pythonlinuxansible
View job →
H
Hp
📍 Texas• $116.2K – $182.4K/yr
1mo ago

Senior Manager, Logistics Security Description - At HP, we create technology that helps people and businesses bring their ideas to life. Our global supply chain plays a vital role in that mission, moving high-value products safely and reliably to customers around the world. We are looking for a collaborative and solutions-oriented Senior Manager, Logistics Security to help protect our logistics network, strengthen transportation risk programs, and lead efforts that reduce loss, improve carrier performance, and support a secure customer experience. Core responsibilities Own logistics security strategy for HP's transportation network. Manage loss, theft, damage, shortage and mis-delivery incidents across carriers, warehouses and distribution partners. Lead transportation claims management, including investigation, documentation, recovery and settlement. Manage relationships with major LSPs, carriers, freight forwarders and 3PLs. Investigate high-value shipment incidents and identify root causes. Develop controls to prevent cargo theft, fraud, unauthorized access and product diversion. Establish security standards for transportation lanes, warehouses and high-risk locations. Track claims and losses using KPIs such as: Claim $ / shipment Loss rate Damage rate Theft incidents Recovery % Claim cycle time Carrier performance Partner with Supply Chain, Logistics, Finance, Legal, Compliance, Security and Insurance teams. Lead corrective/preventive actions with logistics service providers. Review carrier contracts and ensure appropriate liability and insurance coverage. Establish escalation processes for major incidents. Conduct risk assessments of logistics providers and distribution locations. Use dat

financesupply chainlogistics
View job →

As an Overman, the primary responsibility is to ensure the safe and efficient operation of mining activities within the assigned area. Key tasks include taking over shift charge, supervising dump management, face management, drain management, dewatering, haul road maintenance, conducting inspections to ensure bench and OB dump stability, reviewing blasting schedules, and enforcing safety measures such as proper haul road maintenance and dust control. Additionally, the role involves monitoring manpower and equipment usage, maintaining records as per DGMS guidelines, and promptly addressing any safety hazards or incidents to ensure a secure working environment. Source: Adani Group | Job ID: 55067

As an Overman, the primary responsibility is to ensure the safe and efficient operation of mining activities within the assigned area. Key tasks include taking over shift charge, supervising dump management, face management, drain management, dewatering, haul road maintenance, conducting inspections to ensure bench and OB dump stability, reviewing blasting schedules, and enforcing safety measures such as proper haul road maintenance and dust control. Additionally, the role involves monitoring manpower and equipment usage, maintaining records as per DGMS guidelines, and promptly addressing any safety hazards or incidents to ensure a secure working environment. Source: Adani Group | Job ID: 55066

As an Overman, the primary responsibility is to ensure the safe and efficient operation of mining activities within the assigned area. Key tasks include taking over shift charge, supervising dump management, face management, drain management, dewatering, haul road maintenance, conducting inspections to ensure bench and OB dump stability, reviewing blasting schedules, and enforcing safety measures such as proper haul road maintenance and dust control. Additionally, the role involves monitoring manpower and equipment usage, maintaining records as per DGMS guidelines, and promptly addressing any safety hazards or incidents to ensure a secure working environment. Source: Adani Group | Job ID: 51881

AG
Adani Group
📍 Ahmedabad• Full-time
1mo ago

The OT Cybersecurity Specialist is responsible for supporting the implementation, monitoring, and management of cybersecurity measures for the organization’s Operational Technology (OT) systems. This role involves securing industrial control systems (ICS), SCADA systems, and other OT assets from cyber threats, conducting vulnerability assessments, monitoring networks for anomalies, and responding to security incidents. The OT Cybersecurity Specialist will work closely with other IT and cybersecurity teams to ensure a robust and secure OT environment. Source: Adani Group | Job ID: 45584

🔔

Get new incident commander jobs by email

Daily job updates · Unsubscribe anytime