Jobiba hiring network

Incident Commander Jobs

589 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current incident commander jobs. Use filters to narrow by work mode, employment type, experience and date posted.

A
Affirm
📍 Poland• Full-time• Remote• $384K – $576K/yr
19 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions

REMOTEpythonsqlmysql
View job →
A
Affirm
📍 Spain• Full-time• Remote• From €1M/yr
19 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions

REMOTEpythonsqlmysql
View job →
TA
19 days ago

About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw

pythonkuberneteslinux
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $134.4K/yr
19 days ago

At Scale, we believe that the next frontier of artificial intelligence is embodied. The Physical AI team is focused on building general AI that can reason and act in the physical world. By leveraging Scale’s massive, industry-leading data infrastructure, we are partnering with frontier labs to build Foundation Models for Physical AI that will redefine the future of automation. To support our rapid hardware-software iteration cycles and ensure a world-class R&D environment, we are looking for a Safety Coordinator / Lab Lead to anchor our physical testing operations. Role Overview As the Safety Coordinator / Lab Lead , you will play a mission-critical role in scaling our physical testing infrastructure safely and efficiently. This is a high-impact position where your highest-priority responsibility will be owning the end-to-end execution of safety audits and incident documentation . Operating at the intersection of cutting-edge AI foundation models and complex robotics hardware, you will ensure our researchers, engineers, and autonomous systems interact in a secure, compliant, and highly organized environment. Core Responsibilities Priority Focus: Safety Audits & Incident Documentation Rigorous Safety Audits: Design, schedule, and execute routine safety audits across all physical testing environments, robot cells, and hardware workspaces to ensure continuous compliance with internal benchmarks and industrial safety standards. Incident & Near-Miss Documentation: Own the end-to-end incident management pipeline. Act as the primary point of contact for documenting, archiving, and analyzing any lab incidents, mechanical anomalies, or near-misses. Root-Cause Analysis (RCA): Lead structured post-incident investigations to identify systematic risks, authoring comprehensive RCA reports and implementing Corrective and Preventive Actions (CAPA). Data-Driven Risk Mitigation: Treat safety data as a core operational asset—tracking safety metrics and audit trends to proa

awsrestai
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
19 days ago

Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is growing quickly. You will play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features, while also balancing scalability, reusability, and performance. Role Overview: We're looking for a Staff Platform Engineer to serve as the technical backbone of our Engineering organization. You'll own the technical strategy, and delivery of our platform — spanning architecture, DevOps, SRE, security, Dev-ex. This is a hands-on staff level role: you'll set technical direction, drive cross-team alignment, and be the senior escalation point for platform challenges. What you’ll do: Drive platform reliability, scalability, security, and cost efficiency across all environments. Technical Leadership: Provide technical leadership for platform components, Influence the technical strategy and architecture of our cloud platform, from CI/CD pipelines to observability and incident response. Design and implement platform components and reusable integration patterns that minimize custom development efforts, reduce the time spent on repetitive tasks, and ensure that integrations scale across multiple healthcare systems Partner closely with Architecture, DevOps, SRE, and Security teams to deliver cohesive platform solutions Cross-Functional Collaboration: Work closely with product teams, and solutions architects to understand integration needs and ensure the platform meets current and future business requirements. Serve as a senior escalation point for infrastructure and platform incidents Establish frameworks for: AI governance and compliance. Observability of systems. Traceability of decisions and outputs. Ensure enterprise readiness with security, auditability, and reliability in production environments. Security & Compliance : Ensure all p

awsci/cdgit
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
19 days ago

Opportunity Overview: We’re looking for a Manager, Platform Engineering that can lead and grow a high-performing engineering team focused on Developer Experience, DevOps, SRE, and Quality. You will own the systems and processes that enable teams to build, test, release, and operate software with high velocity and reliability, driving engineering efficiency and operational excellence across the organization. What you’ll do: Lead a fast-paced, autonomous team of engineers focused on platform engineering, developer experience, DevOps, SRE, and quality engineering Own and drive the internal developer platform strategy and roadmap, improving how engineering teams build, test, deploy, and operate services Create transparency into engineering efficiency and system health through meaningful metrics across delivery, reliability, and quality Enable teams to move faster by improving CI CD pipelines, environments, tooling, and overall developer workflows Provide technical leadership across platform, infrastructure, and reliability, helping teams build scalable and resilient systems Ensure strong engineering practices across release processes, testing, quality, reliability, and security Define and enforce release guardrails, validation standards, and rollback mechanisms to improve production safety Improve environment stability and consistency across development, QA, and pre production environments Drive test strategy and automation maturity to improve overall product quality and confidence in releases Define and implement observability, monitoring, and alerting standards across systems Improve incident detection, response, and RCA practices, ensuring learnings translate into platform and system improvements Drive cloud infrastructure best practices across AWS, containers, and infrastructure as code Foster a culture of ownership, reliability, and continuous improvement within the team Provide innovative solutions for attracting, developing, and retaining top engineering talent I

DU
DoorDash USA
📍 San Francisco• Full-time• From $1.1M/yr
19 days ago

About the Team DoorDash's Protective Services function safeguards executives, their families, and residences through executive protection, residential security, threat management, and secure transportation across a three-brand global enterprise. The team operates 24/7 and is built on the principle that protection starts before an incident, not after. The people who do this work well combine operational discipline with personal composure, take ownership of outcomes rather than tasks, and represent the program in every interaction with the principal population. About the Role The Protective Services Agent provides close protection for DoorDash executives and their families across all operating environments, including corporate headquarters, domestic and international travel, company events, and residential settings. The role requires a trained executive protection professional who develops and implements security plans, conducts advances and risk assessments, coordinates with law enforcement and venue partners, and serves as the principal's primary security presence. Personnel must be capable of managing incidents and serving as first responders on scene. Operations regularly involve extended duty periods, irregular schedules, and short-notice deployments. Potential travel up to 25% of the time. You are excited about this opportunity because you will… Provide close protection across all principal environments. Serve as the primary security presence for executives and their families across headquarters, travel, events, and residential operations. Maintain continuous situational awareness, assess threats in real time, and execute protective protocols effectively and unobtrusively. Plan and execute security operations. Develop and implement comprehensive security plans for executive movements and events, incorporating site surveys, route planning, risk assessments, staffing requirements, screening procedures, and emergency response strategies. Conduct advances for all pr

awsgitrest
View job →
AG
Adani Group
📍 Ludhiana• Full-time
21 days ago

The EHS Officer shall be responsible for ensuring compliance with all Environment, Health & Safety (EHS) regulations, OISD guidelines, PESO requirements, statutory provisions, and company safety policies at the POL Terminal. The role involves implementation of safety management systems, incident prevention, environmental compliance, emergency preparedness, contractor safety management, and fostering a strong safety culture across the terminal. Responsible for Environment, Health, Fire Fighting System & Safety compliance of POL terminal in all respect. Source: Adani Group | Job ID: 58439

D
Datadog
📍 Remote, France• Full-time• Remote
23 days ago

We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, http://incident.io , Rootly, Jira Cloud Platform, Cortex, or similar and know how to navigate operational challenges of different

REMOTEpythonnode.jsai
View job →

Critical Support Operations Associate, Weekend Coverage Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Critical Support team is Stripe's frontline response function for urgent, high-stakes user issues. We operate at the intersection of hands-on support execution, incident response, and cross-functional coordination. Our mandate is to protect users in their most business-critical moments—resolving executive-triggered escalations, urgent and critical-priority cases, and complex issues that require more than a standard resolution path. We work closely with Internal Risk Management, Product, Engineering, and Legal to drive fast, coordinated outcomes. This is a greenfield team: we are building the operational foundations, workflows, and standards that will define critical support at Stripe for years to come. What you'll do As a Critical Support OA, you will be the first operational responder on Stripe's most urgent and time-sensitive user issues. You will execute against escalation cases across a wide range of complex scenarios—from executive-triggered requests and critical payment failures to urgent user escalations requiring immediate cross-functional coordination. This is a builder role: in addition to resolving cases, you will actively help shape the processes, playbooks, and operational standards that make this team function. You will develop a strong foundation in executive communication, complex problem-solving under pressure, and operational execution in an ambiguous, high-stakes

restaigo
View job →
AG
1mo ago

This role is pivotal in fostering a culture of safety and compliance within the plant. This role is responsible for implementing safety training programs, supporting in regular audits and inspections, and ensuring adherence to all safety regulations. The role will proactively manage safety initiatives, support incident investigations, and drive continuous improvement in safety performance to safeguard personnel and operations. Source: Adani Group | Job ID: 57769

aitraining
View job →

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<

pythonawsazure
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The SRE Leadership Team The SRE Leadership Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great infrastructure is invisible—it just works. Our team champions a culture of continuous learning, data-driven decision-making, and blameless incident response. We work at the intersection of product engineering, architecture, and operations to ensure Auth0 remains the trusted authentication platform for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with a focus on scalability, resilience, and empowering engineers to grow as technical leaders. What You'll Be Doing Lead the SRE team's technical direction , translating organizational vision into actionable roadmaps while driving complex, cross-functional initiatives across product and platform teams Operate at scale through hands-on participation in 24/7 on-call rotations (follow-the-sun weekdays, shared weekends), directly troubleshooting and remediating incidents on critical systems Build infrastructure resilience , designing and implementing monitoring, alerting, and automation improvements that reduce toil and elevate operational efficiency Champion reliability best practices , establishing policies and cultural standards that embed observability, resilience, and software engineering rigor into all engineering efforts Mentor and develop SRE talent , elevati

pythonawsazure
View job →

As Senior Data Scientist for Engineering Systems you will work independently alongside sharp, generous, and pragmatic engineers from Server Query, Atlas Clusters, and Release Quality, among other teams. Together, we tackle problems spanning resource scaling across the Atlas fleet, safe feature rollout to MongoDB clusters, automated incident response and query engine performance. Join the Platform Data Science team and help us research, prototype and ship machine learning features for MongoDB’s core server, query engine and Atlas, our database-as-a-service cloud offering. We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. What You'll Do Partner with Server Query, Atlas Clusters, Release Quality and other engineers to embed algorithmic rigor and optimization into resource scaling, release-safety and monitoring systems across the fleet and inside query engine Deliver production-ready, thoroughly tested statistical and ML algorithms with well-identified limitations that deliver measurable business impact, not just an impressive-sounding methodology Own the full feedback loop: instrument model architecture with the metrics needed to track performance and create dashboards in collaboration with our stellar analytics team, collect feedback from users and metrics to diagnose issues or opportunities, and iterate accordingly Deliver thoughtful, kind code reviews to your peers and act as a core contributor to internal packages, tooling, and team processes that increase developer productivity Measures of Success In 3 months, you’re familiar with our workflow, have an elementary understanding of our product and what teams we work with. You have delivered small-to-medium improvements to our project portfolio In 6 months, you’ve delivered one feature you researched and prototyped from scratch and demonstrated its impact on business metrics of your choice In 12 months, you've established a track record of shipping ML-driven improveme

pythonmongodbaws
View job →

Job description • You will perform the day-to-day data centre/computer operations (such as running the batch programs, printing of reports, systems and data backup, tapes/cartridges preparation and storage, etc). • You will be required to pro-actively monitor the data centre systems' uptime and performance (such as Server, Midrange and/or Mainframe systems, Networking equipment, Telecommunication connections, Email and Security Systems, Data Centre environment, etc) to prevent any down time or low performance. • You will diagnose and log down hardware problems into incident ticketing system, coordinate the problem resolution with the respective vendor or second level support personnel. • You will liaise with the vendors and other technical teams on any data centre related maintenance and installation activities. • You will also assist to provide helpdesk and basic IT troubleshooting support. • You will also assist in other IT related projects implementations. Job Requirements: • "A" Level Holders, ITE Graduates, Diploma and Degree Holders are welcome to apply. Those with at least 2 years of working experience will be considered for senior roles. • Experience in operating Windows Server, UNIX and/or AS/400 systems. • Good in work prioritization and able to maintain the service levels. • A detailed and systematic person, following the data centre operating procedures. • Willing to perform 12-hour Shift Work.

🔔

Get new incident commander jobs by email

Daily job updates · Unsubscribe anytime