Jobiba hiring network

Incident Commander Jobs

589 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current incident commander jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Core Services organization builds and runs the mission-critical online services that product teams rely on in production. We own foundational distributed systems and platform capabilities that enable reliable execution, high-performance services, and large-scale file/data needs across our products. This team is distinct from developer infrastructure and data infrastructure—our focus is production service foundations and core runtime services. About the Role We’re hiring an Engineering Manager, Core Services to help lead teams responsible for highly reliable, high-scale distributed systems that sit on the critical path for OpenAI products. Your team will own foundational production systems that OpenAI’s product engineering teams build on. You’ll collaborate closely with product and infrastructure partners to ship reliable services quickly, and help scale systems and teams as OpenAI grows. You’ll partner closely with senior engineering leaders to scale the org, mature operations, and drive major platform initiatives. This role requires strong technical ability. You’ll be responsible for: Managing and growing a high-performing team of infrastructure engineers. Leading teams building and operating large, critical production platforms, including cluster reliability, scaling, and rollout safety. Building and operating mission-critical distributed systems with strong operational rigor (SLOs, incident response, capacity planning, reliability). Setting technical direction for platform foundations such as workflow/orchestration capabilities, large-scale file/blob/storage services, and core service foundations. Partnering with a broad set of stakeholders, including product engineering, adjacent infrastructure teams, and (where relevant) finance/cost partners. Coaching, mentoring, and developing engineers and emerging leaders. You might thrive in this role if you: Have significant experience leading teams that run mission-critical infrastructure in production

awsrestai
View job →
G
Godaddy
📍 United Kingdom• Full-time
1mo ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About The Team The Commerce Site Reliability Engineering team is responsible for the reliability, scalability, and day-to-day operation of the platforms that power GoDaddy's Commerce ecosystem. We build and operate shared infrastructure, support critical production systems, and partner closely with engineering teams to ensure services remain secure, resilient, and highly available. As a Senior Site Reliability Engineer, you'll join a team that values ownership, operational excellence, and continuous improvement. Engineers are empowered to identify problems, drive meaningful change, and influence how reliability is delivered across the broader Commerce organisation. From improving operational maturity and reducing toil to modernising delivery platforms and strengthening incident response practices, this team plays a key role in enabling engineering teams to move quickly and safely. You'll work closely with engineers across infrastructure, cloud, security, networking, and application teams while helping shape the future of reliability engineering at GoDaddy. This role offers significant opportunity to broaden your impact, develop technical leadership skills, and grow toward Staff and Principal engineering positions over time. What you'll get to do... Lead reliability and operational improvement initiatives across GoDaddy's Commerce platform, helping engineering teams build and operate services safely and at scale. Own critical production systems, drive incident response and post-incident improvements, and continuously raise the bar

typescriptpythonaws
View job →
S
Smartsheet
📍 Bengaluru• Full-time
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. As the Director of Engineering at Smartsheet India, you will build capabilities to empower the world's largest companies to transform their approach to work. You will guide teams that own the grid ecosystem - defining how data linking, synchronization, and grid infrastructure evolve as a cohesive platform. You will ensure architectural decisions are coherent and avoid fragmentation. You will be willing to challenge technical choices. Platform Reliability & Operational Excellence: You will be accountable for the availability and performance of foundational services that other teams depend on. Drive a high bar for on-call health, incident response, and SLA/SLO definition across all the services. You will manage cross-pillar/cross-domain dependencies, negotiate API contracts, and prevent the grid ecosystem from becoming a delivery bottleneck. You will balance the needs of user-facing product features with infrastructural stability and operational health. You will ensure career growth paths are clear for engineers across that spectrum, and develop a strong sense of customer centricity and pillar identity for your teams. You will be comfortable accepting responsibility for impact, service availability, and the effectiveness of your teams. You will be comfortable being at the forefront of AI adoption for delivery and operations, leaning in and helping the team leverage AI for optimum delivery in their ways of working. You are passionate about continuous improvement and have built learning organizations that keep up w

sqlawsagile
View job →
C
Coinbase
📍 Canada• Full-time• Remote• From C$154K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As an Offensive Security Engineer on the Application Security team within Security, you'll pioneer how Coinbase uses frontier AI models to scale vulnerability discovery, red team AI systems, and automate security workflows. This role bridges traditional penetration testing with next-generation AI-augmented offensive security, directly accelerating our ability to protect products and customers. You'll own the development of AI-driven security tooling and collaborate across Vulnerability Management, Offensive Security, and Incident Response to fundamentally shift how we operate. What you'll do: Build, deploy, and maintain custom security scanners that leverage frontier models to detect vulnerabilities at scale across Coinbase's product surface. Lead red teaming efforts against internal AI systems, including jailbreak testing, prompt injection analysis, and tool abuse simulation. Develop AI-driven automation for vulnerability triage, validation, and remediation workflows to accelerate the bug bounty and vulnerability response pipelines. Partner with engineering teams to prioritize, remediate, and verify fixes for critical vulnerabilities discovered through AI-augmented and manual testing. Mentor junior security engineers on integrating AI into offensive security workflows to scale team capabilities. Required Skills and Experience: 3+ years of experience in application s

REMOTEpythonawsai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability Team Focus Areas Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely. Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads. Databases: Building high performance, distributed database systems that power all of OpenAI's product stack. In this

pythonawskubernetes
View job →
O
Okta
📍 San Francisco• Full-time• From $165K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in a global on-call rotation supporting highly available customer-facing systems. Participate in incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with en

pythonsqlpostgresql
View job →
S
Sentry
📍 Toronto• Full-time• Remote• C$162K – C$420K/yr
1mo ago

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About The Role The Security Team is responsible for securing all things Sentry: our customers, our code, and everything in between. We are a small but growing team with broad scope, high trust, and the autonomy to tackle hard security problems with creativity and an engineering mindset. We work at a company with a strong developer culture, building a product that millions of developers genuinely love and rely on. That context shapes everything about how we operate. We take a pragmatic approach to preventing and responding to security risks. In this role not only will you build and contribute to systems which detect malicious activity, you will have the unique opportunity to implement new controls to prevent future incidents. You will work across detection and response and corporate security domains. You'll contribute to practices that keep Sentry secure as we grow: alert triage for corporate and production, detection engineering, deploying preventative controls, identity and access management, investigations and incident response, and more. You'll partner with teams across the company to prevent and respond to security incidents. You will work as a technical collaborator who prioritizes preventative controls, defense in depth, and high signal alerting practices. As Sentry expands our agentic product capabilities and development practices, you'll also find yourself at the frontier of a new set of security approaches and challenges. In this role, you will Maintain, improve, and own detection engineering systems. We own and operate our own detection stack and are building agentic triage with thoughtful security response and orc

REMOTEpythonreactaws
View job →
S
Sentry
📍 San Francisco• Full-time• Remote• $155K – $400K/yr
1mo ago

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About The Role The Security Team is responsible for securing all things Sentry: our customers, our code, and everything in between. We are a small but growing team with broad scope, high trust, and the autonomy to tackle hard security problems with creativity and an engineering mindset. We work at a company with a strong developer culture, building a product that millions of developers genuinely love and rely on. That context shapes everything about how we operate. We take a pragmatic approach to preventing and responding to security risks. In this role not only will you build and contribute to systems which detect malicious activity, you will have the unique opportunity to implement new controls to prevent future incidents. You will work across detection and response and corporate security domains. You'll contribute to practices that keep Sentry secure as we grow: alert triage for corporate and production, detection engineering, deploying preventative controls, identity and access management, investigations and incident response, and more. You'll partner with teams across the company to prevent and respond to security incidents. You will work as a technical collaborator who prioritizes preventative controls, defense in depth, and high signal alerting practices. As Sentry expands our agentic product capabilities and development practices, you'll also find yourself at the frontier of a new set of security approaches and challenges. In this role, you will Maintain, improve, and own detection engineering systems. We own and operate our own detection stack and are building agentic triage with thoughtful security response and orc

REMOTEpythonreactaws
View job →
A
1mo ago

At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale's need to detect and respond to security events across its production and corporate environments is growing as the company scales. We're looking for a Senior Detection and Response Engineer to own detection engineering and to lead incident response when it counts, coordinating the response and driving it to resolution. This is a high-ownership role with real room to shape how detection and response works at Anyscale. You will own the detection pipeline, the response runbooks, and incident response, reporting to the Head of Security and partnering with engineering. This role is based in India. In your first year, success looks like strong detection coverage across our cloud, endpoint, and runtime telemetry, a working correlation and alerting pipeline, and incident response runbooks that have been exercised in practice. What You'll Do Own and build detection coverage across cloud, endpoint, and runtime telemetry. Own a centralized correlation and alerting capability that turns telemetry into actionable detections. Own incident response: runbooks, escalation paths, and coordination during an incident, across corporate and production environments. Drive detection of anomalous activity across the environments

awsazurekubernetes
View job →

Job Details: Job Description: Job Description The Role and Impact As an Industrial Hygienist, you will play a critical role in safeguarding employee health and ensuring workplace safety. You will drive industrial hygiene risk assessments, oversee exposure evaluations, approve chemical usage, and coordinate monitoring of workplace environments. Your expertise will directly contribute to maintaining a safe and compliant work environment while addressing health risks associated with workplace conditions and chemical usage. Business group You will be part of a team dedicated to environmental health and safety within Intel's broader organization. This group focuses on driving compliance with applicable industrial hygiene regulations, developing global policies, and ensuring alignment with Intel's high standards. By fostering a safe and healthy work environment, the team supports Intel's commitment to employee well-being and operational excellence. Key Responsibilities - Conduct qualitative risk assessments and personal protective equipment evaluations for workplace environments. - Evaluate chemical exposure implications and approve their usage based on toxicology data. - Monitor workplace environments to identify and mitigate health risks. - Ensure compliance with local industrial hygiene regulations and Intel's internal standards. - Support site program self-assessments, follow-up plans, and respond to incident investigations. - Participate in project reviews, pre-startup safety reviews, and hazard evaluations. - Conduct surveys of working conditions to assess hazardous exposures and control measure effectiveness. - Develop industrial hygiene policies, programs, and procedural documentation in alignment with global and regional standards. - Investigate employee reports of un

recruitment
View job →

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What you will be doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuous

pythondockerkubernetes
View job →

NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services. What you’ll be doing: Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale. Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation. Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data. Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems. Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation. What we need to see: BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10&#43; years operating production infrastructure or platform services. Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges. <l

pythondockerkubernetes
View job →

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About The Team.... Global Compute runs Optimised Hosting, GoDaddy's global platform for all customer hosting products. Squad R is the engineering team responsible for operating, scaling, and continuously improving the OpenStack-based clouds that power that platform. We treat reliability as an engineering problem: we automate toil away, we plan capacity ahead of demand, and we instrument everything so that we understand our systems before they surprise us. As an SRE III on the team, you'll be a senior technical contributor who others lean on for the hard problems. What you'll get to do... Operate and scale GoDaddy's cloud infrastructure, including our OpenStack-based hosting platform. You'll troubleshoot and improve services spanning compute, networking, and storage in large-scale production environments. Drive the OpenStack migration. Help move customer hosting workloads onto the platform safely — designing and executing migration tooling, validation, and rollback strategies that protect customer experience. Work within a large-scale global hosting environment supporting thousands of servers and customer workloads across multiple regions. Eliminate toil through automation. Build and maintain automation in Python and Puppet to replace manual operational work. Treat repeated manual effort as a bug to be fixed. Strengthen observability. Improve monitoring, alerting, and dashboards so that signal reaches the right engineer at the right time, and so that we can reason about system behavior from data. Participate in on-call and incident response. T

pythondockerlinux
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput, and 99.999 availability. We're looking for a technical leader to help us to continue to scale the service with great people and reliable, cost-effective and efficient infrastructure, processes and tooling. As the Director of Site Reliability Engineering you will oversee the SRE organization focused on Okta platform, Databases, Edge networking, K8s platform, CI/CD, Observability, FinOps, and automation platform & tooling. Job Duties and Responsibilities: Build and lead a high-caliber India-based SRE organization supporting Okta’s production fleet. Partner with global engineering, product, and infrastructure leaders to deliver resilient, scalable, and secure services. Define and execute the India SRE strategy in alignment with global reliability goals. Lead post-incident reviews, drive root-cause analysis, and ensure long-term corrective actions. Participate in incident management, on-call rotations, and blameless RCAs. Implement automation and observability to reduce manual toil and improve operational efficiency. Drive adoption of modern infrastructure practices: infrastructure as code (Terraform), container orchestration (Kubernetes), and AI within Infrastructure org. H

awskubernetesmachine learning
View job →

NVIDIA pioneers computer graphics, gaming, AI, and accelerated computing. We are looking for a Technical Platform Operations Lead to join our team and play an important role in scaling Sales AI applications and platforms. This position offers the opportunity to shape how these solutions operate after launch and help ensure they remain reliable, secure, well governed, widely adopted, and continuously improved. You will collaborate with Sales, Product, Engineering, Data, Security, and IT teams to strengthen platform health, improve the user experience, and increase business impact. What you’ll be doing: Lead end-to-end post-launch operations for Sales AI applications, including availability, performance, support readiness, releases, upgrades, and lifecycle planning. Develop effective processes for incident response, problem management, changes, and issue resolution. Coordinate timely recovery and lasting improvements. Analyze service-level indicators and objectives, adoption metrics, dashboards, alerts, and user feedback to identify risks, performance degradation, and usage gaps. Collaborate with partner teams to translate operational signals and user needs into prioritized improvements and roadmap inputs. Improve adoption and business value through usage analytics, enablement, feedback loops, and user experience enhancements. Establish governance practices for security, access controls, compliance, documentation, and platform support. Develop automation, observability, and self-service capabilities that simplify operations and reduce repetitive work and recurring incidents. Prepare new AI capabilities and releases for production with runbooks, monitoring, rollback plans, support models, and partner enablement. What we need to see: 8&#43; years of experience in technical operations, pl

🔔

Get new incident commander jobs by email

Daily job updates · Unsubscribe anytime