We are seeking a Staff Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Staff Site Reliability Engineer on this new team, you will be responsible for providing technical leadership for the operational foundations that enable deployment at scale of AI applications. You will own the reliability architecture of the platform as it expands across regions and cloud providers, and set the technical direction for how the platform is operated, including capacity planning, multi-cloud expansion, incident response, and SLO discipline. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Bengaluru for our hybrid working model. Position Expectations Own the reliability architecture of the platform across regions and cloud providers Collaborate with the teams building the platform, providing internal support and guidance on operability, capacity, and best practices Set operational standards for the team: on-call quality, incident response, SLO discipline Mentor and technically develop the SRE team Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure Qualifications 10+ years of experience working on software and operating distributed systems, with deep Kubernetes expertise, including designing or evolving multi-cluster platforms Proficiency in Python, Go, or a similar programming language Understand workload isolati
Jobiba hiring network
Incident Commander Jobs
589 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current incident commander jobs. Use filters to narrow by work mode, employment type, experience and date posted.
As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams. Success in this role means smoother launches, clearer roadmaps, stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale. This role can be based remotely on the East Coast What You'll Do Drive Program Planning & Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep work clearly tracked in Jira, and deliver on time Strengthen Production Reliability – Lead change management and launch readiness programs. Partner with SREs and product teams to define and operationalize SLOs/SLIs, and use incident data, metrics, and capacity signals to drive prioritization and continuous improvement Lead Cross-Functional Coordination – Align SRE with Security, Compliance, Cloud platform, and other engineering teams. Coordinate cross-team incident response, ensure clear follow-through, and build trust as the go-to driver of complex, multi-team efforts Build Scalable Systems & Processes – Design lightweight frameworks and communication patterns that help SRE deliver reliably at scale. Work yourself out of the "hero" role by leaving teams better-equipped to execute independently Requirements 8+ years in technical program management, engineering management, or a comparable technical role partnering with software engineering teams Proven track record leading large-scale, cross-team platform initiatives through ambiguity and change Strong knowledge of production change management, software development lifecycle, and reliability metrics (SLOs, SLIs) Skilled at shaping roadmaps and managing dependencies Able to query and interpret metrics, logs, or other data s
At MongoDB, we are disrupting industries and equipping developers with the tools to create extraordinary applications that impact daily lives. As the leading developer data platform and the pioneering database vendor to execute an IPO in over two decades, we offer an environment at the cutting edge of creativity and innovation. The international data management software landscape is vast and expanding swiftly. Data from IDC indicates the global database management systems (DBMS) software market was valued at roughly $82 billion in 2023 and is projected to scale to approximately $200 billion by 2030, demonstrating an exceptional compound annual growth rate. We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. The Role & Team Overview We are seeking an accomplished technical leader to join our Technical Services organization as the head of the EMEA Incident & Escalation Management teams. In this critical capacity, you will guide the teams responsible for ensuring our clients can smoothly operate MongoDB at scale and seamlessly navigate any technical hurdles. This position demands a leader who can maintain a hands-on approach—taking direct ownership of high-profile escalations that require coordinated corporate responses—while concurrently expanding the team, refining processes, and advancing strategic objectives within a global critical situations management framework. Team Overview: As an integral component of the Customer Organization, Technical Services boasts more than 500 global members across worldwide offices. We maintain industry-leading customer satisfaction metrics through a 24x7x365 'follow-the-sun' support framework, with specialized regional teams across the Americas, EMEA, and APAC. The EMEA team operates out of hubs including Dublin, Milan, Barcelona, and Tel Aviv. Core Responsibilities Collaborating with global leadership to architect and execute strategic plans for critical situations management,
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Team The Auth0 Platform Tools team owns the incident management tooling, Slack-based tooling, StatusPage, and local development environments that Auth0 engineers rely on every day. That includes incident.io and the services we have built around it, Statuspage, custom Slack bot applications that automate our incident response and engineering operations workflows, the customer-facing web application behind status.auth0.com, Vivaldi, and Tilt - the tools engineers use to run Auth0 locally. We are seeking an engineer to help build new features across all of these tools. Our stack is primarily TypeScript and Node.js, with a React and Next.js front end, backed by Postgres and Redis, and deployed on Kubernetes on AWS. A significant portion of our incident and engineering operations automation is built on Tines, a no-code automation platform. Prior no-code experience is welcome, but we expect you to learn Tines here and become effective with it. Current initiatives include extending our incident tooling to meet FedRAMP requirements, taking full ownership of the status page, and improving how we communicate incident status to customers. There is real room to improve along the way, from test coverage to resilience to inherited technical debt. We build for two audiences: Auth0 engineers, who depend on our tooling every day, and Auth0's customers, who rely on the status page during incidents. We are looking for an engineer who cares about both and enjoys working wi
The Cyber Deployment Manager partners with customers throughout the full lifecycle—from technical discovery and solution design through implementation, deployment, and adoption. You’ll work closely with Sales and Solutions Engineering during the pre-sales process to understand customer security priorities, assess technical requirements, develop solution architectures, and support demonstrations, workshops, and proofs of concept. After a customer commits, you’ll remain engaged as the technical deployment lead, translating the proposed solution into a production-ready implementation. You’ll guide integrations, establish success criteria, manage technical risks, and help customers operationalize AI across security workflows such as secure code review, vulnerability management, threat detection, incident response, SOC operations, and GRC automation. In this role, you will: Lead technical discovery with security executives, practitioners, architects, and engineering teams. Partner with Sales and Solutions Engineering on solution design, demonstrations, workshops, technical validation, and proofs of concept. Translate customer requirements into clear architectures, deployment plans, success criteria, and implementation milestones. Own the transition from pre-sales solution design into post-sales deployment and adoption. Serve as the primary technical partner during implementation, coordinating customer stakeholders and internal Product, Engineering, Security, and GTM teams. Build and troubleshoot integrations involving APIs, agents, security tools, cloud platforms, data sources, and enterprise workflows. Identify deployment risks, technical blockers, and product gaps, and drive them toward resolution. Help customers establish evaluation frameworks, governance controls, guardrails, monitoring, and human-review processes. Measure adoption and business impact, ensuring deployed solutions deliver meaningful security outcomes. Turn successful customer deployments into reusable
About the Team The Corporate Security team supports the safety, security, and resilience of OpenAI employees globally. The Travel Security function helps employees travel safely and confidently while enabling business activity in a complex and fast-changing operating environment. About the Role As Travel Security Manager, APAC, you will own the regional delivery and continued development of OpenAI's travel security program across Asia-Pacific. You will support the full travel risk management lifecycle, from pre-travel assessment and preparation through active-trip monitoring, incident response, traveler accountability, and post-incident learning. Reporting to the Global Travel Security Manager, you will be the regional lead for travel security matters across APAC. You will work closely with Corporate Security, Protective Intelligence, Finance, People, Legal, Immigration, Global Mobility, Communications, Information Security, and external providers. This is a hands-on role for a seasoned travel security professional. You will independently manage complex cases, advise leaders on risk-based travel decisions, support travelers during incidents and crises, and build scalable regional processes. Broad APAC experience is important, and deeper operating experience in East Asia is particularly valuable. Travel across APAC and occasionally to other regions is expected to comprise approximately 25% of the role. In this role, you will: Own day-to-day travel security operations across APAC, including high-risk travel reviews, traveler briefings, monitoring, escalation, incident response, and traveler accountability. Produce clear, intelligence-led risk assessments, advisories, decision briefs, and practical mitigation options for employees and senior stakeholders. Work with intelligence analysts and external providers to monitor geopolitical, security, health, environmental, and transport risks and turn changing information into timely operational decisions. Lead the travel sec
About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. Our team builds the software, tooling, and operational systems that help manage this fleet at scale. We work across production engineering, distributed systems, capacity management, and operational automation to improve reliability, reduce manual work, and make better use of available compute. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will develop the systems that help manage the GPU fleet powering ChatGPT, including tooling for fleet health, capacity planning, operational automation, and incident response. You will work closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and compute utilization. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference. Develop systems for capacity planning, fleet health monitoring, and resource utilization. Automate operational workflows, including incident detection, diagnosis, and response. Identify and address bottlenecks affecting fleet reliability, scalability, and performance. Partner with infrastructure, research, and product engineering teams to improve the compute platform. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems. Have a background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering. Have built software that automates operational workflows and reduces manual work. Have worked with distributed infrastructure, cluster orchestration, or large-scale int
About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A
About the Team OpenAI’s Security organization exists to enable safe, responsible innovation at scale. As our systems, infrastructure, and research footprint grow, we invest deeply in world-class security capabilities that protect our people, products, and users without slowing progress. This organization safeguards OpenAI’s environments by building advanced detection systems, driving real-time response capabilities, scaling telemetry and logging infrastructure, and delivering actionable threat intelligence to stay ahead of adversaries. About the Role We are seeking a Global Detection and Response Lead to own and scale OpenAI’s cybersecurity detection and response operations. In this role, you will set the strategy and drive execution for security monitoring, incident response, recovery, and post-incident improvements across our global infrastructure. You will be a hands-on leader with deep technical credibility and strong operational instincts. You will build and mentor high-performing teams, partner closely with Infrastructure, Research, Product Security, Enterprise Security, IT, and Engineering, and ensure that detection and response capabilities are embedded by design into the systems that power OpenAI. This is a strategic and practical leadership role requiring deep technical credibility, operational rigor, and the ability to build high-performing teams in a fast-moving environment. In this role, you will: Oversee global detection and response operations, including continuous monitoring, triage, investigation, containment, and remediation of security events across a diverse set of networks and infrastructure. Lead, mentor, and directly manage several small teams of senior engineers across observability, detection and response, and threat intelligence. Hire and scale these functions deliberately and proportionately as OpenAI’s compute footprint and platform ambitions grow. Ensure world-class operational rigor and readiness through management of incident playbooks
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role Millions of people rely on Notion to do their most important work, and protecting that trust is foundational to everything we build. We’re looking for a hands-on Detection Engineer to build and operate the systems and workflows we use to detect and respond to attacks across Notion’s cloud-native environment. You’ll ship high-signal detections, improve the platform that powers them, participate in incident response, and help shape how detection and response engineering scales at Notion. You’ll work closely with Engineering, Corporate Security, and Infrastructure, with broad latitude to identify gaps, prioritize investments, and build what’s needed next. We view detection and response as a software engineering discipline: detections are code, platforms are products, and measurement matters What You'll Achieve Design and maintain high-signal detections across cloud, identity, endpoints, and SaaS environments. Build and improve the detection platform, including rule lifecycle management, tuning, measurement, and rollout safety. Develop tooling and automation that accelerate triage, enrichment, investigation, and detection
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role Millions of people rely on Notion to do their most important work, and protecting that trust is foundational to everything we build. We’re looking for a hands-on Detection Engineer to build and operate the systems and workflows we use to detect and respond to attacks across Notion’s cloud-native environment. You’ll ship high-signal detections, improve the platform that powers them, participate in incident response, and help shape how detection and response engineering scales at Notion. You’ll work closely with Engineering, Corporate Security, and Infrastructure, with broad latitude to identify gaps, prioritize investments, and build what’s needed next. We view detection and response as a software engineering discipline: detections are code, platforms are products, and measurement matters What You'll Achieve Design and maintain high-signal detections across cloud, identity, endpoints, and SaaS environments. Build and improve the detection platform, including rule lifecycle management, tuning, measurement, and rollout safety. Develop tooling and automation that accelerate triage, enrichment, investigation, and detection
The Senior Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, performance, and operability of production systems across our platforms, by applying software engineering practices to operations, with a focus on automation, observability, and incident response.
Seceon is a trusted Ransomware Detection Company that helps organizations identify, prevent, and respond to ransomware attacks before they disrupt critical operations. Powered by AI-driven threat detection, real-time monitoring, and automated incident response, Seceon delivers comprehensive cybersecurity protection across cloud, network, and endpoint environments. Its advanced platform detects suspicious behavior, blocks malicious activity, and minimizes the risk of data loss or business downtime. By combining proactive threat intelligence with continuous security monitoring, Seceon enables businesses of all sizes to strengthen cyber resilience, maintain compliance, and defend against evolving ransomware threats with confidence. https://www.seceon.com/ransomware-detection/
Seceon is a trusted Ransomware Detection Company that helps organizations identify, prevent, and respond to ransomware attacks before they disrupt critical operations. Powered by AI-driven threat detection, real-time monitoring, and automated incident response, Seceon delivers comprehensive cybersecurity protection across cloud, network, and endpoint environments. Its advanced platform detects suspicious behavior, blocks malicious activity, and minimizes the risk of data loss or business downtime. By combining proactive threat intelligence with continuous security monitoring, Seceon enables businesses of all sizes to strengthen cyber resilience, maintain compliance, and defend against evolving ransomware threats with confidence. https://www.seceon.com/ransomware-detection/
Seceon is a trusted Ransomware Detection Company that helps organizations identify, prevent, and respond to ransomware attacks before they disrupt critical operations. Powered by AI-driven threat detection, real-time monitoring, and automated incident response, Seceon delivers comprehensive cybersecurity protection across cloud, network, and endpoint environments. Its advanced platform detects suspicious behavior, blocks malicious activity, and minimizes the risk of data loss or business downtime. By combining proactive threat intelligence with continuous security monitoring, Seceon enables businesses of all sizes to strengthen cyber resilience, maintain compliance, and defend against evolving ransomware threats with confidence. https://www.seceon.com/ransomware-detection/
Get new incident commander jobs by email
Daily job updates · Unsubscribe anytime