The Team: We are Datadog's in-house product experts. The technical solutions team enables Datadog's worldwide growth by educating potential clients and ensuring that existing customers are happy and successful. We share our technical and product expertise with customers through demos, presentations, technical evaluations, and ongoing support. Technical solutions is a growing global team that collaborates constantly to share knowledge and continuously advance our technical skillset. The Opportunity: Join a diverse team of traditional and non-traditional backgrounds, working together to solve complex problems the right way. You will lead a team immersed in a startup-like environment where you will be challenged, but also will immediately witness your contributions to Datadog. You Will: Manage, develop, and mentor a fast-paced team of Solutions Engineers who respond to client requests, reproduce and troubleshoot issues, and dive into the 400+ integrations that Datadog works with Ensure successful onboarding of new Solutions Engineers Help triage Solutions Engineering queues and work with Product and Engineering on urgent matters Review and help prioritize Solutions Engineering escalations Dispatch customer requests throughout the day as part of a queue rotation schedule Oversee demo training and assist in demo certification Assist with incident response during outages/incidents, communicating with customers and providing our internal teams with info that we’re getting first-hand from those customers Build out documentation and knowledge base articles for a variety of technologies You Are: Passionate about people management and/or mentorship with previous experience leading a team Self-motivated, detail-attentive, and have a desire for continuous learning A critical thinker who defaults to a client-centric approach A tinkerer with some programming experience and a basic knowledge of Linux Able to work a rotating schedu
Jobiba hiring network
Incident Commander Jobs
589 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current incident commander jobs. Use filters to narrow by work mode, employment type, experience and date posted.
We are Datadog's in-house product experts. The technical solutions team enables Datadog's worldwide growth by educating potential clients and ensuring that existing customers are happy and successful. We share our technical and product expertise with customers through demos, presentations, technical evaluations, and ongoing support. Technical solutions is a growing global team that collaborates constantly to share knowledge and continuously advance our technical skillset. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Manage, develop, and mentor a fast-paced team of Solutions Engineers who respond to client requests, reproduce and troubleshoot issues, and dive into the 1000+ integrations that Datadog works with Ensure successful onboarding of new Solutions Engineers Help triage Solutions Engineering queues and work with Product and Engineering on urgent matters Review and help prioritize Solutions Engineering escalations, and follow up with customers in Japan Dispatch customer requests throughout the day as part of a queue rotation schedule Oversee demo training and assist in demo certification Assist with incident response during outages/incidents, communicating with customers and providing our internal teams with info that we’re getting first-hand from those customers Build out documentation and knowledge base articles for a variety of technologies Who You Are: Passionate about people management and/or mentorship with previous experience leading a team Self-motivated, detail-attentive, and have a desire for continuous learning A critical thinker who defaults to a client-centric approach A tinkerer with some programming experience and a basic knowledge of Linux Fluent in written and spoken English and Japanese as this role will
We are Datadog's in-house product experts. The technical solutions team enables Datadog's worldwide growth by educating potential clients and ensuring that existing customers are happy and successful. We share our technical and product expertise with customers through demos, presentations, technical evaluations, and ongoing support. Technical solutions is a growing global team that collaborates constantly to share knowledge and continuously advance our technical skillset. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Manage, develop, and mentor a fast-paced team of Support Engineers who respond to client requests, reproduce and troubleshoot issues, and dive into the 1000+ integrations that Datadog works with Ensure successful onboarding of new Support Engineers Help triage Support Engineering queues and work with Product and Engineering on urgent matters Review and help prioritize Support Engineering escalations, and follow up with customers in Korea Dispatch customer requests throughout the day as part of a queue rotation schedule Oversee demo training and assist in demo certification Assist with incident response during outages/incidents, communicating with customers and providing our internal teams with info that we’re getting first-hand from those customers Build out documentation and knowledge base articles for a variety of technologies Who You Are: Passionate about people management and/or mentorship with minimum 3 years of experience leading a team Curious, self-motivated, detail-attentive, and have a desire for continuous learning A critical thinker who defaults to a client-centric approach A tinkerer with some programming experience and a basic knowledge of Linux Business or native-level fluency in English and Korean (
As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams. Success in this role means smoother launches, clearer roadmaps, stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale. This role can be based out of our Dublin or Cork office or remotely in Ireland. What You'll Do Drive Program Planning & Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep work clearly tracked in Jira, and deliver on time Strengthen Production Reliability – Lead change management and launch readiness programs. Partner with SREs and product teams to define and operationalize SLOs/SLIs, and use incident data, metrics, and capacity signals to drive prioritization and continuous improvement Lead Cross-Functional Coordination – Align SRE with Security, Compliance, Cloud platform, and other engineering teams. Coordinate cross-team incident response, ensure clear follow-through, and build trust as the go-to driver of complex, multi-team efforts Build Scalable Systems & Processes – Design lightweight frameworks and communication patterns that help SRE deliver reliably at scale. Work yourself out of the "hero" role by leaving teams better-equipped to execute independently Requirements 8+ years in technical program management, engineering management, or a comparable technical role partnering with software engineering teams Proven track record leading large-scale, cross-team platform initiatives through ambiguity and change Strong knowledge of production change management, software development lifecycle, and reliability metrics (SLOs, SLIs) Skilled at shaping roadmaps and managing dependencies Able to query and interpret
As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental research problems and collaborating with Datadog's product and engineering teams to translate research advances into products. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Conduct research in generative AI and machine learning, building specialized foundation models and trained agents for observability Train multimodal models on large-scale, diverse telemetry data (metrics, logs, traces, topology, events) using distributed training infrastructure Design and build simulated environments and RL training loops for on-policy agent training and evaluation Collaborate with cross-functional teams (Product, Engineering) to integrate capabilities like multimodal world modeling and autonomous agents into Datadog's products Stay at the forefront of foundation models, world models, and RL-based agent research Contribute to r
As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental research problems and collaborating with Datadog's product and engineering teams to translate research advances into products. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Conduct research in generative AI and machine learning, building specialized foundation models and trained agents for observability Train multimodal models on large-scale, diverse telemetry data (metrics, logs, traces, topology, events) using distributed training infrastructure Design and build simulated environments and RL training loops for on-policy agent training and evaluation Collaborate with cross-functional teams (Product, Engineering) to integrate capabilities like multimodal world modeling and autonomous agents into Datadog's products Stay at the forefront of foundation models, world models, and RL-based agent research Contribute to r
As Engineering Manager for Threat Detection, you will lead a high-performing team that powers Datadog's detection program. Threat Detection is the organization responsible for keeping Datadog ahead of an evolving threat environment: closing coverage gaps faster, raising the bar on signal quality, and shipping detections that hold up under the scale and complexity of cloud-native infrastructure. Your team will combine direct detection expertise, platform engineering, and applied AI to ship detections at a pace and scale traditional rule-writing alone cannot match. Examples of what your team will work on include detection-authoring agents, the detection platform that powers every rule in production, coverage analysis, alert triage and response automation, and the evaluation infrastructure that holds these systems to a high bar of fidelity. Detection authorship is a shared responsibility across the organization, and your team will contribute both by building the systems that scale our authoring capacity and by writing detections directly when their domain expertise is the right tool. You will partner closely with our Security Incident & Response Team (SIRT), Cyber Threat Intelligence (CTI), AI Engineering teams, and Datadog's broader Security organization. This is a high-impact leadership role: you will grow a team of security and software engineers responsible for building and executing our detection and AI strategy. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the strategy, roadmap, and execution of Datadog Security's shift to AI-accelerated detection and response. Drive development of high-fidelity detections as a shared responsibility across the organization, ensuring your team's systems and direct contributions raise the bar on coverage and
As a Research Engineer on our team, you will partner with Research Scientists to turn research ideas into working systems, building the data, tooling, and infrastructure that enable rapid iteration, trustworthy evaluation, and a smooth path from prototype to production. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Build and operate multimodal data pipelines, training and evaluation infrastructure, benchmarks, and internal tooling Implement models, run experiments at scale, and profile for reliability, performance, and cost Build simulation environments and replay infrastructure for agent training and evaluation Orchestrate distributed training and distributed RL with Ray, including scheduling, scaling, and failure recovery Establish rigorous automated benchmarks and regression tests for world model predictions, agent performance, and simulation fidelity Collaborate with Research Scientists, Product, and Engineeri
AI agents are transforming the way developers interact with software - and databases are no exception. We're seeking a Senior Software Engineer to join our AI Interfaces team within AI Builder Experience (ABX), where you'll provide technical direction, shape architecture, and build the core products that make it seamless for developers and AI agents to work with MongoDB. Our team owns the surfaces through which humans and agents connect to MongoDB - including the MongoDB MCP Server, Agent Skills, our Intelligent Assistant Platform, and purpose-built agents. In short: if it's how an agent talks to MongoDB, we're building it. This is a new team charting new territory, and as a Senior Software Engineer here, you'll be right at the frontier - building the technologies that let AI applications and agentic workflows work seamlessly with MongoDB at scale. You'll integrate with fast-moving, often unproven technologies, make pragmatic calls in the face of ambiguity, and own high-visibility projects end to end with minimal guidance. We're looking for product-minded engineers who thrive on autonomy and take pride in shipping. MongoDB engineering teams pride themselves on building high-quality software and living our cultural values every day - we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. This position requires participation in a 24/7 on-call rotation to ensure business continuity and incident response capabilities. This role can be based out of our Gurugram office. Position Expectations Work closely with research, product management, product engineering, product design, peers, as well as other teams within the company to define the first version and future evolution of our AI interfaces Design, build, and deliver well-tested core pieces of the platform - including the MongoDB MCP Server, Agent Skills, the Intelligent Assistant Platform, and purpose-built agents - in collaboration with othe
MongoDB’s mission is to empower innovators to create, transform, and disrupt industries by unleashing the power of software and data. We enable organizations of all sizes to easily build, scale, and run modern applications by helping them modernize legacy workloads, embrace innovation, and unleash AI. Our industry-leading developer data platform, MongoDB Atlas, is the only globally distributed, multi-cloud database and is available in more than 115 regions across AWS, Google Cloud, and Microsoft Azure. Atlas allows customers to build and run applications anywhere—on premises, or across cloud providers. With offices worldwide and over 175,000 new developers signing up to use MongoDB every month, it’s no wonder that leading organizations, like Samsung and Toyota, trust MongoDB to build next-generation, AI-powered applications. The Escalation Manager is a critical role within Technical Services. As a member of our global Incident and Escalation Management team, they work internally with our Engineering, Services, Sales and Product Management teams, as well as externally with customers and partners, to coordinate and drive the resolution of critical technical issues and incidents. Transparency is key and is achieved by providing timely and accurate updates to senior management regarding active escalations, as well as important detail on the status of the customer account. Individuals in this role are highly organized, proactive and professional. You are one who excels in fast-paced environments and can assess business impact, mobilize cross-functional teams, and drive technical escalations with urgency and ownership. We are looking for someone who has a customer-focused mindset with excellent communication and expectation-setting abilities. You have a technical background in Support, Services, DevOps, Systems Engineering, or Database environments, and are experienced in incident response or crisis management. You will have strong negotiation and objection-handling skill
Cloud Operations Engineers are responsible for building internal tools and process automation. Day-to-day duties are creating and monitoring systems alert dashboards, reviewing critical event and system logs, accessing customer instances that underpin their production databases, and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace. We are looking for a Lead with strong technical leadership experience as well as technical depth who is looking to collaborate closely with Cloud Operations Engineering Management in building and maintaining a high-performing team that delivers high quality outcomes while fostering psychological safety and professional growth. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Core responsibilities Team leadership: partner with and assist COE Management with the tasks of providing ongoing technical feedback to engineers, support their growth and creating an inclusive team environment Execution and delivery: play a key role in guiding team members through project deliverables ensuring high quality outcomes while also assisting in meeting or resetting timelines when required Time management: between assisting team members with day to day tasks ranging from incident to project management Cross-functional collaboration: work closely with Product, Technical Services and R&D to surface team’s pain points and drive alignment with the goal of providing an excellent user experience to the end customer Coordinate with Lead counterparts within Cloud Operations as well as Technical Services to ensure our uptime guarantees to the MongoDB Atlas customer base Assist and collaborate with the team on scoping, designing, deploying and ongoing maintenance of systems that focus on reducing mean time to resolve customer incidents Detec
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About Okta for AI Agents Okta secures access for 20,000 organizations and billions of users. Okta for AI Agents extends that work to the agentic shift. Deploying an AI agent is not like deploying traditional software. You are putting professional work output into production, and it needs deep integration, continuous tuning, and change management. Every agent needs an identity, a scope, an audit trail, and a way to be shut down when it goes wrong. Most enterprises have not built this yet. We are. We hire builders who see the cracks in enterprise agent identity that everyone else has learned to live with. The Role You embed inside four to five of Okta’s most strategic enterprise customers as their dedicated technical partner for agent identity. You sit alongside their identity, platform, and security engineering teams, write production code in their environment, and own the technical outcome from prototype through production. You are a builder-consultant. You go past architecture diagrams to code, debug, and ship bespoke agent identity solutions inside the customer’s environment. You ship secure agents faster for the customer, and you feed real field insight back to Okta product engineering. Responsibilities Become the customer’s trusted technical voice on agent security. Sit in their standups, design reviews, and incident response. Earn a seat on their architecture review board and security council for agent risk decisions. Architect and deploy with the cust
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Federal Operations Engineering Group Okta's Federal Operations team supports government customers operating in FedRAMP-authorized, IL4, and IL5 environments. We deliver the same 99.999% availability promise to the federal market while meeting the strict security, compliance, and operational requirements that come with it. We're looking for a technical leader who understands both the SRE discipline and the unique demands of federal customer relationships — someone who can hold a technical conversation with an agency ISSO in the morning and unblock an incident bridge with an engineering team in the afternoon. As Senior Manager of Federal SRE Operations, you will own the operational health of Okta's federal environments, lead a team of engineers working within compliance-governed change processes, and serve as a trusted point of contact for federal security stakeholders. What you'll be doing Lead
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from
Get new incident commander jobs by email
Daily job updates · Unsubscribe anytime