Jobs in United States

Incident Commander in United States

159 active opportunities · Updated October 2026

Explore current incident commander jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

S
📍 Bellevue, WA, United States· Full-time
✓ High-confidence listingCompany trend -92%
Quick readStrong listing-quality and freshness signals

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. The Grid Service and Platform Engineering team is looking for a highly motivated and collaborative Software Engineering Manager. This role involves leading the engineering of mission-critical, tier 0 service infrastructure, the foundational data platform that powers Smartsheet at scale. You will oversee services that handle millions requests per day, operate at 99.999% availability, and deliver low-latency, high-throughput performance for millions of customers worldwide. We are an agile team that operates iteratively, focused on building high-quality software and adhering to rigorous operational best practices across complex, cross-functional distributed systems. This full-time position reports to the Director, Engineering and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Manage one or more related teams of 6–10+ software engineers, driving development of tier 0 grid services and platform infrastructure that millions of customers depend on daily. Own and uphold 99.999% service availability targets across critical platform services, embedding reliability engineering, incident management, and on-call rigor into team culture. Help architect and guide technical vision to evolve low-latency, high-throughput service platforms capable of sustaining millions requests per day with predictable, consistent performance under load. Guide and mentor engineers on distributed systems architecture, scalability patterns, and platform best pr

VueAWSAgileScrum
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $295.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer on the Sharing team, you will lead large, multi-team initiatives with long-term technical vision and group-level impact. You will be expected to define and drive the platform-wide strategy powering how millions of users capture, share, and discover content on Roblox. In this role, you will set engineering standards, mentor senior engineers, and serve as a key architect of our long-term technical direction. You Will Drive Strategy & Execution: Own the outcome of complex, business-critical programs spanning several teams, often lasting years. Innovate at Scale: Develop and drive a multi-year technical vision for content creation and sharing, anticipating scale, technology, and business evolution. Elevate Reliability: Lead high-severity incident response across groups; drive durable systemic solutions that improve reliability and velocity. Architect Foundations: Regularly improve shared infrastructure and foundational systems, introducing frameworks that uplift development speed across the organization. Align Teams: Aligns multiple teams on shared technical direction, producing detailed design docs, phased roadmaps, and planning models that balance short an

JavaAWSGitAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $293.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox’s database team develops the next-generation, multi-tenant database platform that elastically scales and underpins every online data workload at Roblox. As a principal engineer on the database team, you will shape the architecture, build and launch critical database capabilities that keep our services fast, reliable and efficient at global scale. You will report to the Technical Director for Storage. You will: Design and implement new engine features —indexing, storage formats, WAL and replication protocols, sharding, and query-planner enhancements—that push latency, throughput, and availability boundaries. Evolve the control plane to deliver elastic scaling, autonomous healing, and zero-downtime schema or tenant moves across global regions. Profile and optimize critical code paths using kernel-level tracing and advanced performance tooling; drive systematic tail-latency reductions. Establish engineering best practices by leading design reviews, performance benchmarks, failure drills, and post-incident retrospectives. Automate everything : develop frameworks for testing, CI/CD, rollout safety, observability, and autoscaling so that the platform operates hands-off at scale. Ment

SQLPostgreSQLMySQLAWS
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $187.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior QA Engineer, you will be the first dedicated QA hire for the Safety Engineering Group, which is responsible for building the infrastructure and policies that make Roblox the safest online community in the world. You will report to the QA Engineering Manager for Universal Apps and collaborate across the Safety organization to own end-to-end test strategies for high-stakes, 24/7 incident response systems and parent-facing safety features. This role is a unique opportunity to shape the quality culture of a company-level mandate from the ground up. You will: Partner with Engineering and Product to define and maintain comprehensive, risk-based test coverage for Safety features and systems. Own the end-to-end test strategy for Safety initiatives, including functional, regression, integration, and edge-case validation. Leverage and expand internal automation frameworks to automate critical user flows and drive API-first validation strategies. Identify high-impact automation opportunities and incorporate industry best practices to design robust, abuse-resistant test scenarios. Develop, track, and report on quality metrics (e.g., defect trends, coverage gaps) to proactively surface relea

AWSGitRestAI
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $192K/yr

Quick readStrong listing-quality and freshness signals

This Engineering Manager will lead the Data Visualizations Explorations team within the Graphing organization, setting product direction, and coaching and developing team members. They will staff and drive projects that build end to end experiences for Datadog’s core users: observability engineers. This includes extending the capabilities of core widgets like Hostmap and Geomap Visualizations and finding innovative ways to leverage existing Datadog data sources. This role also involves close partnerships with other product teams to deeply understand customer needs and deliver compelling data experiences. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Work with Product Management and Design to plan and staff projects for the Data Visualization Explorations team Coach and develop engineers at various levels Ensure strong cross-team communication and design best practices around project management, code review, architecture patterns, and more. Proactively anticipate cross-team dependencies and blockers to goals. Identify opportunities to appropriately reuse or customize features across dashboards, notebooks, and product pages. Deeply understand the needs of our customers and other Datadog products we work with. Participate in customer conversations, review product briefs, read feature requests, and coach team members to adopt these practices as well. Define and maintain high standards for operations practices, including bug triage and remediation, incident response, and gathering and analyzing performance telemetry for our widgets. Who You Are: At least 2 years of people management experience in a software engineering or similar setting Strong TypeScript/JavaScript skills, including familiarity with front

JavaScriptTypeScriptJavaReact
D
📍 Colorado, California, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $96K/yr

Quick readStrong listing-quality and freshness signals

We are Datadog's in-house product experts. The Technical Solutions team enables Datadog's worldwide growth by educating potential clients and ensuring that existing customers are happy and successful. Premier Technical Support Engineers (PSEs) are primarily focused on assisting prospects and customers with any technical questions about Datadog. PSEs engage with Datadog’s Premier customers not only via standard technical support channels, but also get involved via cadence calls, business reviews, and side projects. You will be immersed in a fast-paced environment where you will be challenged, but will also immediately witness your contributions to Datadog and to our customers. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Respond to Premier customer requests (phone / chat / tickets) on our fast paced team while continuing to educate our clients on the use of the platform Develop relationships with our Premier customers, working hand-in-hand to understand their specific environment Reproduce customer issues and assist customers implementing 1,000+ Datadog integrations Handle urgent escalation requests that may result in customer-facing troubleshooting calls, and internal or external incident management Build subject matter expertise in many Datadog product areas Autonomously troubleshoot complex and/or high-priority customer issues without guidance Drive product and engineering conversations based on needs, use cases, and problems learned during client interactions Provide mentorship to junior members of the team and serve as their escalation partner Participate in routine health check meetings with Premier customers Build out and improve documentation and knowledge base articles for a variety of technologies Who You Are: Experienced in mul

LinuxAIGoRust
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $252K/yr

Quick readStrong listing-quality and freshness signals

We're looking for an Engineering Manager II to own and grow the Observability Pipelines engineering org at a pivotal moment in the product's lifecycle. Observability Pipelines is Datadog's on-premise, vendor-agnostic telemetry pipeline product, with a lot still to build as it grows and scales. It sits at the center of a fast-consolidating market, is central to Datadog's data pipeline optimization story for Logs and Metrics customers. This is a build-and-scale opportunity: you'll grow the management and technical leadership layers, co-own the roadmap with Product, and define how this org operates as it continues to expand. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You'll Do: Directly manage the OP org including EM1s across NYC and Paris, set technical direction, and be the connective tissue across a distributed team Build out the management and technical leadership layers as the org continues to grow - today ~20 ICs Partner directly with Product to co-own the roadmap and strategy, helping decide where OP’s engineering investment goes next Set and evolve the operating rhythm across the group: planning cadence, on-call and incident standards, and cross-team alignment Own key cross-org relationships with the SaaS Logs Pipelines team, the BYOC team, and the Vector open-source community Coach managers and senior engineers, and build the succession and growth plans that let the org scale beyond you Who You Are: Experienced managing managers across distributed teams, with a track record of raising the bar on how those teams operate, not just delivering through them Back

D
📍 Denver, Colorado, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $92K/yr

Quick readStrong listing-quality and freshness signals

We are Datadog's in-house product experts. The Datadog Federal Support Engineering team is dedicated to serving as highly trusted technical advisors for our Public Sector customers, who operate within some of the most highly regulated and security-constrained environments. These customers include various government agencies and organizations with critical, sensitive missions. As a Federal Support Engineer 3, this role places you at the forefront of supporting these customers' mission-critical workloads. These complex workloads are often deployed across sophisticated hybrid and multi-cloud architectures, requiring deep expertise in cloud technologies, monitoring, and security best practices. Your primary responsibility is to ensure the complete success of these customers across their entire lifecycle with Datadog. Whether you’re looking to learn from the best or be the best, the Federal Support team is dedicated to furthering personal development and team success. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Engage with public sector customers via multiple channels (ticketing system, live chat, calls, and screensharing tools) to identify and resolve technical support requests. Troubleshoot, investigate, and resolve complex technical issues in highly constrained environments across Datadog's 1000+ integrations, often with limited logs or sanitized data. Handle urgent escalation cases that may result in customer-facing troubleshooting calls, and internal or external incident management Become a subject matter expert in many Datadog product areas Partner with Product, Engineering, and Account teams to to validate bugs and advocate for customer-impacting improvements Provide mentorship to junior members of the team and serve

RestMicroservicesAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $320K/yr

Quick readStrong listing-quality and freshness signals

As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental research problems and collaborating with Datadog's product and engineering teams to translate research advances into products. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Conduct research in generative AI and machine learning, building specialized foundation models and trained agents for observability Train multimodal models on large-scale, diverse telemetry data (metrics, logs, traces, topology, events) using distributed training infrastructure Design and build simulated environments and RL training loops for on-policy agent training and evaluation Collaborate with cross-functional teams (Product, Engineering) to integrate capabilities like multimodal world modeling and autonomous agents into Datadog's products Stay at the forefront of foundation models, world models, and RL-based agent research Contribute to r

GitMachine LearningAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $192K/yr

Quick readStrong listing-quality and freshness signals

As Engineering Manager for Threat Detection, you will lead a high-performing team that powers Datadog's detection program. Threat Detection is the organization responsible for keeping Datadog ahead of an evolving threat environment: closing coverage gaps faster, raising the bar on signal quality, and shipping detections that hold up under the scale and complexity of cloud-native infrastructure. Your team will combine direct detection expertise, platform engineering, and applied AI to ship detections at a pace and scale traditional rule-writing alone cannot match. Examples of what your team will work on include detection-authoring agents, the detection platform that powers every rule in production, coverage analysis, alert triage and response automation, and the evaluation infrastructure that holds these systems to a high bar of fidelity. Detection authorship is a shared responsibility across the organization, and your team will contribute both by building the systems that scale our authoring capacity and by writing detections directly when their domain expertise is the right tool. You will partner closely with our Security Incident & Response Team (SIRT), Cyber Threat Intelligence (CTI), AI Engineering teams, and Datadog's broader Security organization. This is a high-impact leadership role: you will grow a team of security and software engineers responsible for building and executing our detection and AI strategy. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the strategy, roadmap, and execution of Datadog Security's shift to AI-accelerated detection and response. Drive development of high-fidelity detections as a shared responsibility across the organization, ensuring your team's systems and direct contributions raise the bar on coverage and

PythonCI/CDRestAI
A
📍 United States· Full-time
✓ High-confidence listingCompany trend -99%

From $212K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Difference You Will Make: Airbnb is entering a new era—reimagining how fraud, safety, and quality are measured, protected, and elevated across the digital landscape. Reporting to the Director of Advanced Analytics, Fraud & Safety, you will lead a dispersed team of ~8 advanced analysts tasked with turning safety measurement into governed, self-serve decision systems for Policy, Ops, Legal, Product, and Engineering. You will be a pivotal leader safeguarding Airbnb’s global community. You’ll steer a multidisciplinary team that designs, delivers, and scales state-of-the-art analytics for: Safety (physical & digital) Connected-account & circumvention detection Privacy protection & risk management This role isn’t just about reporting what happened; it’s about building the systems that help Airbnb see around corners. You will democratize data access, build always-on scenario simulators for fraud and safety, and turn incident impacts into seamless signals for continuous improvement. Your insights and systems will empower every stakeholder—from legal to operations, policy to product—to make bold, data-driven, and context-aware decisions in real time. A Typical Day: Own end-to-end analytical workflows for platform safety controls: data ingestion, feature engineering, modeling, experimentation, and visualization. A core mandate is to co-own and drive the enterprise-wide “Safety Single Source of Truth” (SSoT) strategy —a unified, structured data asset and taxonomy that underpins decision-making across Product, Operations, Policy, and Legal. Ensure metrics are future-proof, privacy-compliant, and g

PythonSQLGitAI
A
📍 United States· Full-time
✓ High-confidence listingCompany trend -99%

From $156K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Community Support org handles tens of millions of interactions yearly, engaging with Airbnb customers by phone, messaging, chat, or social media channels. The group handles hundreds of issues across categories including Cancellations, Account Issues, Refunds, Payments, Reservations, Extenuating Circumstances, Booking & Listing issues, Safety & Claims. The organization is globally distributed with offices in San Francisco, Dublin, Montreal, Seattle, Singapore, Manila, Gurgaon and an extensive partner network serving all regions. The Difference You Will Make: In this role you will own the strategic design and continuous evolution of risk frameworks, the risk registry, and executive risk narratives for Community Support — translating investigative findings, AI/ML model outputs, operational signals, and emerging threats into decisions and actions that protect Airbnb. A defining element of this role is your close partnership with the Insider Threat program. Together, you will form a complementary unit: you will ensure that investigative outcomes are contextualized within the broader risk ecosystem — informing risk appetite decisions, shaping detection strategy, and driving remediation accountability. You will co-own escalation frameworks, jointly challenge detection model effectiveness, and ensure that the feedback loop between investigations, risk governance, and AI-driven detection is robust and continuously improving. Beyond analysis and governance design, you will serve as the program manager for systemic root cause resolution — ensuring that when investigations, incident

ReactAWSAIGo
A
📍 United States· Full-time· Remote
✓ Quality checkedCompany trend -99%

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: BizTech fosters culture and connection at Airbnb by providing reliable corporate tools, innovative products, and technical support for all teams. We drive technical breakthroughs and strategies that redefine what it means to belong anywhere, delivering greater value for the business and our people. The Global Operations team at BizTech manages production services across Airbnb’s corporate environment, delivering reliable operations through Observability, Incident Management, Core Operations, and AI-enabled automation. We partner across BizTech to scale service quality, efficiency, and resilience. The Difference You Will Make: As a Senior Staff Engineer in Operations, you will lead and mentor a high-performing team to scale our AI-enabled operations model and deliver AIOps solutions that streamline operational workstreams and help BizTech teams focus on their core work with confidence. Ops owns triage and resolution, proactive monitoring across networks, systems, applications, and cloud services via a homegrown observability platform, and drives process excellence through automation and shift-left programs. You will set the technical bar, model operational excellence, and ensure high-quality, reliable service. Your scope includes leading projects ac

PythonAWSCI/CDAI
A
📍 United States· Full-time
✓ High-confidence listingCompany trend -99%

From $212K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Our web and API surfaces handle requests from guests and hosts alongside a growing volume of automated agents: AI assistants, crawlers, and scrapers. We build the systems that bring clarity to this traffic, combining in-house ML and vendor signals to decide in real time how to serve billions of daily requests. Anti-bot and anti-scraping detection is our most adversarial mandate, but the wider challenge is full traffic classification: building evaluation frameworks that tell legitimate automation apart from abusive actors, so high-stakes decisions hold up across the fleet. The Difference You Will Make: You will architect and maintain Airbnb’s end-to-end traffic classification ML systems, balancing high-performance model deployment with rigorous offline data pipelines. Success is measured by your ability to harden edge-traffic policies—targeting reduced bot-incident MTTM—and by establishing rigorous evaluation practices that ensure foundational signal accuracy and evasion-resistance across the fleet. A Typical Day: Own the complete lifecycle of traffic-scoring models, from problem framing to real-time deployment, managing the adversarial feedback loop to ensure high evasion-resistance and directly drive reductions in bot-incident MTTM. Architect robust offline-to-online pipelines that produce certified source-of-truth datasets, establishing rigorous evaluation frameworks—such as stratified benchmarks and leakage-prevention checks—to ensure every model improvement is empirically measurable and defensible. Execute model optimization within strict millisecond latency budgets at the

SQLGitMachine LearningAI
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI’s Network Engineering team within IT and Security advances the mission of deploying artificial general intelligence (AGI) for the benefit of all by delivering secure, scalable, and resilient network services. We build and operate the connectivity that supports OpenAI’s offices, labs, campuses, cloud environments, people, and devices. By combining strong network fundamentals with security, reliability, automation, and user-centered design, we enable impactful AI research, corporate operations, and product innovation. About the Role As a Network Engineer at OpenAI, you will design, operate, and continuously improve the global networks that connect our offices, labs, campuses, PoPs, cloud environments, people, and devices. The role spans strategic platform engineering and responsive production operations: you will shape architecture, standards, roadmaps, lifecycle plans, and automation while supporting incidents, escalations, and time-sensitive delivery. Operational signals will inform what we stabilize, simplify, standardize, or automate next. We work backward from user needs, investigate root causes, own outcomes end-to-end, and move quickly without compromising security. We are looking for a versatile engineer who can make pragmatic reliability and security tradeoffs, communicate clearly, and turn recurring operational work into durable platforms, tooling, and standards. You will partner across IT, Security, AppEng, Research, Applied, workplace teams, carriers, and vendors. In this role, you will: Design, implement, and operate secure, scalable enterprise networks across offices, labs, campuses, PoPs, cloud connectivity, and hybrid environments. Set strategic direction for network services through architecture, standards, roadmaps, lifecycle planning, capacity strategy, and measurable reliability outcomes. Own production operations, including on-call, incident response, escalations, and time-sensitive delivery, while protecting user experience,

PythonAWSAzureCI/CD
🔔

Get new incident commander jobs in United States by email

Daily job updates · Unsubscribe anytime