Jobs in United States

Incident Commander in New York

26 active opportunities · Updated October 2026

Explore current incident commander jobs in New York. Filter by work mode, employment type, experience, department, date posted and distance.

D
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -88.9%

Datadog Incident Response is an end-to-end incident operations solution native to Datadog’s unified observability and security platform. It brings the entire incident lifecycle — from alert to resolution — together in one place so engineers can respond fast with confidence instead of losing time switching between disconnected monitoring, paging, and incident tools. As a Product Marketing Manager (PMM) - Incident Response, you will develop go-to-market strategy for new products and features, create the content that enables our sales and partner teams, and touch on all areas of the business while helping move Datadog forward. We give our Product Marketing Managers the opportunity to collaborate, investigate, and idealize how we can gear our product strategy to yield the highest results. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Drive the go-to-market strategy for Datadog Incident Response products, such as Incident Management, On-Call, and Incident AI as an expert of your product area. Launch new features with compelling messaging and positioning, developing assets such as slide decks, blogs, product demos, webinars, and solution briefs. Develop high-impact sales enablement assets (battlecards, pitch decks, demos) grounded in deep market and competitive insights. Partner with cross-functional teams to execute campaigns across webinars, paid media, organic channels, and sponsored events. Drive customer marketing initiatives including case studies, testimonials, and

L
📍 New York, NY, United States
✓ Quality checkedCompany trend -25%

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Safety is fundamental to the Lyft experience and to the trust riders, drivers, and our broader communities place in our platform. Lyft’s Safety and Customer Care team works through product, technology, data, policy, and operations to prevent harm before it happens, respond compassionately when it does, and continuously learn from every incident in efforts to make Lyft the safest way to get around. We are looking for an experienced safety leader with deep Trust & Safety experience to help turn Lyft’s Trust & Safety strategy into impact at scale, with the final shape of this role informed by the strengths of the leader we bring on. This is a highly visible leadership role within Lyft’s Safety & Customer Care organization. You will report to the Senior Director and General Manager of Trust & Safety and sit on the Trust & Safety leadership team. You will additionally partner deeply with Product, Engineering, Data Science, Research & Design, Legal, Compliance, Risk, Communications, Finance, and other teams across Lyft. The ideal candidate is an exceptional operator and people leader with Trust & Safety expertise, strong strategic judgment, and a track record of leading complex, high-stakes Trust & Safety organizations at scale. You can move fluidly between setting strategy with senior executives, developing leaders, overseeing large global operations and budgets, and diving into an emerging safety issue when the situation demands it. Responsibilities Lead Global Safety Compliance & Risk Oversee Lyft’s global safety compliance, transparency, and risk functions, supporting strong functional leaders and individual contributors on your team and strengthening capabilities as Lyft scales globally. Set direction across compliance readiness, safety transparency repo

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $272K/yr

Quick readStrong listing-quality and freshness signals

We’re looking for a Senior Staff Software Engineer with deep experience in GenAI/ML to join Datadog’s Application Performance Monitoring (APM) team. APM is a product which provides deep visibility into applications, enabling users to identify performance bottlenecks, troubleshoot issues, and optimize services. With distributed tracing, profiling, out-of-the-box dashboards, and seamless correlation with other telemetry data, Datadog APM provides some of the deepest and most structured visibility into the health and performance of applications. This context sets us up for an opportunity to be the world leaders in agentic investigations and incident troubleshooting. You’ll act as a technical leader within the APM group, focused on agentic workflows. You’ll lead efforts to design, train, evaluate, and deploy GenAI/ML models at scale. We’re looking for a product-minded ML engineer with strong technical expertise, excellent communication skills, and a track record of driving impactful initiatives end to end. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Serve as the technical owner for GenAI initiatives within APM, leading design, development, and deployment of ML/AI-powered features across multiple teams. Guide long-term strategy and technical direction for GenAI workflows across APM and related products. Build and benchmark GenAI/ML models using state-of-the-art techniques. Contribute to Datadog’s broader senior engineering community through thought leadership and collaboration on company-wide initiatives. Collaborate with cross-functional teams to build automated investigation and triaging tools. Influence product direction by bringing a strong product mindset to your work, always advocating for the end user. Guide teams through ambiguity, sc

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $234K/yr

Quick readStrong listing-quality and freshness signals

We’re looking for a Staff Software Engineer with deep experience in GenAI/ML to join Datadog’s Application Performance Monitoring (APM) team. APM is a product which provides deep visibility into applications, enabling users to identify performance bottlenecks, troubleshoot issues, and optimize services. With distributed tracing, profiling, out-of-the-box dashboards, and seamless correlation with other telemetry data, Datadog APM provides some of the deepest and most structured visibility into the health and performance of applications. This context sets us up for an opportunity to be the world leaders in agentic investigations and incident troubleshooting. You’ll act as a technical leader within the APM group, focused on agentic workflows. You’ll lead efforts to design, train, evaluate, and deploy GenAI/ML models at scale. We’re looking for a product-minded ML engineer with strong technical expertise, excellent communication skills, and a track record of driving impactful initiatives end to end. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Act as a technical leader within the APM organization, driving GenAI/machine learning projects from concept to production. Build and benchmark GenAI/ML models using state-of-the-art techniques. Collaborate with cross-functional teams to build automated investigation and triaging tools. Influence product direction by bringing a strong product mindset to your work, always advocating for the end user. Guide teams through ambiguity, scaling challenges, and evolving requirements with clear technical direction. Actively mentor engineers and influence engineering culture through leadership in design reviews, technical talks, and working groups. Who You Are: You have a BS/MS/PhD in a scientific field or equiva

Machine LearningAIGoRust
B
📍 New York, New York, United States· Full-time
✓ High-confidence listing

$192K – $240K/yr

Quick readStrong listing-quality and freshness signals

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Senior Software Engineer, Infrastructure (Release Engineering) at Brex, you will design, build, and operate the core systems that power Brex’s release, observability, and incident management processes. You will partner closely with product, platform, and operations teams to ensure releases are safe, fast, and reliable, and that our infrastructure scales securely as Brex grows. Where you’ll work This role will be based in our New York office. We are a hybrid environment that combines the energy and connections

PythonJavaSQLAWS
C-
📍 New York, NY, United States· Full-time
✓ High-confidence listing

$225K – $300K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. Today, CLEAR is well-known as a leader in digital and biometric identification, reducing friction for our members wherever an ID check is needed. We’re looking for a Senior Software Engineer to establish our Observability framework and foundations. You will join us to accelerate building and scaling our innovative systems that support our growing identity platform. You will drive on Observability best practices to find and fix gaps in our observability and our overall systems. You will also lead practices such as load testing, capacity planning, game days, chaos testing, and incident post-mortems. What You Will Do: Embed within the Engineering pillar to deeply understand the product and implement observability across all key flows Facilitate and build load testing cases, ensuring we understand the limits and scaling factors of our services and systems Contribute to observability and support the design of new services and systems, ensuring highly reliable and scalable concepts are implemented Build and lead practices such as game days, chaos engineering, and failure analysis Build long-term capacity plans, with an eye toward reliability and cost-efficiency Who You Are: 6+ experience writing production-grade software in a modern language, such as Java and Python. Strong knowledge of distributed systems concepts (think CAP theorem), microservices architecture, and distributed tracing . Experience with modern observability systems such as Datadog. Experience with performance debugging tools and patterns. You should be able to read a f

PythonJavaGitRest
C-
📍 New York, NY, United States· Full-time
✓ High-confidence listing

$275K – $350K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We are seeking a strategically-minded, technology-focused, and customer-centric Engineering Manager to lead one of our Infrastructure teams here. You will lead a team responsible for building, operating, and scaling the cloud infrastructure and platform systems that underpin CLEAR’s services, ensuring reliability, performance, and security across our environments. A successful candidate brings strong experience in cloud infrastructure, distributed systems, and operational excellence, along with a solid foundation in software engineering. You are an effective communicator who can lead complex infrastructure initiatives from inception through delivery, and thrive in fast-paced environments. This role requires a focus on building resilient, scalable systems, driving automation, and leading and developing high-performing engineering teams. What you'll do: Hire, develop, and grow engineering talent through coaching, mentorship, performance management, and career development planning Set clear goals and expectations, provide regular feedback, and foster accountability across the team Own and execute the roadmap for cloud infrastructure and platform engineering, and reliability initiatives Design, build, and operate a scalable, secure, and highly available cloud platform infrastructure Drive automation across infrastructure provisioning, deployment, and operations to improve efficiency and reduce manual overhead Establish and enforce best practices for system reliability, observability, incident response, and disaster recovery Partner with eng

PythonJavaAWSKubernetes
S
📍 New York, NY, United States· Full-time
✓ High-confidence listingCompany trend -95.9%

$190.4K – $285.6K/yr

Quick readStrong listing-quality and freshness signals

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Lead the technical design and architecture of major platform initiatives, author design documents and build consensus across engineering teams. Define technical roadmaps for complex, multi-quarter projects that span multiple teams. Make critical architectural decisions for company documentation infrastructure, balancing scalability, reliability, and developer experience. Evaluate and set direction for integrating emerging technologies, including AI/LLM capabilities, into company documentation platforms and authoring tools. Establish and evolve engineering standards, best practices and technical guidelines for the team and broader organization. Partner with engineering teams across the company to understand documentation needs and design integrated solutions. Design, build and maintain scalable, reliable and performant services and systems. Contribute high-quality code across the full stack and navigate codebases with different languages and tools. Debug and resolve complex production issues and improve system reliability. Take ownership of system health and incident response. Who you are Minimum requirements Must have a Bachelor's degree or foreign equivalent in Computer Science, Software Engineering, Engineering, or a related field, plus four (4) years of experience in Software Engineering. Must have four (4) years of experience in each of the following: - Working in a full stack environment with a foc

TypeScriptJavaMongoDBAI
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for a Detection & Response Engineer to build the systems that help us identify, investigate, and respond to threats across our platform. This is an engineering role focused on automation. You'll build detections, investigation tooling, and response capabilities that scale with our infrastructure, using AI where it meaningfully improves signal, investigation speed, and operational effectiveness. You'll work closely with infrastructure, platform, and security engineers to ensure every incident makes the platform more resilient. What You'll Work On: Detection Engineering Design and build high-fidelity detections for attacks, abuse, and anomalous behavior across our infrastructure and production systems Continuously improve detections based on telemetry, threat intelligence, and lessons learned from incidents Improve visibility across cloud infrastruc

SQLKubernetesGitLinux
R
📍 New York, NY, United States· Full-time
✓ High-confidence listingCompany trend -99.2%

From $10K/yr

Quick readStrong listing-quality and freshness signals

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role We're looking for a Technical Program Manager who can operate at the intersection of engineering, product, and business — someone deeply technical, trusted instinctively by engineers, and sharp enough to drive clarity and momentum across complex, cross-functional programs. This is a high-agency role with real executive visibility and direct impact on how Ramp's engineering organization scales. We're looking for someone who is energized by complexity, deeply curious about what AI can unlock for engineering teams, and eager to apply it hands-on in their work. You should be someone who experiments with AI tools regularly, thinks about how they change the way software gets built, and brings that perspective into how you run programs. What You’ll Do Lead large-scale technical programs across engineering and adjacent teams—from CI/CD and infrastructure scaling to incident response, and driving other strategic projects across the engineering organization Own Ramp's engineering incident response program, improving processes, running retrospec

CI/CDRestAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $192K/yr

Quick readStrong listing-quality and freshness signals

This Engineering Manager will lead the Data Visualizations Explorations team within the Graphing organization, setting product direction, and coaching and developing team members. They will staff and drive projects that build end to end experiences for Datadog’s core users: observability engineers. This includes extending the capabilities of core widgets like Hostmap and Geomap Visualizations and finding innovative ways to leverage existing Datadog data sources. This role also involves close partnerships with other product teams to deeply understand customer needs and deliver compelling data experiences. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Work with Product Management and Design to plan and staff projects for the Data Visualization Explorations team Coach and develop engineers at various levels Ensure strong cross-team communication and design best practices around project management, code review, architecture patterns, and more. Proactively anticipate cross-team dependencies and blockers to goals. Identify opportunities to appropriately reuse or customize features across dashboards, notebooks, and product pages. Deeply understand the needs of our customers and other Datadog products we work with. Participate in customer conversations, review product briefs, read feature requests, and coach team members to adopt these practices as well. Define and maintain high standards for operations practices, including bug triage and remediation, incident response, and gathering and analyzing performance telemetry for our widgets. Who You Are: At least 2 years of people management experience in a software engineering or similar setting Strong TypeScript/JavaScript skills, including familiarity with front

JavaScriptTypeScriptJavaReact
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $252K/yr

Quick readStrong listing-quality and freshness signals

We're looking for an Engineering Manager II to own and grow the Observability Pipelines engineering org at a pivotal moment in the product's lifecycle. Observability Pipelines is Datadog's on-premise, vendor-agnostic telemetry pipeline product, with a lot still to build as it grows and scales. It sits at the center of a fast-consolidating market, is central to Datadog's data pipeline optimization story for Logs and Metrics customers. This is a build-and-scale opportunity: you'll grow the management and technical leadership layers, co-own the roadmap with Product, and define how this org operates as it continues to expand. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You'll Do: Directly manage the OP org including EM1s across NYC and Paris, set technical direction, and be the connective tissue across a distributed team Build out the management and technical leadership layers as the org continues to grow - today ~20 ICs Partner directly with Product to co-own the roadmap and strategy, helping decide where OP’s engineering investment goes next Set and evolve the operating rhythm across the group: planning cadence, on-call and incident standards, and cross-team alignment Own key cross-org relationships with the SaaS Logs Pipelines team, the BYOC team, and the Vector open-source community Coach managers and senior engineers, and build the succession and growth plans that let the org scale beyond you Who You Are: Experienced managing managers across distributed teams, with a track record of raising the bar on how those teams operate, not just delivering through them Back

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $320K/yr

Quick readStrong listing-quality and freshness signals

As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental research problems and collaborating with Datadog's product and engineering teams to translate research advances into products. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Conduct research in generative AI and machine learning, building specialized foundation models and trained agents for observability Train multimodal models on large-scale, diverse telemetry data (metrics, logs, traces, topology, events) using distributed training infrastructure Design and build simulated environments and RL training loops for on-policy agent training and evaluation Collaborate with cross-functional teams (Product, Engineering) to integrate capabilities like multimodal world modeling and autonomous agents into Datadog's products Stay at the forefront of foundation models, world models, and RL-based agent research Contribute to r

GitMachine LearningAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $192K/yr

Quick readStrong listing-quality and freshness signals

As Engineering Manager for Threat Detection, you will lead a high-performing team that powers Datadog's detection program. Threat Detection is the organization responsible for keeping Datadog ahead of an evolving threat environment: closing coverage gaps faster, raising the bar on signal quality, and shipping detections that hold up under the scale and complexity of cloud-native infrastructure. Your team will combine direct detection expertise, platform engineering, and applied AI to ship detections at a pace and scale traditional rule-writing alone cannot match. Examples of what your team will work on include detection-authoring agents, the detection platform that powers every rule in production, coverage analysis, alert triage and response automation, and the evaluation infrastructure that holds these systems to a high bar of fidelity. Detection authorship is a shared responsibility across the organization, and your team will contribute both by building the systems that scale our authoring capacity and by writing detections directly when their domain expertise is the right tool. You will partner closely with our Security Incident & Response Team (SIRT), Cyber Threat Intelligence (CTI), AI Engineering teams, and Datadog's broader Security organization. This is a high-impact leadership role: you will grow a team of security and software engineers responsible for building and executing our detection and AI strategy. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the strategy, roadmap, and execution of Datadog Security's shift to AI-accelerated detection and response. Drive development of high-fidelity detections as a shared responsibility across the organization, ensuring your team's systems and direct contributions raise the bar on coverage and

PythonCI/CDRestAI
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $156K/yr

Quick readStrong listing-quality and freshness signals

As a Product Manager – IaC Detection, you will define, build, and launch capabilities that proactively detect infrastructure issues in code (e.g. Terraform, Helm) before they can be deployed into production and escalate into production incidents. The Infrastructure Monitoring team has pioneered shift-left detection in the industry with Bits Infrastructure Operations , and we’re looking for a Product Manager to expand this capability to a broader set of use cases Customers (and thus developers) are increasingly standardizing on IaC tools to deploy and maintain ever-growing infrastructure in the cloud. At the same time, SREs and Infra teams struggle with an increasing number of production incidents. By shifting-left and identifying high-impact infra changes before they are deployed, we help reduce production incidents, reduce waste, and free up SRE time to focus on value-added tasks. You will own the roadmap to expand IaC detection to a broader set of use cases, including cost detection, blast radius impact, as well as configuration changes on infrastructure powering applications like nginx, postgres and more. You’ll partner closely with Engineering, Design, and customers to build and iterate on the roadmap, build product market fit, drive customer adoption (including internal usage), and focus on coverage and correctness of the AI system. This is an opportunity to lead an initiative at the intersection of AI, infrastructure operations, and autonomous observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the product roadmap for IaC Detection, enabling customers to proactively detect and catch high-impact infrastructure and configuration changes before they are deployed into production and escalate into incidents. Define the end-to

GitAIRustTerraform
🔔

Get new incident commander jobs in New York, United States by email

Daily job updates · Unsubscribe anytime