Jobs in United States

Incident Commander in United States

159 active opportunities · Updated October 2026

Explore current incident commander jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $234K/yr

Quick readStrong listing-quality and freshness signals

We’re looking for a Staff Software Engineer with deep experience in GenAI/ML to join Datadog’s Application Performance Monitoring (APM) team. APM is a product which provides deep visibility into applications, enabling users to identify performance bottlenecks, troubleshoot issues, and optimize services. With distributed tracing, profiling, out-of-the-box dashboards, and seamless correlation with other telemetry data, Datadog APM provides some of the deepest and most structured visibility into the health and performance of applications. This context sets us up for an opportunity to be the world leaders in agentic investigations and incident troubleshooting. You’ll act as a technical leader within the APM group, focused on agentic workflows. You’ll lead efforts to design, train, evaluate, and deploy GenAI/ML models at scale. We’re looking for a product-minded ML engineer with strong technical expertise, excellent communication skills, and a track record of driving impactful initiatives end to end. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Act as a technical leader within the APM organization, driving GenAI/machine learning projects from concept to production. Build and benchmark GenAI/ML models using state-of-the-art techniques. Collaborate with cross-functional teams to build automated investigation and triaging tools. Influence product direction by bringing a strong product mindset to your work, always advocating for the end user. Guide teams through ambiguity, scaling challenges, and evolving requirements with clear technical direction. Actively mentor engineers and influence engineering culture through leadership in design reviews, technical talks, and working groups. Who You Are: You have a BS/MS/PhD in a scientific field or equiva

Machine LearningAIGoRust
D
📍 United States· Full-time· Remote
✓ High-confidence listing

$220K – $275K/yr

Quick readStrong listing-quality and freshness signals

Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. Discord exists to give people the power to create space to find belonging — to talk regularly with the people they care about and build genuine relationships with friends and communities close to home or around the world. We're looking for an Analytics Manager to lead our Scaled Abuse Countermeasures and Research (SCAR) team — the team that safeguards Discord’s platform integrity. SCAR detects, analyzes, and disrupts high-volume threats through a combination of automated systems, deep research, and active incident response. This role reports to the Head of Safety Intelligence and Automation. What You'll Be Doing Lead and mentor a team of data analysts, scientists, and researchers who investigate active threats, identify platform abuse vectors, and uncover adversarial patterns. Define a strategic roadmap that prioritizes and disrupts the highest impact abuse operations through structured research, rigorous analyses, and live experimentation. Collaborate closely with the safety machine learning team to improve models by identifying the threat signals that translate into long-term, automated countermeasures. Partner cross-functionally with Product, Engineering, Data Science, Policy, Legal, and Revenue, influencing safety-by-design decisions upstream of abuse. Influence capacity toward high-impact infrastructure, such as automated rule engines, ML models, or agent moderation tools. What you should have 2+ years of people management experience leading technical teams, including engineers, data scientists, analysts, applied researchers, or equivalent. 4+ years of experience working in a Trust & Safety dom

PythonSQLRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

PythonAWSAzureGit
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -88.6%

$230K – $260K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role Millions of people rely on Notion to do their most important work, and protecting that trust is foundational to everything we build. We’re looking for a hands-on Detection Engineer to build and operate the systems and workflows we use to detect and respond to attacks across Notion’s cloud-native environment. You’ll ship high-signal detections, improve the platform that powers them, participate in incident response, and help shape how detection and response engineering scales at Notion. You’ll work closely with Engineering, Corporate Security, and Infrastructure, with broad latitude to identify gaps, prioritize investments, and build what’s needed next. We view detection and response as a software engineering discipline: detections are code, platforms are products, and measurement matters What You'll Achieve Design and maintain high-signal detections across cloud, identity, endpoints, and SaaS environments. Build and improve the detection platform, including rule lifecycle management, tuning, measurement, and rollout safety. Develop tooling and automation that accelerate triage, enrichment, investigation, and detection

AWSAzureGCPKubernetes
PE
📍 Westford, Ma 01886, United States
✓ Quality checked

Seceon is a trusted Ransomware Detection Company that helps organizations identify, prevent, and respond to ransomware attacks before they disrupt critical operations. Powered by AI-driven threat detection, real-time monitoring, and automated incident response, Seceon delivers comprehensive cybersecurity protection across cloud, network, and endpoint environments. Its advanced platform detects suspicious behavior, blocks malicious activity, and minimizes the risk of data loss or business downtime. By combining proactive threat intelligence with continuous security monitoring, Seceon enables businesses of all sizes to strengthen cyber resilience, maintain compliance, and defend against evolving ransomware threats with confidence. https://www.seceon.com/ransomware-detection/

PE
📍 238 Littleton Road, Suite #200, United States
✓ Quality checked

Seceon is a trusted Ransomware Detection Company that helps organizations identify, prevent, and respond to ransomware attacks before they disrupt critical operations. Powered by AI-driven threat detection, real-time monitoring, and automated incident response, Seceon delivers comprehensive cybersecurity protection across cloud, network, and endpoint environments. Its advanced platform detects suspicious behavior, blocks malicious activity, and minimizes the risk of data loss or business downtime. By combining proactive threat intelligence with continuous security monitoring, Seceon enables businesses of all sizes to strengthen cyber resilience, maintain compliance, and defend against evolving ransomware threats with confidence. https://www.seceon.com/ransomware-detection/

MT
📍 Manassas, VA - Fab 6, United States
✓ Quality checkedCompany trend +1150%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As the Logistics Security Program Manager for EMEA, this role is an essential part of Micron’s Global Security team. The person leads the management and continuous refinement of the company’s important logistics security efforts across the EMEA area. This position serves as the regional authority, advancing risk-focused security methods that protect high-value shipments, improve supply chain durability, and lower transportation security risks in intricate multimodal logistics networks. As a senior individual contributor, the Logistics Security Program Manager takes charge of regional program initiatives on their own. They apply solid judgment to shifting threat conditions and collaborate with colleagues across functions and external partners to produce security results. The position involves balancing security, operational efficiency, and business continuity while transforming regional risks into scalable, practical controls that advance cargo visibility, shipment protection, and incident readiness. Responsibilities: Act as the EMEA logistics security authority, guiding the creation and implementation of risk-focused security programs for valuable and sensitive shipments involving carriers, freight forwarders, and logistics providers. Develop, apply, and manage shipment security controls, including tracking, telematics, geofencing, chain of custody, tamper detection, monitoring, critical issue handling, recovery processes, and carrier compliance requirements. Conduct carrier, route, l

AISupply ChainLogisticsProcurement
C
📍 Scottsdale, United States
✓ Quality checkedCompany trend +340.2%

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. The PBM Batch Operations role is responsible for providing 24x7 operational support for Pharmacy Benefit Management (PBM) production processing environments. The position monitors, controls, and supports enterprise batch workloads, mainframe systems, iSeries environments, and associated operational processes to ensure critical pharmacy and business applications execute successfully and on schedule. Schedule: WorkDays: TBD Hours: 7 AM to 7PM AZ time Shift Structure: Three 12-hour shifts Additional Requirement: Must be available to work overtime as needed to provide coverage PBM Batch Operations Functional Responsibilities The PBM Operations environment includes responsibility for: Monitoring and supporting IWS (IBM Workload Scheduler) batch processing. Batch job interventions (restart, hold, kill, force complete). Mainframe IPL support. Mainframe console monitoring across multiple LPARs. RxClaim and iSeries batch monitoring. PBM Disaster Recovery support. Vendor escort activities and data center operational support. Data center security ticket processing. MIR3 paging and incident notifications. ServiceNow ticket management. Procedure verification and operationa

C
📍 Scottsdale, United States
✓ Quality checkedCompany trend +340.2%

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. The PBM Batch Operations role is responsible for providing 24x7 operational support for Pharmacy Benefit Management (PBM) production processing environments. The position monitors, controls, and supports enterprise batch workloads, mainframe systems, iSeries environments, and associated operational processes to ensure critical pharmacy and business applications execute successfully and on schedule. Schedule: WorkDays: Wednesday through Saturday Hours: 7:00 PM to 5:00 AM AZ time Shift Structure: 10-hour shifts Additional Requirement: Must be available to work overtime as needed to provide coverage PBM Batch Operations Functional Responsibilities The PBM Operations environment includes responsibility for: Monitoring and supporting IWS (IBM Workload Scheduler) batch processing. Batch job interventions (restart, hold, kill, force complete). Mainframe IPL support. Mainframe console monitoring across multiple LPARs. RxClaim and iSeries batch monitoring. PBM Disaster Recovery support. Vendor escort activities and data center operational support. Data center security ticket processing. MIR3 paging and incident notifications. ServiceNow ticket management. Procedure ver

C
📍 Woonsocket, Woonsocket, United States
✓ Quality checkedCompany trend +340.2%

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. At CVS Health, Site Reliability Engineering (SRE) is fundamental to delivering the reliable, secure, and scalable technology experiences that support millions of patients, customers, pharmacists, and healthcare professionals every day. Our SRE organization drives operational excellence across critical healthcare and retail platforms through innovation, automation, observability, and engineering best practices. The Executive Director, Site Reliability Engineering serves as the strategic leader responsible for the reliability, resilience, and performance of CVS Health's retail and pharmacy technology ecosystem. This executive will define and execute a comprehensive reliability strategy, oversee large global engineering teams, and establish a long-term vision for observability, automation, and operational excellence across thousands of store locations. Working closely with senior business and technology leaders, the Executive Director will champion modern SRE practices, accelerate incident response capabilities, and deliver real-time operational visibility that enables proactive issue prevention and exceptional customer and patient experiences. Key Responsibilities Strategic Leadership & Vision Define and lead the enterprise-wide Site Reliability Engineering strategy supporting CVS Health's retail and pharmacy operations. Align reliability and operational

AWSAzureGCPKubernetes
W
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -7.5%
Quick readStrong listing-quality and freshness signals

🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Join WRITER's security team as a staff detection and response engineer and help protect the AI infrastructure that's transforming how the world works. You'll build sophisticated detection systems that identify attacks targeting our AI platform, training data, and model deployments while creating automated response capabilities that scale with our explosive growth. This isn't just traditional security work – you're defending cutting-edge AI/AGI systems against adversaries who are evolving their tactics as fast as AI itself advances. This role combines hands-on security engineering with strategic thinking to stay ahead of novel threats that don't exist in textbooks yet. You'll be the operational arm of our security function, translating threat intelligence into real-time detections, coordinating incident response across multiple teams, and hunting for sophisticated attacks across GPU clusters and distributed training environments. If you're excited by the challen

Z
📍 Bellevue, Washington, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler’s high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You’ll Do (Role Expectations) Maintain h

PythonKubernetesLinuxAI
N
📍 Remote, United States· Remote
✓ Quality checkedCompany trend -13.7%

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data

PythonSQLAWSAzure
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

NVIDIA is transforming how the world uses AI, cloud, and accelerated computing, and trust is at the center of that mission. Our Attestation and Trust Services team builds the secure cloud services that show customers their NVIDIA platforms are healthy, resilient, and ready for their most important workloads. In this role, you help design and run services that sit at the intersection of hardware, security, and large-scale distributed systems. We partner closely with security, silicon, platform, and cloud teams to bring new ideas into reliable production services that people rely on every day. We care about building systems that last, supporting each other, and creating space for learning and experimentation. If you enjoy solving complex problems, keeping services running smoothly, and collaborating with teammates from many disciplines, we would love to talk with you! What you’ll be doing: Your main focus will be on building and managing our core attestation cloud services. Day-to-day responsibilities include crafting APIs and integrations, boosting reliability, and working alongside NVIDIA teams to convert hardware trust mechanisms and standards into production-ready solutions. You will contribute significantly to shaping how customers verify that NVIDIA platforms are secure and prepared for their workloads. Crafting and evolving attestation cloud services, APIs, and SDK/CLI integration points that confirm the integrity of NVIDIA platforms across data center, AI, networking, and partner environments. Improving reliability and operational maturity through SLOs/SLIs, alerting, runbooks, incident response, and safe rollout practices. Crafting resilient service behavior that handles dependency failures, caching challenges, regional issues, customer-side resilience needs, and graceful degradation. Architecting trust-material distribution for certificate status, re

JavaAWSAzureGCP
MI
📍 Atlanta, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Hexnode Hexnode, the Enterprise software division of Mitsogo Inc., was founded with a mission to simplify the way people work. Operating in over 100 countries, Hexnode UEM empowers organizations in diverse sectors. Fueling the transformation to a seamless ecosystem of connected tools, Hexnode is revolutionizing the enterprise software and cybersecurity landscape. Role Overview We are seeking a AWS Operations Specialist to manage and maintain our cloud infrastructure and device ecosystems. This is a highly operational, execution-focused role—not an architecture position. The ideal candidate has 2 to 4 years of experience executing infrastructure as code, monitoring environments, and following documented playbooks to keep our systems secure and resilient. Because this role handles secure environments, candidates must be US Citizens and capable of passing a comprehensive federal background check. Key Responsibilities Infrastructure Execution: Run, maintain, and execute existing Terraform and Ansible scripts to deploy and update infrastructure. GovCloud Monitoring: Actively monitor our AWS GovCloud dashboards, keeping a close eye on system health, performance metrics, and security baselines. Mobile Device Management: Manage Android Enterprise kiosk configurations, ensuring secure deployments and smooth device operations. Incident Response & Triage: Respond swiftly to operational alerts by strictly following our documented team playbooks. Escalation: Identify anomalies or issues that fall outside established, documented procedures and escalate them accurately to the engineering team. Required Qualifications & Profile Citizenship: Must be a US Citizen (required for GovCloud infrastructure management). Background: Must be able to successfully clear a rigorous federal background investigation. Experience: 2 to 4 years of hands-on experience in a technical operations, DevOps, or SysAdmin role. Technical Familiarity: Comfort executing/running Terraform and

AWSAISwiftGo
🔔

Get new incident commander jobs in United States by email

Daily job updates · Unsubscribe anytime