About the Team OpenAI’s User Operations team shepherds our customers’ adoption of AI and ensures that our customers' product experience is nothing short of exceptional. We are building the very first post-AGI support team. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others, to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. Within Premium Support, Dedicated Support Engineers combine deep technical troubleshooting with an enduring understanding of our most strategic customers’ architectures, critical workloads, and business priorities. Through proactive reliability work, ownership during incidents, and AI-powered support capabilities, we help customers operate successfully as their use of OpenAI grows. About the Role We’re looking for a senior leader to build and scale our Dedicated Support Engineering function globally. You will define its strategy, build the team, and establish how we deliver technically rigorous, proactive support for customers running some of the most complex and consequential workloads on OpenAI. This role combines organizational leadership, technical judgment, and executive customer engagement. You will establish a model in which DSEs develop deep customer context, independently advance difficult investigations, anticipate operational risks, and drive issues through resolution. You will also turn what the team learns into improvements that benefit customers across OpenAI. You should bring experience building technical organizations that maintain long-term accountability for enterprise customers. Leadership in Technical Account Management, enterprise Support Engineering, or a comparable technical customer function is particularly rel
Jobs in United States
Incident Commander in United States
159 active opportunities · Updated October 2026
Showing
15 jobs
Explore current incident commander jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
From $156K/yr
As a Product Manager – IaC Detection, you will define, build, and launch capabilities that proactively detect infrastructure issues in code (e.g. Terraform, Helm) before they can be deployed into production and escalate into production incidents. The Infrastructure Monitoring team has pioneered shift-left detection in the industry with Bits Infrastructure Operations , and we’re looking for a Product Manager to expand this capability to a broader set of use cases Customers (and thus developers) are increasingly standardizing on IaC tools to deploy and maintain ever-growing infrastructure in the cloud. At the same time, SREs and Infra teams struggle with an increasing number of production incidents. By shifting-left and identifying high-impact infra changes before they are deployed, we help reduce production incidents, reduce waste, and free up SRE time to focus on value-added tasks. You will own the roadmap to expand IaC detection to a broader set of use cases, including cost detection, blast radius impact, as well as configuration changes on infrastructure powering applications like nginx, postgres and more. You’ll partner closely with Engineering, Design, and customers to build and iterate on the roadmap, build product market fit, drive customer adoption (including internal usage), and focus on coverage and correctness of the AI system. This is an opportunity to lead an initiative at the intersection of AI, infrastructure operations, and autonomous observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the product roadmap for IaC Detection, enabling customers to proactively detect and catch high-impact infrastructure and configuration changes before they are deployed into production and escalate into incidents. Define the end-to
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. As a Premium Support Engineer at Replit, you’ll be the front line for our highest-value customers — delivering fast, expert, and reliable technical support when it matters most. You’ll handle complex product issues, guide customers through critical incidents, and ensure every interaction meets the highest standard of quality and speed. Replit is at the forefront of AI-driven software development, and how we support customers is constantly evolving. You’ll play a critical role in shaping how Premium Support adapts to new products, new customer expectations, and AI-assisted workflows, operating effectively in ambiguity and driving clarity for your team. You’ll combine deep technical troubleshooting with calm, confident communication to keep builders moving — whether it’s an enterprise team deploying at scale or a top-tier developer relying on Replit to power their business. In this role you will: Provide swift, high-priority support to Premium customers, responding within strict SLAs. Diagnose, reproduce, and resolve complex technical issues across the Replit platform. Escalate and track high-impact issues with Product and Engineering, ensuring timely fixes and transparent communication. Lead customer-facing communications during outages or incidents. Identify recurring issues and collaborate internally to reduce time-to-resolution. Contribute to internal tooling, automation, and documentation that improves team efficiency. Partner with Engineering, Product, Sales and other internal teams to ensure Premium customers receive a consistent, high-quality experience. Help onboard and mentor other support engineers, raising the team’s overall bar for responsiveness and quality. Required skills and experience: 3+ years in techn
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. As a Premium Support Engineer at Replit, you’ll be the front line for our highest-value customers — delivering fast, expert, and reliable technical support when it matters most. You’ll handle complex product issues, guide customers through critical incidents, and ensure every interaction meets the highest standard of quality and speed. Replit is at the forefront of AI-driven software development, and how we support customers is constantly evolving. You’ll play a critical role in shaping how Premium Support adapts to new products, new customer expectations, and AI-assisted workflows, operating effectively in ambiguity and driving clarity for your team. You’ll combine deep technical troubleshooting with calm, confident communication to keep builders moving — whether it’s an enterprise team deploying at scale or a top-tier developer relying on Replit to power their business. In this role you will: Provide swift, high-priority support to Premium customers, responding within strict SLAs. Diagnose, reproduce, and resolve complex technical issues across the Replit platform. Escalate and track high-impact issues with Product and Engineering, ensuring timely fixes and transparent communication. Lead customer-facing communications during outages or incidents. Identify recurring issues and collaborate internally to reduce time-to-resolution. Contribute to internal tooling, automation, and documentation that improves team efficiency. Partner with Engineering, Product, Sales and other internal teams to ensure Premium customers receive a consistent, high-quality experience. Help onboard and mentor other support engineers, raising the team’s overall bar for responsiveness and quality. Required skills and experience: 3+ years in techn
About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight: Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful autonomy (for example future versions of auto-review ). About the Role We’re looking for strong executors with excellent judgment, comfort with ambiguity, and an understanding of frontier model research. You don’t need prior safety or alignment experience, we also welcome people that recently realized that alignment and safety is a critical area to contribute to. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Train and evaluate frontier models to reduce harmful or misaligned agent actions, forming clear hypotheses and executing independently through ambiguity. Mine incidents and build scalable measurement, data-processing, and evaluation systems that turn real failures into repeatable safety signals. Collaborate closely with post-training, capabilities, oversight, and pre-training partners to ship research-backed mitigations into large-scale training and agent systems. You might thrive in this role if you: Have demonstrated strength in research engineering, ML en
About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight : Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful agent autonomy (for example future versions of auto-review ). About the Role This role focuses on oversight and system-level mitigations that enable increasingly capable agents to operate safely and autonomously in real environments. We prioritize building oversight systems that are used in practice today, both internally and externally (see our recent work on action monitoring for codex and former code review ). We also study longer-term questions about how increasingly capable agentis systems can be supervised, constrained, and corrected. We’re looking for a safety&security minded researcher or engineer who can reason rigorously about security boundaries and agent behavior, then build and test practical mitigations. A background in AI control or security is welcome but not required. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and evaluate system-level controls for agent actions like agent-based review. Plan how they fit in a broader syste
$116.2K – $182.4K/yr
Senior Manager, Logistics Security Description - At HP, we create technology that helps people and businesses bring their ideas to life. Our global supply chain plays a vital role in that mission, moving high-value products safely and reliably to customers around the world. We are looking for a collaborative and solutions-oriented Senior Manager, Logistics Security to help protect our logistics network, strengthen transportation risk programs, and lead efforts that reduce loss, improve carrier performance, and support a secure customer experience. Core responsibilities Own logistics security strategy for HP's transportation network. Manage loss, theft, damage, shortage and mis-delivery incidents across carriers, warehouses and distribution partners. Lead transportation claims management, including investigation, documentation, recovery and settlement. Manage relationships with major LSPs, carriers, freight forwarders and 3PLs. Investigate high-value shipment incidents and identify root causes. Develop controls to prevent cargo theft, fraud, unauthorized access and product diversion. Establish security standards for transportation lanes, warehouses and high-risk locations. Track claims and losses using KPIs such as: Claim $ / shipment Loss rate Damage rate Theft incidents Recovery % Claim cycle time Carrier performance Partner with Supply Chain, Logistics, Finance, Legal, Compliance, Security and Insurance teams. Lead corrective/preventive actions with logistics service providers. Review carrier contracts and ensure appropriate liability and insurance coverage. Establish escalation processes for major incidents. Conduct risk assessments of logistics providers and distribution locations. Use dat
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are seeking a mid-level Infrastructure Vulnerability Management Engineer with a strong background in Cloud Security, DevSecOps, and Infrastructure-as-Code (IaC). In this role, you will bridge the gap between security, compliance, DevOps, and Platform engineering teams. You will identify infrastructure misconfigurations, secure multi-cloud environments, and manage continuous vulnerability lifecycles across cloud workloads, containers, and data repositories to satisfy strict regulatory compliance frameworks. You will also serve as a technical infrastructure responder during security incidents, deploying real-time cloud or network countermeasures to protect our production ecosystem. What You'll Do Core Responsibilities Infrastructure Scanning & Triage: Perform continuous security scanning across our cloud posture and workloads. Review, validate, and prioritize flaws and misconfigurations based on CVSS scores, real-world exploitability, and infrastructure network exposure. Posture Management & Visibility : Own and optimize Cloud Security Posture Management (CSPM), Kubernetes Security Posture Management (KSPM), and Data Security Posture Management (DSPM) tools to ensure uniform compliance, prevent data leakage, and maintain hardened baselines. Infrastructure-as-Code (IaC) Security: Configure, tune, and embed automated IaC security scanning tools into CI/CD pipelines to identify architectural risks (e.g., overly permissive IAM, public S3 buckets/Cloud Storage) before they are deployed to production. Workload & Container Security: Manage the continuous vulnerability scanning lifecycle for container images, registries, and Virtual Machines (VMs), partnering with SRE and Platform teams to build aut
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are seeking a mid-level AppSec Vulnerability Management Engineer with a strong software development background. In this role, you will bridge the gap between security, compliance, and engineering teams. You will identify application vulnerabilities, maintain software supply chain security, and drive tracking to satisfy strict regulatory compliance frameworks. You will also serve as a technical responder during security incidents, deploying real-time countermeasures to protect our software ecosystem. What You'll Do Core Responsibilities Vulnerability Scanning & Triage: Perform periodic application security scanning activities. Review results and prioritize flaws based on CVSS scores, real-world exploitability, and system exposure. Compliance-Driven Tracking: Track, document, and manage vulnerabilities according to strict compliance SLAs (e.g., SOC 2, ISO 27001, PCI-DSS). Maintain audit-ready evidence of remediation timelines and exception approvals. Executive Reporting & Alerting: Escalate and report critical exposures directly to the CISO and senior leadership. Maintain dashboards and alerting mechanisms that visualize vulnerability status, risk trends, and compliance posture. Software Supply Chain Security: Ownership of the organization's Software Bill of Materials (SBOM). Continually update SBOM inventories to ensure compliance with modern regulatory requirements and dependency tracking. Help Replit mature through various SLSA levels for supply chain security. Remediation Collaboration: Partner with development teams to provide clear mitigation paths. Review, write, and patch code directly when necessary to resolve security flaws. Tooling Integration: Configure and tune automated security te
From $156K/yr
The Team: As a Security Engineer 2 on the Cyber Threat Intelligence team, you will help Datadog stay ahead of evolving threats by identifying, analyzing, and operationalizing intelligence on threat actors, campaigns, and emerging threats. Working within Security Engineering, you will partner closely with security teams to translate intelligence into actionable security improvements across the company. You will serve as a subject matter expert on how the cyber threat landscape intersects with Datadog and contribute to intelligence-led decision making during both steady-state operations and active security incidents. This role provides opportunities to influence detection, response, and security strategy through technical analysis, collaboration, and intelligence-driven initiatives. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Develop and maintain tooling that automates the collection, processing, analysis, and dissemination of threat intelligence. Assess emerging vulnerabilities, threat activity, and security events to help stakeholders understand potential impact to Datadog. Conduct threat hunting and infrastructure analysis to identify adversary activity relevant to Datadog and improve defensive controls. Partner with security teams to operationalize intelligence into detections, investigations, and response workflows. Coordinate with information-sharing communities to gather, evaluate, and disseminate actionable intelligence. Produce technical briefings, threat reports, and intelligence products for security and engineering stakeholders. Who You Are: Experienced in writing and presenting operational and technical intelligence for threat detection, response, and security stakeholders. Skilled in partnering with detection and response te
As an Engineering Manager II at Datadog, you will own a significant portion of the Logs business — one of Datadog's largest and most strategic products, with major growth opportunities still ahead. Sitting at the table where strategy is defined, you will drive execution across a ~30-person engineering org at an exciting moment of product evolution where AI is pushing us to rethink how we handle observability data. This is a high-scope role for someone energized by hard problems, eager to grow at the intersection of business and engineering excellence, and driven to build a product that radically transforms how engineering teams around the world operate and resolve incidents. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Own the success of the Logs business for your scope — co-define the Logs technical roadmap and product vision with Product, Design, and peer leadership. You will blend an understanding of customers, markets, and revenue drivers with deep engineering rigor to deliver our core vision Act as a key collaborator across Datadog, building strong partnerships with platform and infrastructure teams to ensure a highly reliable, scalable and performant product. Invest in the next generation of engineering leaders — growing your managers, expanding their impact, and building a high-performing engineering organization where people continuously develop and thrive. Foster a culture of ambition and resilience — where people lean into hard problems
MongoDB is seeking a senior leader to build a new communications function within the Customer Organization. This role will own the strategy, operating model, governance and execution framework for a spectrum of situations requiring customer and stakeholder communications, ranging from high-stakes operational situations, incidents, critical advisories, security events, product changes, through to recurring action-required and informational type needs. In this role, the leader will ensure MongoDB delivers transparent, accurate, empathetic, and well-timed communications during key moments of customer need. They will drive cross-functional collaboration across Technical Services, Engineering, Product, Security, Customer Success, Sales, Legal, and Corporate Communications to guarantee a unified and seamless corporate presence across both critical and non-critical events. This is a highly cross-functional leadership role that is explicitly expected to operate in a player-coach model. You will shape the long-term vision for the function, define the processes and measures of success, and personally guide communications during the most complex and visible situations. This role can be based out of one of our offices or remotely in the United States region. What you’ll do Build the function Define the charter, scope, roadmap, and operating model for the Customer Communications function at MongoDB Codify decision rights, review workflows, and escalation criteria for all operational communications. Orchestrate a scalable global model that ensures 24/7 follow-the-sun support. Recruit and mentor a team of specialized Communications Managers known for technical fluency and operational precision. Own critical operational communications <u
Job Requisition ID # 26WD95847 Position Overview Autodesk is seeking a motivated, experienced, and skilled engineer to join our Connected Delivery platform team. This platform is responsible for enabling the discovery, download, and install of all the Autodesk products, services, and third-part apps. You'll be responsible for building a full-stack application that is resilient, scalable, and highly available. As an ideal candidate, you’ll have led teams developing modern web experiences with cloud services in a fast-paced, agile environment. Responsibilities Lead, design, and develop high-quality, secure, performant applications Provide project and team leadership to break down, estimate, and organize work Participate in agile ceremonies of the scrum team Work closely with the product manager and team to understand and elaborate on the requirements Provide guidance to own team and others on software development best practices Work with the team to troubleshoot code-level problems quickly and efficiently Identify risk and propose mitigation strategies associated with the design Participate in code reviews to ensure new code conforms to the highest standards Interact with internal and external customers to identify and resolve product defects Mentor and develop junior developers Respond on a rotation basis to escalated incidents after hours or on weekends to ensure 24/7 availability of our platform Minimum Qualification BS in computer scie
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. As a Support Engineer you are at the forefront of helping those developers create. In a given day you’ll help folks with all sorts of things ranging from general Replit usage questions, identifying and debugging issues on the platform and helping to build and maintain the tools we use. We use tools like Zendesk, Linear, Slack, and Replit itself to get the job done. We also value solving the problems we can ourselves. We have a number of tools we build and maintain ourselves. You will work on a small, global team of support specialists and engineers united by Replit's mission to make the next billion software creators. Together, you'll ensure developers worldwide have the support they need to bring their ideas to life. In this role you will: Work directly with Replit customers via support tickets to solve technical product issues, bugs, and provide technical guidance while meeting daily ticket volume and resolution time targets. Collaborate with the rest of the Support team in telling the story of our users to the rest of Replit. Directly contribute to and maintain the Support team's Knowledge Base. Support customer-facing communications for outages and incidents and swiftly report ongoing bugs for Eng to triage and fix. Required Skills and Experience: Prior experience in software development or relevant technical support experience Proven ability to work efficiently in fast-paced, high-volume support environments with strict productivity metrics. Strong time management and ticket triage skills with comfort in performance tracking and regular reporting. Excellent communication skills with initiative to ask questions when encountered with the unknown. Experience with support tools and ticketing systems (e.g. Zendesk) Nic
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary CVS Health's Adjudication & Client Experience Engineering organization is seeking a motivated and highly skilled Senior Analyst - Software Development Engineering to join our Application Production Support team. This role will support critical business applications by providing production support, troubleshooting technical issues, and contributing to ongoing application enhancements and stability improvements. As a Sr. Analyst, you will work closely with Lead Engineers, Software Development Engineers, Product Owners, QA teams, and business stakeholders to investigate and resolve production incidents, perform root cause analysis, and implement code fixes. You will be responsible for supporting enterprise applications built on Java, Angular, APIs, and Cloud platforms while ensuring the reliability and performance of systems that serve our PBM (Pharmacy Benefit Management) business. This role is ideal for a hands-on engineer who enjoys solving production challenges, developing software solutions, and collaborating within a fast-paced environment. The successful candidate will contribute to application support activities, system enhancements, and continuous improvement initiatives while growing their technical and business domain expertise. Required Qualifications 5-8 years of experience in software development, application s
Other cities to consider
More places hiring for this role
Get new incident commander jobs in United States by email
Daily job updates · Unsubscribe anytime