About the Team OpenAI's Industrial Compute organization is building and operating the infrastructure foundation for the next generation of AI. Infrastructure Operations works across facilities, hardware, network operations, incident management, data center engineering, delivery teams, and external partners to bring capacity online safely, understand its operational state, and improve it over time. As OpenAI's data center portfolio grows across first-party and partner-delivered capacity, the organization needs clear goals, trusted data, repeatable processes, and systems that make ownership, risk, readiness, and performance visible. This role will help build the operating mechanisms that allow Infrastructure Operations to scale with rigor. About the Role We are seeking a Technical Program Manager to own the systems, data, reporting, governance, and program-management backbone for Infrastructure Operations. Reporting to the Delivery & Operations Lead, you will translate strategy into executable goals and operating cadences, turn operational needs into software and data solutions, and create the mechanisms that keep a rapidly evolving organization aligned and accountable. This role will also own the current 1P+3P delivery-tracking layer within Operations: milestones, delivery timelines, quantity forecasts, risks, decisions, and executive reporting. You will partner closely with 1P Delivery Program Management, Compute TPMs, Data Center Engineering, construction, commissioning, and operations leaders to ensure that delivery information becomes complete, usable input for readiness, handover, and ongoing operations. You will own program health and the operating system around it: the goals, data definitions, workflows, reporting, decision paths, and follow-through that help functional DRIs execute. The ideal candidate is comfortable in ambiguity, technically fluent enough to implement real systems, and relentless about converting scattered information into durable mechan
Jobs in United States
Incident Commander in United States
159 active opportunities · Updated October 2026
Showing
15 jobs
Explore current incident commander jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We are the first line of defense against fraud and abuse on the Plaid platform. Our mission is to ensure the safety and integrity of our platform for consumers and customers. As a Fraud and Abuse Operations Analyst , you will be responsible for responding to fraud and abuse events, investigating claims, and triaging incidents. We also partner with product and engineering teams to inform and improve fraud mitigation strategies. Responsibilities: Safeguard Plaid's Platform: Participate in the abuse on-call rotation, directly protecting our users and customers by responding to and resolving fraud and abuse events. Your timely actions will be instrumental in maintaining trust and security. Drive Investigations and Mitigate Risks: Investigate fraud and abuse claims from diverse sources, partnering with senior teammates on complex cases. Your findings will inform decisions and strategies, directly impacting Plaid's ability to prevent future incidents and minimize financial losses. Proactively perform threat modeling of abuse surfaces and continuously survey external fraud trends, adversary techniques, tooling, and emerging threat vectors Support Incident Response: Help triage and manage fraud and abuse ev
From $295.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. About the role: As a Principal Security Engineer on the Detection and Response (D&R) team at Roblox, you'll play a key role designing and developing effective custom security data pipeline systems, detection strategies and automations for response workflows to defend our critical assets from threat actors. You will also lead real-time incident response, actively investigate events and analyze threat actor techniques to prioritize emerging threats to ensure Roblox is equipped to mitigate and react to critical challenges. You will play a vital part to ensure the safety of our community and enterprise by proactively fostering a high-performing, inclusive security culture. This is a hybrid in-office role. You Will: Be a D&R authority! You will deliver robust detection & response capabilities: build new threat detection systems (keeping false positives low) while also automating processes with scripts, playbooks and orchestration tooling. Implement ETL pipelines : Design and develop customized data processing pipelines. Conduct security operations : Actively monitor security events and participate in on-call rotations to lead real-time incident response to contain and mitigate potent
We're looking for a Product Analyst to support our Point of Sale (POS) product team. You'll work closely with the POS Product Owner to translate business and venue needs into clear requirements, support day-to-day POS issues, and help drive continuous improvement of the POS platform across all venues. This role is a strong fit for someone who has spent real time in the trenches with POS systems, whether building them, configuring them, or supporting the people who use them every day, and who wants to move deeper into product-facing work. What You'll Do Partner with the POS Product Owner to gather, document, and prioritize requirements for POS enhancements, integrations, and fixes Analyze POS support tickets and incident trends to identify recurring issues, root causes, and opportunities for product improvement Translate business and venue operations needs into clear user stories, acceptance criteria, and functional specifications Support testing and validation of new POS features and releases, including UAT coordination with venue stakeholders Serve as an escalation point for complex POS issues, working with engineering, support, and vendor teams to drive resolution Monitor POS system performance and data quality, flagging anomalies that could affect transactions, inventory, or reporting Maintain and update product documentation, release notes, and training materials for interna
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. The Logistics Security Intelligence Analyst supports Micron’s global logistics security program by producing timely, actionable intelligence on shipment risk, cargo theft trends, route exposure, carrier performance, alert activity, and logistics security incidents. This individual contributor role helps strengthen shipment visibility, support incident response, and enable data-driven decisions for valuable and sensitive shipments across Micron’s transportation network. Responsibilities: Collect, analyze, and report on logistics security data related to high-value and high-risk shipments, including shipment value, route risk, carrier performance, tracking status, alert activity, and incident history. Monitor internal, vendor, industry, open-source, and law-enforcement sources for cargo theft trends, route disruptions, regional security developments, and emerging threats. Prepare intelligence summaries, dashboards, route profiles, regional threat updates, incident trend reports, and briefing materials for logistics security leaders and multi-functional collaborators. Support lane, route, carrier, provider, and regional risk assessments by identifying risk indicators, documenting findings, and helping translate analysis into practical control recommendations. Analyze shipment monitoring alerts such as route deviation, unauthorized stop, signal loss, seal breach, geofence violation, cargo separation, and other logistics security events. Support blocking issue and incident response a
Overview The Associate Director of Service Engineering leads the reliability, availability, and operational excellence of Natera’s lab-facing platforms. This role ensures that clinical systems, laboratory equipment workflows, and data pipelines operate with high reliability, scalability, and compliance in a regulated healthcare environment. You will lead a team responsible for production stability, incident response, and service health, partnering closely with Production Engineering, Lab Operations, Bioinformatics, Infrastructure, Facilities, and Compliance to support mission-critical genetic testing and diagnostics. Key Responsibilities Leadership & Team Development Lead, mentor, and scale a team of Service Engineers / SREs supporting clinical production systems Establish clear expectations around ownership, on-call readiness, and operational excellence Drive hiring, onboarding, performance management, and career growth Foster a blameless, learning-oriented culture focused on patient impact and reliability Service Reliability & Production Operations Own reliability and availability for production services supporting laboratory operations, reporting, and customer delivery Define and manage SLAs, SLOs, and operational KPIs aligned with clinical and business priorities Lead major incident response, ensuring rapid triage, clear communication, and thorough post-incident reviews Oversee on-call rotations, escalation paths, and operational playbooks Ensure operational readiness and go-live support for new assays, pipelines, and platform capabilities Technical Strategy & Execution Partner with Engineering and Development teams to design resilient, fault-tolerant systems Drive best practices for monitoring, alerting, logging, and observability across lab and cloud platforms Reduce operational toil through automation, tooling, and process improvements Advocate for reliability, performance, and scalability requirements early in t
Job Details: Job Description: The Role As a Cloud Application Development Engineer within Intel Manufacturing Foundry Cloud Services (imFCS) , you will design, develop, deploy, and support cloud-native applications that power semiconductor manufacturing, engineering automation, and AI-driven factory operations. You will build scalable, secure, and resilient platforms that improve engineering productivity, enable advanced analytics, and accelerate Intel Foundry's digital transformation. Key Responsibilities Develop and maintain cloud-native applications and services supporting manufacturing and engineering workflows. Support 24x7 manufacturing operations through on-call rotations, incident response, and root-cause analysis. Design and implement scalable microservices, APIs, and containerization systems utilizing Kubernetes orchestrator and cloud native applications. Architect secure cloud solutions spanning application, data, networking, identity, and observability domains. Build and maintain CI/CD pipelines, Infrastructure-as-Code, and DevOps automation. Provide technical leadership for contractor and partner development teams across multiple geographic regions, driving architecture, implementation, and operational excellence for cloud-native solutions. Collaborate with manufacturing, engineering, and software teams to deliv
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Safety is fundamental to the Lyft experience and to the trust riders, drivers, and our broader communities place in our platform. Lyft’s Safety and Customer Care team works through product, technology, data, policy, and operations to prevent harm before it happens, respond compassionately when it does, and continuously learn from every incident in efforts to make Lyft the safest way to get around. We are looking for an experienced safety leader with deep Trust & Safety experience to help turn Lyft’s Trust & Safety strategy into impact at scale, with the final shape of this role informed by the strengths of the leader we bring on. This is a highly visible leadership role within Lyft’s Safety & Customer Care organization. You will report to the Senior Director and General Manager of Trust & Safety and sit on the Trust & Safety leadership team. You will additionally partner deeply with Product, Engineering, Data Science, Research & Design, Legal, Compliance, Risk, Communications, Finance, and other teams across Lyft. The ideal candidate is an exceptional operator and people leader with Trust & Safety expertise, strong strategic judgment, and a track record of leading complex, high-stakes Trust & Safety organizations at scale. You can move fluidly between setting strategy with senior executives, developing leaders, overseeing large global operations and budgets, and diving into an emerging safety issue when the situation demands it. Responsibilities Lead Global Safety Compliance & Risk Oversee Lyft’s global safety compliance, transparency, and risk functions, supporting strong functional leaders and individual contributors on your team and strengthening capabilities as Lyft scales globally. Set direction across compliance readiness, safety transparency repo
We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<
# Best Email Phishing Simulation Services Email remains one of the most common ways cybercriminals target organizations. Attackers use convincing messages to trick employees into clicking malicious links, opening harmful attachments, sharing credentials, or transferring sensitive information. Even with firewalls, antivirus solutions, and security monitoring in place, one successful phishing email can create a serious security incident. **Email Phishing Simulation Services** help organizations measure employee awareness and strengthen their ability to recognize and report suspicious messages. ## What Is Email Phishing Simulation? Email phishing simulation is a controlled security awareness exercise that recreates realistic phishing scenarios without exposing employees to actual malicious activity. Authorized security professionals design simulated phishing emails based on common attack techniques and send them to selected users according to an approved testing plan. The objective is not to blame employees. Instead, it is to understand how users respond to realistic threats and identify areas where additional security awareness training may be required. A well-designed simulation can test responses to credential-harvesting emails, fake password-reset notifications, suspicious invoices, delivery alerts, account warnings, and other social engineering scenarios. ## Why Phishing Simulations Are Important Technology alone cannot completely prevent phishing attacks. Cybercriminals continually improve their techniques, making fraudulent messages increasingly difficult to distinguish from legitimate communications. Regular phishing simulations can help organizations: * Measure employee phishing awareness. * Identify users who may require additional training. * Test reporting and response procedures. * Improve recognition of suspicious emails. * Reduce the likelihood of credential theft. * Strengthen the organization's security culture. * Track awareness improvements
From $82K/yr
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: When a significant incident impacts, or threatens to impact, our community or company, Global Crisis Management works to coordinate a swift and effective company-wide response to the crisis in order to mitigate harm and best support those affected. Our team is composed of dedicated professionals who possess both a breadth and depth of institutional knowledge striving to always do the right thing in intense, and often ambiguous, situations. The Difference You Will Make: As a Disaster Response Coordinator within the Global Crisis Management (GCM) team, you will play a pivotal role in a dynamic team dedicated to mitigating the impact of unexpected events by preparing, acting decisively, and doing what is right for all of our stakeholders. A Typical Day: Global Incident Response Review global incident notifications and determine the response strategy based on the threat severity and business impact Conduct proactive research to monitor threats and events globally for potential impact to Airbnb stakeholders Use internal tools to promptly inform both our internal and external communities as needed Draft and coordinate the delivery of preparedness and response messaging for both external and internal audiences Recognize the need to make critical business decisions that consider the tradeoffs and benefits of potential options Cross-functional Collaboration and Partnership Effectively coordinate disaster response with cross-functional teams Serve as point of contact for the GCM Disaster Response team’s responsibilities and actions Provide clear situation reports and critica
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. We are seeking an experienced Manager of Security Operations, reporting to the Director of Security, to lead our exceptionally talented Security Engineering team. Vanta’s Security Operations team provides essential security operational services, including configuring, monitoring, and maintaining our security tools and infrastructure. You’ll be responsible for leading the Security Operations team and assisting with incident response, setting our detection and response strategy, and assisting with investigations. You’ll also work cross-functionally to ensure we maintain compliance and improve our secure maturity. What you’ll do as a Manager, Security Operations at Vanta: Lead and grow a team of the best security operations analysts in the world, with a view of security that is AI-first, human-centric, and trust-based. Help define the strategy for Vanta’s security operations team, and empower the team to implement robust security protocols and stay ahead of emerging threats. Leverage AI to improve efficiency of team processes, and improve the maturity of the overall security program. Lead and drive incident response from detection, remediation, to prevention Identify, develop, and implement new processes in our security operations program Identify new technologies and emerging threats for our organization, and plan a technology-driven approach to addressing new risks How to be successful in this role: Strong leadership experience in security and an ability to lead a global team from a foundation of transparency and trust. Strong security operations experience, with emphasis on implementing security controls in a SaaS and cloud env
What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.
We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, incident.io, Rootly, Jira Cloud Platform, Cortex, or similar and know how to navigate operational challenges of different s
From $272K/yr
We’re looking for a Senior Staff Software Engineer with deep experience in GenAI/ML to join Datadog’s Application Performance Monitoring (APM) team. APM is a product which provides deep visibility into applications, enabling users to identify performance bottlenecks, troubleshoot issues, and optimize services. With distributed tracing, profiling, out-of-the-box dashboards, and seamless correlation with other telemetry data, Datadog APM provides some of the deepest and most structured visibility into the health and performance of applications. This context sets us up for an opportunity to be the world leaders in agentic investigations and incident troubleshooting. You’ll act as a technical leader within the APM group, focused on agentic workflows. You’ll lead efforts to design, train, evaluate, and deploy GenAI/ML models at scale. We’re looking for a product-minded ML engineer with strong technical expertise, excellent communication skills, and a track record of driving impactful initiatives end to end. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Serve as the technical owner for GenAI initiatives within APM, leading design, development, and deployment of ML/AI-powered features across multiple teams. Guide long-term strategy and technical direction for GenAI workflows across APM and related products. Build and benchmark GenAI/ML models using state-of-the-art techniques. Contribute to Datadog’s broader senior engineering community through thought leadership and collaboration on company-wide initiatives. Collaborate with cross-functional teams to build automated investigation and triaging tools. Influence product direction by bringing a strong product mindset to your work, always advocating for the end user. Guide teams through ambiguity, sc
Other cities to consider
More places hiring for this role
Get new incident commander jobs in United States by email
Daily job updates · Unsubscribe anytime