About the Team OpenAI’s User Operations team shepherds our customers’ adoption of AI and ensures that our customers' product experience is nothing short of exceptional. We are building the very first post-AGI support team. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others, to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. About the Role We are looking for a hands-on lead to build and run OpenAI's Incidents & Escalations function within User Operations. This is a player-coach role with a meaningful hands-on operating component. You will set the operating model and also step into active incidents and urgent escalations when needed, coordinating with on-call teams, driving clear ownership, supporting communications, and ensuring issues move through resolution and post-incident closure. During active incidents, you will coordinate with the relevant on-call teams and cross-functional responders across Engineering, Infrastructure, Support Delivery, Product, and Go-To-Market. You will help keep teams aligned, maintain timelines, clarify ownership, escalate when needed, and ensure internal, executive, customer-facing, and external communications are accurate and timely, including status page updates when required. For escalations, you will build and run the processes for tracking, triaging, mitigating, and resolving critical customer and user issues. After incidents and escalations, you will own the follow-through: retrospectives, root cause identification, action item tracking, trend analysis, and process improvements that reduce repeat issues over time. You will also help define the long-term operating model for incidents and escalations across Support D
Jobs in United States
Incident Commander in San Francisco
56 active opportunities · Updated October 2026
Showing
15 jobs
Explore current incident commander jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
From $183.3K/yr
About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. What you'll do There is no MSSP and no tier-1 queue here. Detection & Response engineers own their detections end to end: you write them, you tune them, and your team is paged when they fire. The security team is spread across the globe with a follow-the-sun pager rotation so nobody is paged at 3am local. The adversaries are real. The business is growing fast and the threat surface is growing with it. Defining the necessary telemetry is part of the job. Detection engineering Build and tune detections across endpoint, identity, SaaS, and cloud , treating them as software: version-controlled, peer-reviewed, and shipped through the same CI/CD practices the rest of engineering uses. Track detection quality as measured quantities : coverage against MITRE ATT&CK, precision, time-to-detect. We don’t build-and-forget here. Response & automation Own incident response: triage, contain, remediate, and write the retrospective that turns the incident into a systemic fix. Build automation that removes toil from investigations, and partner closely with the US-based team so context carries across time zones instead of getting lost at handoff. Telemetry & partnershi
About the Team OpenAI's Industrial Compute organization is building and operating the infrastructure foundation for the next generation of AI. Infrastructure Operations works across facilities, hardware, network operations, incident management, data center engineering, delivery teams, and external partners to bring capacity online safely, understand its operational state, and improve it over time. As OpenAI's data center portfolio grows across first-party and partner-delivered capacity, the organization needs clear goals, trusted data, repeatable processes, and systems that make ownership, risk, readiness, and performance visible. This role will help build the operating mechanisms that allow Infrastructure Operations to scale with rigor. About the Role We are seeking a Technical Program Manager to own the systems, data, reporting, governance, and program-management backbone for Infrastructure Operations. Reporting to the Delivery & Operations Lead, you will translate strategy into executable goals and operating cadences, turn operational needs into software and data solutions, and create the mechanisms that keep a rapidly evolving organization aligned and accountable. This role will also own the current 1P+3P delivery-tracking layer within Operations: milestones, delivery timelines, quantity forecasts, risks, decisions, and executive reporting. You will partner closely with 1P Delivery Program Management, Compute TPMs, Data Center Engineering, construction, commissioning, and operations leaders to ensure that delivery information becomes complete, usable input for readiness, handover, and ongoing operations. You will own program health and the operating system around it: the goals, data definitions, workflows, reporting, decision paths, and follow-through that help functional DRIs execute. The ideal candidate is comfortable in ambiguity, technically fluent enough to implement real systems, and relentless about converting scattered information into durable mechan
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We are the first line of defense against fraud and abuse on the Plaid platform. Our mission is to ensure the safety and integrity of our platform for consumers and customers. As a Fraud and Abuse Operations Analyst , you will be responsible for responding to fraud and abuse events, investigating claims, and triaging incidents. We also partner with product and engineering teams to inform and improve fraud mitigation strategies. Responsibilities: Safeguard Plaid's Platform: Participate in the abuse on-call rotation, directly protecting our users and customers by responding to and resolving fraud and abuse events. Your timely actions will be instrumental in maintaining trust and security. Drive Investigations and Mitigate Risks: Investigate fraud and abuse claims from diverse sources, partnering with senior teammates on complex cases. Your findings will inform decisions and strategies, directly impacting Plaid's ability to prevent future incidents and minimize financial losses. Proactively perform threat modeling of abuse surfaces and continuously survey external fraud trends, adversary techniques, tooling, and emerging threat vectors Support Incident Response: Help triage and manage fraud and abuse ev
What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.
About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A
$230K – $260K/yr
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role Millions of people rely on Notion to do their most important work, and protecting that trust is foundational to everything we build. We’re looking for a hands-on Detection Engineer to build and operate the systems and workflows we use to detect and respond to attacks across Notion’s cloud-native environment. You’ll ship high-signal detections, improve the platform that powers them, participate in incident response, and help shape how detection and response engineering scales at Notion. You’ll work closely with Engineering, Corporate Security, and Infrastructure, with broad latitude to identify gaps, prioritize investments, and build what’s needed next. We view detection and response as a software engineering discipline: detections are code, platforms are products, and measurement matters What You'll Achieve Design and maintain high-signal detections across cloud, identity, endpoints, and SaaS environments. Build and improve the detection platform, including rule lifecycle management, tuning, measurement, and rollout safety. Develop tooling and automation that accelerate triage, enrichment, investigation, and detection
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Join WRITER's security team as a staff detection and response engineer and help protect the AI infrastructure that's transforming how the world works. You'll build sophisticated detection systems that identify attacks targeting our AI platform, training data, and model deployments while creating automated response capabilities that scale with our explosive growth. This isn't just traditional security work – you're defending cutting-edge AI/AGI systems against adversaries who are evolving their tactics as fast as AI itself advances. This role combines hands-on security engineering with strategic thinking to stay ahead of novel threats that don't exist in textbooks yet. You'll be the operational arm of our security function, translating threat intelligence into real-time detections, coordinating incident response across multiple teams, and hunting for sophisticated attacks across GPU clusters and distributed training environments. If you're excited by the challen
$192K – $240K/yr
Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Senior Software Engineer, Infrastructure (Release Engineering) at Brex, you will design, build, and operate the core systems that power Brex’s release, observability, and incident management processes. You will partner closely with product, platform, and operations teams to ensure releases are safe, fast, and reliable, and that our infrastructure scales securely as Brex grows. Where you’ll work This role will be based in our San Francisco office. We are a hybrid environment that combines the energy and connect
About the Team The Applied organization brings OpenAI’s most advanced technology to the world through products like ChatGPT and the APIs that power a growing ecosystem of developer and enterprise applications. Data Engineering builds and operates the trustworthy, secure, and reliable data systems that power decisions across OpenAI. About the Role We’re looking for a Data Engineering Manager to lead the Growth & Revenue data engineering team. This leader will own the data strategy and execution for the data subject areas spanning growth accounting across all product surfaces, product partnerships, checkout, billing, payments, revenue, and monetization, helping OpenAI understand how people adopt, engage with, and pay for our products. You will partner closely with several Data Science, Business, and Engineering partners to connect product behavior to trustworthy subscriber, payment, and revenue measurement. In this role, you will: Build, manage, and grow a high-performing, inclusive team across the Growth & Revenue data subject areas. Define the data strategy for all the data subject areas you own. Deliver durable, well-modeled data products that connect product behavior, subscription state, checkout events, payment outcomes, and revenue. Establish trusted metric definitions and data quality standards so product, growth, finance, and executive leaders can make fast, consistent decisions. Partner with Data Science and Product teams to support experimentation, causal measurement, funnel analysis, and scalable self-serve analytics. Partner with Finance and Financial Engineering to ensure analytical revenue views reconcile to financial truth and production billing systems. Raise operational excellence for critical pipelines, including reliability, observability, privacy, governance, and incident response. Set a clear roadmap, make principled tradeoffs, and communicate progress and risk across technical and business stakeholders. You might thrive in this role if yo
About the Team OpenAI’s Legal team helps advance our mission by tackling novel legal issues in AI. Our team brings together professionals across technology, privacy, intellectual property, corporate, employment, tax, regulatory, and litigation. Our regulatory compliance work turns legal requirements into practical programs that support responsible AI development and deployment. About the Role As a Legal Program Manager focused on regulatory compliance, you will build and manage cross-functional programs that translate counsel’s guidance into practical, sustainable operations. Your initial focus may include content moderation and/or frontier AI governance, with the mix shaped by team priorities and your strengths. You will partner with internal and external counsel, other legal program managers, and technical and business teams to coordinate implementation, evidence collection, reporting, and ongoing compliance. You’ll build repeatable systems that scale across regulations, products, and jurisdictions, helping teams navigate emerging requirements with clarity and sound judgment. This full-time role is based in San Francisco, CA, or New York, NY. In this role, you will: Lead regulatory compliance programs end to end: define scope, owners, milestones, dependencies, risks, and escalation paths, and drive execution with counsel and cross-functional partners. Translate counsel’s regulatory guidance into repeatable workflows, controls, and documentation. Depending on your portfolio, this may include content moderation disclosures, transparency reporting, reporting and appeals workflows, or frontier AI model launch readiness, evaluation and risk-management evidence, and incident reporting. Build strong partnerships across Product, Engineering, User Operations, Governance, Risk and Compliance (GRC), Global Affairs, Communications, and Go-to-Market to align program priorities and deliverables. Support regulatory inquiries, audits and investigations with counsel, organizing ev
About the Team OpenAI's Environmental, Health & Safety (EHS) team partners across the company to enable safe, responsible growth. This role will be a senior EHS partner to our robotics operations as our operating footprint expands. You will work closely with Robotics, Engineering, Operations, Facilities, Workplace, Construction, Security, People, Legal, and external partners to integrate safety into how our facilities are designed, built, staffed, and operated. About the Role We are seeking a Senior Manager, EHS - Robotics to provide senior, hands-on safety leadership across multiple sites. This role is for an experienced EHS leader who can operate independently in a fast-moving technical environment, anticipate risk before it becomes an incident, and build practical safety systems that scale with the business. Our robotics operations are scaling quickly, with a growing mix of construction, commissioning, workforce expansion, and around-the-clock operations. This creates a complex and evolving risk environment and requires strong preventive planning, consistent field presence, and sustained incident-management oversight. This leader will own site-level EHS strategy and execution for the robotics portfolio while partnering with the broader EHS team on company-wide standards and programs. The successful candidate will look around corners: planning for future ramps, identifying requirements early, influencing design and operating decisions, and creating durable systems that enable teams to move quickly without compromising safety. In this role you will: Serve as a senior EHS partner across robotics R&D operations, owning site-level safety strategy, priorities, and execution as the footprint scales. Anticipate EHS requirements for new operations, equipment, processes, construction phases, staffing ramps, and 24/7 operations, and translate them into practical plans before work begins. Lead proactive risk identification and control across robotics activities, incl
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Join WRITER's security team as a staff detection and response engineer and help protect the AI infrastructure that's transforming how the world works. You'll build sophisticated detection systems that identify attacks targeting our AI platform, training data, and model deployments while creating automated response capabilities that scale with our explosive growth. This isn't just traditional security work – you're defending cutting-edge AI/AGI systems against adversaries who are evolving their tactics as fast as AI itself advances. This role combines hands-on security engineering with strategic thinking to stay ahead of novel threats that don't exist in textbooks yet. You'll be the operational arm of our security function, translating threat intelligence into real-time detections, coordinating incident response across multiple teams, and hunting for sophisticated attacks across GPU clusters and distributed training environments. If you're excited by the challen
About the Team The Core Services organization builds and runs the mission-critical online services that product teams rely on in production. We own foundational distributed systems and platform capabilities that enable reliable execution, high-performance services, and large-scale file/data needs across our products. This team is distinct from developer infrastructure and data infrastructure—our focus is production service foundations and core runtime services. About the Role We’re hiring an Engineering Manager, Core Services to help lead teams responsible for highly reliable, high-scale distributed systems that sit on the critical path for OpenAI products. Your team will own foundational production systems that OpenAI’s product engineering teams build on. You’ll collaborate closely with product and infrastructure partners to ship reliable services quickly, and help scale systems and teams as OpenAI grows. You’ll partner closely with senior engineering leaders to scale the org, mature operations, and drive major platform initiatives. This role requires strong technical ability. You’ll be responsible for: Managing and growing a high-performing team of infrastructure engineers. Leading teams building and operating large, critical production platforms, including cluster reliability, scaling, and rollout safety. Building and operating mission-critical distributed systems with strong operational rigor (SLOs, incident response, capacity planning, reliability). Setting technical direction for platform foundations such as workflow/orchestration capabilities, large-scale file/blob/storage services, and core service foundations. Partnering with a broad set of stakeholders, including product engineering, adjacent infrastructure teams, and (where relevant) finance/cost partners. Coaching, mentoring, and developing engineers and emerging leaders. You might thrive in this role if you: Have significant experience leading teams that run mission-critical infrastructure in production
About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability Team Focus Areas Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely. Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads. Databases: Building high performance, distributed database systems that power all of OpenAI's product stack. In this
Other cities to consider
More places hiring for this role
Get new incident commander jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime