Jobs in United States

Incident Response Analyst in San Francisco

55 active opportunities · Updated October 2026

Explore current incident response analyst jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -73.6%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently. You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. EXAMPLE INITIATIVES You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services. Building AI-assisted tooling for incident triage and response. RESPONSIBILITIES Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking. Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code. Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution. Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations. Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, con

KubernetesGitMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -73.6%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As the Engineering Manager for Baseten's Cloud Platform team, you will directly manage a team of cloud platform engineers responsible for building the systems and processes that keep our infrastructure scalable, reliable, and efficient — from automated deployments and monitoring to performance optimization and incident response. You are a people-first leader with a strong cloud infrastructure background. You set a high bar for reliability and operational excellence, engage credibly in technical discussions and code reviews, and know how to build a culture of ownership and accountability. You'll spend most of your time close to the work: unblocking your team, shaping technical direction on day-to-day decisions, and developing your engineers. At Baseten, we work closely with our users to understand their struggles operationalizing ML — you'll keep your team connected to that mission and translate user learnings into better infrastructure. RESPONSIBILITIES Recruit, hire, and grow a high-performing team of cloud platform engineers; provide ongoing coaching, feedback, and career development through regular 1:1s. Set clear performance expectations, hold a high bar, and create an environment where engineers do their best work. Foster a culture of ownership, accountability, and continuous improvement. Drive day-to-day technical decisions through design reviews, code reviews, and architectural discussions; translate th

KubernetesCI/CDGitMachine Learning
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -73.6%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: As a Software Engineer at Baseten, you will own one of the most critical surfaces of our business: pricing, billing, and revenue infrastructure. As we launch more and more products— billing is no longer just operational plumbing. It is a strategic lever for growth. This role will establish clear ownership of billing as a function and create leverage for Finance, Sales, and GTM teams while maintaining a seamless customer experience. RESPONSIBILITIES: Own Baseten’s end-to-end billing and revenue infrastructure, including pricing, invoicing, metering, and reporting foundations. Build and evolve our billing platform and integrations (including Orb), ensuring correctness, auditability, and a high-trust experience for customers and internal teams. Partner closely with Finance, Sales, GTM, and Forward Deployed Engineering to turn real-world workflows into reliable internal tooling and automation (quoting, approvals, renewals, usage reconciliation, revenue reporting). Design systems that scale with new products, packaging, and go-to-market motions, making billing a strategic lever for growth. Drive reliability and operational excellence for revenue-critical workflows: monitoring, alerting, incident response, backfills, and clear runbooks. Lead from the front on high-impact projects: clarify requirements, propose crisp technical approaches, ship iteratively, and raise the bar on quality and velocity. Debug and resolve

Machine LearningAIGoRust
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -70%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Security Engineering is the engineering function inside the Plaid security org that focuses on developing the industry-leading security systems and infrastructure. Security Engineering owns most of Plaid’s security-related infrastructure: secure data storage, key management systems, internal identity platform, internal authentication systems, internal permission management, and internal authorization service. We develop solutions across data encryption, key management, access control, and data loss prevention to protect sensitive consumer data. We believe in the Zero Trust security model and are always looking for ways to improve our authentication and access control platforms. About the role You will develop security capabilities to secure Plaid infrastructure and to secure sensitive data access. You will own, maintain, and build Plaid’s security infrastructure and services like Key Management System and Secure Token Service. You will consult with product engineers to ensure Plaid services meet security standards. You will help educate and support other engineering teams to improve security in their own products and services. You will assist with Plaid’s incident response and security awareness pro

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI’s Network Engineering team within IT and Security advances the mission of deploying artificial general intelligence (AGI) for the benefit of all by delivering secure, scalable, and resilient network services. We build and operate the connectivity that supports OpenAI’s offices, labs, campuses, cloud environments, people, and devices. By combining strong network fundamentals with security, reliability, automation, and user-centered design, we enable impactful AI research, corporate operations, and product innovation. About the Role As a Network Engineer at OpenAI, you will design, operate, and continuously improve the global networks that connect our offices, labs, campuses, PoPs, cloud environments, people, and devices. The role spans strategic platform engineering and responsive production operations: you will shape architecture, standards, roadmaps, lifecycle plans, and automation while supporting incidents, escalations, and time-sensitive delivery. Operational signals will inform what we stabilize, simplify, standardize, or automate next. We work backward from user needs, investigate root causes, own outcomes end-to-end, and move quickly without compromising security. We are looking for a versatile engineer who can make pragmatic reliability and security tradeoffs, communicate clearly, and turn recurring operational work into durable platforms, tooling, and standards. You will partner across IT, Security, AppEng, Research, Applied, workplace teams, carriers, and vendors. In this role, you will: Design, implement, and operate secure, scalable enterprise networks across offices, labs, campuses, PoPs, cloud connectivity, and hybrid environments. Set strategic direction for network services through architecture, standards, roadmaps, lifecycle planning, capacity strategy, and measurable reliability outcomes. Own production operations, including on-call, incident response, escalations, and time-sensitive delivery, while protecting user experience,

PythonAWSAzureCI/CD
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%

From $2M/yr

Quick readStrong listing-quality and freshness signals

About the team OpenAI’s mission is to build safe artificial general intelligence (AGI) which benefits all of humanity. This long-term undertaking brings the world’s best scientists, engineers, and business professionals into one lab together to accomplish this. In pursuit of this mission, our Go To Market (GTM) team is responsible for helping customers learn how to leverage and deploy our highly capable AI products across their business. The Cybersecurity specialist sales team partners with Account Directors, Technical Success, Marketing, and Partnerships to drive cybersecurity adoption that help bring AI to as many users as possible. About the role Our Sales team has a unique mission to help cybersecurity customers understand the deep impact that highly capable AI models can bring to their businesses, operations, employees, and customers. This role is a mixture of technical understanding, industry expertise, vision, partnership, and value-driven strategy. As an Account Director focused on Cybersecurity, you will own executive-level relationships with leading cybersecurity firms and help them safely and effectively deploy OpenAI’s technology across their organizations. You’ll work with customers to identify and scale high-impact use cases across areas such as security operations, threat intelligence, risk management, incident response, vulnerability management, workforce enablement, customer support, and enterprise knowledge management. You’ll be a key driver of opportunities through the entire sales cycle, from pipeline generation to closure and successful deployment. You’ll work with researchers, engineers, and solution strategists to help customers transform their operations and evolve the cybersecurity industry with AI. This role is based in San Francisco, Seattle or New York. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. We are open to US-based remote candidates. In this role, you’ll: Support Accou

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI’s Cyber team works to make frontier AI a decisive advantage for defenders. The Cyber Blue Team is an operator-led group focused on turning real defensive problems into better models, useful products, safe Codex workflows, and integrations with the security tools defenders already use. Our ambition is simple: Raise attacker cost. Lower defender toil. Prove it by defending OpenAI; scale it through the ecosystem. We are not setting out to build another SIEM or autonomous SOC. We want to build the AI reasoning and workflow layer that helps security teams investigate threats, create and validate detections, improve their controls, and respond with greater speed and confidence. About the Role We are looking for a Product Manager to help build a new generation of AI-powered cyber defense products. You will work closely with security practitioners, researchers, engineers, designers, internal security teams, customers, and technology partners to turn emerging model capabilities into products that solve meaningful defensive problems. This is an early-stage product role. The work will span product discovery, prototyping, evaluation, development, launch, and iteration. You will help the team identify where AI can create the most value for defenders and translate those opportunities into clear, usable, and trustworthy product experiences. Initial areas of focus may include: Detection engineering and detection-content development Threat hunting and investigation Security validation and control testing AI-agent and MCP runtime defense Integrations with security platforms and enterprise workflows Safe, governed assistance for incident response The specific roadmap will continue to evolve based on model progress, practitioner needs, internal learnings, and customer feedback. In This Role, You Will Work with security practitioners to understand high-value defensive workflows, recurring pain points, and opportunities for AI to materially improve outcomes. Help sh

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the team The AI Deployment Engineering team is responsible for helping developers and enterprises safely and effectively deploy OpenAI technologies in production. We act as trusted technical advisors and thought partners for customers, working side by side with their teams to identify high-value use cases, design practical architectures, and move from prototype to durable deployment. Cybersecurity is one of the most urgent domains where AI can help. Security teams are under pressure to reason across code, logs, infrastructure, tickets, alerts, and vulnerability data faster than ever. As frontier models become more capable, organizations need deep technical guidance on how to evaluate, validate, and safely deploy AI systems in security-critical workflows. About the role We are looking for a Cyber AI Deployment Engineer to partner with customers and help them apply OpenAI models, APIs, Codex, and agentic workflows to real cybersecurity use cases. You will work with CISOs, security executives, application security leaders, SOC teams, security engineering teams, and hands-on practitioners to identify where AI can create measurable security outcomes. This is a customer-facing technical role for someone who can move fluidly between executive strategy, practitioner-level cyber depth, and hands-on solution design. You will help customers evaluate and deploy workflows such as secure code review, vulnerability triage, threat modeling, remediation, SOC and incident response workflows, detection engineering, cloud security, GRC automation, and security validation. You will collaborate closely with Sales, Solutions Engineering, Product, Engineering, Research, and Security to turn customer needs into safe deployment patterns, reusable field assets, and product feedback. This role is based in our San Francisco HQ. We offer relocation support to new employees. In this role, you will: Deeply embed with strategic customers as the technical lead for AI-enabled cybersecurity work

JavaScriptPythonJavaAWS
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team The Support team is central to ensuring that our customers' experience with our products is nothing short of exceptional. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. Given OpenAI’s breakneck shipping cadence and growth – and the expectation that it will only accelerate – our ability to architect automation systems and agentic workflows for scale is central to our ability to maintain exceptional support quality in the face of AGI. About the Role As a Support Vendor Manager, you will own the health, performance, and long-term scalability of multiple support partner and vendor relationships. This is a vendor leadership role first and foremost: you will drive commercial and operational accountability (SLAs, QBRs, escalation paths, remediation plans), while also building the operating model that enables support to scale without linear headcount growth. You’ll collaborate closely with User Operations teams (e.g., Trust & Safety, Fraud & Risk), Systems/Tooling, Data partners, and Product/PM stakeholders as we launch new workflow and launch and scale new programs. You’ll be responsible for: End-to-end vendor leadership: Own day-to-day oversight, relationship health, and executive-level accountability for multiple support vendors/BPOs. Performance management & remediation: Define and manage SLA/KPI performance expectations, run WBRs/QBRs, identify performance gaps, and drive structured turnaround plans with clear owners and timelines. Escalation and risk management: Serve as the primary escalation point for vendor issues, including incident response, surge events, quality regress

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun

AWSAzureGCPKubernetes
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -86%

$272K – $320K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role Build the most advanced AI Meeting Notes product — and expand it into broader “AI data capture” features that help teams turn conversations into durable context, tasks, and knowledge. Our mission is to 10x the rate of business context & data that enters Notion — optimized for agents — so teams get superhuman memory across workstreams and customers. Notion workspaces that use AI Meeting Notes already enter 6x more data on a daily basis, so we’re well on our way. What You'll Achieve Ship end-to-end product experiences across capture → transcript → summary → follow-ups (full-stack ownership). Make meeting & data capture feel effortless and magical (e.g., speaker identification via audio waveforms, richer in-meeting UX, smarter organization). Improve summary quality that teams trust: structure, factuality, and citations that make downstream agents and humans more capable. Raise the bar on reliability & observability across the pipeline (SLOs, debugging workflows, incident response) for realtime systems. Build agentic meeting workflows that turn discussions into tasks, follow-ups, and organized knowledge — so “w

RestAIGoRust
P
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and performance of Postman’s AI-powered API and agentic systems in production. This role focuses on monitoring, availability, incident response, and automation to support AI services and tools trusted by millions of developers globally. What You’ll Do Develop and manage reliability metrics (SLOs) for AI-driven API services and agentic AI platform features Implement comprehensive observability and monitoring systems for real-time performance and fault detection Design and drive automated failover, recovery, and incident response strategies for high-availability AI infrastructure Optimize resource utilization, particularly GPU/accelerator efficiency, ensuring cost-effective AI system operation Collaborate closely with engineering, platform, and product teams to align reliability efforts with broader organizational goals Lead efforts to build internal tooling and automation focused on AI system stability and operational excellence Drive continuo

AIGoRustDevOps
F
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -85.5%

From $165.4K/yr

Quick readStrong listing-quality and freshness signals

About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. What you'll do Identity & access Advance our identity posture: SSO coverage, phishing-resistant MFA rollout, SCIM lifecycle automation, and least-privilege access across the SaaS and cloud estate. Build the detections and guardrails that catch account takeover, MFA fatigue attacks, and session token theft before they turn into incidents. Endpoint & device lifecycle Write and ship device policy as code — configuration profiles, remediation scripts, and enforcement rules across macOS and Windows — with staged rollout and rollback built in from day one. Maintain and improve our EDR stack's detection and response coverage across the fleet. SaaS posture Reduce SaaS risk at scale through SSPM tooling and automation , including detection of risky OAuth grants, shadow IT, and configuration drift across our critical SaaS applications. Own security configuration for the SaaS tools hundreds of Flexporters use daily (Google Workspace, Slack, and similar), and keep pace as we add AI agents and MCP integrations to that surface. Automation & enablement Automate the parts of corporate security that don't need a human — device provisioning, access reviews, vendor securi

PythonRestAgileAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team At OpenAI, our User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We operate at the front line of real-world safety and risk management, translating user and operational signals into timely decisions, effective interventions, and improvements to our systems. This role sits on a team focused on building operational capacity for new, ambiguous, and fast-moving areas of work. The team defines what needs to be built, creates the operating model to support it, and works with partner teams to make the work scalable and durable over time. About the Role We are seeking a Device Safety & Risk Operations Specialist to build the safety operating model for a new category of consumer hardware. This is a senior individual-contributor role for someone who can turn emerging product risks and incomplete requirements into practical workflows, controls, launch plans, and durable systems. You will define how product-safety incidents, critical escalations, regulated cases, and privacy-sensitive issues should be identified, investigated, escalated, resolved, and learned from. You will also establish operational requirements for case management, data access, decision logging, quality assurance, monitoring, and cross-functional response. You will stand up priority workflows through launch and early operations, then help transition them into durable homes across USRO and partner teams. The right person combines deep operational judgment with strong technical and hardware product fluency. They can move from executive-level risk framing to detailed workflow design, tabletop exercises, launch readiness, frontline guidance, and post-launch improvement. Location / work model: San Francisco, CA; hybrid, 3 days/week in-office. Please note: This role may involve exposure to sensitive or concerning material. Strong discretion, judgment, and resilience are essential. In This Role, You Will:

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team At OpenAI, our User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We operate at the front line of real-world safety and risk management, translating user and operational signals into timely decisions, effective interventions, and improvements to our systems. This role sits on a team focused on building operational capacity for new, ambiguous, and fast-moving areas of work. The team defines what needs to be built, creates the operating model to support it, and works with partner teams to make the work scalable and durable over time. About the Role We are seeking a Device Safety & Risk Operations Specialist to build the safety operating model for a new category of consumer hardware. This is a senior individual-contributor role for someone who can turn emerging product risks and incomplete requirements into practical workflows, controls, launch plans, and durable systems. You will define how product-safety incidents, critical escalations, regulated cases, and privacy-sensitive issues should be identified, investigated, escalated, resolved, and learned from. You will also establish operational requirements for case management, data access, decision logging, quality assurance, monitoring, and cross-functional response. You will stand up priority workflows through launch and early operations, then help transition them into durable homes across USRO and partner teams. The right person combines deep operational judgment with strong technical and hardware product fluency. They can move from executive-level risk framing to detailed workflow design, tabletop exercises, launch readiness, frontline guidance, and post-launch improvement. Location / work model: San Francisco, CA; hybrid, 3 days/week in-office. Please note: This role may involve exposure to sensitive or concerning material. Strong discretion, judgment, and resilience are essential. In This Role, You Will:

AWSRestAIGo
🔔

Get new incident response analyst jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime