Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng
Jobiba hiring network
Senior Incident Response Analyst React Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior incident response analyst react jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Lead the technical design and architecture of major platform initiatives, author design documents and build consensus across engineering teams. Define technical roadmaps for complex, multi-quarter projects that span multiple teams. Make critical architectural decisions for company documentation infrastructure, balancing scalability, reliability, and developer experience. Evaluate and set direction for integrating emerging technologies, including AI/LLM capabilities, into company documentation platforms and authoring tools. Establish and evolve engineering standards, best practices and technical guidelines for the team and broader organization. Partner with engineering teams across the company to understand documentation needs and design integrated solutions. Design, build and maintain scalable, reliable and performant services and systems. Contribute high-quality code across the full stack and navigate codebases with different languages and tools. Debug and resolve complex production issues and improve system reliability. Take ownership of system health and incident response. Who you are Minimum requirements Must have a Bachelor's degree or foreign equivalent in Computer Science, Software Engineering, Engineering, or a related field, plus four (4) years of experience in Software Engineering. Must have four (4) years of experience in each of the following: - Working in a full stack environment with a foc
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About The Team The Commerce Site Reliability Engineering team is responsible for the reliability, scalability, and day-to-day operation of the platforms that power GoDaddy's Commerce ecosystem. We build and operate shared infrastructure, support critical production systems, and partner closely with engineering teams to ensure services remain secure, resilient, and highly available. As a Senior Site Reliability Engineer, you'll join a team that values ownership, operational excellence, and continuous improvement. Engineers are empowered to identify problems, drive meaningful change, and influence how reliability is delivered across the broader Commerce organisation. From improving operational maturity and reducing toil to modernising delivery platforms and strengthening incident response practices, this team plays a key role in enabling engineering teams to move quickly and safely. You'll work closely with engineers across infrastructure, cloud, security, networking, and application teams while helping shape the future of reliability engineering at GoDaddy. This role offers significant opportunity to broaden your impact, develop technical leadership skills, and grow toward Staff and Principal engineering positions over time. What you'll get to do... Lead reliability and operational improvement initiatives across GoDaddy's Commerce platform, helping engineering teams build and operate services safely and at scale. Own critical production systems, drive incident response and post-incident improvements, and continuously raise the bar
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in a global on-call rotation supporting highly available customer-facing systems. Participate in incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with en
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What you will be doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuous
NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services. What you’ll be doing: Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale. Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation. Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data. Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems. Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation. What we need to see: BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10+ years operating production infrastructure or platform services. Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges. <l
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operati
Location Details: At GoDaddy, the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operat
We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Instacarts Detection Engineering team sits at the core of our Security organization, building and operating the systems that identify, surface, and respond to threats across one of North America's largest grocery technology platforms. We own the full detection lifecycle, from telemetry collection and signal design to automated response, across a complex, cloud-native environment spanning endpoint, cloud, container, and SaaS. As a Senior Detection Engineer II, you'll be a technical anchor on the team: developing high-fidelity detection logic, hunting for novel attacker techniques, and raising the bar for how we think about coverage, quality, and scale. You'll work closely with Engineering, Red Team, Incident Response, Fraud, and Trust & Safety to ensure our detections reflect real-world adversary behavior; not just signatures. We operate with a detectio
We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Instacarts Detection Engineering team sits at the core of our Security organization, building and operating the systems that identify, surface, and respond to threats across one of North America's largest grocery technology platforms. We own the full detection lifecycle, from telemetry collection and signal design to automated response, across a complex, cloud-native environment spanning endpoint, cloud, container, and SaaS. As a Senior Detection Engineer II, you'll be a technical anchor on the team: developing high-fidelity detection logic, hunting for novel attacker techniques, and raising the bar for how we think about coverage, quality, and scale. You'll work closely with Engineering, Red Team, Incident Response, Fraud, and Trust & Safety to ensure our detections reflect real-world adversary behavior; not just signatures. We operate with a detectio
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Site Reliability Engineer (SRE) - Security and Data Systems Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software engineering and systems administration, where you'll be responsible for building and maintaining highly reliable, scalable, and secure infrastructure. You will be a key contributor, applying your expertise to automate manual processes and proactively solve complex problems before they become incidents, handling incidents, and includes on-call shifts. * This position requires the ability to access U.S. National Security information. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15 ) upon hire. Responsibilities Platform & Reliability: Design, build, and maintain the core infrastructure that underpins our security SaaS offerings, ensuring high availability, performance, and scalability. This includes building and operating the tooling for our Snowflake data systems. Automation: Develop robust automation using code to eliminate toil and ensure consistency across our environments. You'll be a key driver in automating everything from infrastructure provisioning to application deployment and incident response. Security &
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Site Reliability Engineer (SRE) - Security and Data Systems Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software engineering and systems administration, where you'll be responsible for building and maintaining highly reliable, scalable, and secure infrastructure. You will be a key contributor, applying your expertise to automate manual processes and proactively solve complex problems before they become incidents, handling incidents, and includes on-call shifts. Responsibilities Platform & Reliability: Design, build, and maintain the core infrastructure that underpins our security SaaS offerings, ensuring high availability, performance, and scalability. This includes building and operating the tooling for our Snowflake data systems. Automation: Develop robust automation using code to eliminate toil and ensure consistency across our environments. You'll be a key driver in automating everything from infrastructure provisioning to application deployment and incident response. Security & Compliance: Work closely with our security teams to embed a security-first mindset into all our processes and infrastructure. You will be responsible for ensuring our systems and data platforms are compliant with industry standards. Incident Response: Participate in on-call rotations and be a primary responder for critical inci
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What You Will Be Doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuously improve reliability and efficiency. Build and maintain runbooks, standard operating procedures, and comprehensive documentation for storage services and automation.<
Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. Remote (US East Coast preferred, for timezone coverage) About the team Cloud Infrastructure owns the platform every Synthesia product runs on — AWS, Kubernetes, MongoDB, Temporal, our observability stack, and the vendor and cost relationships underneath them. We're a small, high-leverage team scaling toward a domain-ownership model: small groups that both build and operate the systems they're accountable for. The role We're hiring a dedicated SRE to take real ownership of operational excellence across Cloud Infrastructure. Today, too much critical operational knowledge — vendor relationships, cost management, and incident response — lives with one or two people. Your mission is to take genuine ownership of those domains, make them resilient to any single person, and raise the bar on how reliably we run. This is not simply a ticket-queue or keep-the-lights-on role. You'll own domains end to end: understand them deeply, operate them well, and build the automation and tooling that make them boring . We deliberately pair operational and engineering work so the role grows rather than narrows. What you'll own Incident management & operational excellence — take custody of the incident process: on-call quality, resp
About the Role Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Senior Platform Engineer, you'll own medium-to-high-complexity platform projects end-to-end and help shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll partner with Staff engineers and product teams to make Kubernetes effortless across the engineering org, building self-service automation and scalable AWS infrastructure that lets product teams ship faster, safer, and with less cognitive load. If you're excited about building the systems that other engineers will rely on every day, this role is for you. Key Responsibilities Lead high-impact platform projects — design and ship capabilities that move the needle on developer experience, reliability, or security, and set the bar for quality, testing, and safe deployment practices. Build the AI-augmented platform. Design tooling and workflows that help engineers get more out of AI-assisted development — think infra primitives that are easy to reason about, automated review, and policy-as-code that keeps the guardrails strong as AI shifts how code gets written. Own Infrastructure-as-Code for Kubernetes, AWS, and GCP using Terraform, Helm, Kustomize, and emerging tooling — and make it consumable enough that an LLM can safely PR against it. Evolve our CI/CD backbone (Argo CD / Workflows / Rollouts, GitHub Actions) to make deploys faster, safer, and easier to reason about. Instrument and operate. Drive observability with Datadog and Amplitude, own dashboards and SLOs, and use the data to push reliability forward. Participate in on-call, lead incident response when needed, and turn postmortems into durable platform improvements. Reduce toil and tech debt with pragmatic remediation
Get new senior incident response analyst react jobs by email
Daily job updates · Unsubscribe anytime