Jobiba hiring network

Incident Commander Jobs

589 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current incident commander jobs. Use filters to narrow by work mode, employment type, experience and date posted.

R
Roblox
📍 San Mateo• Full-time• From $295.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer on the Sharing team, you will lead large, multi-team initiatives with long-term technical vision and group-level impact. You will be expected to define and drive the platform-wide strategy powering how millions of users capture, share, and discover content on Roblox. In this role, you will set engineering standards, mentor senior engineers, and serve as a key architect of our long-term technical direction. You Will Drive Strategy & Execution: Own the outcome of complex, business-critical programs spanning several teams, often lasting years. Innovate at Scale: Develop and drive a multi-year technical vision for content creation and sharing, anticipating scale, technology, and business evolution. Elevate Reliability: Lead high-severity incident response across groups; drive durable systemic solutions that improve reliability and velocity. Architect Foundations: Regularly improve shared infrastructure and foundational systems, introducing frameworks that uplift development speed across the organization. Align Teams: Aligns multiple teams on shared technical direction, producing detailed design docs, phased roadmaps, and planning models that balance short an

javaawsgit
View job →
R
Roblox
📍 San Mateo• Full-time• From $293.8K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox’s database team develops the next-generation, multi-tenant database platform that elastically scales and underpins every online data workload at Roblox. As a principal engineer on the database team, you will shape the architecture, build and launch critical database capabilities that keep our services fast, reliable and efficient at global scale. You will report to the Technical Director for Storage. You will: Design and implement new engine features —indexing, storage formats, WAL and replication protocols, sharding, and query-planner enhancements—that push latency, throughput, and availability boundaries. Evolve the control plane to deliver elastic scaling, autonomous healing, and zero-downtime schema or tenant moves across global regions. Profile and optimize critical code paths using kernel-level tracing and advanced performance tooling; drive systematic tail-latency reductions. Establish engineering best practices by leading design reviews, performance benchmarks, failure drills, and post-incident retrospectives. Automate everything : develop frameworks for testing, CI/CD, rollout safety, observability, and autoscaling so that the platform operates hands-off at scale. Ment

sqlpostgresqlmysql
View job →
R
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. We are seeking a highly motivated Analyst to support and continuously improve the tooling ecosystem that powers Roblox's Safety, Privacy, Trust & Safety, and Support Operations teams. In this role, you will work at the intersection of operations and technology, helping scale tooling, improve operational workflows, and ensure agents across global vendor sites have the tools they need to operate efficiently. You will join a growing team of administrators, incident managers, and product support specialists in India, providing comprehensive tooling and admin support to our global operations teams. You will contribute to building and maintaining a robust suite of agent tooling by supporting systematic changes, identifying improvement opportunities, and delivering exceptional service to internal stakeholders. Work Schedule : This role is currently aligned to a 2:00 PM to 11:00 PM IST schedule. As part of supporting a 24×7 operational environment, working hours may change based on business requirements, and you should be flexible to work across rotational shifts, including weekends and holidays, as needed. YOU WILL Administer and configure applications – including Zendesk, Absorb, Jira,

awsgitagile
View job →
R
Roblox
📍 San Mateo• Full-time• From $187.8K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior QA Engineer, you will be the first dedicated QA hire for the Safety Engineering Group, which is responsible for building the infrastructure and policies that make Roblox the safest online community in the world. You will report to the QA Engineering Manager for Universal Apps and collaborate across the Safety organization to own end-to-end test strategies for high-stakes, 24/7 incident response systems and parent-facing safety features. This role is a unique opportunity to shape the quality culture of a company-level mandate from the ground up. You will: Partner with Engineering and Product to define and maintain comprehensive, risk-based test coverage for Safety features and systems. Own the end-to-end test strategy for Safety initiatives, including functional, regression, integration, and edge-case validation. Leverage and expand internal automation frameworks to automate critical user flows and drive API-first validation strategies. Identify high-impact automation opportunities and incorporate industry best practices to design robust, abuse-resistant test scenarios. Develop, track, and report on quality metrics (e.g., defect trends, coverage gaps) to proactively surface relea

awsgitrest
View job →
T
Twilio
📍 - Ireland• Full-time• Remote
1mo ago

Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as our next Software Engineer on Twilio’s platform engineering observability team. About the job This position is needed to help our platform engineering observability team. Twilio is undergoing a large-scale observability transformation—and you can help shape the foundation. Observability is a strategic pillar and a key enabler for faster incident response, deeper customer-centric insights, and more cost-effective platform operations. As a Software Engineer on the Platform Observability team, you’ll play a critical role in re-architecting how telemetry flows and is utilized through Twilio—making it structured, accessible, affordable, and actionable. Over the next 3 years, Twilio is rebuilding nearly every component of our observability platform, from data collection to real-time analytics. You will drive core initiatives that shift Twilio from fragmented tooling and wasteful data sprawl to a unified, OpenTelemetry-first observabil

REMOTEpythonjavaaws
View job →
T
Twilio
📍 - US• Full-time• Remote
1mo ago

Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Operator, Global Security Operations Center (GSOC). About the job This position is needed to maintain a broad situational awareness over internal and external physical events that could impact employee safety or pose a threat to business assets or operations across Twilio’s global footprint. GSOC Operators monitor a variety of data sources including internal access control systems, external incident aggregators, travel safety applications, and general open source reporting materials. GSOC Operators serve as the primary contact for employees with physical safety and security questions or concerns. When an incident occurs, the GSOC Operator may dispatch security resources, escalate to crisis management teams, send broad communications to Twilio employees, or execute a mass notification to account for employee safety. Responsibilities: In this role, you’ll: Monitoring and Surveillance Utilize various physical security

REMOTErestaigo
View job →
F
Fin
📍 England• Full-time
1mo ago

Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin is transforming customer service through AI, helping businesses deliver fast, accurate, and reliable support at scale. Trust is foundational to that mission. The Cloud Security team is responsible for protecting the platforms that power Fin. We partner closely with infrastructure and product engineering teams to secure cloud environments, detect emerging threats, respond to incidents, and build the security foundations that enable teams to move quickly with confidence. The team owns critical cloud security capabilities including detection engineering, cloud security monitoring, incident response, cloud security controls, and the security tooling that protects Fin's production environments. The team is responsible for securing the cloud platforms and production systems that underpin every Fin customer interaction. The mission of the team is to help Fin build and operate trusted AI-powered customer service experiences by making security a natural part of how our cloud

F
Fin
📍 Dublin• Full-time
1mo ago

Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin is transforming customer service through AI, helping businesses deliver fast, accurate, and reliable support at scale. Trust is foundational to that mission. The Cloud Security team is responsible for protecting the platforms that power Fin. We partner closely with infrastructure and product engineering teams to secure cloud environments, detect emerging threats, respond to incidents, and build the security foundations that enable teams to move quickly with confidence. The team owns critical cloud security capabilities including detection engineering, cloud security monitoring, incident response, cloud security controls, and the security tooling that protects Fin's production environments. The team is responsible for securing the cloud platforms and production systems that underpin every Fin customer interaction. The mission of the team is to help Fin build and operate trusted AI-powered customer service experiences by making security a natural part of how our cloud

A
1mo ago

We're looking for an Engineering Manager who combines strong technical judgment with people leadership to help build the systems and team practices that keep Asana resilient at scale. This role is a great fit for someone who enjoys turning broad reliability problems into clear priorities, and can balance foundational engineering work with iterative delivery. You'll help shape what Platform Reliability means at Asana while building a team that ships durable systems, strong operational practices, and high-trust partnerships. You will define the reliability roadmap for a rapidly growing global platform, transforming reliability into a core architectural advantage. You'll partner closely with platform engineering, infrastructure, and product teams in Warsaw, Reykjavik and San Francisco to protect Asana under real-world load, improve how traffic and failure modes are handled, and ensure reliability is designed in rather than added later. This is a role for someone who can lead through influence, coach engineers, and raise the bar for execution and collaboration. This role is based in our Warsaw office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do, and your recruiter can share more about the in-office requirements. We offer a Contract of Employment (UoP) for our employees in Poland. What you'll achieve Build and lead a new Platform Reliability team, hiring and developing engineers while setting a clear standard for collaboration, ownership, technical excellence and growth. Partner with technical leaders to define the roadmap for core reliability systems such as load shedding, rate limiting, circuit breakers, traffic controls, and other platform guardrails. Establish and evangelize best-in-class operating practices for incident response, postmortems, and proactive risk management through SLOs a

awskubernetesrest
View job →
A
Asana
📍 Warsaw• Full-time• $372K – $432K/yr
1mo ago

Asana’s rapid growth brings new challenges in keeping our systems fast, reliable, and resilient. As our product evolves, we’re making a major investment in reliability – and building a brand new SRE team in Warsaw is a key part of that strategy. This is your chance to help shape it from day one. This isn’t a traditional “ops” role – we’re looking for strong software engineers who are passionate about building reliable, distributed systems. You’ll work closely with a small SRE team in San Francisco, infrastructure engineers in Reykjavik, and an established infrastructure team in Warsaw. Warsaw will be a significant hub for our future infrastructure engineering and operations. As one of the first engineers here, you’ll have a real say in how we build reliable infrastructure, manage incidents, and support the rest of the company. This role is based in our Warsaw office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do, and your recruiter can share more about the in-office requirements. We offer a Contract of Employment (UoP) for our employees in Poland. What you’ll achieve Influence the future of Asana’s SRE practice, especially as we grow the Warsaw team. Lead reliability-focused projects across our stack – from infrastructure to tooling to incident response. Define and implement Asana’s incident management process – we’re investing here, and you’ll help shape how it works. Build internal platforms and frameworks that help other teams improve the reliability of their services. Be part of (and help shape) a sustainable on-call rotation – shared across teams in Warsaw, San Francisco, and Reykjavik. On average, we handle ~1 page per day, but it’s not constant, and we care about keeping things sane. Work with our stack: AWS, Kubernetes (EKS), Datadog, MySQL (RDS), ElasticSearch (OpenSearch), Redis

typescriptpythonsql
View job →
D
Datadog
📍 New York• Full-time• From $192K/yr
1mo ago

This Engineering Manager will lead the Data Visualizations Explorations team within the Graphing organization, setting product direction, and coaching and developing team members. They will staff and drive projects that build end to end experiences for Datadog’s core users: observability engineers. This includes extending the capabilities of core widgets like Hostmap and Geomap Visualizations and finding innovative ways to leverage existing Datadog data sources. This role also involves close partnerships with other product teams to deeply understand customer needs and deliver compelling data experiences. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Work with Product Management and Design to plan and staff projects for the Data Visualization Explorations team Coach and develop engineers at various levels Ensure strong cross-team communication and design best practices around project management, code review, architecture patterns, and more. Proactively anticipate cross-team dependencies and blockers to goals. Identify opportunities to appropriately reuse or customize features across dashboards, notebooks, and product pages. Deeply understand the needs of our customers and other Datadog products we work with. Participate in customer conversations, review product briefs, read feature requests, and coach team members to adopt these practices as well. Define and maintain high standards for operations practices, including bug triage and remediation, incident response, and gathering and analyzing performance telemetry for our widgets. Who You Are: At least 2 years of people management experience in a software engineering or similar setting Strong TypeScript/JavaScript skills, including familiarity with front

javascripttypescriptjava
View job →
D
Datadog
📍 Colorado• Full-time• From $96K/yr
1mo ago

We are Datadog's in-house product experts. The Technical Solutions team enables Datadog's worldwide growth by educating potential clients and ensuring that existing customers are happy and successful. Premier Technical Support Engineers (PSEs) are primarily focused on assisting prospects and customers with any technical questions about Datadog. PSEs engage with Datadog’s Premier customers not only via standard technical support channels, but also get involved via cadence calls, business reviews, and side projects. You will be immersed in a fast-paced environment where you will be challenged, but will also immediately witness your contributions to Datadog and to our customers. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Respond to Premier customer requests (phone / chat / tickets) on our fast paced team while continuing to educate our clients on the use of the platform Develop relationships with our Premier customers, working hand-in-hand to understand their specific environment Reproduce customer issues and assist customers implementing 1,000+ Datadog integrations Handle urgent escalation requests that may result in customer-facing troubleshooting calls, and internal or external incident management Build subject matter expertise in many Datadog product areas Autonomously troubleshoot complex and/or high-priority customer issues without guidance Drive product and engineering conversations based on needs, use cases, and problems learned during client interactions Provide mentorship to junior members of the team and serve as their escalation partner Participate in routine health check meetings with Premier customers Build out and improve documentation and knowledge base articles for a variety of technologies Who You Are: Experienced in mul

linuxaigo
View job →

Observability Pipelines (OP) is Datadog's on-premise, vendor-agnostic telemetry pipeline product. As an Engineering Manager on the team, you'll own people management and engineering execution for one of OP's core missions, spanning areas like Integrations (ingesting from and routing to the many source and destination systems customers rely on), streaming insights, cost control, or pipeline capabilities, reliability and scalability. You'll partner directly with Product to help shape the roadmap, and work closely with your peer EMs and senior ICs to define how OP operates and grows. This is an opportunity to build your management craft while having real influence over the technical direction of a fast-growing product area. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Own people management and engineering execution Establish a strong operating rhythm for the team Drive high standards for on-call rotations and incident response Partner with Product on the roadmap, balancing product priorities with technical realities Lead, coach, and grow the careers of engineers on your team Who You Are: Experienced managing engineers directly, comfortable owning a team’s operating rhythm end-to-end, from planning through execution and stakeholder communication to incident and on-call ownership Have a technical background in distributed systems and data infrastructure Have experience with on-premises or customer-installed software concepts A product-minded partner to have on the team — you enjoy working with Product on strategy Experience with high-performance or Rust-based data pipeline systems Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications o

aigorust
View job →
D
Datadog
📍 New York• Full-time• From $252K/yr
1mo ago

We're looking for an Engineering Manager II to own and grow the Observability Pipelines engineering org at a pivotal moment in the product's lifecycle. Observability Pipelines is Datadog's on-premise, vendor-agnostic telemetry pipeline product, with a lot still to build as it grows and scales. It sits at the center of a fast-consolidating market, is central to Datadog's data pipeline optimization story for Logs and Metrics customers. This is a build-and-scale opportunity: you'll grow the management and technical leadership layers, co-own the roadmap with Product, and define how this org operates as it continues to expand. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You'll Do: Directly manage the OP org including EM1s across NYC and Paris, set technical direction, and be the connective tissue across a distributed team Build out the management and technical leadership layers as the org continues to grow - today ~20 ICs Partner directly with Product to co-own the roadmap and strategy, helping decide where OP’s engineering investment goes next Set and evolve the operating rhythm across the group: planning cadence, on-call and incident standards, and cross-team alignment Own key cross-org relationships with the SaaS Logs Pipelines team, the BYOC team, and the Vector open-source community Coach managers and senior engineers, and build the succession and growth plans that let the org scale beyond you Who You Are: Experienced managing managers across distributed teams, with a track record of raising the bar on how those teams operate, not just delivering through them Back

aigorust
View job →
D
Datadog
📍 Denver• Full-time• From $92K/yr
1mo ago

We are Datadog's in-house product experts. The Datadog Federal Support Engineering team is dedicated to serving as highly trusted technical advisors for our Public Sector customers, who operate within some of the most highly regulated and security-constrained environments. These customers include various government agencies and organizations with critical, sensitive missions. As a Federal Support Engineer 3, this role places you at the forefront of supporting these customers' mission-critical workloads. These complex workloads are often deployed across sophisticated hybrid and multi-cloud architectures, requiring deep expertise in cloud technologies, monitoring, and security best practices. Your primary responsibility is to ensure the complete success of these customers across their entire lifecycle with Datadog. Whether you’re looking to learn from the best or be the best, the Federal Support team is dedicated to furthering personal development and team success. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Engage with public sector customers via multiple channels (ticketing system, live chat, calls, and screensharing tools) to identify and resolve technical support requests. Troubleshoot, investigate, and resolve complex technical issues in highly constrained environments across Datadog's 1000+ integrations, often with limited logs or sanitized data. Handle urgent escalation cases that may result in customer-facing troubleshooting calls, and internal or external incident management Become a subject matter expert in many Datadog product areas Partner with Product, Engineering, and Account teams to to validate bugs and advocate for customer-impacting improvements Provide mentorship to junior members of the team and serve

restmicroservicesai
View job →
🔔

Get new incident commander jobs by email

Daily job updates · Unsubscribe anytime