The IT Systems Engineering team builds the systems and processes that keep Asana running: endpoint management, identity, and the enterprise applications every Asana employee uses every day. As Manager of IT Applications and Client Platform Engineering, you’ll lead four engineers: two on client platform, two on applications. You report directly to the Head of IT Systems Engineering. This is a player/coach role: you set direction, run planning, and step in on escalations and architecture decisions while your engineers own most of the day-to-day. Your primary depth is in endpoint engineering; you have enough SaaS administration background to lead the apps side without micromanaging it. We’re looking for someone who knows endpoint work deeply and is ready to grow as a manager. Your title history matters less than technical foundation and how you work with people. This role is based in our San Francisco office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you’re interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Lead and develop your team. Manage and mentor four engineers (two on CPE, two on Apps). Set priorities, clear blockers, and help each person grow. Own the client platform program. Set strategy and lead execution for Asana’s endpoint fleet: macOS, Windows, and mobile. This means MDM infrastructure, zero-touch deployment, software catalog, and OS lifecycle. We’re primarily an Apple shop with a growing Windows footprint. Strengthen endpoint security. Work with the Corporate Security team to harden our endpoint posture, respond to incidents, and build initiatives that connect security priorities and endpoints. Oversee the applications engineering function. Your Apps engineers manage the
Jobiba hiring network
Incident Commander Jobs
589 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current incident commander jobs. Use filters to narrow by work mode, employment type, experience and date posted.
As a Staff Engineer on the Data Platform Experience team, you'll help shape how Datadog engineering teams build, operate, and evolve products on the Observability Data Platform. You'll lead the design and delivery of shared platform capabilities that reduce developer friction, improve operational visibility, and enable engineering teams to move faster with confidence. This role combines deep distributed systems expertise with technical leadership across multiple teams, influencing platform strategy while remaining hands-on in the code. You'll have the opportunity to solve company-wide challenges spanning cost intelligence, operational tooling, platform health, and developer experience. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead strategic engineering initiatives that improve how product teams build, operate, and evolve services on the Observability Data Platform. Design and build scalable platform capabilities for cost intelligence, including cloud cost allocation, trend analysis, and optimization recommendations. Develop operational intelligence and self-service tooling that helps engineering teams understand platform health, troubleshoot incidents, and improve operational efficiency. Drive reusable platform services and developer workflows that increase engineering autonomy while reducing operational complexity across multiple products. Provide technical leadership across teams by influencing architecture, mentoring engineers, and raising engineering standards through hands-on technical contributions. Participate in the team's on-call rotation and continuously improve platform reliability, observability, and operational excellence. Who You Are: You have experience designing and building large-scale SaaS or cloud platforms with deep expertise i
Data Semantics is Datadog’s authority on semantic knowledge, providing shared infrastructure that powers both Datadog’s product experiences and AI capabilities. As Datadog continues its investment in OpenTelemetry-native observability, semantic interoperability, and AI-powered workflows, this team sits at the center of some of the company’s most strategic platform initiatives. As a Staff Engineer, you will serve as a technical leader for the team, balancing stewardship of critical production systems with the exploration of new platform capabilities that improve how telemetry is modeled, understood, and consumed across Datadog. You will work closely with engineering and product partners to define standards, drive technical direction, and deliver solutions that scale across Datadog’s observability platform. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the technical direction of the Data Semantics team while remaining deeply hands-on in design, implementation, and delivery. Build and scale semantic infrastructure that bridges OpenTelemetry, Datadog-native telemetry, cloud-provider telemetry, and customer-defined data models. Drive platform initiatives focused on schema evolution, telemetry standardization, data insights, and semantic interoperability across Datadog products. Partner with engineering and product teams across the platform to define standards, align stakeholders, and deliver high-leverage platform capabilities. Mentor engineers through design reviews, technical guidance, operational excellence, and long-term career development. Participate in on-call rotations and lead investigation and resolution efforts for complex production incidents affecting critical platform services. Who You Are: You have significant experience designing, opera
As a Staff Engineer on the Data Platform Experience team, you'll help shape how Datadog engineering teams build, operate, and evolve products on the Observability Data Platform. You'll lead the design and delivery of shared platform capabilities that reduce developer friction, improve operational visibility, and enable engineering teams to move faster with confidence. This role combines deep distributed systems expertise with technical leadership across multiple teams, influencing platform strategy while remaining hands-on in the code. You'll have the opportunity to solve company-wide challenges spanning cost intelligence, operational tooling, platform health, and developer experience. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead strategic engineering initiatives that improve how product teams build, operate, and evolve services on the Observability Data Platform. Design and build scalable platform capabilities for cost intelligence, including cloud cost allocation, trend analysis, and optimization recommendations. Develop operational intelligence and self-service tooling that helps engineering teams understand platform health, troubleshoot incidents, and improve operational efficiency. Drive reusable platform services and developer workflows that increase engineering autonomy while reducing operational complexity across multiple products. Provide technical leadership across teams by influencing architecture, mentoring engineers, and raising engineering standards through hands-on technical contributions. Participate in the team's on-call rotation and continuously improve platform reliability, observability, and operational excellence. Who You Are: You have experience designing and building large-scale SaaS or cloud platforms with deep expertise i
Datadog is seeking a strategic, results oriented Director to lead and scale our Technical Support Engineering (TSE) organization across LATAM. This leader will be responsible for strengthening and expanding our regional support presence, building high-performing teams, and ensuring operational excellence as we continue our rapid growth.You will lead a team of managers and support engineers in our LATAM region, partnering cross-functionally to deliver exceptional customer outcomes and consistent global standards. A t Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead, develop, and scale a multi-country Technical Support Engineering organization across LATAM Evaluate and launch new support locations within LATAM as customer growth requires Act as the owner for your team’s metrics and performance - partnering closely with upper management and HR on performance management and employee issues Provide strategic guidance and oversight to Managers on regional projects, ensuring their successful and timely completion. Lead regional hiring efforts to attract and retain top technical talent in competitive LATAM markets. Ensure that all quarterly hiring targets for the region and functional area are met Ensure the successful onboarding and development of Technical Support Engineers Drive cross-functional projects or initiatives to improve team productivity, process, or procedure Collaborate with internal teams and customers on high-priority escalations/incidents and act as a resource to resolve escalations from team members as necessary Conduct regular 1:1’s with team members to provide constructive feedback and skills development. Liaise with other groups throughout the organization (Sales, Customer Success, Engineering, Product) to tackle urgent
MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. We are looking for a talented Senior Software Engineer to join our team as we execute on a multi-year roadmap and prepare to launch and rapidly scale services that handle petabytes of data. Come do some of the most interesting work of your career as we Think Big and Go Far for our customers! This role can be based out of our New York City office (hybrid working model) or remotely in the North America region. What you’ll do Design, build, and operate control plane services powering an elastic and multi-tenant storage layer for thousands of database instances. Solve problems around maintaining high availability and performance during load spikes, hardware failures, cloud provider outages, and other disruptions. Contribute to a culture of operational excellence through dashboards, playbooks, and on-call improvements. Lead complex technical projects from planning through successful deployment with clear stakeholder updates. Partner closely with peers across database, cloud, and infrastructure engineering teams as well as project management to investigate incidents and develop long-term roadmaps. Mentor junior engineers and foster a collaborative team environment. We’re looking for someone with 5+ years of professional software development experience building, deploying, and operating multi-tenant cloud services with a focus on operational excellence. Experience with large backend/compiled codebases, such as Rust or C/C++. Experience with containerization and orchestration platforms (e.g. Kubernetes). Experience with observability tooling (e.g. time series metrics, dashboards). Experience with distributed systems
MongoDB is helping define the next era of application development as organizations modernize legacy systems, build AI-native experiences, and power mission critical workloads across cloud, hybrid, and on-premises environments. From MongoDB Atlas, our AI ready developer data platform, to our industry leading server technology and enterprise software offerings, MongoDB provides a flexible and unified software stack that helps customers build faster, scale with confidence, and support a wide range of modern application needs. MongoDB’s standing in the market reflects that momentum: the company continues to be recognized as a leader in cloud database management while also serving organizations that require the performance, flexibility, and control of self-managed deployments. For candidates, that means the opportunity to join a company with strong market relevance, a builder-first culture, and a clear strategic focus on shaping the future of modern, intelligent applications wherever they run. We’re looking to speak to candidates based in Palo Alto, California for our hybrid working model. Life in MongoDB Technical Services As part of the Support organization, you will be directly interfacing with our largest customers and their most difficult problems. There will almost certainly be sweat. But don’t worry – you’ll have time to prepare before you take on your first escalation. The limits of your understanding will constantly be pushed, and you’ll be challenged to grow. Underlying it all is the fuel that drives our team: customer obsession. We keep mission-critical deployments online, recover from high-pressure incidents with grace and humility, become trusted advisors, and work so closely with our customers that they would swear we were part of their team. These are just a few of the things you can expect to do as a Technical Services Engineer: Become well-versed in core aspects of MongoDB Gain deeper expertise in specialized areas of the product Combine technical
We are seeking a Senior Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Senior Site Reliability Engineer on this new team, you will be responsible for enabling deployment at scale of AI applications and improving the performance, scalability, and reliability of the distributed systems infrastructure for this new product. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Position Expectations Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing Identify and configure key metrics to detect incidents and quantify service health, availability, and performance Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure Mentor early-career SREs and contribute to the team’s operational practices as it grows Qualifications Strong background in software development and operating distributed systems 6+ years of experience building and operating distributed systems, with proficiency in Python, Go, or a similar programming language Experience operating Kubernetes in production and debugging below the abstraction layer, including scheduling, cluster networking, and node-level issues Expertise in cloud infrastructure platforms, in
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Auth0 provides an unparalleled authentication experience for hundreds of millions of users worldwide. Our commitment to reliability is a key foundation of our product and our dedication to exceeding customer availability expectations is a core engineering focus. As a Senior Site Reliability Engineer, you'll join our SRE team based in Europe to ensure our production systems are not only operational but also resilient, scalable, and ready for exponential growth. This isn't just about keeping the lights on; it's about directly contributing to the platform's core resiliency and robustness. You'll be a hands-on builder, crafting solutions that make our system more reliable by design. What you’ll do: Design and build custom software in Go to enhance the platform's reliability, resiliency, and redundancy. Partner with engineering teams to embed reliability principles, improving the availability, performance, and observability of our services. Use your deep understanding of infrastructure and observability principles to identify opportunities for improvement within the product and implement solutions. Contribute to our follow-the-sun on-call rotation, providing rapid, effective response to critical incidents and using your expertise to troubleshoot, mitigate or accurately escalate production issues. Because our team is globally distributed, your on-call shifts will only occur during your standard local working hours. Develop and refine our SRE tooling and proc
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Auth0 Platform Observability team owns the observability tooling that monitors the Auth0 Platform, and we are looking for an Observability Engineer to help ensure that our Product and Platform Engineers can monitor and observe our platform while continuing to rapidly ship software that our customers love. Our engineers maintain and automate observability tooling for our entire platform, including metrics, logs, and traces. We are looking for engineers passionate about monitoring, observing, measuring uptime and availability, and ensuring platform stability. If you have experience within the Site Reliability Engineering (SRE) field or working as a Development Operations (DevOps) engineer, and you have a passion for Observability tooling, this position will allow you to further your learning and development in these areas. As a Senior Engineer on this team, you will act as a core technical leader. You will work cross-functionally to help integrate services with our instrumentation libraries, support product teams, and actively investigate incidents to identify our observability gaps. Responsibilities: Proven ability to champion observability best practices, acting as an educator who can effectively correct anti-patterns and teach other engineering teams how to build robust, standardized instrumentation. Be an expert in running services in production environments Contribute to the process of designing services for high growth and high availability. Provisi
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Role We are looking for a highly curious, self-directed Full-stack engineer who thrives on solving complex technical challenges for our Digital Technology team. You are a versatile developer comfortable navigating everything from front-end interfaces (React/Next.js) and CMS environments (AEM) to cloud infrastructure (Vercel/Cloudflare). You possess an 'ownership' mindset: you don’t wait for assignments, you identify gaps and help drive plans, troubleshoot production incidents through to resolution, and drive technical alignment between frontend, backend, and design teams. You have a sharp eye for great UX/UI. You enjoy partnering closely with designers to refine interactive experiences, ensuring that technical implementations not only work flawlessly but also feel intuitive and polished for the end user. You are a proactive innovator who experiments with new technologies, including AI-driven tools, to optimize our systems and improve developer workflows. You will: Participate in cross-team initiatives from end to end, including code reviews, design reviews, operational robustness discuss
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The AI-Core Team The team is building a scalable Agentic AI platform for agents that run real engineering work, agents that write code, triage incidents, investigate alerts, remediate vulnerabilities, upgrade libraries, patch flaky tests, automate runbooks, and other workloads across Okta's infra. Owning this platform means owning the hard cross cutting problems: agent identity, delegated access, secure isolated execution, orchestration, and governance, at the scale and reliability bar that Infra demands. The work is novel and high ownership, and the candidate should be drawn to problems the industry hasn't solved yet. What You'll Work On Design and implement backend APIs and services that make up the agentic platform Build the agent identity and machine-to-machine authentication system, including credential management and delegated access flows Build the agent knowledge base and memory layer so agents retain context within and across sessions Buil
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. At Okta, our motto is "Always On" and nowhere do we embrace that more than in Technical Operations. We strive to build the most reliable and performant systems on the planet through the skillful use of automation. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. You will work on: ● Mentoring, managing, and leading a team of SRE’s with a broad range of expertise and experience. ● Being an evangelist and advocate for security best practices, leading initiatives and projects to strengthen our security posture for our most critical infrastructure. ● Responding to production incidents, driving us to remediation as quickly as possible and determining how we can prevent them in the future. ● Triaging and troubleshooting complex production issues to ensure reliability and performance. ● Working closely with our stakeholders across the organization to ensure our new capabilities are aligned to our competing constraints of reliability, security, and delivery velocity. ● Partnering directly with recruiting and people ops to hire and retain the best talent in the world. ● Keep sharp eyes on our metrics, including vulnerability scanning and security posture, 
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is seeking an experienced Security Engineer to build and implement detection and response improvements and adapt to emerging threats to protect employees and infrastructure. In this role you will have the opportunity to solve challenging problems and provide a meaningful impact on our overall security posture. We are looking for a candidate with a passion for both security and innovation. What you'll do: Build alerts and automation workflows to improve capabilities to detect and response to external and internal security threats Manage our logging pipelines and infrastructure and onboard new logging sources to improve our detection coverage Develop and maintain internal tooling to expand and automate team detection and response capabilities Respond to alerts generated from our tooling and run incidents as part of an on-call rotation Co
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! Server Platform is at the core of Figma's infrastructure foundation! This area owns the foundational layer of infrastructure that underpins all other infrastructure and application services at Figma. The Foundation team under Server Platform is Infrastructure's infrastructure. This team enables the efficient operation and rapid development of reliable infrastructure at Figma through opinionated platforms. We are building the V2 of Infrastructure and setting key initiatives that will set the tone for infrastructure for the next half decade. We're looking for thoughtful and technical leaders with experience focused on designing the next generation edge and network to handle our growing customer traffic and services. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: Design, build and operate scalable distributed systems and services that power Figma's innovative tools for design and collaboration Develop and secure our internal and edge network running on AWS Help design and deploy an internal service mesh network which can support Figma's growing portfolio of services and features Develop intelligent detection and auto-scaling countermeasures for DDoS threats Debug and resolve production issues across services and multiple layers of the stack, reducing mean time to resolution for critical incidents We’d love to hear from you if you have: 4+ years of experience building infrastructure components and services at scale in a cl
Get new incident commander jobs by email
Daily job updates · Unsubscribe anytime