Jobiba hiring network

Incident Commander Jobs

589 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current incident commander jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the team The AI Deployment Engineering team is responsible for helping developers and enterprises safely and effectively deploy OpenAI technologies in production. We act as trusted technical advisors and thought partners for customers, working side by side with their teams to identify high-value use cases, design practical architectures, and move from prototype to durable deployment. Cybersecurity is one of the most urgent domains where AI can help. Security teams are under pressure to reason across code, logs, infrastructure, tickets, alerts, and vulnerability data faster than ever. As frontier models become more capable, organizations need deep technical guidance on how to evaluate, validate, and safely deploy AI systems in security-critical workflows. About the role We are looking for a Cyber AI Deployment Engineer to partner with customers and help them apply OpenAI models, APIs, Codex, and agentic workflows to real cybersecurity use cases. You will work with CISOs, security executives, application security leaders, SOC teams, security engineering teams, and hands-on practitioners to identify where AI can create measurable security outcomes. This is a customer-facing technical role for someone who can move fluidly between executive strategy, practitioner-level cyber depth, and hands-on solution design. You will help customers evaluate and deploy workflows such as secure code review, vulnerability triage, threat modeling, remediation, SOC and incident response workflows, detection engineering, cloud security, GRC automation, and security validation. You will collaborate closely with Sales, Solutions Engineering, Product, Engineering, Research, and Security to turn customer needs into safe deployment patterns, reusable field assets, and product feedback. This role is based in our San Francisco HQ. We offer relocation support to new employees. In this role, you will: Deeply embed with strategic customers as the technical lead for AI-enabled cybersecurity work

javascriptpythonjava
View job →
O
1mo ago

About the Team Our London-based team builds the backend systems that help ChatGPT scale reliably. We work on infrastructure close to the product, partnering with engineering teams to improve the performance, resilience, and operability of critical user-facing systems. Our work combines backend software engineering with distributed systems and production reliability. We build shared capabilities, improve high-traffic workflows, and make it easier to introduce new product functionality without compromising performance or availability. About the Role This role is for software engineers who want to build and evolve backend systems operating at significant scale. You’ll write production code, design shared infrastructure, and solve technical challenges involving performance, distributed systems, and system reliability. You’ll also own how those systems behave in production: how changes are rolled out, how issues are detected and diagnosed, and how recurring operational problems can be addressed through better software and system design. This is a strong fit for backend engineers who enjoy complex systems problems and want a direct connection between the infrastructure they build and the experience of ChatGPT users. In this role, you will: Design, build, and maintain backend systems supporting high-traffic ChatGPT experiences. Develop shared services, APIs, and infrastructure that help product teams build and launch new capabilities safely. Improve the performance, scalability, and efficiency of production systems as usage and product complexity grow. Build and improve systems for asynchronous processing and other large-scale backend workloads. Lead architectural improvements and infrastructure migrations while maintaining correctness, compatibility, and safe rollout and rollback. Strengthen monitoring, alerting, and diagnostics to detect problems early and reduce customer impact. Participate in on-call, incident response, and root-cause analysis, and turn operational lea

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team: The Database Systems team specializes in high-performance distributed databases. Our team built Rockset, the real-time search, analytics, and vector database that powers all vector search and retrieval augmented generation (RAG) at OpenAI. In addition to retrieval, as an online database, Rockset powers core functionality across all of OpenAI's product lines and many critical internal use cases. About the Role : We are looking for engineers passionate about distributed systems, close-to-the-metal performance optimization (our core engine is written in C++), and building scalable database infrastructure from the ground up. As an engineer on the Database Systems team, you'll contribute to the core database engine, driving improvements across ingestion, query execution, indexing, and storage. You'll partner with teams across OpenAI to unlock new product capabilities and help scale online database reliability and throughput as usage grows by orders of magnitude. In this role you will: Design, build, and operate high-performance distributed systems Identify and resolve performance bottlenecks to scale infrastructure to the next order of magnitude Define long-term technical direction and guide system evolution Collaborate with product, engineering, and research teams to deliver scalable and reliable infrastructure Dig deep into complex production issues across the stack Contribute to incident response, postmortems, and best practices for system reliability You might thrive in this role if you: Have significant experience building, scaling, and optimizing distributed systems at scale Are curious about database internals, storage engines, or low-latency query systems Enjoy debugging challenging performance issues in complex, high-throughput systems Have experience operating production clusters at scale (e.g., Kubernetes or other orchestration systems) Think rigorously about scalability, correctness, and reliability Thrive in fast-paced environments with high

awsazuregcp
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Support team is central to ensuring that our customers' experience with our products is nothing short of exceptional. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. Given OpenAI’s breakneck shipping cadence and growth – and the expectation that it will only accelerate – our ability to architect automation systems and agentic workflows for scale is central to our ability to maintain exceptional support quality in the face of AGI. About the Role As a Support Vendor Manager, you will own the health, performance, and long-term scalability of multiple support partner and vendor relationships. This is a vendor leadership role first and foremost: you will drive commercial and operational accountability (SLAs, QBRs, escalation paths, remediation plans), while also building the operating model that enables support to scale without linear headcount growth. You’ll collaborate closely with User Operations teams (e.g., Trust & Safety, Fraud & Risk), Systems/Tooling, Data partners, and Product/PM stakeholders as we launch new workflow and launch and scale new programs. You’ll be responsible for: End-to-end vendor leadership: Own day-to-day oversight, relationship health, and executive-level accountability for multiple support vendors/BPOs. Performance management & remediation: Define and manage SLA/KPI performance expectations, run WBRs/QBRs, identify performance gaps, and drive structured turnaround plans with clear owners and timelines. Escalation and risk management: Serve as the primary escalation point for vendor issues, including incident response, surge events, quality regress

awsrestai
View job →
O
1mo ago

About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun

awsazuregcp
View job →
O
1mo ago

About the Team Corporate Security helps protect OpenAI's people, offices, events, and operations through practical and risk-informed security programs. The team partners closely with Workplace, Legal, People/HR, Resilience & Safety, Executive Operations, GSOC, regional leaders, vendors, and building management. About the Role OpenAI is seeking a Paris-based Security Operations Manager to lead corporate and physical security operations for Paris, provide primary support for Brussels, and coordinate regional support for Munich, Zurich, and other European cities as needed. This person will be a trusted security partner for local teams, regional stakeholders, vendors, building management, and global Corporate Security peers. In this role, you will: Lead day-to-day physical security operations for the Paris office and supporting Brussels operations. Coordinate regional security support for Munich, Zurich, and other European cities where business activity, events, travel, or executive visits require coverage. Manage practical security controls, including access management, visitor workflows, vendor coordination, incident response, emergency readiness, and site security documentation. Partner with Workplace, building management, Legal, People/HR, EHS/Resilience, GSOC, Executive Operations, and local leadership to keep security effective, locally appropriate, and employee-friendly. Support event and executive visit planning with clear escalation paths, stakeholder alignment, and proportionate risk mitigation. Improve vendor performance, operational standards, reporting, and remediation tracking. You might thrive in this role if you have: 8+ years of experience in corporate security, physical security operations, protective services, emergency management, public service, military, law enforcement, or a closely related operational field. Experience leading physical security operations in a modern office environment, ideally in technology, professional services, financial

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world. Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure. About the Role We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology. You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more. There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future. In this role, you will: Build, lead, and grow high-performing infrastructure engineering teams. Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search. Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach. Champion pragmatic use of agent technology to amplify execution velocity. Reduce operational toil and incident frequency through better abstractions, gua

awsrestai
View job →
N
Notion
📍 San Francisco• Full-time• $272K – $320K/yr
1mo ago

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role Build the most advanced AI Meeting Notes product — and expand it into broader “AI data capture” features that help teams turn conversations into durable context, tasks, and knowledge. Our mission is to 10x the rate of business context & data that enters Notion — optimized for agents — so teams get superhuman memory across workstreams and customers. Notion workspaces that use AI Meeting Notes already enter 6x more data on a daily basis, so we’re well on our way. What You'll Achieve Ship end-to-end product experiences across capture → transcript → summary → follow-ups (full-stack ownership). Make meeting & data capture feel effortless and magical (e.g., speaker identification via audio waveforms, richer in-meeting UX, smarter organization). Improve summary quality that teams trust: structure, factuality, and citations that make downstream agents and humans more capable. Raise the bar on reliability & observability across the pipeline (SLOs, debugging workflows, incident response) for realtime systems. Build agentic meeting workflows that turn discussions into tasks, follow-ups, and organized knowledge — so “w

restaigo
View job →

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and performance of Postman’s AI-powered API and agentic systems in production. This role focuses on monitoring, availability, incident response, and automation to support AI services and tools trusted by millions of developers globally. What You’ll Do Develop and manage reliability metrics (SLOs) for AI-driven API services and agentic AI platform features Implement comprehensive observability and monitoring systems for real-time performance and fault detection Design and drive automated failover, recovery, and incident response strategies for high-availability AI infrastructure Optimize resource utilization, particularly GPU/accelerator efficiency, ensuring cost-effective AI system operation Collaborate closely with engineering, platform, and product teams to align reliability efforts with broader organizational goals Lead efforts to build internal tooling and automation focused on AI system stability and operational excellence Drive continuo

aigorust
View job →

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. About the Role As a Site Reliability Engineer at Ema, you will own the stability, availability, and operational health of our agentic AI platform across customer environments. You'll work closely with Engineering and DevOps to provision infrastructure, drive deployment excellence, and keep production running at the quality bar our enterprise customers expect — 99.9%+ uptime, proactive incident response, and continuous improvement. What You'll Do Infrastructure & Deployment Design and provision cloud infrastructure (GCP, Azure, AWS) tailored to customer environments, with security, scalability, and compliance built in Execute on-call SaaS deployments with minimal downtime; automate and optimize deployment workflows end-to-end Production Stability & Observability Monitor logs, alerts, and metrics to maintain SLA commitments and catch issues before they escalate Diagnose and resolve production incidents with speed and rigor; drive root cause analysis and permanent fixes Collaborate with DevOps to enhance monitoring dashboards and alerting frameworks; deliver clear system health reporting to internal and customer stakeholders Documentation & Knowledge Management Maintain de

awsazuregcp
View job →
S
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Cloud Support Engineer (CSE) Job Description Snowflake seeks a Senior Cloud Support Engineers who combine technical expertise, customer empathy, and an AI-first mindset . You'll leverage and refine AI tools to accelerate troubleshooting, improve knowledge bases, and reduce time to resolution—safely and responsibly. Experience in 24x7 technical support, escalation handling, on-call rotations, and incident management is ideal. Key Responsibilities: As a Senior Cloud Support Engineer , you manage customer cases and are accountable for accelerating their time-to-resolution, ensuring platform stability, and driving a world-class support experience. Operating as a full-stack technical support resource, this role blends deep hands-on troubleshooting with the customer empathy and communication skills of a trusted advisor. Customer Value and Incident Ownership Own the Customer Experience: Manage customer issues from initial triage through resolution and follow-up, ensuring clear communication and timely updates. Deliver Support Value with AI: Leverage AI assistants and diagnostics to accelerate triage and root-cause analysis while maintaining accuracy and safety standards. Outcome-Based Support: Focus on business impact, ensuring resolutions fix issues, reduce recurrence, and

sqlawsazure
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Customer Experience Engineering Team builds the internal and external technologies that scale Snowflake’s global support and sales organizations. We empower our technical experts by providing the advanced tools they need to resolve complex issues and drive customer success. Our team specializes in software engineering, data-driven decisions, ML, and LLM-based solutions . We build production-grade systems to automate manual processes and augment the capabilities of our technical staff. Our current focus includes: LLMs : Developing and deploying LLM and agent-based architectures for streamlining troubleshooting Scalable Evaluations : Implementing large-scale evaluations to ensure the quality and reliability of our internal and external tools Process Automation : Designing intelligent workflows that eliminate bottlenecks and allow our experts to focus on the most technical aspects of the Snowflake platform Incident discovery: using embeddings, LLMs, clustering, and agents to detect potential widespread issues more quickly Now, the team is growing, and we are looking for a Software Engineer to join us. In this role, you will work closely with the state of the art LLM models, fine-tune them, develop agents, apply various clusterings, summarizations, embeddings, and so on. Ev

pythonjavakubernetes
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Customer Experience Engineering Team builds the internal and external technologies that scale Snowflake’s global support and sales organizations. We empower our technical experts by providing the advanced tools they need to resolve complex issues and drive customer success. Our team specializes in software engineering, data-driven decisions, ML, and LLM-based solutions . We build production-grade systems to automate manual processes and augment the capabilities of our technical staff. Our current focus includes: LLMs : Developing and deploying LLM and agent-based architectures for streamlining troubleshooting Scalable Evaluations : Implementing large-scale evaluations to ensure the quality and reliability of our internal and external tools Process Automation : Designing intelligent workflows that eliminate bottlenecks and allow our experts to focus on the most technical aspects of the Snowflake platform Incident discovery: using embeddings, LLMs, clustering, and agents to detect potential widespread issues more quickly Now, the team is growing, and we are looking for a Software Engineer to join us. In this role, you will work closely with the state of the art LLM models, fine-tune them, develop agents, apply various clusterings, summarizations, embeddings, and so on. Ev

pythonjavakubernetes
View job →
PE
Private Employer
📍 Dublin• Full-time• Hybrid
1mo ago

Shape the Future with Dun & Bradstreet At Dun & Bradstreet, we believe data has the power to create a better tomorrow. As a global leader in business decisioning data and analytics, we help companies worldwide grow, manage risk, and innovate. Since 1841, businesses have trusted us to turn uncertainty into opportunity. We’re a diverse, global team that values creativity, collaboration, and bold ideas. Are you ready to make an impact and help shape what’s next? Join us! Explore opportunities at dnb.com/careers. The AI Operations Engineer is responsible for supporting the reliability and operational intelligence of cloud-hosted and AI-enabled services. This role focuses on observability, CI/CD-integrated reliability, alerting, and automated remediation to reduce noise, detect issues early, and improve incident response.

ci/cdairust
View job →
PE
Private Employer
📍 Uttar Pradesh, India• Full-time
1mo ago

Technical Lead - Backend About Us: Paytm is India's leading mobile payments and financial services distribution company. Pioneer of the mobile QR payments revolution in India, Paytm builds technologies that help small businesses with payments and commerce. Paytm’s mission is to serve half a billion Indians and bring them to the mainstream economy with the help of technology. About the role: We are looking for a Tech Lead for the Gold Tech team to drive technical direction, architecture, and delivery of Paytm Digital Gold — a high-scale bullion platform serving consumer and merchant use cases at lakhs+ rpm. This role combines hands-on engineering with technical leadership — you will architect solutions, unblock the team, own critical money-path reliability, and partner with Product and Engineering leadership on roadmap and execution. We are building reliable, high-availability financial products and need a Tech Lead who can own modules end-to-end, raise the bar on code quality, and mentor junior engineers. Key Responsibilities: ● Define and evolve technical architecture for Gold services — scalability, reliability, security, and maintainability. ● Lead design and implementation of high-availability, high-volume transactional systems. ● Own end-to-end delivery of major initiatives — breakdown, estimation, risk management, and production rollout. ● Set engineering standards: code quality, review practices, testing strategy, observability, and incident response. ● Drive cross-service integration design. ● Guide the team on Kafka event design, scheduler orchestration, CDC pipelines, and data consistency patterns. ● Mentor and grow engineers (SSE and below); conduct reviews, pair on complex problems, and build team capability. ● Contribute hands-on to critical modules; unblock the team on complex bugs, performance issues, and production fires. ● Represent Gold Tech in architecture reviews, tech debt prioritization, and platform-wide initiatives. Skills Required: ●

javanode.jssql
View job →
🔔

Get new incident commander jobs by email

Daily job updates · Unsubscribe anytime