At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Network Engineer Location: MPK/Bellevue/Dublin Snowflake's Enterprise Technology Network Services team is looking for a Senior Network Engineer to lead the design, operation, and optimization of our Zero Trust and secure-access platform. This role is Zscaler-centric — you will own the health, performance, and roadmap of our ZIA/ZPA deployment — while working across a modern, multi-cloud network stack that supports a global workforce of 10,000+ users. You'll be the escalation point for the most complex connectivity issues and a driver of automation and observability across the environment. What You'll Do Own and operate the Zscaler platform (ZIA, ZPA, ZDX, ZCC) end-to-end, including policy frameworks, app-segmentation models, PAC/traffic-forwarding standards, App Connector topology, and NSS/log-streaming design. Troubleshoot secure-access incidents like tunnel flapping, broker/connector health, SSL inspection edge cases, DNS/DTLS failures, and lead root-cause analysis for systemic issues. Manage Palo Alto firewalls, Panorama and GlobalProtect VPN, while planning migration toward Zscaler solutions. Support Aruba (Central) Wireless, Ekahau and Cisco Catalyst switches globally. Design, install, and configure network devices and ISP circuits at new offices. Build automatio
Jobiba hiring network
Incident Commander Jobs
589 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current incident commander jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Requirements: 5+ years of experience writing high-quality production code Experience building high-performance distributed systems at a large scale (the more battle scars, the better) Strong cloud skills Strong knowledge of low-level operating system foundations (Linux kernel, file systems, containers, etc.) Experience with performance engineering (tell us a story of when you shaved off a few milliseconds!) Ability to work in-person in our NYC or SF office. Prior experience with Rust is nice to have, but not required. Ability to participate in on-call rotation and respond to production incidents.
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: At Modal, we sell cloud services atop which our customers run their critical production systems. As a rapidly growing new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform, customer base, and our team. This role is for people who are deep systems thinkers, love stacking nines, and thrive from making others move faster at scale. Responsibilities include: Identifying architectural changes to improve reliability and performance. Fostering a culture of reliability across Modal’s engineering organization. Defining and implementing operational processes such as deployments, upgrades, etc. Operating systems like Kubernetes, Postgres, Redis, etc. Participating in on-call rotations, and responding to production incidents. Requirements: 5+ years of experience writing high-quality production code. 2+ years of
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We’re looking for strong engineers with experience building developer tools that users love to work with. Our ideal candidate is someone with a demonstrated drive to build beautiful interfaces that enhance developer productivity. Requirements: 5+ years of experience developing high-quality Python libraries with broad user-bases, ideally including some experience maintaining open-source software. Knowledge of advanced Python features, especially async programming. A strong product sense that manifests as a focus on developer ergonomics and productivity. A high level of customer empathy, good communication skills, and an openness to working directly with our users to help solve their problems. Ability to participate in on-call rotation and respond to production incidents. Ability to work in-person in our NYC or Stockholm office. Any of the following would be a plus:
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Requirements: 5+ years of experience writing high-quality production code Experience building high-performance distributed systems at a large scale (the more battle scars, the better) Strong cloud skills Strong knowledge of low-level operating system foundations (Linux kernel, file systems, containers, etc.) Experience with performance engineering (tell us a story of when you shaved off a few milliseconds!) Ability to work in-person in our Stockholm office. Prior experience with Rust is nice to have, but not required. Ability to participate in on-call rotation and respond to production incidents.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Cloud Platform Engineer, you'll envision and build robust systems and processes that ensure our infrastructure is scalable, reliable, and efficient. This can range from automating deployments and monitoring systems to optimizing performance and managing incidents. We all work closely with our users, learning from their past struggles in operationalizing ML, onboarding them onto our platform, and turning our learnings into ideas for improving Baseten. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Infrastructure team: Multi-cloud capacity management Inference on B200 GPUs Multi-node inference Fractional H100 GPUs for efficient model serving RESPONSIBILITIES Build and maintain scalable infrastructure to support the deployment and operation of machine learning models. Establish standards and best practices for reliability and performance across the infrastructure. Automate processes when relevant, particularly for managing CI/CD pipelines. Own products and projects end-to-end, functioning as both an engineer and a project manager, with a focus on user empathy, project specification, and end-to-end execution. Collaborate with cross-functional teams to understand project requirements and translate them into technical solutions. Mentor junior team members and contribute to knowledge sharing within the organization. Navigate ambiguity and exercise good judgment on tradeoffs and
About Supabase Supabase is an open source Firebase alternative. We give developers a Postgres database, authentication, instant APIs, edge functions, and real-time subscriptions — all in one platform. We are building the infrastructure layer for the next generation of applications. Corporate IT at Supabase reports into the Security organization. Identity and endpoint hygiene are treated as security controls, not administrative overhead. You will work with a small, senior team with direct access to engineering leadership and a mandate to automate everything. About the Role You will work directly with our IDM/MDM Lead to own the day-to-day operations of our identity and endpoint stack — Okta, Slack, Iru (MDM), and the integrations that tie them together. This role is equal parts identity management and endpoint operations, with a strong expectation that you automate what you repeat and document what you automate. This role provides follow-the-sun IT and identity coverage alongside our IDM/MDM Lead on the West Coast. Fully remote, with a strong preference for candidates based in EST or APAC. What You’ll Own Identity & Access Management Administer Okta day-to-day: user provisioning, group management, SSO application configuration, and MFA policy enforcement. Own joiner-mover-leaver (JML) workflows — ensure access is granted on day one, adjusted on role change, and fully revoked on departure with no manual gaps. Maintain and improve Okta lifecycle automation, reducing manual provisioning toil and closing the window between HR events and access changes. Audit access regularly: identify stale accounts, over-provisioned roles, and orphaned app assignments before they become incidents. Support FIDO2/WebAuthn and YubiKey deployment for privileged access across the organization. Endpoint Management & MDM Administer Iru (formerly Kandji) MDM for macOS fleet: device enrollment, configuration profiles, compliance baselines, and policy enforcement. Ensure all managed endpo
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! As a Senior Security Operations Engineer you will: Serve as trusted advisor to team’s leadership and partner teams by clearly articulating business risks associated with security issues Harden our cloud-native environments (AWS, OCI, GCP) by introducing secure by default designs and features into network, tooling, and processes Own and drive resolutions for enabling engineers to design, build, and use infrastructure securely at scale by deploying secure architectures using infrastructure-as-code and reusable code libraries Manage IAM / RBAC for cloud infrastructure, and partner with IT on streamling authentication/authorization to ensure unified access control across the board Deploy and operationalize some of the security services and tools (eg: SIEM, SOAR, domain monitoring, endpoint tooling, cloud security tooling) Respond to security incidents and harden environments post-incidents. Support control monitoring and remediation for compliance initiatives Gather and analyze security metrics to address security issues with cross-team dependencies Be a problem solver who is empathetic to developer concerns and will employ construc
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! As a Senior Security Operations Engineer you will: Serve as trusted advisor to team’s leadership and partner teams by clearly articulating business risks associated with security issues Harden our cloud-native environments (AWS, OCI, GCP) by introducing secure by default designs and features into network, tooling, and processes Own and drive resolutions for enabling engineers to design, build, and use infrastructure securely at scale by deploying secure architectures using infrastructure-as-code and reusable code libraries Manage IAM / RBAC for cloud infrastructure, and partner with IT on streamling authentication/authorization to ensure unified access control across the board Deploy and operationalize some of the security services and tools (eg: SIEM, SOAR, domain monitoring, endpoint tooling, cloud security tooling) Respond to security incidents and harden environments post-incidents. Support control monitoring and remediation for compliance initiatives Gather and analyze security metrics to address security issues with cross-team dependencies Be a problem solver who is empathetic to developer concerns and will employ construc
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 This role bridges infrastructure and product engineering: you'll build genuine partnerships across product, AI, and enterprise teams so that ownership is shared and velocity is never blocked by platform constraints. As Director, you'll set a forward-looking cloud vision, proactively align with stakeholders across the business, and ensure the platform scales for multi-shard, multi-region growth while meeting security and compliance commitments (SOC 2, business continuity/disaster recovery, and enterprise security frameworks). The Role: Cost Efficiency: Drive significant annual infrastructure savings by migrating data workloads to EKS and self-hosting key services such as OpenSearch. Ingress Convergence: Deprecate legacy frontend ALBs and consolidate to a single EKS-managed ALB per shard, unblocking faster deployments across the org. Coverage & Bench Depth: Eliminate single points of ownership across Networking, OpenSearch, and Terraform through cross-training and targeted hiring into coverage gaps. Stakeholder Alignment: Stand up a recurring alignment cadence with Product, AI, Enterprise, and Security so infra planning is driven by demand, not ad-hoc interrupts. Automation: Ship automated shard buildout via Backstage to remove manual toil from enterprise scaling. Reliability: Cut P0/P1 incidents attributed to Cloud Platform (DNS, ALB misconfiguration) through hardened ingress patterns and Terraform-policy guardrails, including blocking unauthenticated public endpoints. Roadmap Ownership: Deliver a roadmap covering Agent enablement, centralized IaC, and deployment rollout acceleration, tied to AI and
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 About This Role ClickUp is looking for a Senior Database Reliability Engineer to join our Database Operations team. You’ll be responsible for the performance, integrity, security, and availability of our PostgreSQL databases running on Linux in AWS. This role focuses on day-to-day database administration, operational excellence, and ensuring data is consistently reliable and well-managed across environments. Key Responsibilities Administer, monitor, and maintain PostgreSQL databases (250GB+) in production environments (AWS RDS, Aurora, EC2) Ensure database availability, performance, and data integrity through proactive monitoring and maintenance Perform routine database administration tasks including patching, upgrades, vacuuming, reindexing, and statistics management Execute and manage backup, restore, and disaster recovery procedures, including regular testing of recovery plans Handle user access management, roles, and database security to ensure compliance with best practices Perform capacity planning, storage management, and growth forecasting Troubleshoot database issues, including performance bottlenecks, locking/contention, and failed jobs Support application teams with query tuning, schema changes, and release deployments Manage database migrations, upgrades, and change requests with minimal downtime Maintain documentation for database configurations, standards, and operational procedures Participate in on-call rotations and provide support for production incidents Qualifications 7+ years of experience in a senior database administrator or Database Engineering role Strong hands-on PostgreSQL ad
About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role Join our growing security team and help drive security detection and response initiatives across Ramp. This will include a focus on maturing our security detection and alerting capabilities across our federal and public sector environments. Please note that this role will require you to be comfortable with working in-person at our NYC HQ (located near Madison Square Park) at least 2 days/week What You’ll Do Respond and assist with security requests and incidents submitted by Ramp team members Review logging, alerting, and audit sources to identify potential security incidents and perform initial triage on identified incidents Contribute to the creation, upkeep, and tuning of runbooks and security alerts to effectively handle, triage, and improve security alerts Work closely with the Ramp Security Engineers to improve security alerting and automated remediation Utilize log ingestion platform for security analytics and identification of tactics, techniques and patterns of attackers Design and implement automation to detect and respond t
At Linear, we're building the product development system for teams and agents. AI is fundamentally changing how software gets built, and we’re shaping the tools this new era requires. Founded in 2019, Linear has become the platform of choice for more than 40,000 companies (including OpenAI, Coinbase, and Ramp) to plan, build, and ship their products. Today, our team is distributed across North America, Europe, and Australia, and we’re continuing to grow internationally. What unites us is relentless focus, fast execution, and a deep care for software craftsmanship. As a small team, we’re all generalists and constantly picking up new challenges. When it comes to code, we’re looking to work with experienced people who can pick a problem and solve it. We use TypeScript and build scalable systems so we can continuously make progress on a solid foundation. We don’t expect you to have a background in everything we use, but we do expect strong JavaScript fundamentals and a background working with React and TypeScript. Location & work mode Linear is a remote-first company, with optional co-working offices in San Francisco, New York, and London. This role is open to candidates based in North America and Europe. You can work from most timezones within these regions. We value deep focus and async collaboration, with intentional moments to connect in person through team off-sites, optional co-working, and occasional travel. What you'll do Build new user-facing features with everything from database models to GraphQL resolvers and UI components Optimize our data synchronization stack by applying better serialization protocols Add real-time collaborative editing to our content editor Improve performance by profiling and tweaking virtualized list rendering Add analytics, monitoring, and alerts to our service so that we can better respond to operational incidents Open-source any non-trivial innovations that come out of our work on the product Redefine best-in-class software deve
At Linear, we're building the product development system for teams and agents. AI is fundamentally changing how software gets built, and we’re shaping the tools this new era requires. Founded in 2019, Linear has become the platform of choice for more than 40,000 companies (including OpenAI, Coinbase, and Ramp) to plan, build, and ship their products. Today, our team is distributed across North America, Europe, and Australia, and we’re continuing to grow internationally. What unites us is relentless focus, fast execution, and a deep care for software craftsmanship. As a small team, we’re all generalists that work across the full stack (built in Typescript end-to-end). We’re looking for experienced engineers that thrive in an environment of autonomy and individual responsibility to help us build the future of product development. Location & work mode Linear is a remote-first company, with optional co-working offices in San Francisco, New York, and London. This role is open to candidates based in the US and Europe. You can work from anywhere within those regions. We value deep focus and async collaboration, with intentional moments to connect in person through team off-sites, optional co-working, and occasional travel. What you'll do Build new user-facing features with everything from database models to GraphQL resolvers and UI components Optimize our data synchronization stack by applying better serialization protocols Add real-time collaborative editing to our content editor Improve performance by profiling and tweaking virtualized list rendering Add analytics, monitoring, and alerts to our service so that we can better respond to operational incidents Open-source any non-trivial innovations that come out of our work on the product Redefine best-in-class software development processes so that we can build a purpose-built product. What we're looking for 5+ years of experience building customer-facing products at a high-quality software company Strong React and Typ
About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. Protecting Flexport's global operations and assets, enabling secure trade worldwide. The opportunity: Join Flexport's dynamic Compliance Program as the Manager, Supply Chain Protection & Loss Intelligence. This critical role serves as the single point of accountability for both supply chain protection and loss prevention across Flexport's entire footprint — offices, fulfillment centers, and partner sites worldwide. The Manager will be responsible for protecting Flexport employees, visitors, assets, and facilities through the design and execution of security programs, while simultaneously driving a proactive loss prevention strategy across the Omni fulfillment network. Day-to-day, this means managing a team of 50+ contracted security officers across multiple locations, overseeing electronic security infrastructure (CCTV, access control), conducting investigations, and managing the security ticketing system. Strategically, it means building the analytical foundation to identify loss trends and security gaps before they become incidents, authoring the policies and SOPs that standardize how Flexport protects its people and assets globally, and ensuring every new fa
Get new incident commander jobs by email
Daily job updates · Unsubscribe anytime