Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! As a Senior Security Operations Engineer you will: Serve as trusted advisor to team’s leadership and partner teams by clearly articulating business risks associated with security issues Harden our cloud-native environments (AWS, OCI, GCP) by introducing secure by default designs and features into network, tooling, and processes Own and drive resolutions for enabling engineers to design, build, and use infrastructure securely at scale by deploying secure architectures using infrastructure-as-code and reusable code libraries Manage IAM / RBAC for cloud infrastructure, and partner with IT on streamling authentication/authorization to ensure unified access control across the board Deploy and operationalize some of the security services and tools (eg: SIEM, SOAR, domain monitoring, endpoint tooling, cloud security tooling) Respond to security incidents and harden environments post-incidents. Support control monitoring and remediation for compliance initiatives Gather and analyze security metrics to address security issues with cross-team dependencies Be a problem solver who is empathetic to developer concerns and will employ construc
Jobiba hiring network
Senior Cloud Operations Engineer Jobs
7,292 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Core Engineering Team Okta powers authentication and authorization for thousands of organizations worldwide. We make access to applications safe, secure, and seamless for billions of logins worldwide. Within Okta, the Core team builds software and frameworks and works with infrastructure teams to deliver 99.99% uptime for our core authentication and authorization products. The Software Engineering Manager Opportunity Okta's Core Engineering team is responsible for building and evolving shared infrastructure and services that lay the foundation for what other engineering teams build on. We're in charge of common shared services like search, cache, configuration management, frameworks for async job management, and email pipeline, to name a few. We're cloud native, where redundancy, multi-tenancy, scale, resource optimization, and resiliency are first-class citizens. With Okta's mantra of 'Always On!' there's never a dull moment. Our biggest asset is our team of passionate engineers and technically minded managers. What You’ll Be Doing Manage a distributed team, including setting expectations and removing blockers, creating a collaborative working environment, hiring and recruitment, providing coaching and career management discussions Collaborate with managers, architects, product owners, project managers, test partners, security and operations engineers to implement best practices related to resilience Communicate and organize cross-team projects with hi
PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. Senior Product Manager for AI and Automation PagerDuty is redefining how modern engineering and operations teams work. PagerDuty’s Automation Platform includes Workflows, Actions, Connectors, and a growing agentic layer built on Skills and Tools. This is the backbone of how teams eliminate toil, respond to incidents autonomously, and ultimately enable AI-native SRE agents. As Senior Product Manager for AI and Automation, you will own product strategy and execution across our Operations Cloud SaaS and on-premises automation products and lead the roadmap for the agentic automation experience we’re building for autonomous SRE agents. This is a high-visibility, high impact role that sits at the intersection of developer tooling, enterprise operations, and frontier AI product design. You will report directly to the Senior Director of Product Management for the AI & Automation group and will partner tightly with engineering, design, GTM, and enterprise customers. Key Responsibilities Define and drive the multi-year roadmap for Workflows and Actions, covering both cloud-delivered SaaS and on-p
PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. Senior Product Manager for AI and Automation PagerDuty is redefining how modern engineering and operations teams work. PagerDuty’s Automation Platform includes Workflows, Actions, Connectors, and a growing agentic layer built on Skills and Tools. This is the backbone of how teams eliminate toil, respond to incidents autonomously, and ultimately enable AI-native SRE agents. As Senior Product Manager for AI and Automation, you will own product strategy and execution across our Operations Cloud SaaS and on-premises automation products and lead the roadmap for the agentic automation experience we’re building for autonomous SRE agents. This is a high-visibility, high impact role that sits at the intersection of developer tooling, enterprise operations, and frontier AI product design. You will report directly to the Senior Director of Product Management for the AI & Automation group and will partner tightly with engineering, design, GTM, and enterprise customers. Key Responsibilities Define and drive the multi-year roadmap for Workflows and Actions, covering both cloud-delivered SaaS and on-p
PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. About the role PagerDuty’s Operations Cloud runs on a platform that ingests billions of signals and turns them into real-time action for thousands of customers. We’re looking for a Senior AI/ML Engineer who lives at the intersection of two disciplines: large-scale distributed systems and applied AI. In this role you will design and ship AI systems that run in production at PagerDuty’s scale — powering Incident Management AI Agents, event intelligence, and the LLM-powered capabilities embedded across our platform. You’ll own the full lifecycle, from framing the problem to serving reliably at scale. We are looking for a candidate who is genuinely passionate about building with modern AI — LLMs, agents, and retrieval — but grounded in the realities of building resilient, high-throughput systems. What you’ll do Design and build AI-powered features — LLM agents, retrieval, and event intelligence — that operate on high-volume, real-time event streams, from problem framing through production deployment and monitoring. Architect and own the systems behind them: agent and prompt orchestration, retrieval pipelin
About Prophecy The leader in AI-native data preparation and analysis, Prophecy is revolutionizing how the world’s top enterprises turn data chaos into reliable insights. We introduce the AI-native data lifecycle (generate, refine, deploy) where our industry leading AI agents and humans work hand-in-hand in visual and document interfaces to analyze, transform and prepare data, to ship trusted insights at enterprise scale. Don’t miss the rocket ship—join Prophecy and build the next data revolution. Position Summary This is a high-impact opportunity to be a senior DevOps engineer in a fast-growing startup, based in Prophecy’s India engineering center. You will own and evolve the foundations that keep our engineering org fast, secure, and cost-efficient at scale — spanning cloud cost management, security DevOps, and CI/CD and engineering operations. You will work with a team of dynamic engineers who take pride in solving complex problems, and you will have the autonomy to set direction in your areas of ownership. The Impact You Will Have Cloud cost (FinOps) Own and evolve our cloud cost optimization program across AWS, Azure, GCP, Databricks, Snowflake and Bigquery building on the programmatic monitoring and controls Analyze billing, asset-inventory, and utilization data across departments to identify wasteful spend and provide actionable insights to optimize it. Develop and maintain cost optimization strategies, roadmaps, and forecasting models Partner with engineering teams to design cost-efficient architectures without compromising scalability or reliability. Build and maintain automation for infrastructure provisioning, scaling, and cost control. Security DevOps Partner with engineering and security to drive our security-hardening program across workstreams such as identity & access governance, secrets & credential lifecycle, cloud access, and CI/CD hardening. Implement and automate guardrails: secrets management, least-privilege access, cr
About the Team At OpenAI, our User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We translate real-world user and operational signals into timely decisions, practical interventions, and improvements to our products and systems. This role will take on new, ambiguous, or underdeveloped operational risks and help mature them into scalable capabilities. We work across USRO and partner closely with Product, Engineering, Data Science, Product Policy, Legal, Safety, Support, and external vendors or partnership stakeholders. About the Role We are seeking a Senior Operations Analyst to take on complex, ambiguous safety and risk problems and turn them into practical operational solutions that can scale. This is a senior individual-contributor role for a versatile operator who is comfortable moving between queues, investigation, analysis, workflow design, hands-on execution, and cross-functional leadership. Depending on team needs, the role may focus on emerging-risk incubation, cloud deployment partnerships, or other new operational areas. You will be expected to move quickly, work hands-on, and create structure without waiting for perfect requirements or a large support team. The work starts with the problem, not a prescribed process. You may investigate unstructured user signals, stand up a lightweight workflow, build an AI-assisted tool, improve an existing operation, or help a new launch become operationally ready. The goal is to produce durable systems that other people can run, not simply complete a series of individual tasks. The portfolio will change with company priorities and may span established harm areas, emerging-risk incubation, cloud deployments and partnerships, device safety, or new product launches. Some hires may focus primarily on cloud deployment operations, including launch readiness, partner coordination, safety workflows, and operational monitoring. You will ty
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer to join our Cloud Infrastructure & Operations team. This is a remote role based in the Netherlands, reporting to the Senior Director, Software Engineering. As a Staff SRE, you will leverage your expertise in Linux/UNIX System Administration to build scalable infrastructure and manage platforms like Kubernetes using automation and high security standards. You will troubleshoot complex Linux networking and security issues, manage firewall technologies, and ensure secure access across our global platforms and applications. What you’ll do (Role Expectations) Create and maintain highly scalable solutions based on KVM LINUX, Kubernetes, and Public Cloud Providers Analyze and troubleshoot systems performance and issues across the OS and Applications Maintain platform security and observability using nftables and robust monitoring tools Manage and deploy systems and s
PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. About the role PagerDuty’s Operations Cloud runs on a platform that ingests billions of signals and turns them into real-time action for thousands of customers. We’re looking for an early-career AI/ML Engineer who is excited to grow at the intersection of two disciplines: large-scale distributed systems and machine learning. In this role you will help build and ship AI systems that run in production at PagerDuty’s scale — powering Incident Management AI Agents, event intelligence, and the LLM-powered capabilities embedded across our platform. You’ll work alongside senior engineers on real production problems, learning how AI features go from a prototype to something that serves reliably at scale. We are looking for a candidate who is genuinely excited about building with modern AI — LLMs, agents, and retrieval — eager to learn how resilient, high-throughput systems are built, and motivated to grow into an engineer who is strong in both. What you’ll do Contribute to AI-powered features — LLM agents, retrieval, and event intelligence — that operate on high-volume, real-time data, with support and guidanc
About the Role Redwood is scaling public cloud infrastructure and AI features across multiple product lines, and we need a FinOps Lead to bring rigor, visibility, and accountability to that spend. This is a senior individual-contributor role with the authority to drive cross-functional cost governance directly. This role owns the translation of raw cloud cost data and AI spent into the models, forecasts, and governance mechanisms that let engineering, product, and executive leadership make informed decisions. This is not a bill-monitoring role. You will build the cost attribution infrastructure that ties cloud/AI spend to specific product lines and, ultimately, to ROI involving architecture and engineering teams to make the right trade offs and decisions in line with the strategic roadmap. The role will report to the Senior Director, Product Engineering Operations, and work closely with Cloud Engineering, the CPO's org, and engineering leadership to create cost visibility and defensibility informing leadership on efficiency strategies. Your ability to be an effective communicator, collaborative team player, and analytical thinker will be keys to success in this role. Responsibilities Cost Visibility, Attribution & Optimization Own and continuously improve cost models that attribute AWS (and other public cloud) spend by product line, team, and environment Drive tagging governance and hygiene, define standards, audit compliance, and close attribution gaps that prevent accurate cost-per-product reporting Build shared cost allocation models to drive transparency into per-team cost drivers where there are prevalent savings plans, reserved instances and network charges Build toward feature-level cost attribution that connects infrastructure spend to product ROI, not just aggregate bill totals Create and manage optimization programs with achievable savings targets, including working across engineering and finance to rightsize, clean up and modernize infrastructur
PagerDuty, Inc. (NYSE: PD) is the global leader in AI-first digital operations. By automatically detecting, diagnosing, and remediating issues, the PagerDuty Platform orchestrates AI agents and automated workflows with context from over 750 integrations. Trusted by approximately two-thirds of the Fortune 100 and nearly half of the Fortune 500, PagerDuty is the industry standard for organizations scaling resilient, autonomous operations. Notable customers include Chipotle, Cloudflare, Docusign, Fox, Nvidia, Salesforce, Spotify, Zoom and more. We are growing rapidly and hiring top talent with leading AI skills across engineering, sales, product, marketing, and beyond as we build the leading digital operations platform. About the Role PagerDuty is seeking a Principal Product Manager, Platform Security to own the strategy and execution of how we secure, harden, and defend our Operations Cloud platform. This role sits within our Product Development organization and reports to the Sr Director of Product, Platform & Partners. This is a senior individual contributor role. You'll bring the same rigor to security that a great PM brings to a product: deep customer empathy, structured threat modeling, clear risk tiering frameworks, and a bias toward measurable outcomes. You'll also be the connective tissue between Product, Engineering, IT, and Legal, ensuring security strategy translates into engineering execution and customer trust. This role owns the full lifecycle from recommendation to implementation to operations. You'll be the decision-maker on risk acceptance, control exceptions, and incident escalation in real-time. The ideal candidate has operated at the intersection of product management and security engineering in a later-stage B2B SaaS environment. You've owned security architecture decisions end-to-end, built security infrastructure, and have the credibility to influence both product roadmaps and engineering practices without formal authority. What You'l
As a Senior Security Engineer focused on Datadog’s Cloud SIEM product, you will help shape the future of security operations by transforming real-world security expertise into scalable detection, investigation, and response capabilities. You will develop high-impact threat detection content, improve AI-assisted security workflows, and help defenders identify and respond to threats across cloud-native and enterprise environments. Working closely with Product, Engineering, and Security Research teams, you will influence the evolution of Datadog Security products while advancing detection coverage across emerging technologies and attack surfaces. This role offers the opportunity to contribute to open source initiatives, publish security research, and help define the next generation of agentic security operations capabilities. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Research attacker techniques, defensive strategies, and emerging threats, translating findings into scalable security capabilities that protect customers at cloud scale. Design and improve AI-powered investigation, threat hunting, and response workflows that support Datadog’s agentic SOC capabilities. Own the lifecycle of threat detections and automated security workflows, from research and design through deployment, measurement, and continuous improvement. Develop high-fidelity detection content across cloud platforms, SaaS applications, identity systems, endpoints, networks, and other modern attack surfaces. Partner with Product, Engineering, Security Research, and customers to influence roadmap decisions and improve security outcomes across the platform. Mentor security engineers and drive improvements through automation, tooling, rapid prototyping, and data-driven optimization. Who Yo
As a Senior Security Engineer focused on Datadog’s Cloud SIEM product, you will help shape the future of security operations by transforming real-world security expertise into scalable detection, investigation, and response capabilities. You will develop high-impact threat detection content, improve AI-assisted security workflows, and help defenders identify and respond to threats across cloud-native and enterprise environments. Working closely with Product, Engineering, and Security Research teams, you will influence the evolution of Datadog Security products while advancing detection coverage across emerging technologies and attack surfaces. This role offers the opportunity to contribute to open source initiatives, publish security research, and help define the next generation of agentic security operations capabilities. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Research attacker techniques, defensive strategies, and emerging threats, translating findings into scalable security capabilities that protect customers at cloud scale. Design and improve AI-powered investigation, threat hunting, and response workflows that support Datadog’s agentic SOC capabilities. Own the lifecycle of threat detections and automated security workflows, from research and design through deployment, measurement, and continuous improvement. Develop high-fidelity detection content across cloud platforms, SaaS applications, identity systems, endpoints, networks, and other modern attack surfaces. Partner with Product, Engineering, Security Research, and customers to influence roadmap decisions and improve security outcomes across the platform. Mentor security engineers and drive improvements through automation, tooling, rapid prototyping, and data-driven optimization. Who Yo
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . As a Software Dev Senior Engineer , you will own the reliability, scalability, and operational excellence of our Cloud-based services. You will define and enforce reliability standards, drive the adoption of SRE practices across engineering teams, and build the systems and tooling that keep our production infrastructure healthy. We follow a DevOps model: Development and Operations teams are integrated, and the SRE function acts as the reliability layer — setting Service Level Objectives, managing error budgets, and continuously reducing toil through engineering. Key Responsibilities: Define, publish, and continuously refine Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs ) for all critical services, partnering with product and engineering leadership. Own the error budget framework: track consumption, enforce error budget policies, and drive reliability investments when budgets are at risk. Lead the design and implementation of comprehensive observability platforms — metrics, structured logging, and distributed tracing — to ensure full visibility into pro
NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel
Get new senior cloud operations engineer jobs by email
Daily job updates · Unsubscribe anytime