Jobiba hiring network

Lead Cloud Operations Engineer Jobs

6,753 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

D
1mo ago

We're on a mission to build the best platform in the world to defend the enterprise from code-to-cloud-to-runtime. Used by thousands of companies globally, Datadog security products uniquely leverage Datadog’s unified security and observability platform so Security, DevOps and SRE can collaborate rapidly and seamlessly to deliver better detection, prioritization and remediation. Our product and engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. In this competitive market, the Group Product Manager for Code Security will play a mission-critical role in providing product and strategy leadership to grow Datadog’s market share through differentiation, innovation and compelling customer value. This leader will lead a talented and growing team of product managers and work with world class engineers to build and grow multiple Code Security products that play an essential role for our customers’ code security programs, and growing Datadog into a security industry leader. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Run and grow multiple Code Security products to meet revenue and business targets with the goal of building a multi-hundred million dollar annual business. Lead and own product strategy and roadmap for accountable security products, fully aligned to revenue and business goals and with compelling differentiation and customer value. Ensure predictable roadmap execution across direct and partner teams to achieve product and business outcomes required to meet the revenue and business goals. Analyze and develop pricing and packaging strategies to maximize revenue through attaching deep understanding of market dynamics and other strategic leverage points. Drive GTM strategy with GTM partner teams

aigorust
View job →
D
Datadog
📍 New York• Full-time• From $192K/yr
1mo ago

As the Engineering Manager for Commercial Audit, you will lead a high-performing team responsible for scaling Datadog’s security and compliance posture through automation, tooling, and engineering excellence. Our GRC (Governance, Risk, and Compliance) function is a critical partner to the broader Security and Engineering organizations, ensuring that Datadog not only meets rigorous global regulatory standards but does so in a way that is efficient, scalable, and integrated into our cloud-native infrastructure. You will manage a team of engineers and analysts who are transitioning to a GRC engineering direction to treat compliance as a software problem, leveraging AI, custom tooling, CI/CD pipelines, and cloud-native services to turn complex regulatory requirements into actionable, automated controls. You will lead the strategy, roadmap, and execution of Datadog’s Commercial Audit initiatives. This is a high-impact leadership role where you will grow a team of engineers and analysts responsible for directly maintaining our compliance programs and related audits (e.g., SOC2, PCI, HIPAA, ISO) while looking to improve efficiency and effectiveness through platforms and tooling. You will act as a bridge between technical engineering, legal, and compliance, enabling the organization to move fast while maintaining a secure and compliant environment. You will champion a culture of "compliance-as-code," identifying opportunities to automate evidence collection, streamline control testing, and reduce manual toil for both your team and our partner engineering teams. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the strategy, roadmap, and execution of Datadog’s commercial security compliance efforts, shifting from manual audit processes to automated, scalable

pythonawsazure
View job →
M
1mo ago

About the Role We are seeking a Staff Enterprise Architect, Data to lead the strategy, design, and modernization of our enterprise data landscape. This role operates at the intersection of data architecture, engineering, and AI enablement, defining solutions to integrate our Data Lake and Data Warehouse across multi-cloud platforms. Over the next 12-18 months, you will enable self-service data access and natural language query capabilities for business users. You will architect Master Data Management and data lineage frameworks ensuring AI models operate on high-quality, governed data. You will also evaluate and implement AI-powered tools to automate data quality monitoring and enhance data security. We're looking to speak with candidates based in the San Francisco Bay Area for our hybrid working model. Key Responsibilities Data Strategy & Roadmap Design semantic layer architecture standardizing business metrics enterprise-wide. Define governance guardrails ensuring natural language queries access validated master data sources Develop Master Data strategy for Customer and Product domains (phases 1-2), Finance and People to follow. Define golden record requirements, stewardship models, and system-of-record hierarchy. Partner with business owners on master data governance Define cross-cloud data integration strategy and reference architecture. Specify patterns (federation, replication, abstraction layer) balancing performance, cost, and data freshness. Document trade-offs and recommend implementations for batch and near-real-time use cases Develop 12-24 month data architecture roadmaps for Finance, Sales, Product, and People. Identify capability gaps and recommend technology investments with business value and effort estimates Systems Design & Solution Leadership Evaluate AI-powered data observability platforms for quality monitoring, pipeline failure prediction, and data classification. Define requirements, lead vendor POCs, and establish integration patterns

pythonsqlmongodb
View job →
PE
Private Employer
📍 India• Full-time
1mo ago

About the role We are seeking an experienced and detail-oriented Information Security and Cloud Security Auditor to join our team. The ideal candidate will have 3-7 years of expertise in data security and privacy control implementation, internal auditing, third-party risk management, cybersecurity governance, and cloud security (banking sector preferred). This role will be responsible for conducting comprehensive IT and cloud security audits, ensuring compliance with regulatory requirements, and enhancing our information security policies and procedures. Key Responsibilities -Conduct IT and cloud security audits across various domains, including IT General Controls (ITGC), Information Security Controls, Cloud Security, Network Security, Vulnerability Management, and Vendor Risk Assessments. -Assess compliance with relevant laws, regulations, industry standards, and organizational policies across both on-premises and cloud environments. -Develop and enhance information security and cloud security policies and procedures in alignment with industry best practices. -Maintain thorough documentation of audit findings, risk assessments, security measures, working papers, audit program checklists, and evidence for internal and external reporting. -Validate ITGC, cloud security, and application-specific controls and manage audit documentation, risk assessments, evidence gathering, and control testing. -Follow up on audit findings and ensure timely closure of non-compliance and security issues identified during audits. -Manage and oversee third-party risk assessments and audits, ensuring robust security controls are implemented across traditional and cloud-based service providers. -Lead and participate in the development, migration, and implementation of security controls and policies for network and cloud security solutions. -Conduct risk-based security assessments of internal, vendor, and third-party hosted environments, covering both traditional IT infrastructure and cloud

awsazuregcp
View job →
S
Snowflake
📍 Bellevue• Full-time
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Build the future of the AI Data Cloud. Join the Snowflake Stellar Design System team. Snowflake Design is a world-class team of curious problem solvers on a mission to evolve and improve the way the world works with data. Great design is a strategic differentiator for Snowflake — but that differentiation doesn’t come automatically. We’re looking for talented, collaborative design engineers to join the team and drive the mission forward while doing career-defining work. As a Staff Design Engineer, you are the cornerstone of delivering the ultimate user experience. You will be entrusted with the highest level of polish, creativity, and interaction design, pushing the boundaries of what's possible in the most innovative areas of our product. This role offers the unique opportunity to have broad creative license, empowering you to lead and create groundbreaking systems and experiences that marry aesthetics with functionality, making data more accessible and actionable for our customers. WHAT YOU WILL DO: Embed with feature teams to up-level experiences Setting the front-end architecture vision, not just contributing to it Driving adoption and influencing teams outside your direct org Owning the technical roadmap for a system or component pillar Design and build exceptional expe

typescriptreactnodejs
View job →

Job Details: Job Description: Join an enthusiastic team of engineers in Intel's Networking Solutions Group (NSG) focused on enabling next generation of programmable Infrastructure Processing Units (IPUs) with our lead customers as part of the Customer Experience Support (CES) organization. Intel brings decades of leadership in networking, virtualization, packet processing, storage, and security to a new class of IPU products that accelerate host networking functions and support emerging use cases such as security, virtualization, storage, load balancing, and data path optimization. Working closely with major cloud service providers and Intel development teams, you will help deliver customized IPU based solutions that enhance isolation, security, performance, storage and system management for our customers. A big part of the day-to-day job is to help customers manage feature request processes, enable solutions, and debug issues. Projects and responsibilities include but are not limited to: • Gain our customers' trust, understand their needs, and build POCs to meet them. Work closely with internal and external partners to understand use cases and requirements. • Be the go-to technical resource for customers building complex Datacenters, AI infrastructure as well as helping them understand performance characteristics for solutions. • Prepare and deliver technical content to customers including presentations, workshops, etc. • Contribute across the full IPU lifecycle, including board and platform bring up, low-level device initialization, OS driver and kernel configuration, system management, feature enablement, use case testing, debugging, and verification. • Defines systems implementation and integration solutions and plans to ensure optimum performance and reliability across hardware, firmware and software w

dockergitlinux
View job →

Distinguished Technologist - Principal Architect, Platform Security and Firmware Systems Description - HP is seeking a senior technical leader to serve as a lead architect for hardware rooted Platform Root of Trust for HP's commercial PC platform. In this role, you will be responsible for envisioning, defining, and driving the architecture of hardware, firmware, software, and cloud-integrated security capabilities that protect HP commercial PCs from current and emerging threats. You will also work across the hardware and security architecture teams, collaborate with HP's Security Research Labs, Firmware (BIOS) team, our security business that creates and deploys industry leading enterprise security and manageability solutions, product management, manufacturing, supply chain, customer enablement, and field organizations to create cohesive system solutions that deliver unique customer value. This position requires deep expertise in embedded systems, platform firmware, hardware, security, device manageability, and PC system architecture. The role also requires the ability to translate complex technical capabilities into customer-relevant value propositions and to influence internal and external stakeholders without relying solely on formal authority. Responsibilities Serve as principal architect for HP Endpoint Security Controller roadmap, associated hardware and firmware based capabilities, cryptographic identity, attestation, secure communication, event logging, and future controller generations. Define and evolve HP Sure Start architecture and related firmware resilience capabilities, including hardware-enforced firmware authenticity, recovery, update, and protection mechanisms. Drive secure BIOS and firmware configuration management strategies, including HP Sure Admin ecosystem capabilities, key management, zero-touch provision

pythonjavac#
View job →
C
Coder
📍 United States• Full-time• Remote
1mo ago

As an Engineering Manager on Coder’s Core Workspaces team, you’ll lead engineers building and evolving the systems behind our agentic development experience. You’ll help make agents more capable, reliable, and useful across real development environments. You’ll guide technical direction while growing the team and keeping execution sharp. You’ll work closely with Engineering, Product, and Design across the agent harness, integrations, and developer workflows. What you’ll do here Lead and grow a team within our Workspaces organization. Set technical direction across the agent harness, integrations, and workflows. Stay close to the code and contribute to architecture and implementation decisions. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve reliability, performance, and operability across agentic systems. Coach engineers, raise the technical bar, and create clarity around priorities and tradeoffs. What we’re looking for Experience managing and growing software engineering teams. Strong hands-on engineering experience with React and TypeScript . Experience with Go . Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS . Strong technical judgment and comfort working through ambiguity. A track record of helping engineers grow while maintaining a high execution bar. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP , agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute. Experience building abstractions across multiple model providers. Deep experience with AWS, Kube

REMOTEtypescriptreactaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability Team Focus Areas Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely. Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads. Databases: Building high performance, distributed database systems that power all of OpenAI's product stack. In this

pythonawskubernetes
View job →
C
Coder
📍 United Kingdom• Full-time
1mo ago

As an Engineering Manager on Coder’s Agentic Engineering team, you’ll lead engineers building and evolving the systems behind our agentic development experience. You’ll help make agents more capable, reliable, and useful across real development environments. You’ll guide technical direction, grow the team, and keep execution sharp. You’ll work closely with Engineering, Product, and Design across the agent harness, integrations, and developer workflows. What you’ll do here Lead and grow a team within our Agentic Engineering organization. Set technical direction across the agent harness, integrations, and workflows. Stay close to the code and contribute to architecture and implementation decisions. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve reliability, performance, and operability across agentic systems. Coach engineers, raise the technical bar, and create clarity around priorities and tradeoffs. What we’re looking for Experience managing and growing software engineering teams. Strong hands-on engineering experience with Go. Experience with React and TypeScript. Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS. Strong technical judgment and comfort working through ambiguity. A track record of helping engineers grow while maintaining a high execution bar. Our tech stack Backend: Go, Postgres Frontend: TypeScript, React Infrastructure: AWS, Kubernetes Observability: Prometheus, Grafana CI/CD: GitHub Actions Bonus tacos if you have (Tacos? If you need an ice-breaker, ask how we say thanks by giving tacos!) Experience building coding agents, developer tools, or cloud development enviro

typescriptreactaws
View job →
D
Datadog
📍 Portugal• Full-time• Remote
1mo ago

This role is part of Datadog’s Security Agent team, which powers critical security capabilities across Workload Protection, Vulnerability Management, Cloud Security products, and other emerging security offerings. As a Staff Software Engineer, you will lead the design and development of low-level Linux instrumentation and runtime security technologies that help customers detect threats, monitor system activity, and protect cloud-native workloads at scale. You will work on complex technical challenges involving eBPF, Linux kernel internals, performance-sensitive systems, and large-scale data collection while influencing technical direction across multiple product teams. This role offers significant ownership, broad organizational impact, and the opportunity to shape the future of Datadog’s security platform. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the architecture and development of security agent capabilities that power runtime threat detection and workload protection across Datadog Security products. Design and build reusable eBPF-based monitoring functionality for process, file, and network visibility within Linux environments. Drive end-to-end delivery of new features, from technical strategy and design through implementation, testing, and rollout. Establish and evolve testing methodologies that improve platform coverage, detection quality, reliability, and performance. Partner with product, security, infrastructure, and engineering teams to deliver shared platform capabilities used across multiple Datadog products. Provide technical leadership by influencing engineering direction, mentoring peers, and helping resolve complex cross-functional challenges. Who You Are: You have significant experience building software in Linux environments,

REMOTElinuxaigo
View job →
D
Datadog
📍 Remote; Spain, Remote• Full-time• Remote
1mo ago

This role is part of Datadog’s Security Agent team, which powers critical security capabilities across Workload Protection, Vulnerability Management, Cloud Security products, and other emerging security offerings. As a Staff Software Engineer, you will lead the design and development of low-level Linux instrumentation and runtime security technologies that help customers detect threats, monitor system activity, and protect cloud-native workloads at scale. You will work on complex technical challenges involving eBPF, Linux kernel internals, performance-sensitive systems, and large-scale data collection while influencing technical direction across multiple product teams. This role offers significant ownership, broad organizational impact, and the opportunity to shape the future of Datadog’s security platform. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the architecture and development of security agent capabilities that power runtime threat detection and workload protection across Datadog Security products. Design and build reusable eBPF-based monitoring functionality for process, file, and network visibility within Linux environments. Drive end-to-end delivery of new features, from technical strategy and design through implementation, testing, and rollout. Establish and evolve testing methodologies that improve platform coverage, detection quality, reliability, and performance. Partner with product, security, infrastructure, and engineering teams to deliver shared platform capabilities used across multiple Datadog products. Provide technical leadership by influencing engineering direction, mentoring peers, and helping resolve complex cross-functional challenges. Who You Are: You have significant experience building software in Linux environments,

REMOTElinuxaigo
View job →
D
1mo ago

This role is part of Datadog’s Security Agent team, which powers critical security capabilities across Workload Protection, Vulnerability Management, Cloud Security products, and other emerging security offerings. As a Staff Software Engineer, you will lead the design and development of low-level Linux instrumentation and runtime security technologies that help customers detect threats, monitor system activity, and protect cloud-native workloads at scale. You will work on complex technical challenges involving eBPF, Linux kernel internals, performance-sensitive systems, and large-scale data collection while influencing technical direction across multiple product teams. This role offers significant ownership, broad organizational impact, and the opportunity to shape the future of Datadog’s security platform. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the architecture and development of security agent capabilities that power runtime threat detection and workload protection across Datadog Security products. Design and build reusable eBPF-based monitoring functionality for process, file, and network visibility within Linux environments. Drive end-to-end delivery of new features, from technical strategy and design through implementation, testing, and rollout. Establish and evolve testing methodologies that improve platform coverage, detection quality, reliability, and performance. Partner with product, security, infrastructure, and engineering teams to deliver shared platform capabilities used across multiple Datadog products. Provide technical leadership by influencing engineering direction, mentoring peers, and helping resolve complex cross-functional challenges. Who You Are: You have significant experience building software in Linux environments,

linuxaigo
View job →
D
Datadog
📍 New York• Full-time• From $154K/yr
1mo ago

We’re looking for a Senior Technical Product Marketing Manager to join our security product marketing team and help bring Datadog’s rapidly growing security offerings to market. In this high-impact role, you’ll collaborate closely with Product, Sales, Sales Engineering, and Enablement to translate complex technical capabilities into compelling narratives that drive awareness, adoption, and differentiation. As the first hire in this space, you'll own key go-to-market efforts, lead technical positioning for strategic initiatives, and mentor others on content strategy and enablement best practices. This is a unique opportunity to shape how Datadog tells its security story to a global market. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Define and execute the technical marketing plan for Datadog’s security product line, from launches to scaled adoption. Partner cross-functionally with Product, PMM, Sales Engineering, and Enablement to craft differentiated messaging, inform roadmap decisions, and build cross-product solution narratives. Create and deliver high-impact sales tools including battlecards, investigation flows, objection handling guides, and competitive workshops. Lead competitive strategy by synthesizing market insights and producing content that positions Datadog as a differentiated leader in cloud-native security. Act as a technical subject matter expert and trusted advisor — coaching field teams, reviewing enablement content, and influencing internal strategy. Represent Datadog in customer briefings, industry events, and webinars, serving as a go-to voice on observability and security. Who You Are: 8+ years of experience in technical product marketing, developer relations, product management, solutions engineering, or related roles in

kubernetesaigo
View job →
M
1mo ago

As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams. Success in this role means smoother launches, clearer roadmaps, stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale. This role can be based out of our Dublin or Cork office or remotely in Ireland. What You'll Do Drive Program Planning & Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep work clearly tracked in Jira, and deliver on time Strengthen Production Reliability – Lead change management and launch readiness programs. Partner with SREs and product teams to define and operationalize SLOs/SLIs, and use incident data, metrics, and capacity signals to drive prioritization and continuous improvement Lead Cross-Functional Coordination – Align SRE with Security, Compliance, Cloud platform, and other engineering teams. Coordinate cross-team incident response, ensure clear follow-through, and build trust as the go-to driver of complex, multi-team efforts Build Scalable Systems & Processes – Design lightweight frameworks and communication patterns that help SRE deliver reliably at scale. Work yourself out of the "hero" role by leaving teams better-equipped to execute independently Requirements 8+ years in technical program management, engineering management, or a comparable technical role partnering with software engineering teams Proven track record leading large-scale, cross-team platform initiatives through ambiguity and change Strong knowledge of production change management, software development lifecycle, and reliability metrics (SLOs, SLIs) Skilled at shaping roadmaps and managing dependencies Able to query and interpret

mongodbawsazure
View job →
🔔

Get new lead cloud operations engineer jobs by email

Daily job updates · Unsubscribe anytime