MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in New York City for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workf
Jobiba hiring network
Lead Cloud Operations Engineer Jobs
6,876 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current lead cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in Cork for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workflows Pref
About the Team OpenAI’s Governance, Risk, and Compliance team helps ensure security and privacy are grounded in how our products and systems actually operate. Assurance Operations partners with Security, Engineering, Infrastructure, Product, Privacy, and Legal to make controls provable, risk decisions explicit, and audit readiness a result of well-designed systems. About the Role We are hiring a technical, product-minded GRC builder who can own consequential audits while improving the control and evidence systems behind them. You will build a reusable common control framework, use Codex to automate assurance work, validate changing system scope, and turn repeated audit friction into measurable improvements. We are looking for someone who questions inherited assumptions, solves novel problems creatively, works closely with engineers, and makes the next audit easier by improving the underlying system. You’ll be responsible for: Lead external, internal, customer, and certification audit work from scoping through evidence review, fieldwork, remediation, and closeout. Build a common control framework linking risk, control intent, implementation, owner, system, environment, evidence, and applicable frameworks. Validate actual scope and ownership instead of assuming last year's controls, product boundaries, or evidence remain accurate. Use Codex to build and test evidence checks, control mappings, request triage, owner workflows, monitoring, and remediation reporting. Partner with engineers on cloud architecture, identity, logging, data flows, software changes, vulnerabilities, and control effectiveness. Design maintainable, permission-aware tools that preserve source provenance, human review, and evidence integrity. Reduce repeated requests and operational burden for control owners through measurable workflow improvements. Define roadmaps, decision rights, milestones, success metrics, and clear cross-functional escalations. We’re looking for someone with: Direct ownership
Become a part of our caring community Job Description Summary The Lead Solutions Architect provides architecture leadership for CenterWell Home Health programs and platforms, shaping conceptual and reference architectures, governing solution designs, and aligning delivery teams to cloud and data strategies. The scope includes high-priority initiatives as well as interoperability and provider-data integrations that span CenterWell and Humana Insurance. As Lead Solution Architect, you'll be the senior individual contributor on a team with broad accountability across CenterWell's dispensing pharmacy portfolio — mail order, specialty, retail, and associated platforms. You'll own the architectural vision for complex, multi-system initiatives, shape how technology decisions get made, and act as a connective force between business strategy, engineering execution, and enterprise standards. You will operate within CenterWell IT – Cross-CenterWell Architecture. You will collaborate with product, engineering, EA Activation, security, data, and operations. You will engage governance forums to enable Integrated Health across CenterWell and Humana Insurance. Key Activities Quickly conduct structured knowledge transfer with existing architects and relevant stakeholders to capture critical in-flight designs and decisions. Review current initiatives and establish an architectural roadmap aligned with organizational priorities. Develop or refine reference architectures and design patterns for core platforms and solutions. Collaborate with governance and compliance teams to validate designs against enterprise standards and regulatory requirements. Define integration strategies and solution blueprints for key systems and data flows. Establish architecture review processes and decision forums to support de
About the Role: We are looking for a Senior DevOps Engineer to join our DevOps team at K Health. You will own and evolve the infrastructure underpinning a healthcare AI platform serving patients and enterprise health system partners. This is a high-ownership role: you will architect and operate cloud environments across K Health and its enterprise partners, lead complex infrastructure migrations, drive disaster recovery programs, and help build the next generation of AI-powered operations tooling. You will also mentor junior engineers and collaborate closely with product and engineering teams across the company. This is a hybrid role based in New York City (4 days/week in office) and includes participation in a daytime on-call rotation. What you will do: Own the design, implementation, and evolution of our GKE-based Kubernetes infrastructure across K Health and enterprise partner environments. Build and maintain our Terraform modular infrastructure library, including reusable modules with automated testing, across GCP, Cloudflare, and AWS. Architect, build, and maintain GitLab CI/CD shared pipeline templates used by all engineering teams (build, test, security scanning, deployment). Own and maintain self-hosted infrastructure software running in-cluster, including GitLab, ArgoCD, Langfuse, DependencyTrack, NGINX Ingress, and others. Implement and support security and compliance controls across infrastructure and the software supply chain - secrets management, pipeline secret detection, container scanning, SOC2 and HIPAA. Drive disaster recovery readiness: design failover scenarios, author runbooks, and lead periodic DR tests. Lead development of AI-powered operations tooling and agentic infrastructure. Monitor, troubleshoot, and improve production system reliability; respond to incidents during on-call shifts. Mentor junior DevOps engineers and establish team-wide engineering standards. What we are looking for: 5+ years of experience in DevOps, platform engineering,
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet's customers in Europe, the Middle East, and Africa expect our Customer Trust team to understand their compliance challenges, speak their language, and respond quickly to security assessments. We're looking for a Sr. Security Engineer I to lead Customer Trust operations across EMEA—responding to security questionnaires, managing vendor risk assessments, and building trusted relationships with enterprise customers in these regions. You will be responsible for questionnaire triage, completion, and queue management for EMEA customers. You'll work closely with EMEA sales teams, understand regional compliance requirements (GDPR, EU data protection, sector-specific frameworks), and ensure Smartsheet maintains a strong reputation for responsiveness and technical credibility in these high-value markets. You will work remotely from the UK and report to our Sr. Director, GRC Engineering, based in the US You Have 5+ years of experience in customer trust, vendor risk management, security assessment, or customer-facing security roles at SaaS or cloud platform companies. Proven experience completing and responding to customer security questionnaires, vendor assessments, and RFIs. Strong understanding of GDPR, EU data protection, and regional compliance requirements: Familiarity with data residency, data processing agreements, DPIAs, and how cloud services operate within EU regulatory frameworks. Knowledge of GRC frameworks: Working knowledge of SOC 2, ISO 27001, and compliance standards commonly referenced in EMEA asse
About the Team OpenAI’s Compute Strategy team is responsible for securing and scaling the core resources that power our research and products. We partner across engineering, finance, legal, and operations to identify, negotiate, and execute strategic partnerships that expand OpenAI’s capacity for compute, power, and data center infrastructure. Our mandate spans energy procurement, real estate development, colocation, cloud service providers, silicon and strategic supply chain, and infrastructure financing—ensuring OpenAI can grow with speed, resilience, and cost-efficiency. About the Role We are hiring several Business Development Lead, Compute Strategy positions focused on compute infrastructure. Each hire will bring deep expertise in one or more focus areas while collaborating across the broader infrastructure stack. In this role, you will source opportunities, structure partnerships, and negotiate high-value agreements across OpenAI’s infrastructure ecosystem. You will work directly with external partners and suppliers while collaborating internally with engineering, legal, finance, and operations to ensure we have the resources needed to support state-of-the-art AI systems. This role requires technical fluency, commercial judgment, and disciplined execution. Your work will directly shape how quickly, reliably, and efficiently OpenAI can bring new compute capacity online. Each hire will focus on building partnerships and executing deals in one or more of the following areas: Energy and Power: securing scalable and sustainable energy supply. Land and Real Estate: identifying and securing strategic sites. Colocation : evaluating and contracting for third-party data center capacity. Cloud Service Providers (CSPs): structuring partnerships with hyperscalers and specialized AI cloud providers. Silicon: building semiconductor partnerships to secure advanced silicon and resilient long-term supply. Fiber & Equipment: securing fiber & critical data center equipmen
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta Workflows is the secure, no-code automation platform that empowers organizations to build identity-centric workflows across cloud applications — all without writing code. Our intuitive drag-and-drop interface allows enterprises to automate complex business processes at scale, enhancing productivity, enforcing security, and simplifying IT operations. Customers like Netflix, MGM, and NTT rely on Workflows to automate high-impact identity scenarios with speed and confidence. As we continue to scale, we’re investing in extensibility, developer experience, and performance. Join us to help shape the future of cloud automation and low-code development. Location: Bengaluru, Karnataka, India Work Mode: Hybrid (2-3 days Onsite per week) Note: "This role requires in-person onboarding and travel to our Bengaluru, IN office during the first week of employment." Position Description We’re hiring a Staff Full-Stack Engineer to join the Flow builder team within Okta Workflows. This team owns the core no-code canvas that enables both internal teams and our customers to build powerful automation experiences with ease. As a Staff Engineer, you’ll lead initiatives that span front-end and back-end services — delivering performant, secure, and scalable features. You’ll help define architecture, drive implementation, and collaborate closely with Design, PM, and Platform teams. You’ll also work directly with our technical architects to help shape what
About the Role Join Peloton’s Global Network Services team as a Senior Engineer, sitting at the unique intersection of high-scale enterprise networking and high-stakes media production. In this role, you will architect and support the global infrastructure that powers our corporate offices, warehouses, and flagship New York Broadcast Studio. Acting as the key bridge between Global Cloud Engineering and Studio Operations, you will deploy highly resilient, scalable architectures that deliver live streaming content seamlessly to millions of members worldwide. Your Daily Impact Global Architecture & Deployment: Design, optimize, and secure Peloton's global network footprint, integrating Cisco, Meraki, Aruba, and Palo Alto Networks across hybrid on-premises and AWS environments. Studio & Broadcast Reliability: Lead deep-dive traffic analysis and troubleshooting for our NY studio, ensuring 24/7 uptime for live broadcasts, real-time video streaming, and OTT media delivery. Operational Leadership & Lifecycle Management: Manage the end-to-end network project lifecycle—from traffic shaping and SD-WAN optimization to establishing SOPs, security policies (with InfoSec), and handling Tier-3 disaster recovery. You Bring To Peloton Broad & Deep Network Expertise: 8+ years in Network Engineering, including 6+ years in complex SaaS environments and 4+ years designing public cloud networking (specifically AWS). Advanced Protocol & Security Mastery: Expert knowledge of L2/L3 protocols (BGP, OSPF, EIGRP), security protocols (IPsec, 802.1x, RADIUS), SD-WAN, and SDN/SDDC full-stack solutions. Media & Streaming Specialization: 3+ years optimizing IP networks specifically for live video delivery, utilizing multicast/unicast technologies and streaming protocols like HLS, RTMP, and SRT. Education & Elite Certifications: A Bachelor’s degree in Engineering or Computer Science, backed by active CCIE or JNCIE certifications (required). Collaborative Mindset: A curious
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta Workflows is the secure, no-code automation platform that empowers organizations to build identity-centric workflows across cloud applications — all without writing code. Our intuitive drag-and-drop interface allows enterprises to automate complex business processes at scale, enhancing productivity, enforcing security, and simplifying IT operations. Customers like Netflix, MGM, and NTT rely on Workflows to automate high-impact identity scenarios with speed and confidence. As we continue to scale, we’re investing in extensibility, developer experience, and performance. Join us to help shape the future of cloud automation and low-code development. Position Description We’re hiring a Staff Full-Stack Engineer to join the Integration Builder team within Okta Workflows. This team owns the core no-code surface that enables both internal teams and third-party developers (ISVs) to build powerful integrations and automation experiences with ease. As a Staff Engineer, you’ll lead initiatives that span front-end and back-end services — delivering performant, secure, and scalable features. You’ll help define architecture, drive implementation, and collaborate closely with Design, PM, and Platform teams. You’ll also work directly with our technical architects to help shape what we build — and how we build it. This is a high-impact role in a growing, strategic product area with strong executive visibility. Role Details: Design, build, and maintain end-to-end feature
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! About the Role We are seeking a seasoned Manager, Software Engineering with 12+ years of experience to lead our Database Engineering and Cloud Infrastructure team. In this role, you will lead a team of high-performing engineers responsible for architecting, scaling, and optimizing multi-cloud relational and in-memory database platforms. You will bridge technical execution, engineering leadership, and strategic infrastructure planning across AWS and Azure environments. Key Responsibilities Technical Leadership & Architecture Lead the architectural design and operations of enterprise-grade, multi-cloud relational databases across AWS (RDS PostgreSQL, MySQL, Aurora) and Azure (Database for PostgreSQL/MySQL, Azure SQL Managed Instance). Drive high-availability architecture strategies, including Multi-AZ deployments, auto-failover groups, read replica scaling, and cross-region disaster recovery (DR). Oversee zero-downtime operations, including major-version engine upgrades, schema migrations, and blue/green deployment strategies. In-Memory Infrastructure & Open-Source Strategy Manage scale operations for in-memory datastores (AWS ElastiCache, Azure Cache for Redis), focusing on cluster mode operations, eviction policies, and persistence tuning. Spearhead open-source caching initiatives and migration pathways from Redis to Valkey (e.g., AWS ElastiCache for Valkey) using zero-downtime tools like RedisShake to ensure open-source license compliance and optimize cloud spend. Aut
Join Delphi - Where Innovation meets transformation At Delphi, we believe in creating an environment where our people thrive. Our hybrid work model empowers you to choose where you work—whether it's from the office, your home, or a mix of both—so you can prioritize what matters most. We are committed to supporting your personal goals, family, and overall well-being while driving transformative results for our clients. We welcome exceptional talent from anywhere across the globe. Interviews and onboarding are conducted virtually, reflecting our digital-first mindset. Rooted in the region, we specialize in delivering tailored, impactful solutions in Data, Advanced Analytics and AI, Infrastructure, Cloud Security, and Application Modernization. Whether it’s enabling predictive analytics , transforming operations with automation, or driving customer engagement with intelligent platforms, we are the trusted partner for organizations ready to embrace a smarter, more efficient future. We are looking for a highly skilled and detail-oriented Consultant – Quality Assurance to ensure the delivery of high-quality, secure, and reliable software solutions. The ideal candidate will have strong expertise in test strategy, automation frameworks, security testing, and AI-driven QA practices. This role requires close collaboration with cross-functional teams to drive quality across the SDLC while implementing modern automation and security standards. What You’ll Do: • Design, develop, and execute comprehensive test strategies, test plans, and test cases • Perform functional, regression, integration, system, and UAT testing • Lead and implement test automation frameworks to improve efficiency and coverage • Conduct Vulnerability Assessment & Penetration Testing (VAPT) – Expert Level (Preferred) to identify security risks • Integrate AI/ML-driven testing solutions for pre
Join Delphi - Where Innovation meets transformation At Delphi, we believe in creating an environment where our people thrive. Our hybrid work model empowers you to choose where you work—whether it's from the office, your home, or a mix of both—so you can prioritize what matters most. We are committed to supporting your personal goals, family, and overall well-being while driving transformative results for our clients. We welcome exceptional talent from anywhere across the globe. Interviews and onboarding are conducted virtually, reflecting our digital-first mindset. Rooted in the region, we specialize in delivering tailored, impactful solutions in Data, Advanced Analytics and AI, Infrastructure, Cloud Security, and Application Modernization. Whether it’s enabling predictive analytics , transforming operations with automation, or driving customer engagement with intelligent platforms, we are the trusted partner for organizations ready to embrace a smarter, more efficient future. We are looking for a hands-on Senior QA Consultant to lead end-to-end testing of AI and Generative AI applications across enterprise environments. This role combines strong expertise in traditional QA engineering with modern AI evaluation and validation practices. The ideal candidate will drive quality assurance initiatives for RAG pipelines, multi-agent systems, OCR and Speech-to-Text solutions while ensuring production grade quality outcomes in regulated industries such as Finance, Healthcare, and Insurance. The successful candidate will lead a small QA pod, collaborate closely with engineering, AI/ML, product, and business teams, and act as the client-facing QA owner for enterprise AI engagements. This role requires strong technical leadership, automation expertise, AI evaluation capabilities, and excellent stakeholder management skills. Experience Requirements • 7–10 years of experien
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Digital Experience (DX) business unit to deliver the next generation of intelligent, cloud-native fleet management solutions for our customers. You will contribute to our mission of simplifying operations by building scalable SaaS platforms that leverage AI, security, and automation . This role involves working across the full application lifecycle, from Pure IAM to Pure1 Manage , collaborating with cross-functional teams to translate business needs into resilient, production-ready systems. WHAT YOU'LL DO Own the end-to-end design, development, and operation of mission-critical processing services , ensuring seamless, secure, and compliant data flow between edge devices and the Pure1 cloud platform. Partner with Product and Architecture teams to translate complex requirements into scalable, resilient architectural designs that drive significant organizational impact from initial concept through to production deployment. Drive continuous innovation by experimenting with new technologies, platform ecosystems, and architectural patterns to improve system performance, security, and cost-effectiveness for petabytes of real-time data. Maintain a quality-first mindset throughout the entire software development lifecycle, emphasizing comprehensive unit testing, thorough code reviews, and robust Continuous Integration/Continuous Deployment (CI/CD) pipelines. Lead the resolution of complex inter-operab
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As a hands-on senior leader for our India SecOps team, you will shape and safeguard Everpure’s security posture at the intersection of detection engineering, threat hunting, attack surface management, and incident response. Positioned as a strategic cornerstone in Bangalore, you will empower an elite engineering team, optimize critical SecOps pipelines, and partner cross-functionally across global engineering and infrastructure groups. By driving execution excellence and high team morale, you ensure our enterprise platform and global telemetry remain resilient against evolving threats. WHAT YOU'LL DO Scale & Lead SecOps Operations: Architect, mentor, and grow the India SecOps team to foster an environment of high morale, technical excellence, and rapid execution across detection engineering and incident response. Proactively Manage & Remediate Attack Surface: Own end-to-end Attack Surface Management (ASM) across cloud environments, SaaS applications, endpoints, and secrets management to measurably minimize enterprise exposure and mitigate risk. Optimize Telemetry & Incident Response: Mature SIEM and SOAR automation pipelines to drastically reduce mean time to detect, contain, and respond (MTTD/MTTC/MTTR) while continuously elevating alert fidelity and signal confidence. Drive Cross-Functional Alignment & RCA Postmortems: Lead continuous validation through purple-teaming and incident postmortems alo
Get new lead cloud operations engineer jobs by email
Daily job updates · Unsubscribe anytime