Jobiba hiring network

Lead Cloud Operations Engineer Jobs

6,876 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

DC
Diligent Corporation
📍 New York• Full-time• From $131K/yr
16 days ago

Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production

pythonawsazure
View job →
J
Jamf
📍 Us Remote• Full-time• Remote• From $113.3K/yr
1mo ago

At Jamf, we believe in an open, flexible culture based on respect and trust. Our track record and thriving work environment all stem from the freedom we grant ourselves to get the job done right. We take pride in helping tens of thousands of customers around the globe succeed with Apple. The secret to our success lies in our connectivity, while operating with a high degree of flexibility. Work-life balance remains our priority while feeling connected is important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: The Senior Software Engineer is responsible for building the tools required to help organizations succeed with Apple. Lead others on the agile team to break down problems and apply the appropriate designs and practices to build Jamf products. Subject matter expert in various Jamf components and product offerings. Mentor and coach others while delivering new components and features with high quality and reliability. You may be required to work periodically at a Jamf office or collaborative work location with other Jamf employees in your area for certain events or moments that matter. What you can expect to do in this role : Break down customer problems into work you and the team can execute on. Independently complete tasks from start to finish with high quality. Ability to communicate technical concepts to stakeholders. Use your knowledge of Engineering best practices to ask the right questions, solve problems and build great software with a high level of quality. Produce designs for new and existing features. Clearly communicate technical concepts with others in the organization (Technical Communication, Support, Product and Cloud). Performs all job responsibilities in alignment with the core values, mission and purpose of the organization. Adheres to the highest moral, ethical and legal standards to deliver and environment that promotes respect, innovation and creativity

REMOTErestagileai
View job →
M
Mongodb
📍 New York City; United States• Full-time• From $126K/yr
1mo ago

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. We are looking for a talented Senior Software Engineer to join our team as we execute on a multi-year roadmap and prepare to launch and rapidly scale services that handle petabytes of data. Come do some of the most interesting work of your career as we Think Big and Go Far for our customers! This role can be based out of our New York City office (hybrid working model) or remotely in the North America region. What you’ll do Design, build, and operate control plane services powering an elastic and multi-tenant storage layer for thousands of database instances. Solve problems around maintaining high availability and performance during load spikes, hardware failures, cloud provider outages, and other disruptions. Contribute to a culture of operational excellence through dashboards, playbooks, and on-call improvements. Lead complex technical projects from planning through successful deployment with clear stakeholder updates. Partner closely with peers across database, cloud, and infrastructure engineering teams as well as project management to investigate incidents and develop long-term roadmaps. Mentor junior engineers and foster a collaborative team environment. We’re looking for someone with 5+ years of professional software development experience building, deploying, and operating multi-tenant cloud services with a focus on operational excellence. Experience with large backend/compiled codebases, such as Rust or C/C++. Experience with containerization and orchestration platforms (e.g. Kubernetes). Experience with observability tooling (e.g. time series metrics, dashboards). Experience with distributed systems

mongodbawsazure
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Opportunity: Okta Access Gateway (OAG) Enterprises run on a mix of modern cloud services and mission-critical on-premises systems (such as Oracle E-Business Suite, SAP, PeopleSoft, and custom legacy web apps). Okta Access Gateway (OAG) solves the enterprise hybrid cloud challenge by extending Okta’s cloud identity, Adaptive MFA, and Zero Trust security policies to on-premises and legacy applications without requiring custom code changes or traditional VPNs. As the Engineering Manager for Okta Access Gateway in Toronto, you will lead and grow a team of software engineers building the next generation of our hybrid access and gateway infrastructure. You will partner closely with Product Management, Architecture, Security, and Quality teams to deliver high-throughput, mission-critical security software deployed across multi-cloud and enterprise datacenters globally. What You’ll Do People Leadership & Team Growth Lead, mentor, and empower an engineering team, fostering an inclusive, high-performance, and psychologically safe engineering culture. Drive career progression, goal setting, regular 1:1s, and continuous feedback to help engineers grow their technical and leadership skills. Attract, interview, and hire diverse engineering talent to scale Okta’s engineering presence in Toronto. Delivery & Operational Excellence Own the end-to-end execution and delivery of key product roadmap initiatives, balancing feature velocity, technical debt, and softwar

javaawsazure
View job →
J
Jumio
📍 India• Full-time• Remote
16 days ago

Role Purpose We’re looking for a Staff/Senior Machine Learning Engineer with deep expertise in computer vision and biometrics to lead the design and scaling of face recognition systems in production. You’ll build and train models, and own ML systems end-to-end on AWS. The final job level for this role will be determined following the interview process. What You’ll Do Lead the design and development of computer vision systems for biometrics (face attributes, detection, quality, and recognition) Rigorous fairness analysis and benchmarking of biometric models across various datasets and operating conditions. Architect, train, and optimize models using PyTorch, Tensorflow, and/or JAX Own and evolve end-to-end ML pipelines, from data ingestion to deployment. Design automated pipelines (Airflow) for data ingestion and cleaning. You will be responsible for curating balanced training sets and generating synthetic data to address both quality and diversity gaps. Production Engineering: Own the path to production. Optimize models for low-latency inference (quantization, distillation, TensorRT/ONNX) and manage deployment on AWS. Mentor ML engineers, conduct code/design reviews, and drive technical best practices across the Computer Vision team. What We’re Looking For Experience: 5+ years of industry experience in Machine Learning, with at least 3 years dedicated to Biometrics or Face Analysis. Deep expertise in computer vision and biometrics, especially face recognition. Fairness & Ethics: You understand the sources of algorithmic bias in Computer Vision and have practical experience measuring and mitigating disparate impact. Strong Engineering: Expert proficiency in Python (both machine learning and vision libraries such as Pillow, OpenCV, PyTorch, etc). You write clean, modular, production-ready code. Systems Architecture: Experience designing end-to-end ML pipelines (Data to Train to Deploy) and working with workflow orchestrators like Airflow. Cloud Native: Hands-on ex

REMOTEpythonawsrest
View job →
D
Datadog
📍 Spain; Paris, France• Full-time
1mo ago

As a Staff Engineer on Datadog's Compute – Disruption and Workload Placement team, you'll help define how our Kubernetes fleet scales to meet the demands of rapidly growing AI and cloud-native workloads. You'll work on the systems that ensure engineering teams have the right compute capacity, in the right region, at the right time across AWS, Google Cloud, and Azure. This is a highly technical, high-impact role where you'll shape the future of capacity orchestration, influence platform architecture, and solve infrastructure challenges that directly support Datadog's continued growth. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead the technical direction of capacity management and workload placement for Datadog's Kubernetes platform spanning 100,000+ virtual machines across multiple cloud providers. Design and build systems that optimize how engineering workloads are scheduled and deployed across regions while balancing capacity constraints, reliability, and performance. Partner across infrastructure teams to evolve multi-region and multi-cloud capacity orchestration as Datadog continues to scale. Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency. Use data and capacity signals to influence infrastructure decisions, forecast growth, and improve workload placement strategies. Who You Are: You have significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems. You are an experienced software engineer with strong programming skills, ideally in Go or a comparable systems programming language. You have hands-on experience with at least one major cloud provider (AWS, Google Cloud, or Azure) and understand distributed cloud infrastructure. Yo

awsazurekubernetes
View job →
D
Datadog
📍 Lisbon• Full-time
1mo ago

As a Staff Engineer on Datadog's Compute – Disruption and Workload Placement team, you'll help define how our Kubernetes fleet scales to meet the demands of rapidly growing AI and cloud-native workloads. You'll work on the systems that ensure engineering teams have the right compute capacity, in the right region, at the right time across AWS, Google Cloud, and Azure. This is a highly technical, high-impact role where you'll shape the future of capacity orchestration, influence platform architecture, and solve infrastructure challenges that directly support Datadog's continued growth. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead the technical direction of capacity management and workload placement for Datadog's Kubernetes platform spanning 100,000+ virtual machines across multiple cloud providers. Design and build systems that optimize how engineering workloads are scheduled and deployed across regions while balancing capacity constraints, reliability, and performance. Partner across infrastructure teams to evolve multi-region and multi-cloud capacity orchestration as Datadog continues to scale. Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency. Use data and capacity signals to influence infrastructure decisions, forecast growth, and improve workload placement strategies. Who You Are: You have significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems. You are an experienced software engineer with strong programming skills, ideally in Go or a comparable systems programming language. You have hands-on experience with at least one major cloud provider (AWS, Google Cloud, or Azure) and understand distributed cloud infrastructure. Yo

awsazurekubernetes
View job →
D
1mo ago

As a Staff Engineer on the Data Platform Experience team, you'll help shape how Datadog engineering teams build, operate, and evolve products on the Observability Data Platform. You'll lead the design and delivery of shared platform capabilities that reduce developer friction, improve operational visibility, and enable engineering teams to move faster with confidence. This role combines deep distributed systems expertise with technical leadership across multiple teams, influencing platform strategy while remaining hands-on in the code. You'll have the opportunity to solve company-wide challenges spanning cost intelligence, operational tooling, platform health, and developer experience. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead strategic engineering initiatives that improve how product teams build, operate, and evolve services on the Observability Data Platform. Design and build scalable platform capabilities for cost intelligence, including cloud cost allocation, trend analysis, and optimization recommendations. Develop operational intelligence and self-service tooling that helps engineering teams understand platform health, troubleshoot incidents, and improve operational efficiency. Drive reusable platform services and developer workflows that increase engineering autonomy while reducing operational complexity across multiple products. Provide technical leadership across teams by influencing architecture, mentoring engineers, and raising engineering standards through hands-on technical contributions. Participate in the team's on-call rotation and continuously improve platform reliability, observability, and operational excellence. Who You Are: You have experience designing and building large-scale SaaS or cloud platforms with deep expertise i

javakubernetesai
View job →
D
Datadog
📍 Spain; Paris, France• Full-time
1mo ago

As a Staff Engineer on the Data Platform Experience team, you'll help shape how Datadog engineering teams build, operate, and evolve products on the Observability Data Platform. You'll lead the design and delivery of shared platform capabilities that reduce developer friction, improve operational visibility, and enable engineering teams to move faster with confidence. This role combines deep distributed systems expertise with technical leadership across multiple teams, influencing platform strategy while remaining hands-on in the code. You'll have the opportunity to solve company-wide challenges spanning cost intelligence, operational tooling, platform health, and developer experience. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead strategic engineering initiatives that improve how product teams build, operate, and evolve services on the Observability Data Platform. Design and build scalable platform capabilities for cost intelligence, including cloud cost allocation, trend analysis, and optimization recommendations. Develop operational intelligence and self-service tooling that helps engineering teams understand platform health, troubleshoot incidents, and improve operational efficiency. Drive reusable platform services and developer workflows that increase engineering autonomy while reducing operational complexity across multiple products. Provide technical leadership across teams by influencing architecture, mentoring engineers, and raising engineering standards through hands-on technical contributions. Participate in the team's on-call rotation and continuously improve platform reliability, observability, and operational excellence. Who You Are: You have experience designing and building large-scale SaaS or cloud platforms with deep expertise i

javakubernetesai
View job →
S
Sentry
📍 San Francisco• Full-time• Remote
12 days ago

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role A strong and reliable platform is essential to scaling Sentry for the future. Our Platform organization is responsible for everything that powers Sentry—from cloud infrastructure and streaming systems to storage, deployment, and security. We own the core services and technical foundations that enable every product and engineering team at Sentry to move fast and build with confidence. We're looking for a passionate and pragmatic Senior Staff Software Engineer to help lead this evolution. In this role, you’ll report directly to the VP of Engineering and collaborate with teams across the company to shape the future of Sentry’s platform. What You’ll Do Architect the future of Sentry by translating business needs and product strategy into clear, scalable technical blueprints. Partner with product and engineering leaders to align technical roadmaps with company goals. Lead cross-cutting initiatives across the Platform org—owning them end-to-end and driving meaningful outcomes. Promote engineering excellence by mentoring platform engineers, sharing best practices, and setting high standards for system design, scalability, and operational quality. Review major architectural proposals and help ensure consistency, maintainability, and long-term technical health across the company. You’ll Love This Job If You... Enjoy designing and building platforms that help teams move faster and scale safely. Thrive on solving complex, multi-dimensional problems across product, infrastructure, and organizational layers. Want to make architectural decisions that shape Sentry’s long-term success. Bring new ideas, tools, and frameworks t

REMOTEawsgcpkubernetes
View job →
J
Jumio
📍 Austria• Full-time• Remote
16 days ago

Machine Learning Engineer We’re looking for a Machine Learning Engineer with deep expertise in computer vision and biometrics to lead the design and scaling of face recognition systems in production. You’ll build and train models, and own ML systems end-to-end on AWS. The final job level for this role will be determined following the interview process. What You’ll Do Lead the design and development of computer vision systems for biometrics (face attributes, detection, quality, and recognition) Rigorous fairness analysis and benchmarking of biometric models across various datasets and operating conditions. Architect, train, and optimize models using PyTorch, Tensorflow, and/or JAX Own and evolve end-to-end ML pipelines, from data ingestion to deployment. Design automated pipelines (Airflow) for data ingestion and cleaning. You will be responsible for curating balanced training sets and generating synthetic data to address both quality and diversity gaps. Production Engineering: Own the path to production. Optimize models for low-latency inference (quantization, distillation, TensorRT/ONNX) and manage deployment on AWS. Mentor ML engineers, conduct code/design reviews, and drive technical best practices across the Computer Vision team. What We’re Looking For Experience: Multiple years of industry experience in Machine Learning, with at least 3 years dedicated to Biometrics or Face Analysis. Deep expertise in computer vision and biometrics, especially face recognition. Fairness & Ethics: You understand the sources of algorithmic bias in Computer Vision and have practical experience measuring and mitigating disparate impact. Strong Engineering: Expert proficiency in Python (both machine learning and vision libraries such as Pillow, OpenCV, PyTorch, etc). You write clean, modular, production-ready code. Systems Architecture: Experience designing end-to-end ML pipelines (Data to Train to Deploy) and working with workflow orchestrators like Airflow. Cloud Native:

REMOTEpythonawsrest
View job →
J
16 days ago

Machine Learning Engineer IV – (Computer Vision) We’re looking for a Staff/Senior Machine Learning Engineer with deep expertise in computer vision and biometrics to lead the design and scaling of face recognition systems in production. You’ll build and train models, and own ML systems end-to-end on AWS. The final job level for this role will be determined following the interview process. What You’ll Do Lead the design and development of computer vision systems for biometrics (face attributes, detection, quality, and recognition) Rigorous fairness analysis and benchmarking of biometric models across various datasets and operating conditions. Architect, train, and optimize models using PyTorch, Tensorflow, and/or JAX Own and evolve end-to-end ML pipelines, from data ingestion to deployment. Design automated pipelines (Airflow) for data ingestion and cleaning. You will be responsible for curating balanced training sets and generating synthetic data to address both quality and diversity gaps. Production Engineering: Own the path to production. Optimize models for low-latency inference (quantization, distillation, TensorRT/ONNX) and manage deployment on AWS. Mentor ML engineers, conduct code/design reviews, and drive technical best practices across the Computer Vision team. What We’re Looking For Strong industry experience in Machine Learning, dedicated to Biometrics or Face Analysis. Deep expertise in computer vision and biometrics, especially face recognition. Fairness & Ethics: You understand the sources of algorithmic bias in Computer Vision and have practical experience measuring and mitigating disparate impact. Strong Engineering: Expert proficiency in Python (both machine learning and vision libraries such as Pillow, OpenCV, PyTorch, etc). You write clean, modular, production-ready code. Systems Architecture: Experience designing end-to-end ML pipelines (Data to Train to Deploy) and working with workflow orchestrators like Airflow. Cloud Native: Hands-on exper

REMOTEpythonawsrest
View job →
J
1mo ago

At Jamf, we believe in an open, flexible culture based on respect and trust. Our track record and thriving work environment all stem from the freedom we grant ourselves to get the job done right. We take pride in helping tens of thousands of customers around the globe succeed with Apple. The secret to our success lies in our connectivity, while operating with a high degree of flexibility. Work-life balance remains our priority while feeling connected is important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. This role is offered as a hybrid in Brno, Czech Republic . We are only able to accept applications for those based in the Czech Republic or who have sponsorship to live and work in the Czech Republic. What you'll do at Jamf : At Jamf, we empower people to be their best selves and do their best work. You will join the Netopyre team that contributes to Jamf's Security Cloud vectoring strategy - owning the Jamf Trust macOS app end to end and the vectoring pieces of Jamf Trust iOS. The technical core is Apple's Network Extension frameworks: we build and ship a Content Filter extension for real-time traffic inspection and a packet tunnel provider at the heart of our ZTNA client, running across four shipping app variants, roughly twenty shared Swift packages, and two release trains. Writing clean, well-tested Swift is part of our philosophy, and we stay current with platform evolution - Swift concurrency, Network Extension APIs, and the latest tooling. Collaboration is at the heart of how we operate - we approach challenges as collective, brainstorming and problem-solving together. You will contribute to technical direction, hold the quality bar on agent-produced code and deliberately grow the engineers around you. What you can expect to do in this role: Lead projects end to end as an epic owner - design, coordination and delivery, not just implementation. Carry a full s

restaiswift
View job →
C
1mo ago

The Role We are looking for a Senior Partner Deployed Engineer to join the Customer Solutions team, focused on building Coder’s partner ecosystem across EMEA. This is a new function at Coder, modeled on the Forward Deployed Engineer role pioneered by companies like Palantir and now the fastest-growing technical role at frontier AI companies. The difference: instead of embedding with a single customer to deploy a platform, you embed with strategic partners to help them understand, position, and deliver Coder’s platform across their entire customer base. You are part Field CTO, part Industry Strategist, part Technical Specialist, and part Partner Relations Lead. You will be the technical authority within our EMEA partner ecosystem: setting the vision for how Coder fits into each partner’s AI, cloud, and modernization offerings, co-creating the GTM sales plays that partner sellers take to market, building the demonstrations and workshops that generate pipeline, and producing the Partner Relations content that establishes Coder’s technical brand in the ecosystem. This is not a support role. You are equally comfortable holding a strategic roadmap conversation with a partner CTO and debugging a Kubernetes deployment in a partner’s lab in the same afternoon. You are energized by the challenge of building a new category through partnerships, motivated by the multiplied impact of enabling an entire ecosystem rather than a single customer, and capable of operating with full autonomy in a fast-moving startup environment. What You’ll Do Serve as the strategic technical thought partner to EMEA partner leadership, including practice leads, CTOs, and solutions architects at global systems integrators, regional cloud consultancies, and hyperscaler field teams. Own the technical relationship and set the vision for how Coder fits into each partner’s AI and modernization portfolio. Co-create sales plays tailored to each partner’s customer base, vertical focus, and services capabilitie

awsazurekubernetes
View job →

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are looking for a highly skilled PSIRT Engineer to lead the vulnerability response program for Replit’s cloud-native AI platform. You will own the lifecycle of security vulnerabilities affecting our products and services—from intake to validation, remediation coordination, and public disclosure. This role requires strong technical ability to reproduce vulnerabilities , deep understanding of web/app/cloud exploit classes, and experience operating bug bounty and coordinated disclosure programs. You will work closely with Engineering, Cloud Security, SecOps, SRE, and IT teams to ensure vulnerabilities are fixed quickly and communicated responsibly. What You’ll Do Vulnerability Intake, Triage & Validation Manage intake from bug bounty platforms (HackerOne preferred), customer reports, automated scanners, pentest reports, and coordinated disclosure channels. Independently validate, reproduce, severity-score, and document findings. Identify duplicates and maintain a clean vulnerability records pipeline. Assess relevance and exploitability using OWASP, cloud misconfiguration patterns, and identity/authentication/authorization risks (Oauth, OIDC). Remediation Coordination & SLA Management Work with Engineering, SecOps, IT, SRE, and Cloud Security to confirm product impact and drive remediation. Provide detailed reproduction steps, proof-of-concepts, and technical analyses. Track SLAs, remediation progress, regression testing, and systemic improvements. Support SOC 2, ISO 27001, and pentest evidence needs as part of vulnerability lifecycle governance. Bug Bounty & Vulnerability Disclosure Program Management Design and evolve the bug bounty program, including scope, rules, and reward structures. Man

pythongcpci/cd
View job →
🔔

Get new lead cloud operations engineer jobs by email

Daily job updates · Unsubscribe anytime