Clear all

Jobiba hiring network

Senior Cloud Operations Engineer Jobs

7,292 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
Okta
📍 Bengaluru• Full-time
16 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Company Description: Okta - we are the World’s Identity Company. We don’t just protect logins; we secure the digital life of the Fortune 100. Built from the ground up in the cloud, Okta securely and simply connects people to their applications from any device, anywhere, at any time. Okta integrates with existing directories and identity systems, as well as thousands of on-premises, cloud and mobile applications, and runs on a secure, reliable and extensively audited cloud-based platform. Who are we looking for: We are looking for a Senior Software Engineer in Test. An individual who takes ownership and builds viable solutions. A team oriented individual who can demonstrate working independently, as an individual contributor, in a distributed working environment. Detail oriented and methodical in approaching tasks with excellent research and analytical skills. An ideal candidate will be someone who is passionate about automation & appling those automation skills in testing, cloud native, large-scale, mission-critical software in a fast-paced agile environment while partnering with cloud Infrastructure and operations teams. Okta engineering strongly believes in automated testing, and an iterative process to build high-quality next generation software. This role is mainly focused on automation of Cloud Infrastructure testing and supporting SREs, including bespoke solutions rolled out by Developer Productivity teams. The Quality Engineering team work

pythonsqlaws
View job →
O
Okta
📍 Toronto• Full-time• From C$136K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Platform Network Engineering Team Auth0 by Okta is an easy-to-implement authentication and authorization platform designed by developers for developers. We make access to applications safe, secure, and seamless for over 100 million daily logins worldwide. Our modern approach to identity enables this Tier 0 global service to deliver convenience, privacy, and security so customers can focus on innovation. The Senior Software Engineer Opportunity You will be part of the Platform Network engineering team responsible for all connectivity of Auth0. You will play a key engineering role as we evolve our network architecture to meet the demands of enormous growth and support the hundreds of millions of users who rely on us to provide uninterrupted access. You will get to work with engineers throughout the engineering organization. What you’ll be doing Implement internal and edge networking infrastructure and design solutions that work at global scale and with multi-cloud and multi-region constraints. Carry cross-team initiatives from end to end: code reviews, design reviews, operational robustness, security hygiene, etc. Design and develop new services, tools, and automation to expose network functionality to other Okta engineering and operations teams. Research and implement solutions addressing cross-cutting concerns such as routing, failover, and scaling. Participate in the team’s on-call rotation. What you’ll bring to the role Have 3+ years of

awsazurekubernetes
View job →
G
Godaddy
📍 British Columbia• Full-time• From C$107K/yr
1mo ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operati

pythonawsci/cd
View job →

Location Details: At GoDaddy, the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operat

pythonawsci/cd
View job →
N
Nvidia
📍 Remote, United States• Remote
1mo ago

NVIDIA is looking for an experienced software engineer with infrastructure experience to become a senior member of the Cloud Foundations Automation - Development Team. We build and manage the automation ecosystem supporting NVIDIA's GPU Cloud and NVIDIA SuperPod deployments. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most hard-working and dedicated people on the planet working for us. If you're creative and autonomous, we want to hear from you! What you'll be doing: Developing software to enable efficient network design, deployment and day 2 management. Building product focused software solutions, used by internal and external customers. Helping us as we transform our workflows and organization into a centrally orchestrated configuration management framework, operating at scale across geographies. Owning and driving integrations with various service APIs such as Cloud Service Providers, to automate creation of environments and auto populate data sources in turn. Building on open source software, designing and implementing data structures and UI interfaces to automate processes from equipment purchase to device config generation to deployment to operations. Streamlining deployment mechanisms and life cycle operations Developing modern service architectures around streaming data and event pipelines. Working with infrastructure domain experts on true, zero touch deployment solutions and utilizing best of breed high performance computing management solutions. Be a proactive problem solver, looking out for new opportunities to improve our services and customer experience. Communicate readily with your peers across the organization, b

REMOTEpythonkubernetesai
View job →
H
Hasbro
📍 Boston• Full-time• From $119.6K/yr
15 days ago

We take play seriously. We’re looking for curious adventurers ready to find their party, fueled by imagination and drive to build what’s never been built before. At Hasbro and Wizards of the Coast, you’ll collaborate with passionate teams to reimagine our iconic brands and create experiences that spark joy, connection, and community through the magic of play. This is your chance to shape legendary play that lasts a lifetime. The Senior Network Engineer leads Wizards of the Coast's enterprise network — datacenters, corporate offices, studios, and AWS cloud. Our stack runs on a Juniper/Mist campus fabric, a Palo Alto Networks security edge, and cloud-native AWS connectivity. You'll set technical direction, drive complex initiatives end to end, and mentor the broader team. Come help us build the future of network operations. What You'll Do: Own the architecture, build, and roadmap for our Juniper/Mist campus and branch infrastructure, and lead its evolution across sites and business units. Lead the Palo Alto Networks security stack, from policy architecture to secure-by-design standards across teams. Own end-to-end AWS cloud network implementation — VPC, Transit Gateway, Direct Connect/VPN, Route 53 — and hybrid connectivity, including BGP, OSPF, and SD-WAN traffic engineering at scale. Drive automation and AI adoption across network operations, from config deployment to AI-powered monitoring, observability, and root-cause analysis. Be the go-to for critical issues, and mentor less-experienced engineers through code review, troubleshooting, and skill-building. What You'll Bring: 10+ years in enterprise network engineering — routing, switching, wireless — with deep expertise in Juniper (EX/QFX/SRX, Mist) and/or Cisco (Catalyst, Nexus), plus hands-on work with intelligent ops tools like Mist/Marvis at scale. Expert-level BGP, OSPF, SD-WAN, and load balancing chops. You've built these solutions from scratch, not just maintained them! Extensive Palo Alto Net

M
Mongodb
📍 New York City• Full-time• From $126K/yr
1mo ago

The MongoDB Cloud Services Team is a diverse group of contributors working together to help our users manage MongoDB at global scale. The Cloud Team is responsible for MongoDB Atlas: our database as a service offering, and fastest growing product, which allows users to deploy fault-tolerant, globally distributed MongoDB clusters in just minutes. The Backup Team delivers essential infrastructure to help our customers in their hour of need - providing the ability to quickly restore a massive, distributed database to any point in time at the click of a button. The Backup Team’s mission is to make MongoDB backup more reliable, faster, and also cheaper. This team is responsible for the Backup Agent (Go), the extensive server-side infrastructure (Java) which manages 100s of TB of data and processes billions of operations per day, and the user interface (Javascript) that customers use to manage their backups. Common project themes are performance, scaling, and ease of use. We are looking to speak to candidates who are based in New York for our hybrid working model. We're looking for someone who is Skilled at writing large-scale, distributed backend systems in a compiled language (Java, C#, Go, etc.) Fond of chasing down tough problems in a distributed systems environment Cool under pressure - has wrangled production crises, and secretly finds this a little fun Experienced with Linux, and able to correlate application performance problems with underlying hardware limits Comfortable working across the stack of a modern web application Always striving to expand their knowledge Curious, collaborative and intellectually honest Responsibilities Work closely with product teams, considering the user’s perspective while helping the team achieve success Collaborate with team members over best practices and core concepts Hold yourself accountable to your actions, maintaining the balance between accomplishing goals with research & development Own our

javascriptjavamongodb
View job →
N
Nvidia
📍 Remote, United States• Remote
11 days ago

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data

REMOTEpythonsqlaws
View job →
G
Guidepoint
📍 Mumbai• Full-time
15 days ago

Overview: The role is responsible for managing and supporting the operations of Global Network Engineering in a hybrid/cloud environment and provides senior-level expertise to the network engineering team. We're looking for an expert in on-premises networking, cloud infrastructure networking (Azure/AWS), and telecommunications (VOIP) in global environments, with knowledge and focus on zero-trust networking methodologies. This position requires off-hours on-call availability. This is a remote position with a 3 PM - 12 AM Shift. Candidates may need to visit the Mumbai office as and when needed in the general shift. What You'll Do: Manage physical network management and support covering Guidepoint offices, including, but not limited to, firewalls, routers, switches, etc. Responsible for managing the Enterprise WiFi Operations ZScaler Network management Monitor global traffic latency, issues, and application performance Update and design next-generation network models Network detection and response Intrusion detection and prevention technologies Azure Cloud Traffic integrating with On-prem traffic Azure Web Security CASB models and methodologies What You Have: 8+ years of experience with core networking in an enterprise environment, ideally global network design and security methodologies. 5+ years of working experience with advanced Azure cloud networking technologies. Must have substantial experience with managing WiFi Operations. CCNP Certified or equivalent experience 3+ years of hands-on experience, preferably with Sonic-wall/Netgear. Solid engineering knowledge of routing, switching (VLANs), Azure networking, firewalls, and traffic management. Update network security based on VNETs and Subnets, with traffic monitoring, etc. Knowledgeable in Azure Application Gateways, Front Door, and other security tools. Nice to have: Must know telecommunication technologies (VOIP, SIP Trunking, Microsoft Teams) Network security experience (vulnerability manage

awsazureai
View job →
N
1mo ago

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys

pythonjavakubernetes
View job →

Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. This position is based at our office in Chennai, India. Appian was built on a culture of in-person collaboration, which we believe is a key driver of our mission to be the best. You will be the product manager working closely with the team whose mission is to strengthen and optimize site infrastructure by delivering essential upgrades, resource efficiency, and scalable solutions. You will be responsible for the direction and roadmap of a component of the Appian Cloud data plane that ensures reliable, high-performance operations for all Appian Cloud customer sites. This role is specifically focused on the cloud-native persistence and messaging layer. You will oversee the backend sub-systems—including technologies like S3 and Redis—that power the platform's internal data plane and core services. This component of the software is not directly user-facing but has strong implications on the scalability and reliability requirements our customers expect. What you will be doing: Prioritize and Define: Work on an agile team to prioritize, define, and ensure the success of infrastructure and managed services features for a high-level strategic roadmap. Stakeholder Collaboration: Prioritize what we should build and when by collaborating with stakeholders on product vision and strategy, while taking customer feedback into account. Technical Discussions: Define how infrastructure features will work through close collaboration with engineers in design sessions an

redisawskubernetes
View job →
PE
Private Employer
📍 New York• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Substrate is the team responsible for Palantir’s core production infrastructure — 100s of K8s clusters — from on-prem to the major cloud hyperscalers, whether they are internet-connected or air-gapped, small hardware footprint or large. As a Senior Software Engineer on Substrate, you will design and build Palantir’s managed Kubernetes product offerings across all these environments. You and your team will be responsible for bootstrapping and operating the entire fleet of K8s clusters with zero manual steps by building industry leading tooling and contributing to core CNCF components. You will also be responsible for ensuring scale, stability and security across a matrix of compliance regimes and hosting infrastructure types. Your team culture emphasizes engineering rigor and operational excellence at scale. This means issues in production should be pre-empted and deeply root-caused, and investments in automation and self-healing systems are key. If you’re excited about infrastructure at scale and working with Kubernetes, this is the right role for you.

kubernetesaigo
View job →
PE
Private Employer
📍 United Kingdom• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Substrate is the team responsible for Palantir’s core production infrastructure — 100s of K8s clusters — from on-prem to the major cloud hyperscalers, whether they are internet-connected or air-gapped, small hardware footprint or large. As a Senior Software Engineer on Substrate, you will design and build Palantir’s managed Kubernetes product offerings across all these environments. You and your team will be responsible for bootstrapping and operating the entire fleet of K8s clusters with zero manual steps by building industry leading tooling and contributing to core CNCF components. You will also be responsible for ensuring scale, stability and security across a matrix of compliance regimes and hosting infrastructure types. Your team culture emphasizes engineering rigor and operational excellence at scale. This means issues in production should be pre-empted and deeply root-caused, and investments in automation and self-healing systems are key. If you’re excited about infrastructure at scale and working with Kubernetes, this is the right role for you.

kubernetesaigo
View job →
PE
Private Employer
📍 Seattle• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Substrate is the team responsible for Palantir’s core production infrastructure — 100s of K8s clusters — from on-prem to the major cloud hyperscalers, whether they are internet-connected or air-gapped, small hardware footprint or large. As a Senior Software Engineer on Substrate, you will design and build Palantir’s managed Kubernetes product offerings across all these environments. You and your team will be responsible for bootstrapping and operating the entire fleet of K8s clusters with zero manual steps by building industry leading tooling and contributing to core CNCF components. You will also be responsible for ensuring scale, stability and security across a matrix of compliance regimes and hosting infrastructure types. Your team culture emphasizes engineering rigor and operational excellence at scale. This means issues in production should be pre-empted and deeply root-caused, and investments in automation and self-healing systems are key. If you’re excited about infrastructure at scale and working with Kubernetes, this is the right role for you.

kubernetesaigo
View job →
PE
Private Employer
📍 Washington• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Substrate is the team responsible for Palantir’s core production infrastructure — 100s of K8s clusters — from on-prem to the major cloud hyperscalers, whether they are internet-connected or air-gapped, small hardware footprint or large. As a Senior Software Engineer on Substrate, you will design and build Palantir’s managed Kubernetes product offerings across all these environments. You and your team will be responsible for bootstrapping and operating the entire fleet of K8s clusters with zero manual steps by building industry leading tooling and contributing to core CNCF components. You will also be responsible for ensuring scale, stability and security across a matrix of compliance regimes and hosting infrastructure types. Your team culture emphasizes engineering rigor and operational excellence at scale. This means issues in production should be pre-empted and deeply root-caused, and investments in automation and self-healing systems are key. If you’re excited about infrastructure at scale and working with Kubernetes, this is the right role for you.

kubernetesaigo
View job →
🔔

Get new senior cloud operations engineer jobs by email

Daily job updates · Unsubscribe anytime