Jobiba hiring network

Cloud Operations Engineer Jobs

2,329 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in Cork for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workflows Pref

mongodbawsazure
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities, such as reducing costs and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, global traffic routing, and resilient cloud infrastructure problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new network topologies, multi-cloud ecosystems, and agentic engineering models. Position Overview: The Staff Software Engineer - Infrastructure will play a foundational role in architecting, evolving, and securing Okta's global core network fabric and multi-cloud platform layers. This position focuses on building highly resilient, edge management in AWS and GCP cloud, executing enterprise-wide cloud migrations (AWS to GCP), enforcing absolute Zero-Trust primitives (mTLS / TLS 1.3), and integrating cutting-edge AI automation layers (Agentic SRE) into our day-to-day operations. As a Staff Engineer on this team, you will act as a key technical anchor in India, working closely with global counterparts to maintain Okta's high-availability SLAs while protecting the platform against active global DDoS attacks. Key Responsibilities: Global Ingress & Routerless Evolution (CFSaaS): Design and engineer Next-Gen traffic

pythonreactaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside our cloud partners, infrastructure providers, and internal engineering teams, we operate hyperscale AI campuses that support the training and deployment of frontier AI models. The Site Operations team serves as OpenAI's on-site operational presence, helping ensure campuses operate safely, efficiently, and in alignment with Industrial Compute standards. We work closely with Hardware Operations, Infrastructure Delivery, Network Operations, Security, Facilities, Construction, and our infrastructure partners to support day-to-day site execution and maintain operational readiness. As Industrial Compute continues to expand globally, Site Operations plays a critical role in ensuring each campus is prepared to support reliable AI infrastructure at scale. About the Role We are seeking a Site Operations Technician to support the daily operation of Industrial Compute campuses. This role acts as OpenAI's on-site technical representative, helping coordinate activities across hardware operations, facilities, construction, logistics, security, and external service providers. You will perform routine site inspections, support asset tracking, coordinate vendor activities, assist with operational readiness, document site conditions, and help ensure infrastructure issues are identified and resolved quickly. The ideal candidate enjoys working in highly technical environments, is detail-oriented, and thrives in fast-paced operational settings where no two days are the same. Key Responsibilities Perform routine walkthroughs of Industrial Compute facilities to verify operational readiness and identify potential issues. Monitor site conditions and report abnormalities involving hardware spaces, network rooms, utilities, logistics areas, and common infrastructure. Support coordination of vendors, contractors, and partner organizations performing work o

awsrestai
View job →
IT
15 days ago

About the Role We are seeking an experienced Azure DevOps Engineer to design, implement, and maintain CI/CD pipelines, cloud infrastructure, and automation solutions on Microsoft Azure. This role bridges development and operations, ensuring reliable, secure, and scalable delivery of applications and infrastructure. Location: Hyderabad-India-Onsite Duration: Fulltime Responsibilities Design, build, and maintain CI/CD pipelines using Azure DevOps (Pipelines, Repos, Artifacts, Boards) Architect and manage Azure cloud infrastructure using Infrastructure as Code (ARM templates, Bicep, or Terraform) Automate build, test, and deployment processes across multiple environments Implement and manage containerization and orchestration (Docker, Azure Kubernetes Service) Monitor system performance, availability, and security using Azure Monitor, Log Analytics, and Application Insights Collaborate with development, QA, and security teams to streamline release management Implement Azure security best practices, identity management (Azure AD/Entra ID), and network architecture Manage cost optimization and governance across Azure subscriptions Troubleshoot production issues and support incident response Document infrastructure, pipelines, and operational procedures Required Qualifications Microsoft Certified: Azure Solutions Architect Expert, Azure Administrator Associate (required) 3+ years of hands-on experience with Azure DevOps and Azure cloud services Strong experience with Infrastructure as Code (Bicep, ARM templates, or Terraform) Proficiency scripting in PowerShell, Bash, or Python Experience with Git version control and branching strategies Solid understanding of networking, security, and identity concepts in Azure Experience with containerization (Docker) and orchestration (Kubernetes/AKS) Familiarity with monitoring and logging tools (Azure Monitor, Application Insights) Preferred Qualifications Additional certifications: Azure DevOps Engineer Expert Experience with multi-

pythonazuredocker
View job →

Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. This position is based at our office in Chennai, India. Appian was built on a culture of in-person collaboration, which we believe is a key driver of our mission to be the best. You will be the product manager working closely with the team whose mission is to strengthen and optimize site infrastructure by delivering essential upgrades, resource efficiency, and scalable solutions. You will be responsible for the direction and roadmap of a component of the Appian Cloud data plane that ensures reliable, high-performance operations for all Appian Cloud customer sites. This role is specifically focused on the cloud-native persistence and messaging layer. You will oversee the backend sub-systems—including technologies like S3 and Redis—that power the platform's internal data plane and core services. This component of the software is not directly user-facing but has strong implications on the scalability and reliability requirements our customers expect. What you will be doing: Prioritize and Define: Work on an agile team to prioritize, define, and ensure the success of infrastructure and managed services features for a high-level strategic roadmap. Stakeholder Collaboration: Prioritize what we should build and when by collaborating with stakeholders on product vision and strategy, while taking customer feedback into account. Technical Discussions: Define how infrastructure features will work through close collaboration with engineers in design sessions an

redisawskubernetes
View job →
KH
K Health
📍 Tel Aviv• Full-time
15 days ago

About the Role: We are looking for a Senior DevOps Engineer to join our DevOps team at K Health. You will own and evolve the infrastructure underpinning a healthcare AI platform serving patients and enterprise health system partners. This is a high-ownership role: you will architect and operate cloud environments across K Health and its enterprise partners, lead complex infrastructure migrations, drive disaster recovery programs, and help build the next generation of AI-powered operations tooling. You will also mentor junior engineers and collaborate closely with product and engineering teams across the company. This is a hybrid role based in New York City (4 days/week in office) and includes participation in a daytime on-call rotation. What you will do: Own the design, implementation, and evolution of our GKE-based Kubernetes infrastructure across K Health and enterprise partner environments. Build and maintain our Terraform modular infrastructure library, including reusable modules with automated testing, across GCP, Cloudflare, and AWS. Architect, build, and maintain GitLab CI/CD shared pipeline templates used by all engineering teams (build, test, security scanning, deployment). Own and maintain self-hosted infrastructure software running in-cluster, including GitLab, ArgoCD, Langfuse, DependencyTrack, NGINX Ingress, and others. Implement and support security and compliance controls across infrastructure and the software supply chain - secrets management, pipeline secret detection, container scanning, SOC2 and HIPAA. Drive disaster recovery readiness: design failover scenarios, author runbooks, and lead periodic DR tests. Lead development of AI-powered operations tooling and agentic infrastructure. Monitor, troubleshoot, and improve production system reliability; respond to incidents during on-call shifts. Mentor junior DevOps engineers and establish team-wide engineering standards. What we are looking for: 5+ years of experience in DevOps, platform engineering,

pythonsqlpostgresql
View job →

Cloud Platform Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations is currently seeking a Cloud Platform Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking a skilled platform engineer to help build and operate the cloud platform services that host critical enterprise applications and software toolchains. In this role, the selected candidate will focus on the shared platform capabilities that enable teams to deploy, run, and maintain containerized and cloud-hosted solutions in a consistent and supportable manner. As both an individual contributor and technical leader, this position will help define and implement platform standards for Kubernetes, container hosting, deployment automation, configuration management, and operational support. This role is focused on platform reliability, repeatability, scalability, and service enablement, rather than custom application software development. Position Responsibilities: Design, implement, and maintain cloud platform services supporting Kubernetes, containers, ingress, storage integration, secrets management, and service connectivity Build and sustain reusable deployment patterns for Commercial-Off-The-Shelf (COTS), Open Source Software (OSS), and internally customized applications Develop and maintain automation for platform provisioning, upgrades, patching, and lifecycle support Manage cluster lifecycle activities including: Cluster upgrades Node management <

awsazuredocker
View job →

Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations organization is currently seeking a Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking an experienced cloud infrastructure professional to help design, build, and sustain the foundational cloud environment supporting critical program needs. In this role, the selected candidate will help establish and operate secure, scalable, and resilient cloud infrastructure environments in Microsoft Azure to enable enterprise applications, software toolchains, and digital engineering workloads. As both an individual contributor and technical leader, this position will work across network, computer, storage, identity, security, and automation domains to deliver repeatable cloud infrastructure patterns and operational excellence. This role is focused on infrastructure operations, sustainment, automation, and reliability, rather than application software development. Position Responsibilities: Design, implement, and maintain Microsoft Azure-based infrastructure solutions including networking, compute, storage, identity integration, and supporting services Develop and maintain Infrastructure as Code (IaC) and configuration automation solutions using Terraform, Ansible, PowerShell, and Bash Implement cloud policies to enforce security, ensure regulatory compliance, and manage user access Build repeatable landing zones and cloud infrastructure patterns that support mul

azureterraformansible
View job →
PE
Private Employer
📍 Seattle• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Mission Manager is Palantir’s PaaS for enabling US Government customers and vendors to run software securely and compliantly in the most sensitive environments, but without the overhead — whether connected, disconnected, cloud, or edge. Built on the strength of Palantir’s Apollo platform, it provides the critical infrastructure needed to rapidly onboard and deploy applications into a secure Kubernetes-based ecosystem, freeing our customers to focus on building and powering mission-critical systems. The Mission Manager offering is still in its earliest days, and by joining us now, you’ll define the strategy for how we develop and scale it — witnessing firsthand the impact of your work on critical missions and the new capabilities you unlock. You’ll drive this by building elegant, robust APIs powered by Kubernetes controllers, bridging the gap between a raw Kubernetes cluster and a fully featured, infrastructure-agnostic runtime that can meet the operational demands of hundreds of specialized microservices. You’ll undertake this challenge alongside an energized team with a wide array of backgrounds and skillsets, all united by an ambitious vision for what’s possible.

kubernetesmicroservicesai
View job →
PE
Private Employer
📍 Washington• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Mission Manager is Palantir’s PaaS for enabling US Government customers and vendors to run software securely and compliantly in the most sensitive environments, but without the overhead — whether connected, disconnected, cloud, or edge. Built on the strength of Palantir’s Apollo platform, it provides the critical infrastructure needed to rapidly onboard and deploy applications into a secure Kubernetes-based ecosystem, freeing our customers to focus on building and powering mission-critical systems. The Mission Manager offering is still in its earliest days, and by joining us now, you’ll define the strategy for how we develop and scale it — witnessing firsthand the impact of your work on critical missions and the new capabilities you unlock. You’ll drive this by building elegant, robust APIs powered by Kubernetes controllers, bridging the gap between a raw Kubernetes cluster and a fully featured, infrastructure-agnostic runtime that can meet the operational demands of hundreds of specialized microservices. You’ll undertake this challenge alongside an energized team with a wide array of backgrounds and skillsets, all united by an ambitious vision for what’s possible.

kubernetesmicroservicesai
View job →
PE
Private Employer
📍 New York• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Mission Manager is Palantir’s PaaS for enabling US Government customers and vendors to run software securely and compliantly in the most sensitive environments, but without the overhead — whether connected, disconnected, cloud, or edge. Built on the strength of Palantir’s Apollo platform, it provides the critical infrastructure needed to rapidly onboard and deploy applications into a secure Kubernetes-based ecosystem, freeing our customers to focus on building and powering mission-critical systems. The Mission Manager offering is still in its earliest days, and by joining us now, you’ll define the strategy for how we develop and scale it — witnessing firsthand the impact of your work on critical missions and the new capabilities you unlock. You’ll drive this by building elegant, robust APIs powered by Kubernetes controllers, bridging the gap between a raw Kubernetes cluster and a fully featured, infrastructure-agnostic runtime that can meet the operational demands of hundreds of specialized microservices. You’ll undertake this challenge alongside an energized team with a wide array of backgrounds and skillsets, all united by an ambitious vision for what’s possible.

kubernetesmicroservicesai
View job →
PE
Private Employer
📍 Washington• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Palantir is at the forefront of some of the most critical and challenging problems in the world. We develop alongside our customers everyday. Our customers span from the cloud to the frontline. As we adapt to solve their most pressing issues in latency, performance, and compute cost, we are building a team of software engineers relentlessly focused on low-level optimization and novel compute architectures. This is a team of developers creating software for the far-edge, including streaming ETL pipelines, inference platforms, and various timing critical applications. This role requires an experienced software engineer who is well versed in low-level development in compiled, native languages such as Rust and C/C++. A successful candidate can optimize software for constrained embedded devices or across large-scale distributed systems. You should have strong knowledge of computer architecture and OS internals.

aic++rust
View job →
PE
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We are seeking a Senior Software Engineer to join a customer-facing product engineering team focused on developing advanced Command and Control (C2) and Agentic Autonomy Software for Autonomous Systems for use in operational and tactical missions. This role involves building, integrating, and deploying state-of-the-art software solutions that combine sensors, actuators, unmanned vehicles, and intelligent decision-making systems including self-hosted and commercial LLMs to enable users to employ and manage robotic and autonomous systems across a variety of mission sets. You will work alongside other experts in sensing, artificial intelligence, physics and simulations, user experience, and edge system engineering to build and deliver solutions that redefine complex military and civilian mission scenarios in the real world. As a part of this team, you will contribute directly to the development of cloud and edge software to command and control autonomous systems, including interfacing with onboard sensors (radar, cameras, RF, and other modalities) and real-time kinetic and non-kinetic systems for individual and swarms of systems. You will work alongside other experts in sensing, artificial intelligence, and systems engineering to shape solutions that redefine complex area defense scenarios.

PE
Private Employer
📍 Seattle• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Apollo is Palantir’s autonomous software management and deployment platform. It enables seamless, continuous delivery of mission-critical software (Foundry, Gotham, AIP) across a vast range of environments: on-prem, public cloud, disconnected (air-gapped) networks, and highly regulated settings (including IL-5 and FedRAMP). As a Software Engineer on the Apollo team, you’ll build and operate a large-scale distributed system to allow the remote operation and maintenance of Kubernetes clusters. Our mission is to extract the entire state of a cluster into a portable, high-performance artifact within minutes, enabling full and almost instant cluster reconstruction from the ground up—all while pushing the limits of speed, reliability, and scale. You’ll design and implement backup and restore solutions for Kubernetes, leveraging proprietary compression infrastructure tailored to Palantir’s unique deployment models. You’ll also build and optimize our container artifact store, which is based on the OCI (Open Container Initiative) distribution spec—the industry standard for storing and distributing container images and artifacts. You’ll own the backbone of every environment Apollo supports, from hyperscalers to Army trucks. If you’re excited by challenges at the intersection of container technologies like OCI and docker, storage, and distributed systems, you’ll find opportunities here to dive deep into storage formats and low-level optimizations, where milliseconds matter. As we increasingly automate cluster creation and management on diverse hardware, you’ll play a key role in scaling Palantir’s presence at the edge and solving tough distributed systems proble

dockerkubernetesrest
View job →
PE
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We are seeking a Senior Software Engineer to join a customer-facing product engineering team focused on developing advanced Command and Control (C2) and Agentic Autonomy Software for Autonomous Systems for use in operational and tactical missions. This role involves building, integrating, and deploying state-of-the-art software solutions that combine sensors, actuators, unmanned vehicles, and intelligent decision-making systems including self-hosted and commercial LLMs to enable users to employ and manage robotic and autonomous systems across a variety of mission sets. You will work alongside other experts in sensing, artificial intelligence, physics and simulations, user experience, and edge system engineering to build and deliver solutions that redefine complex military and civilian mission scenarios in the real world. As a part of this team, you will contribute directly to the development of cloud and edge software to command and control autonomous systems, including interfacing with onboard sensors (radar, cameras, RF, and other modalities) and real-time kinetic and non-kinetic systems for individual and swarms of systems. You will work alongside other experts in sensing, artificial intelligence, and systems engineering to shape solutions that redefine complex area defense scenarios.

🔔

Get new cloud operations engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More cloud operations engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.