Jobiba hiring network

Cloud Operations Engineer Jobs

2,329 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

BE
15 days ago

About Backblaze Backblaze is the object storage leader in the open cloud movement, fueling customer success with cloud storage built purposefully to unlock budgets, unburden administrators, and unleash innovators. Together with our partners, we’re helping customers break free from the restrictive, overpriced legacy solutions that hold them back, and blaze forward with the full power of the open cloud in their hands. Founded in 2007, we scaled the business with less than $3 million in outside funding until 2021, when we did a traditional IPO on the Nasdaq stock exchange. Today, Backblaze generates over $136M ARR and is the leading specialized storage cloud, managing over three billion gigabytes of data storage for 500K+ customers in 175+ countries, including businesses, developers, IT professionals, and individuals. But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us. We’re seeking a Sr. Reliability Engineer ll (DBA) to join our team! About the Role We are seeking a Site Reliability Engineer (SRE) with a DBA (Database Administration) focus to help ensure the stability, scalability, and reliability of our production database systems - primarily Vitess (distributed MySQL) and Cassandra - alongside the rest of our services and infrastructure. This role operates within procedures and runbooks established by our senior DBA SREs, and focuses on building automation, maintaining observability, and supporting incident response to keep customer-facing systems performing at their best. The SRE will collaborate with engineering, product, and operations teams to embed reliability practices into day-to-day development and operations while contributing to tools and processes that improve efficiency and reduce manual effort Key Responsibilities Database Administration Operating and maintaining high-availability database systems — primarily Vitess (distributed MySQL) and Cassandra — against established architecture and runbooks. Op

REMOTEpythonsqlmysql
View job →
P
Pagerduty
📍 Atlanta• Full-time• From $98K/yr
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office, you'll help build and operate the foundational infrastructure that powers PagerDuty's real-time digital operations platform. Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You'll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases. Key Responsibilities ● Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems. ● Contribute to the reliability and scalability of PagerDuty's core platform by hardening existing systems and supporting the rollout of new infrastructure

pythonawsazure
View job →
P
Pagerduty
📍 Atlanta• Full-time• $113K – $171.6K/yr
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. As a Site Reliability Engineer II on the Core Infrastructure team in our Atlanta office, you'll help build and operate the foundational infrastructure that powers PagerDuty's real-time digital operations platform. Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You'll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases. Key Responsibilities ● Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems. ● Contribute to the reliability and scalability of PagerDuty's core platform by hardening existing systems and supporting the rollout of new infrastructur

pythonawsazure
View job →
P
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. PagerDuty is looking for a Machine Learning Engineer who is passionate about collaborating with data scientists, product managers and engineers alike. As part of our team, you will help us accelerate the development and extension of products powered by Gen AI and many other shapes of Machine Learning. You’ll be contributing hands-on to the development of the services and pipelines that enable multiple ML/AI features in our product. You will have the opportunity to collaborate with multiple organizations, taking input and guidance from your senior stakeholders and helping bring our initiatives to reality. You’ll succeed by showcasing excellent capacity to manage time, demonstrating emotional intelligence as you navigate stakeholder relationships, and by continuously improving your technical skill set. Key Responsibilities Build and improve the capabilities that enable and accelerate the production of machine learning (ML) and generative AI (genAI) based solutions Partner with data scientists, effectively sharing engineering context and collaborating to support larger initiatives Incorporate the best ava

pythonawskubernetes
View job →
P
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. PagerDuty is looking for a Machine Learning Engineer who is passionate about collaborating with data scientists, product managers and engineers alike. As part of our team, you will help us accelerate the development and extension of products powered by Gen AI and many other shapes of Machine Learning. You’ll be contributing hands-on to the development of the services and pipelines that enable multiple ML/AI features in our product. You will have the opportunity to collaborate with multiple organizations, taking input and guidance from your senior stakeholders and helping bring our initiatives to reality. You’ll succeed by showcasing excellent capacity to manage time, demonstrating emotional intelligence as you navigate stakeholder relationships, and by continuously improving your technical skill set. Key Responsibilities Build and improve the capabilities that enable and accelerate the production of machine learning (ML) and generative AI (genAI) based solutions Partner with data scientists, effectively sharing engineering context and collaborating to support larger initiatives Incorporate the best ava

pythonawskubernetes
View job →
N
Nvidia
📍 Remote, United States• Remote
12 days ago

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data

REMOTEpythonsqlaws
View job →
BE
Backblaze External Website
📍 Argentina• Full-time• Remote
15 days ago

Backblaze is the object storage leader in the open cloud movement, fueling customer success with cloud storage built purposefully to unlock budgets, unburden administrators, and unleash innovators. Together with our partners, we’re helping customers break free from the restrictive, overpriced legacy solutions that hold them back, and blaze forward with the full power of the open cloud in their hands. Founded in 2007, we scaled the business with less than $3 million in outside funding until 2021, when we did a traditional IPO on the Nasdaq stock exchange. Today, Backblaze generates over $100m in revenue and is the leading specialized storage cloud - managing over three billion gigabytes of data storage for 500K+ customers in 175+ countries, including businesses, developers, IT professionals, and individuals. But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us. We are seeking an AI Workflow Engineer ! About the Role: Backblaze is running a company-wide AI transformation program, and this role is where strategy becomes something you can actually use. As AI Workflow Engineer, you will be the hands-on builder inside the AI Program Office - embedded alongside the Director, AI Enablement, and deployed into non-GTM, non-Engineering business teams to prototype, build, and ship AI-powered workflow solutions that make a measurable difference in how work gets done. This is not a research role or a strategy role. You will spend your time building: prompt architectures, automation flows, lightweight agents, API integrations, and demos that help a business team see - concretely - what AI can do for them. You will move across functions, which means you need to understand how businesses actually operate. The ideal background is someone who started in a business-facing role - finance, operations, marketing, customer success, HR - and developed serious hands-on AI and automation skills on top of that foundation. You will work closely with

REMOTEgitrestai
View job →
N
1mo ago

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys

pythonjavakubernetes
View job →
M
8 days ago

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. About Our Team Our team builds and enables scalable, cloud-native software platforms that power critical manufacturing and business operations. We use modern full stack engineering practices and AI-driven technologies to create innovative solutions that improve automation, decision making, and operational efficiency. Position Overview We are seeking an AI-Centric Full Stack Software Engineer to design, build, and support modern software solutions with a strong focus on AI-enabled applications and intelligent automation. This role goes beyond simply using AI coding assistants. You will have strong understanding of Prompt Engineering, Vibe Coding, Rework Rate Reduction and leverage Custom Agents, and integrations built around the Model Context Protocol (MCP). You will apply strong software engineering fundamentals to assemble, integrate, and operationalize AI capabilities into real production systems being accountable for Full-Stack solutions. The goal of this role is to significantly shorten development and feedback cycles by automating routine engineering work, augmenting human decision-making, and embedding intelligence directly into our development workflows and the applications we deliver! Responsibilities Design, develop, test, deploy, and maintain scalable full stack applications and platform services. Own software solutions end-to-end, including system design, architecture, implementation, testing, observability, deployment, and lifecycle support. Develop an

angularsqlkubernetes
View job →
M
8 days ago

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. About Our Team Our team builds and enables scalable, cloud-native software platforms that power critical manufacturing and business operations. We use modern full stack engineering practices and AI-driven technologies to create innovative solutions that improve automation, decision making, and operational efficiency. Position Overview We are seeking an AI-Centric Full Stack Software Engineer to design, build, and support modern software solutions with a strong focus on AI-enabled applications and intelligent automation. This role goes beyond simply using AI coding assistants. You will have strong understanding of Prompt Engineering, Vibe Coding, Rework Rate Reduction and leverage Custom Agents, and integrations built around the Model Context Protocol (MCP). You will apply strong software engineering fundamentals to assemble, integrate, and operationalize AI capabilities into real production systems being accountable for Full-Stack solutions. The goal of this role is to significantly shorten development and feedback cycles by automating routine engineering work, augmenting human decision-making, and embedding intelligence directly into our development workflows and the applications we deliver! Responsibilities Design, develop, test, deploy, and maintain scalable full stack applications and platform services. Own software solutions end-to-end, including system design, architecture, implementation, testing, observability, deployment, and lifecycle support. Deve

angularsqlkubernetes
View job →

Job Details: Job Description: The Role As a Cloud Application Development Engineer within Intel Manufacturing Foundry Cloud Services (imFCS) , you will design, develop, deploy, and support cloud-native applications that power semiconductor manufacturing, engineering automation, and AI-driven factory operations. You will build scalable, secure, and resilient platforms that improve engineering productivity, enable advanced analytics, and accelerate Intel Foundry's digital transformation. Key Responsibilities Develop and maintain cloud-native applications and services supporting manufacturing and engineering workflows. Support 24x7 manufacturing operations through on-call rotations, incident response, and root-cause analysis. Design and implement scalable microservices, APIs, and containerization systems utilizing Kubernetes orchestrator and cloud native applications. Architect secure cloud solutions spanning application, data, networking, identity, and observability domains. Build and maintain CI/CD pipelines, Infrastructure-as-Code, and DevOps automation. Provide technical leadership for contractor and partner development teams across multiple geographic regions, driving architecture, implementation, and operational excellence for cloud-native solutions. Collaborate with manufacturing, engineering, and software teams to deliv

pythondockerkubernetes
View job →
A
Asana
📍 Vancouver• Full-time
1mo ago

Asana is seeking an experienced IT Systems Engineer (Identity) to join our growing IT Team and support the critical systems that keep IT moving forward. The ideal candidate will be focused on increasing IT efficiency through automation, and improving the end-user experience. Based in our Vancouver office, you will be primarily responsible for the support and implementation of our internal and cloud based IT identity infrastructure and applications. As an Identity Engineer, you will support and implement our internal and cloud-based IT identity infrastructure and applications to increase efficiency through automation. You will collaborate with cross-functional teams to improve the end-user experience while maintaining a crisp and detail-oriented focus on our identity systems. By leveraging your passion for technology and automation, you will play a critical role in scaling Asana's global IT operations. This role is based in our Vancouver office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Drive efficiency by implementing automated SSO on-boarding and off-boarding processes. Empower the organization through expert administration and support of Google Workspace, Slack, and Asana. Design and deliver security policies in partnership with the Security team to protect our global user base. Scale IT services by identifying and executing on high-impact automation opportunities. Standardize and automate routine IT tasks to improve overall system reliability. Resolve complex technical escalations with a focus on building long-term preventative solutions. Foster a robust IT environment by collaborating with Support and Security teams

pythonrestai
View job →
AG
1mo ago

To lead the Generative AI engineering initiatives at Adani AI Labs, focusing on developing intelligent agents and sovereign cloud AI solutions. This role drives the architectural design, implementation, and scaling of Generative AI applications that optimize operations across Adani’s industrial sectors, aligning with the vision to deliver secure, scalable, and intelligent digital transformation. The Engineering Manager will oversee a team of backend engineers while remaining actively involved in technical execution. This is an exciting opportunity to lead mission-critical product engineering initiatives and build state-of-the-art AI-driven systems within a dynamic and innovative environment at Adani. Source: Adani Group | Job ID: 44640

pythongitai
View job →
PE
Private Employer
📍 United Kingdom• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role The Palantir platform is deployed in numerous critical mission environments including combat zones and classified networks—from the back of a Humvee to a command post to the cloud. This means operating in multiple cloud environments, on-prem air-gapped networks, and at the edge—at scale. We are looking for Edge Infrastructure Engineers to build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure. This role demands a deep focus on low-level systems, including the deployment and management of physical bare metal servers in both traditional data centers and edge environments. You will be responsible for physical network engineering and the development of robust infrastructure that ensures performance of the Palantir platform. In addition to ensuring performance and reliability, you will play a critical role in building and scaling new environments in a forward-deployed capacity, including onsite. Edge Infrastructure Engineers combine hardware-level engineering experience with the drive to improve existing systems and the creativity to develop novel solutions for evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate for and participate in sensible, scalable systems design, sharing responsibility for diagnosing, resolving, and preventing production issues across our most demanding deployments.

PE
Private Employer
📍 New York• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Substrate is the team responsible for Palantir’s core production infrastructure — 100s of K8s clusters — from on-prem to the major cloud hyperscalers, whether they are internet-connected or air-gapped, small hardware footprint or large. As a Senior Software Engineer on Substrate, you will design and build Palantir’s managed Kubernetes product offerings across all these environments. You and your team will be responsible for bootstrapping and operating the entire fleet of K8s clusters with zero manual steps by building industry leading tooling and contributing to core CNCF components. You will also be responsible for ensuring scale, stability and security across a matrix of compliance regimes and hosting infrastructure types. Your team culture emphasizes engineering rigor and operational excellence at scale. This means issues in production should be pre-empted and deeply root-caused, and investments in automation and self-healing systems are key. If you’re excited about infrastructure at scale and working with Kubernetes, this is the right role for you.

kubernetesaigo
View job →
🔔

Get new cloud operations engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More cloud operations engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.