Jobiba hiring network

Cloud Operations System Administrator Jobs

2,329 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Role Description As a Software Engineer on the Metadata team, you’ll build and operate the large-scale distributed databases that every Dropbox service depends on. Metadata systems are mission-critical, in the live path for all user operations and must meet stringent requirements for latency, durability, and transactional consistency. You’ll design and evolve the core infrastructure that manages Dropbox’s databases at scale, enabling fast, reliable access to data for millions of users and hundreds of internal services. This work spans distributed systems, replication, caching, and transactional database systems. You’ll collaborate closely with engineers across Infrastructure and Product teams to ensure the metadata layer meets business needs and continues to scale with Dropbox’s growth. This is an opportunity to leverage your expertise in distributed systems and grow into broader technical leadership. Our Engineering Career Framework is viewable by anyone outside the company and describes what’s expected for our engineers at each of our career levels. Check out our blog post on this topic and more here . Responsibilities Design and maintain distributed database systems providing low-latency, strongly consistent data access Implement and optimize replication, consensus, and caching mechanisms to meet availability and performance goals Operate production systems, including participating in the on-call rotation, ensuring high availability and data durability Collaborate with infrastructure and product teams to assess current and future use cases and requirements, supporting the development of a mid- to long-term roadmap that reflects these needs Contribute to system design reviews, postmortems, and reliability improvements Write high-quality, efficient code in Go and Rust for performance-critical systems On-call work may be necessary occasionally to help address bugs, outages, or other operational issues, with the goal of maintaining a stable

REMOTEpythonjavaai
View job →
D
1mo ago

Role Description As a Software Engineer on the Metadata team, you’ll build and operate the large-scale distributed databases that every Dropbox service depends on. Metadata systems are mission-critical, in the live path for all user operations and must meet stringent requirements for latency, durability, and transactional consistency. You’ll design and evolve the core infrastructure that manages Dropbox’s databases at scale, enabling fast, reliable access to data for millions of users and hundreds of internal services. This work spans distributed systems, replication, caching, and transactional database systems. You’ll collaborate closely with engineers across Infrastructure and Product teams to ensure the metadata layer meets business needs and continues to scale with Dropbox’s growth. This is an opportunity to leverage your expertise in distributed systems and grow into broader technical leadership. Our Engineering Career Framework is viewable by anyone outside the company and describes what’s expected for our engineers at each of our career levels. Check out our blog post on this topic and more here . Responsibilities Design and maintain distributed database systems providing low-latency, strongly consistent data access Implement and optimize replication, consensus, and caching mechanisms to meet availability and performance goals Operate production systems, including participating in the on-call rotation, ensuring high availability and data durability Collaborate with infrastructure and product teams to assess current and future use cases and requirements, supporting the development of a mid- to long-term roadmap that reflects these needs Contribute to system design reviews, postmortems, and reliability improvements Write high-quality, efficient code in Go and Rust for performance-critical systems On-call work may be necessary occasionally to help address bugs, outages, or other operational issues, with the goal of maintaining a stable and high-quality experienc

REMOTEsqlmysqlredis
View job →
S
Stripe
📍 South San Francisco• Full-time• $274.5K – $334.6K/yr
1mo ago

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Partner with the Tax team stakeholders to design and implement features of our financial systems that include Oracle Cloud Applications (Financials, TRCS, Fusion Tax), Stripe homegrown Tax Solutions and many other tools. Identify opportunities to automate Tax processes, work on transformation initiatives, implement SOX controls and maintain our financial Applications. Lead Finance systems (primarily Tax area including but not limited to Direct and Indirect Taxes) projects and strategic initiatives - including implementation of new tools and business process changes - from requirements, tool selection, system design, implementation, UAT, training, go-live & post go-live support. Serve as a technology liaison with Tax stakeholders to understand their pain points, opportunities for process improvements, automation & efficiency gains, create strategic roadmap and provide them with technology solutions. Demonstrate an understanding of increasingly complex sales tax, transfer pricing and Tax provisioning concepts and effectively apply knowledge to develop and maintain tax systems and data flows. Design solutions and make architecture decisions by benchmarking with industry best practices to help us scale for the future. Partner with Engineering product teams to develop & deliver technical solutions to optimize financial workflow and integration with other Stripe internal sy

I
Instacart
📍 United States - Remote• Full-time• Remote• From $132K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview About the Role In this role as a Customer Experience Systems Sr. Associate, you'll be the architect behind streamlined workflows, ensuring our customer relationship tools are not just functional but truly resonate with the needs of our teams. You'll have the reins, managing everything from the nitty-gritty of technical requirements to the bird's-eye view of our system's efficiency, especially within Salesforce Service Cloud. Melding your knack for collaboration with your Salesforce savvy, you'll shape how cases flow, how teams communicate, and how each user journey tells a story of success, all while aligning with our broader business strategies. About the Team The Customer Experience team is the empathetic voice of Instacart, engaging directly through real-time calls and chats with our valued customers, shopper

REMOTEaigosalesforce
View job →
D
Discord
📍 San Francisco Bay Area• Full-time• $192K – $216K/yr
1mo ago

Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. We are seeking an experienced Oracle ERP Fusion Technical Developer to join Discord’s Business Systems team. In this role, you will own technical incident resolution, report development, integrations, system enhancements, and technical delivery across our Oracle Fusion ERP environment. You will partner closely with Finance, Accounting, Business Systems, and integration teams to ensure operational stability, scalable solutions, and successful delivery of business-critical initiatives. What you'll be doing Provide L2 (Incident and Problem Management) and L3 (Bug Fix and Enhancement) technical support across the Oracle Fusion ERP environment. Develop, maintain, and enhance BI Publisher (BIP) and OTBI reports, dashboards, and data extracts across GL, AP, AR and Procurement modules Own Oracle Integration Cloud (OIC) break-fix support, monitoring, troubleshooting, and new integration development. Partner with functional analysts and business stakeholders to translate requirements into scalable technical solutions. Create and maintain technical design documents, integration specifications, and deployment documentation. Support month-end, quarter-end, and year-end financial operations by resolving technical issues and performance bottlenecks. Execute root cause analysis , technical testing, deployment validation, and regression testing for quarterly Oracle releases. Collaborate with cross-functional teams on enterprise integrations and long-term solution architecture. Follow and support ITGC, SOX, and change management requirements for all technical changes. What you should have 5+ years of hands-on Oracle Fusi

KH
K Health
📍 Tel Aviv• Full-time
15 days ago

About the Role: We are looking for a Senior DevOps Engineer to join our DevOps team at K Health. You will own and evolve the infrastructure underpinning a healthcare AI platform serving patients and enterprise health system partners. This is a high-ownership role: you will architect and operate cloud environments across K Health and its enterprise partners, lead complex infrastructure migrations, drive disaster recovery programs, and help build the next generation of AI-powered operations tooling. You will also mentor junior engineers and collaborate closely with product and engineering teams across the company. This is a hybrid role based in New York City (4 days/week in office) and includes participation in a daytime on-call rotation. What you will do: Own the design, implementation, and evolution of our GKE-based Kubernetes infrastructure across K Health and enterprise partner environments. Build and maintain our Terraform modular infrastructure library, including reusable modules with automated testing, across GCP, Cloudflare, and AWS. Architect, build, and maintain GitLab CI/CD shared pipeline templates used by all engineering teams (build, test, security scanning, deployment). Own and maintain self-hosted infrastructure software running in-cluster, including GitLab, ArgoCD, Langfuse, DependencyTrack, NGINX Ingress, and others. Implement and support security and compliance controls across infrastructure and the software supply chain - secrets management, pipeline secret detection, container scanning, SOC2 and HIPAA. Drive disaster recovery readiness: design failover scenarios, author runbooks, and lead periodic DR tests. Lead development of AI-powered operations tooling and agentic infrastructure. Monitor, troubleshoot, and improve production system reliability; respond to incidents during on-call shifts. Mentor junior DevOps engineers and establish team-wide engineering standards. What we are looking for: 5+ years of experience in DevOps, platform engineering,

pythonsqlpostgresql
View job →
PE
Private Employer
📍 Seattle• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Apollo is Palantir’s autonomous software management and deployment platform. It enables seamless, continuous delivery of mission-critical software (Foundry, Gotham, AIP) across a vast range of environments: on-prem, public cloud, disconnected (air-gapped) networks, and highly regulated settings (including IL-5 and FedRAMP). As a Software Engineer on the Apollo team, you’ll build and operate a large-scale distributed system to allow the remote operation and maintenance of Kubernetes clusters. Our mission is to extract the entire state of a cluster into a portable, high-performance artifact within minutes, enabling full and almost instant cluster reconstruction from the ground up—all while pushing the limits of speed, reliability, and scale. You’ll design and implement backup and restore solutions for Kubernetes, leveraging proprietary compression infrastructure tailored to Palantir’s unique deployment models. You’ll also build and optimize our container artifact store, which is based on the OCI (Open Container Initiative) distribution spec—the industry standard for storing and distributing container images and artifacts. You’ll own the backbone of every environment Apollo supports, from hyperscalers to Army trucks. If you’re excited by challenges at the intersection of container technologies like OCI and docker, storage, and distributed systems, you’ll find opportunities here to dive deep into storage formats and low-level optimizations, where milliseconds matter. As we increasingly automate cluster creation and management on diverse hardware, you’ll play a key role in scaling Palantir’s presence at the edge and solving tough distributed systems proble

dockerkubernetesrest
View job →
PE
Private Employer
📍 New York• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Apollo is Palantir’s autonomous software management and deployment platform. It enables seamless, continuous delivery of mission-critical software (Foundry, Gotham, AIP) across a vast range of environments: on-prem, public cloud, disconnected (air-gapped) networks, and highly regulated settings (including IL-5 and FedRAMP). As a Software Engineer on the Apollo team, you’ll build and operate a large-scale distributed system to allow the remote operation and maintenance of Kubernetes clusters. Our mission is to extract the entire state of a cluster into a portable, high-performance artifact within minutes, enabling full and almost instant cluster reconstruction from the ground up—all while pushing the limits of speed, reliability, and scale. You’ll design and implement backup and restore solutions for Kubernetes, leveraging proprietary compression infrastructure tailored to Palantir’s unique deployment models. You’ll also build and optimize our container artifact store, which is based on the OCI (Open Container Initiative) distribution spec—the industry standard for storing and distributing container images and artifacts. You’ll own the backbone of every environment Apollo supports, from hyperscalers to Army trucks. If you’re excited by challenges at the intersection of container technologies like OCI and docker, storage, and distributed systems, you’ll find opportunities here to dive deep into storage formats and low-level optimizations, where milliseconds matter. As we increasingly automate cluster creation and management on diverse hardware, you’ll play a key role in scaling Palantir’s presence at the edge and solving tough distributed systems proble

dockerkubernetesrest
View job →
M
8 days ago

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. About Our Team Our team builds and enables scalable, cloud-native software platforms that power critical manufacturing and business operations. We use modern full stack engineering practices and AI-driven technologies to create innovative solutions that improve automation, decision making, and operational efficiency. Position Overview We are seeking an AI-Centric Full Stack Software Engineer to design, build, and support modern software solutions with a strong focus on AI-enabled applications and intelligent automation. This role goes beyond simply using AI coding assistants. You will have strong understanding of Prompt Engineering, Vibe Coding, Rework Rate Reduction and leverage Custom Agents, and integrations built around the Model Context Protocol (MCP). You will apply strong software engineering fundamentals to assemble, integrate, and operationalize AI capabilities into real production systems being accountable for Full-Stack solutions. The goal of this role is to significantly shorten development and feedback cycles by automating routine engineering work, augmenting human decision-making, and embedding intelligence directly into our development workflows and the applications we deliver! Responsibilities Design, develop, test, deploy, and maintain scalable full stack applications and platform services. Own software solutions end-to-end, including system design, architecture, implementation, testing, observability, deployment, and lifecycle support. Develop an

angularsqlkubernetes
View job →
M
8 days ago

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. About Our Team Our team builds and enables scalable, cloud-native software platforms that power critical manufacturing and business operations. We use modern full stack engineering practices and AI-driven technologies to create innovative solutions that improve automation, decision making, and operational efficiency. Position Overview We are seeking an AI-Centric Full Stack Software Engineer to design, build, and support modern software solutions with a strong focus on AI-enabled applications and intelligent automation. This role goes beyond simply using AI coding assistants. You will have strong understanding of Prompt Engineering, Vibe Coding, Rework Rate Reduction and leverage Custom Agents, and integrations built around the Model Context Protocol (MCP). You will apply strong software engineering fundamentals to assemble, integrate, and operationalize AI capabilities into real production systems being accountable for Full-Stack solutions. The goal of this role is to significantly shorten development and feedback cycles by automating routine engineering work, augmenting human decision-making, and embedding intelligence directly into our development workflows and the applications we deliver! Responsibilities Design, develop, test, deploy, and maintain scalable full stack applications and platform services. Own software solutions end-to-end, including system design, architecture, implementation, testing, observability, deployment, and lifecycle support. Deve

angularsqlkubernetes
View job →

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We're looking for an Information Security Engineer who can think broadly about how infrastructure is built, attacked, defended, and improved. You may have deep experience in one area—such as endpoint security, Active Directory, Data Loss Prevention, cloud security, identity, or detection and response—but you’re motivated by solving security problems across the environment. As an Information Security Engineer focused on infrastructure security, you'll help protect Palantir's global endpoints, servers, identity systems, applications, networks, and data flows. You'll work across teams to reduce attack surface, improve defensive controls, investigate security events, and turn security findings into durable architectural and operational improvements. This is not just a corporate security role; the adversaries we face are sophisticated. We need someone who understands both the details and the broader system: how a weakness in one part of the environment can create risk somewhere else, how controls interact, and how to build security improvements that hold up in practice.

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We're looking for an Information Security Engineer who can think broadly about how infrastructure is built, attacked, defended, and improved. You may have deep experience in one area—such as endpoint security, Active Directory, Data Loss Prevention, cloud security, identity, or detection and response—but you’re motivated by solving security problems across the environment. As an Information Security Engineer focused on infrastructure security, you'll help protect Palantir's global endpoints, servers, identity systems, applications, networks, and data flows. You'll work across teams to reduce attack surface, improve defensive controls, investigate security events, and turn security findings into durable architectural and operational improvements. This is not just a corporate security role; the adversaries we face are sophisticated. We need someone who understands both the details and the broader system: how a weakness in one part of the environment can create risk somewhere else, how controls interact, and how to build security improvements that hold up in practice.

M
Mongodb
📍 Toronto• Full-time• From C$144K/yr
1mo ago

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams. We are looking to speak to candidates who can work East Coast hours. The ideal candidate should Have 6+ years of experience in software development and operating distributed systems Proficiency in Python, Go, or a similar language Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Expectations Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products This includes, but is not limited to, contributing to open-source projects, or engineering software-based

pythonmongodbaws
View job →
M
Mongodb
📍 Boston; Miami; New York City; Pittsburgh; Raleigh; United States• Full-time• From $126K/yr
1mo ago

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of our Boston, New York City, Raleigh, Miami, Pittsburgh or remotely in the United States while physically based in an Eastern or Central time zone location. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure g

pythonmongodbaws
View job →
M
1mo ago

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams. We are looking to speak to candidates who can work East Coast hours. The ideal candidate should Have 6+ years of experience in software development and operating distributed systems Proficiency in Python, Go, or a similar language Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Expectations Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products This includes, but is not limited to, contributing to open-source projects, or engineering software-based

pythonmongodbaws
View job →
🔔

Get new cloud operations system administrator jobs by email

Daily job updates · Unsubscribe anytime