Jobiba hiring network

Cloud Operations System Administrator Jobs

2,329 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About The Team.... Global Compute runs Optimised Hosting, GoDaddy's global platform for all customer hosting products. Squad R is the engineering team responsible for operating, scaling, and continuously improving the OpenStack-based clouds that power that platform. We treat reliability as an engineering problem: we automate toil away, we plan capacity ahead of demand, and we instrument everything so that we understand our systems before they surprise us. As an SRE III on the team, you'll be a senior technical contributor who others lean on for the hard problems. What you'll get to do... Operate and scale GoDaddy's cloud infrastructure, including our OpenStack-based hosting platform. You'll troubleshoot and improve services spanning compute, networking, and storage in large-scale production environments. Drive the OpenStack migration. Help move customer hosting workloads onto the platform safely — designing and executing migration tooling, validation, and rollback strategies that protect customer experience. Work within a large-scale global hosting environment supporting thousands of servers and customer workloads across multiple regions. Eliminate toil through automation. Build and maintain automation in Python and Puppet to replace manual operational work. Treat repeated manual effort as a bug to be fixed. Strengthen observability. Improve monitoring, alerting, and dashboards so that signal reaches the right engineer at the right time, and so that we can reason about system behavior from data. Participate in on-call and incident response. T

pythondockerlinux
View job →

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet's expansion into Asia-Pacific government markets depends on mastering two critical compliance frameworks: Australia's Information Security Registered Assessors Program (IRAP) and Japan's Information System Security Management and Assessment Program (ISMAP). Both are essential gatekeepers for federal/national government adoption in their respective markets. We're seeking a Sr. Security Engineer II to own both programs end-to-end—managing IRAP assessments and continuous assurance in Australia, and ISMAP registration and compliance in Japan. This is a specialized, high-impact role for someone with deep expertise in both frameworks and comfort operating within government compliance ecosystems across two distinct cultures and regulatory environments. You'll be the authority on APAC government security standards at Smartsheet, the trusted interface with assessors and government agencies, and the architect of compliance processes that keep us ahead of evolving requirements. This role is critical for capturing high-value government business across the Asia-Pacific region. You will: Own IRAP (Australia) program strategy and execution: Lead the overall roadmap for obtaining and maintaining IRAP authorizations, including assessment coordination, remediation, and authorization maintenance. Own ISMAP (Japan) program strategy and execution: Lead registration and compliance for Japan's government cloud certification program, managing the certification assessment process and maintaining registry status. Coordinate with region

REMOTEawsazuregcp
View job →
M
Mongodb
📍 Sydney• Full-time
1mo ago

The Storage Layer Services team is currently re-architecting the MongoDB Cloud Storage Layer. This is a relatively new team in MongoDB that sits at the heart of the next generation MongoDB Cloud Storage Architecture, and the team is working to build performant multi-tenant distributed storage services both to enhance our existing MongoDB cloud storage architecture and to power more of our customers' use cases more efficiently. Engineering at MongoDB is globally distributed, with a mix of folks being fully remote, hybrid, or in-office. We have a small but growing team that calls Sydney home, and we are looking for a Staff Engineer to join the team working closely with other teams in Sydney and North America. Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies great engineering fundamentals to deliver core features of a popular database, join us! Let’s change what’s possible for application developers, system architects, and database operators. We are looking to speak to candidates who are based in Sydney for our hybrid working model. You’re an ideal candidate if: You have 10+ years of experience in programming, debugging, and performance tuning highly concurrent and/or distributed systems. Especially if you have worked in a systems language (C, C++, Rust, etc) for a number of those years You have a track record as an effective technical leader. You love helping teams be successful at solving vaguely defined problems in iterative and measurable ways. You put the customer first, and don’t hesitate to cross team boundaries in search of the right solution You have a solid grasp of related systems fundamentals, such as cache management, log-based recovery, transactions or performance profiling You’re comfortable reasoning about highly concurrent, asynchronous services — backpressure, tail latency, and the failure modes of replicated state machines You’ve worked on large, highly availabl

mongodbawsazure
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
16 days ago

Opportunity Overview: We’re looking for a Manager, Platform Engineering that can lead and grow a high-performing engineering team focused on Developer Experience, DevOps, SRE, and Quality. You will own the systems and processes that enable teams to build, test, release, and operate software with high velocity and reliability, driving engineering efficiency and operational excellence across the organization. What you’ll do: Lead a fast-paced, autonomous team of engineers focused on platform engineering, developer experience, DevOps, SRE, and quality engineering Own and drive the internal developer platform strategy and roadmap, improving how engineering teams build, test, deploy, and operate services Create transparency into engineering efficiency and system health through meaningful metrics across delivery, reliability, and quality Enable teams to move faster by improving CI CD pipelines, environments, tooling, and overall developer workflows Provide technical leadership across platform, infrastructure, and reliability, helping teams build scalable and resilient systems Ensure strong engineering practices across release processes, testing, quality, reliability, and security Define and enforce release guardrails, validation standards, and rollback mechanisms to improve production safety Improve environment stability and consistency across development, QA, and pre production environments Drive test strategy and automation maturity to improve overall product quality and confidence in releases Define and implement observability, monitoring, and alerting standards across systems Improve incident detection, response, and RCA practices, ensuring learnings translate into platform and system improvements Drive cloud infrastructure best practices across AWS, containers, and infrastructure as code Foster a culture of ownership, reliability, and continuous improvement within the team Provide innovative solutions for attracting, developing, and retaining top engineering talent I

S
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. About Snowflake Founded by industry experts and backed by strategic investors, our disruptive built-for-the-cloud architecture was designed to push the limitations of conventional data warehousing. Our teams breed ambition, challenge ordinary thinking, push the pace of innovation, and work hard to provide our customers with the best experience possible. As an Enterprise Product Manager, your key role will be to optimize and strategically evolve Workday Finance Order to Cash (OTC) IT Systems at Snowflake. You will bridge diverse business stakeholders and skilled development teams, translating complex business requirements into precise, actionable functional specifications for technical implementations. This role offers a unique opportunity to contribute to the entire solution lifecycle for critical OTC systems, from design to deployment. Your expertise will directly impact the efficiency, scalability, and accuracy of Snowflake's revenue-generating processes, ensuring a robust OTC ecosystem. What you will be doing As a Staff Enterprise Product Manager, you will independently contribute to and lead the following areas: Process Optimization: Optimize and strategically evolve Workday Finance Order to Cash (OTC) systems, bridging business needs with technical solutions. Operation

sqlagilescrum
View job →
EI
16 days ago

About the job Senior Backend Engineer |100% Remote We are searching for a seasoned Sr Backend Engineer who will be responsible for developing and maintaining our backend systems, ensuring their efficiency, scalability, and reliability. The ideal candidate has a strong background in PHP and Java, with optional experience in Golang and Node.js. You should be well-versed in working with databases such as MongoDB, Redis, and Postgres, and have a solid understanding of cloud technologies, specifically AWS. Responsibilities Design, develop, and maintain backend systems and APIs to support our application's functionality. Collaborate with cross-functional teams, including front-end developers, product managers, and designers, to deliver high-quality solutions. Write clean, scalable, and well-documented code that adheres to industry best practices and coding standards. Perform code reviews and provide constructive feedback to peers to ensure code quality and consistency. Optimize and improve the performance of existing backend systems. Troubleshoot and debug production issues, providing timely resolutions. Stay up-to-date with emerging technologies and industry trends, identifying opportunities for innovation and improvement. Collaborate with DevOps teams to ensure smooth deployment and operation of backend services in the AWS cloud environment. Requirements Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent work experience). 5-10 years of professional experience as a Backend Engineer. Strong proficiency in Java/Golang is a must. Experience with Node js is a plus. In-depth knowledge of database technologies, including MongoDB, Redis, and Postgres. Solid understanding of cloud computing platforms, particularly AWS. Familiarity with containerization technologies such as Docker and orchestration tools like Kubernetes. Proficiency in writing efficient and optimized SQL queries. Experience with version control systems, such as Git. Excellent pr

javanode.jssql
View job →
S
16 days ago

SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . Position Summary The Senior Software Engineer will manage, support, and customize SonicWall's Oracle Agile PLM platform and support its modernization toward Oracle Fusion Cloud PLM. Beyond core PLM, this role owns EDI and system-to-system integrations between Oracle and other enterprise systems — ERP, CRM, the data platform, and external/partner systems — and contributes to SonicWall's agentic AI initiatives by building AI-assisted automations within the PLM and integration landscape. This is a hands-on, full-stack engineering role with a high degree of autonomy, operating within our established delivery, change, and security processes. The ideal candidate can operate and extend the PLM platform, deliver robust Oracle-to-external integrations, and apply modern AI tooling to reduce manual effort and accelerate product and transactional data flows across the enterprise. Key Responsibilities PLM platform: Manage, support, and customize Oracle Agile PLM (Product Collaboration, PQM, PG&C), and support the roadmap toward Oracle Fusion Cloud PLM. Deliver product changes: Interface with users to de

pythonjavareact
View job →
G
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. About the Role Our Senior Ecosystem Sales Manager (ESM) will be based in Mumbai or Delhi. This is a senior, market-facing role responsible for building and scaling GitLab's partner business in India across the full breadth of the ecosystem: hyperscalers, large India-centric system integrators (SIs), and born-in-the-cloud / AI-native partners. We are looking for someone with the range and seniority to credibly own strategic partner relationships across all of these partner types, and to help grow our business through these partners. This is a senior individual-contributor role for someone who has already built a career scaling par

REMOTEawsazuregcp
View job →
M
Mongodb
📍 Toronto• Full-time• From C$137K/yr
1mo ago

We are hiring a Senior Software Engineer to join our Server Security team. The Server Security team is a development-focused group within MongoDB's core engineering organization. Operating "close to the bottom of the stack," the team builds features that enable database users to secure their data globally. You will work on critical components including: Cryptography: Queryable Encryption , at-rest data encryption, and fundamental cryptographic principles. Identity & Access: Authentication and authorization systems, TLS, and X.509 certificate management Network Security: High-performance, low-latency networking protocols (PKI, Hashing, CRLs) System Integrity: Resilience, observability, and compliance assurance within a large-scale distributed database Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies distributed systems fundamentals to deliver core features of a popular database, join us! Let’s change what’s possible for application developers, system architects, and database operators. The role As a Senior Engineer, you will apply distributed systems fundamentals to deliver core security features. You will be a leader in improving MongoDB's security posture by owning features and leading investigations into complex areas of the codebase. What you’ll do: Build and test new security features in a large, feature-rich C++ codebase Work across engineering, cloud services, and support teams to coordinate feature rollouts and changes Stand for code quality and security best practices, assisting fellow engineers in writing well-reasoned, secure code Use strong diagnostic intuition to solve thorny technical issues related to distributed systems, concurrency, and OS internals This role can be remote or hybrid anywhere in the USA or Canada. We will prioritize candidates who are already located in one of these countries. Candidate Profile We are looking for a highly technical engineer w

javamongodbaws
View job →
M
Mongodb
📍 New York City• Full-time• From $126K/yr
1mo ago

We are hiring a Senior Software Engineer to join our Server Security team. The Server Security team is a development-focused group within MongoDB's core engineering organization. Operating "close to the bottom of the stack," the team builds features that enable database users to secure their data globally. You will work on critical components including: Cryptography: Queryable Encryption , at-rest data encryption, and fundamental cryptographic principles. Identity & Access: Authentication and authorization systems, TLS, and X.509 certificate management Network Security: High-performance, low-latency networking protocols (PKI, Hashing, CRLs) System Integrity: Resilience, observability, and compliance assurance within a large-scale distributed database Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies distributed systems fundamentals to deliver core features of a popular database, join us! Let’s change what’s possible for application developers, system architects, and database operators. The role As a Senior Engineer, you will apply distributed systems fundamentals to deliver core security features. You will be a leader in improving MongoDB's security posture by owning features and leading investigations into complex areas of the codebase. What you’ll do: Build and test new security features in a large, feature-rich C++ codebase Work across engineering, cloud services, and support teams to coordinate feature rollouts and changes Stand for code quality and security best practices, assisting fellow engineers in writing well-reasoned, secure code Use strong diagnostic intuition to solve thorny technical issues related to distributed systems, concurrency, and OS internals This role can be remote or hybrid anywhere in the USA or Canada. We will prioritize candidates who are already located in one of these countries. Candidate Profile We are looking for a highly technical engineer w

javamongodbaws
View job →
M
Mongodb
📍 New York City• Full-time• From $111K/yr
1mo ago

The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency requests around the globe, and comply with various data sovereignty requirements. The SRE Team’s mission is to build this increasingly complex infrastructure, while continually lowering the operational burden associated with it, and increasing our internal visibility into the health of the system. We are strong believers in infrastructure-as-code and self-healing systems. The SRE Team is fully integrated with all the other engineering teams, and the teams work closely together with a soft and traversable boundary between their areas of responsibility. We are looking to speak to candidates who are based in New York City for our hybrid working model. Responsibilities Design and build the infrastructure for a global cloud service that comprises hundreds of thousands of MongoDB clusters, processes a billion metrics per day, and replicates tens of billions of database writes to our backup service Design, implement, and troubleshoot the automation and monitoring of services that seamlessly spans the globe - including several cloud providers Become an expert in infrastructure performance, helping us optimize from the application level all the way through the firmware Build for resilience. Our goal is that nobody’s pager goes off, ever. Are we there yet? No. Are we really close? Very. While we work on that - participate in a weekly on-call rotation Improve our infrastructure capabilities, optimizing for cost, simplicity, and maintainability Requirements 3+ years of experience running a mission critical service at scale in a Linux environment Firm grasp of at least one modern programming language, beyond basic scripting Familiarity with web and network protocols and standards (HTTP, TLS, DNS, etc) Bachelor’s deg

mongodbawsazure
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system

pythonsqlaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Release Engineer team is responsible for building and maintaining the systems that power software delivery—from CI/CD pipelines and artifact management to release automation and fleet telemetry. We ensure software across bootloaders, firmware, operating systems, and cloud services is built reproducibly, validated rigorously, and released safely at scale. About the Role As a Release Engineer, you’ll design, build, and operate release infrastructure that enables reliable, secure, and traceable software delivery across complex multi-component systems. You’ll partner closely with embedded, cloud, and QA teams to ensure that every build—from development to OTA deployment—is fast, verifiable, and production-ready. We’re looking for engineers who take pride in automation, build reproducibility, and system reliability—and who enjoy building the connective tissue that allows hardware and software to ship together seamlessly. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and operate CI/CD pipelines for multi-component builds (bootloader, firmware, OS images, backend, companion apps) using hermetic toolchains. Define versioning and branching strategies; automate promotions, changelogs, and artifact retention. Integrate unit, integration, and hardware-in-the-loop (HIL) test results; quarantine flaky tests, auto-bisect failures, and block unsafe promotions. Build A/B OTA update flows with verity and health checks; run staged rollouts and canaries; implement safe rollback and roll-forward strategies. Implement code signing for binaries and firmware, generate SBOMs, run vulnerability scanning, and attach build attestations and provenance. Manage dashboards and alerts for build health, promotion latency, failure rates, and fleet update telemetry. You might thrive in this role if you: Have experience building and operating buil

pythonawsci/cd
View job →
O
1mo ago

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches

awsazuregcp
View job →

What We Do At GoGuardian, we’re helping build a future where all learners are ready and inspired to solve the world’s greatest challenges. Our award-winning system of learning solutions is purpose-built for K-12 and trusted by school leaders to promote effective teaching and equitable engagement while helping empower educators to keep students safe. What It’s Like to Work at GoGuardian We are an outcomes-focused learning company with a steadfast focus on improving learning environments, one classroom at a time. Working with us means joining a remote team of diverse, committed, mission-driven employees who are inspired by our vision, dedicated to our customers, and ready to roll up their sleeves. Guardians put their heads together to solve problems, learn together from experiments that fail, and stand together by their work with full accountability. We balance our diligence with an inclusive culture that invites everyone to bring their whole self to work. Join us and learn why “I love the people here” is one of the most frequent comments we hear from Guardians. What We Do At GoGuardian, we’re helping build a future where all learners are ready and inspired to solve the world’s greatest challenges. Our award-winning system of learning solutions is purpose-built for K-12 and trusted by school leaders to promote effective teaching and equitable engagement while helping empower educators to keep students safe. The Role We’re looking for a Senior Site Reliability Engineer (SRE) to help design, scale, and maintain the infrastructure that powers our core products and services. In this role, you’ll collaborate with engineering teams to drive operational excellence, optimize system performance, and ensure high availability across production environments. This position sits on Tech Foundation, a team that manages core cloud infrastructure, shared data services, and developer tooling to empower our product teams to deliver software efficiently and secu

javascripttypescriptpython
View job →
🔔

Get new cloud operations system administrator jobs by email

Daily job updates · Unsubscribe anytime