Jobiba hiring network

Cloud Operations System Administrator Jobs

2,329 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

B
11 days ago

This is where your work makes a difference. At Baxter, we believe every person—regardless of who they are or where they are from—deserves a chance to live a healthy life. It was our founding belief in 1931 and continues to be our guiding principle. We are redefining healthcare delivery to make a greater impact today, tomorrow, and beyond. Our Baxter colleagues are united by our Mission to Save and Sustain Lives. Together, our community is driven by a culture of courage, trust, and collaboration. Every individual is empowered to take ownership and make a meaningful impact. We strive for efficient and effective operations, and we hold each other accountable for delivering exceptional results. Here, you will find more than just a job—you will find purpose and pride. Your Role at Baxter Provides enterprise-level technical leadership for cloud shared services and related connected-care ecosystems. Collaborates with engineering, product, and business leaders to shape long-term architectural strategy, establish technical standards, and guide the evolution of secure, cloud-native services that support multiple products, regions, and business domains. Serves as a senior technical leader for distributed systems, GraphQL and API architecture, multi-region cloud strategy, service interoperability, scalability, security, resiliency, observability, SOC 2 readiness, and operational excellence. Partners closely with executive leadership, product management, cybersecurity, quality, regulatory, operations, and engineering teams to align technology investments, architectural decisions, and platform capabilities with business objectives and sustained growth. <span style="color:

node.jsazurekubernetes
View job →

Join Delphi - Where Innovation meets transformation At Delphi, we believe in creating an environment where our people thrive. Our hybrid work model empowers you to choose where you work—whether it's from the office, your home, or a mix of both—so you can prioritize what matters most. We are committed to supporting your personal goals, family, and overall well-being while driving transformative results for our clients. We welcome exceptional talent from anywhere across the globe. Interviews and onboarding are conducted virtually, reflecting our digital-first mindset. Rooted in the region, we specialize in delivering tailored, impactful solutions in Data, Advanced Analytics and AI, Infrastructure, Cloud Security, and Application Modernization. Whether it’s enabling predictive analytics , transforming operations with automation, or driving customer engagement with intelligent platforms, we are the trusted partner for organizations ready to embrace a smarter, more efficient future. We are seeking a versatile and accomplished Senior Full Stack Developer who thrives at the intersection of elegant user experiences and powerful backend systems. This role is for a technical polyglot who can architect complete solutions—from React-powered interfaces to AI-driven Python backends—and lead teams in building cohesive, production-ready applications that deliver exceptional value to our clients. The ideal candidate is equally comfortable crafting performant Next.js applications and designing scalable FastAPI services. You'll be the technical bridge between frontend and backend teams, ensuring seamless integration, consistent best practices, and end-to-end delivery excellence across our consulting engagements. Architect and deliver full-stack solutions using React/Next.js on the frontend and Python/FastAPI on the backend Build responsive, accessible React applications that consume and visualize data from Python-based APIs and services Develop robust

javascripttypescriptpython
View job →
DC
16 days ago

Join Delphi - Where Innovation meets transformation At Delphi, we believe in creating an environment where our people thrive. Our hybrid work model empowers you to choose where you work—whether it's from the office, your home, or a mix of both—so you can prioritize what matters most. We are committed to supporting your personal goals, family, and overall well-being while driving transformative results for our clients. We welcome exceptional talent from anywhere across the globe. Interviews and onboarding are conducted virtually, reflecting our digital-first mindset. Rooted in the region, we specialize in delivering tailored, impactful solutions in Data, Advanced Analytics and AI, Infrastructure, Cloud Security, and Application Modernization. Whether it’s enabling predictive analytics , transforming operations with automation, or driving customer engagement with intelligent platforms, we are the trusted partner for organizations ready to embrace a smarter, more efficient future. We are looking for a hands-on Senior QA Consultant to lead end-to-end testing of AI and Generative AI applications across enterprise environments. This role combines strong expertise in traditional QA engineering with modern AI evaluation and validation practices. The ideal candidate will drive quality assurance initiatives for RAG pipelines, multi-agent systems, OCR and Speech-to-Text solutions while ensuring production grade quality outcomes in regulated industries such as Finance, Healthcare, and Insurance. The successful candidate will lead a small QA pod, collaborate closely with engineering, AI/ML, product, and business teams, and act as the client-facing QA owner for enterprise AI engagements. This role requires strong technical leadership, automation expertise, AI evaluation capabilities, and excellent stakeholder management skills. Experience Requirements • 10–15 years of experie

pythonjavasql
View job →
DC
16 days ago

Join Delphi - Where Innovation meets transformation At Delphi, we believe in creating an environment where our people thrive. Our hybrid work model empowers you to choose where you work—whether it's from the office, your home, or a mix of both—so you can prioritize what matters most. We are committed to supporting your personal goals, family, and overall well-being while driving transformative results for our clients. We welcome exceptional talent from anywhere across the globe. Interviews and onboarding are conducted virtually, reflecting our digital-first mindset. Rooted in the region, we specialize in delivering tailored, impactful solutions in Data, Advanced Analytics and AI, Infrastructure, Cloud Security, and Application Modernization. Whether it’s enabling predictive analytics , transforming operations with automation, or driving customer engagement with intelligent platforms, we are the trusted partner for organizations ready to embrace a smarter, more efficient future. We are looking for a hands-on Senior QA Consultant to lead end-to-end testing of AI and Generative AI applications across enterprise environments. This role combines strong expertise in traditional QA engineering with modern AI evaluation and validation practices. The ideal candidate will drive quality assurance initiatives for RAG pipelines, multi-agent systems, OCR and Speech-to-Text solutions while ensuring production grade quality outcomes in regulated industries such as Finance, Healthcare, and Insurance. The successful candidate will lead a small QA pod, collaborate closely with engineering, AI/ML, product, and business teams, and act as the client-facing QA owner for enterprise AI engagements. This role requires strong technical leadership, automation expertise, AI evaluation capabilities, and excellent stakeholder management skills. Experience Requirements • 7–10 years of experien

pythonjavasql
View job →
A
16 days ago

Who We Are Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. Addepar unifies portfolio, market and client data in a total portfolio view and delivers AI-powered insights within investment and client workflows. More than 1,400 firms in nearly 60 countries use Addepar to manage and advise on nearly $9 trillion in assets. Its open platform integrates with nearly 650 software, data and consulting partners to power end-to-end investment operations across firms of all sizes and complexity. Addepar supports clients worldwide with offices in New York City, Salt Lake City, London, Edinburgh, Pune, Dubai, Geneva, Singapore and São Paulo. The Role We are currently seeking a Staff Software Engineer, Infrastructure to join the AI Platform team that powers seamless insights and interaction through natural language and data intelligence across our AI products. As a Staff Software Engineer, you’ll architect, build, and operate the backend and platform systems that power AI Platform. You’ll work across service design, distributed systems, cloud infrastructure, event-driven processing, observability, CI/CD, and production reliability, helping shape the technical direction of a platform that supports scalable, client-facing AI experiences. This role requires a strong software engineering foundation combined with deep infrastructure and systems thinking. We are looking for an engineer who can write high-quality production code, make sound architectural tradeoffs, and own platform capabilities end-to-end — not someone focused only on scripting, cloud configuration, or infrastructure tooling in isolation. You will collaborate closely with frontend, product, and AI/ML engineers to deliver reliable, secure, and scalable systems that align with Addepar’s standards of performance, resilience, and trust. Applicants must have legal authorization to work in the country where this role is based o

pythonjavasql
View job →
P
Point72
📍 Bengaluru• Full-time
16 days ago

JOB TITLE Senior Storage Engineer A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO The Senior Storage Engineer will be responsible for monitoring, implementation, and management of the organization's File/Block storage infrastructure. This role requires to cover one of the 2 possible shifts in India as well as a deep understanding of File/Block storage technologies and cloud-based solutions, with a focus on ensuring data availability, scalability, and security. The ideal candidate will have extensive experience in storage engineering and a proven ability to manage complex storage environments. • Provide advanced troubleshooting and support for complex storage issues, minimizing downtime and ensuring seamless operations. • Oversee and optimize File/Block storage systems, ensuring high availability and performance. • Conduct capacity planning and forecasting to ensure adequate storage resources are available to meet future demands. • Monitor storage performance and work with US resource to find recommendations for improvements to enhance efficiency and reduce costs. • Implement integration and automation solutions and provide feedback for solution to US Team to streamline operations and improve storage management. • Implement robust security measures to protect data integrity and ensure compliance with industry regulations and standards. • Maintain comprehensive documentation of storage configurations and processes and provide training and guidance to junior team members. WHAT’S

pythonawsazure
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng

pythonsqlpostgresql
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
29 days ago

About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal and VMs dependable & scalable foundations for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Software Engineer, Host Assurance to build and operate the services, APIs, and host software that establish and maintain trust in our compute infrastructure. You will own production software from design and implementation through testing, rollout, observability, and operation. Your work will support capabilities such as machine identity, certificate issuance and enrollment, secure bootstrap, and host attestation across bare-metal and VM environments. Success in this role requires strong technical judgment, the ability to reason across software and host-system boundaries and learn unfamiliar parts of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—inc

REMOTEawsrestai
View job →
M
Mongodb
📍 Gurugram• Full-time
1mo ago

We are seeking an Engineering Manager to join our growing Gurugram Product & Technology team to provide technical direction, direct architecture, and implement core parts of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As an Engineering Manager on this new team, you will be responsible for leading and growing an engineering team, taking on challenging, high-visibility projects that improve and enhance the performance, scalability, and reliability of the distributed systems infrastructure for this new product. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Position Expectations Provide technical leadership and mentorship to a team of engineers, fostering a culture of innovation, quality, and continuous improvement Drive the architectural vision for the platform, ensuring it runs equally well on public clouds, private cloud environments and on-premise Work with product managers, program managers, design & analytics teams and other teams to define, prioritize and deliver new features that delight our users and drive platform improvements Take responsibility for the planning and execution of major features, raise delivery risks Own the monitoring, operations, and maintenance of the systems your team develops Enable the team to operate efficiently by removing technical obstacles, coordinating with other teams on dependencies, and prioritizing the team's overall well-being Contribute to planning for organizational growth, including allocation of engineering resources, participate in hiring and assignment of projects Qualifications 8+ years of experience of building distributed systems, and/or foundatio

pythonjavamongodb
View job →
O
Okta
📍 San Francisco• Full-time• From $165K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in a global on-call rotation supporting highly available customer-facing systems. Participate in incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with en

pythonsqlpostgresql
View job →
A
Anyscale
📍 Remote• Full-time• Remote
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we're building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Anyscale is looking for a Software Engineer to join the ML Developer Experience (MLDevX) team. MLDevX owns the experience layer of the Anyscale platform: the interfaces through which users and coding agents discover, configure, run, observe, debug, and productionize AI workloads. Every user journey crosses this layer through the CLI, SDKs, APIs, UI, Workspaces, MCP, or the workflows and integrations built on top of them. Together, these form the user’s primary interface into Anyscale, turning distributed computing from a systems problem back into a coding problem. We build the common contracts, tools, control-plane services, and architecture that power these surfaces. You will work across the stack from developer tooling to cloud infrastructure and the Ray runtime. Manage long-running operations and make failures across jobs, tasks, actors, nodes, and GPUs easier to diagnose. The systems you build must scale with the platform, remain predictable through failures, and be intuitive for developers, programmable for applications, and operable by coding agents. This is a high impact individual-contributor role with end-to-end ownership. You will work directly with users and field teams to identify high

REMOTEmachine learningaigo
View job →
S
1mo ago

Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. We're looking for an engineer to own the deployment and operational infrastructure of Multigres, our distributed Postgres platform. You'll be responsible for building and maintaining the Multigres Operator, ensuring reliable cloud deployments, and creating the tooling that powers our Kubernetes-based infrastructure. What You’ll Be Responsible for: Build and maintain the Multigres Operator - Maintain our Go-based Kubernetes operator that orchestrates distributed Postgres deployments Architect cloud deployment infrastructure - Design and implement robust deployment patterns for EKS and other Kubernetes platforms Manage storage and networking layers - Work with CSI drivers, persistent volumes, and cross-cloud networking to ensure data reliability and connectivity Develop deployment tooling - Create internal tools and automation for provisioning, scaling, and managing Multigres clusters Ensure operational excellence - Build monitoring, alerting, and diagnostic capabilities into the deployment layer Collaborate across teams - Work with database engineers, SRE, and product teams to deliver seamless deployment experiences You Might Be a Good Fit If You have: Strong systems programming skills - Proficiency in Go and experience building production-grade operators or controllers Deep Kubernetes expertise - Hands-on experience with Kubernetes internals, custom resources, and cloud-managed Kubernetes services (EKS, GKE, AKS) Database operations knowledge - Understanding of database deployment patterns, backup/restore, replication, and high availability Distributed systems experience - Familiarity with consensus protocols, failure scenarios, and designing for resilience Cloud infrastructure background - Experience with cl

kubernetesrestai
View job →
C
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! The Opportunity Cohere seeks a Chief Information Security Officer who can help shape Cohere’s security strategy & the broader conversation around securing AI at scale. You know how to build trust across organizations and communicate trade-offs clearly in environments where speed, innovation, and security must coexist. You will build the playbook and lead the evolution of Cohere’s global security program across corporate systems, cloud infrastructure, AI development, and enterprise operations. As a visible leader for Cohere’s security vision internally and externally you will navigate risk, represent Cohere in industry discussions, and reinforcing our commitment to building AI responsibly. In this role you will: Define and Scale Cohere’s Security Strategy: Define and execute Cohere’s overarching information security strategy in alignment with business priorities and long-term company growth. Build a Modern Risk, Governance & Compliance Program: Lead Cohere in identifying, assessing, and mitigating security risk across all business functions. Secure AI Systems and Technical Infrastructure: Lead the security architecture st

M
Mongodb
📍 Montreal• Full-time• From C$130K/yr
1mo ago

MongoDB’s mission is to empower innovators to create, transform, and disrupt industries by unleashing the power of software and data. We enable organizations of all sizes to easily build, scale, and run modern applications by helping them modernize legacy workloads, embrace innovation, and unleash AI. Our industry-leading developer data platform, MongoDB Atlas, is the only globally distributed, multi-cloud database and is available in more than 115 regions across AWS, Google Cloud, and Microsoft Azure. Atlas allows customers to build and run applications anywhere—on premises, or across cloud providers. With offices worldwide and over 175,000 new developers signing up to use MongoDB every month, it’s no wonder that leading organizations, like Samsung and Toyota, trust MongoDB to build next-generation, AI-powered applications. We are looking for passionate technologists to join our Pre-Sales organization to ensure that our growth is grounded and guided by strong technical alignment with our platform and the needs of our customers. MongoDB Pre-Sales Solution Architects are responsible for guiding our customers and users to design and build reliable, scalable systems using our data platform. Our team is made up of seasoned technical sales professionals, software architects, entrepreneurs, and developers who take direct responsibility for customer success, including the design of their software, deployment, and operations. You'll work closely with our sales executives, helping customers solve business problems by leveraging our solutions, playing a key role in winning deals and driving the business forward. You'll be a trusted advisor to a wide range of users from startups to the world's largest enterprise IT organizations. This role can be based out of our Toronto office or remotely in the Quebec region. As an ideal candidate, you will have: Ideally 8 to 11 years of related experience in a customer facing role, with 5 to 7 years of experience in pre-sales with ent

pythonjavanode.js
View job →

The Team This role can sit in our NYC HQ on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas platform. As a senior SRE, you will be expected to be able to design & build complex systems, operate with autonomy and act as owner for everything you do. The SRE Atlas team works alongside the various Atlas software engineering teams to provide expertise about running systems at scale, build new tooling and automation and perform essential maintenance of the Atlas fleet. This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions that have the ability to impact our customer’s most crucial workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This role requires engineers to have a customer-first mindset to ensure that everything we do results in a stronger product and a better experience for all Atlas customers. The ideal candidate should Have 5+ years of experience running critical systems at scale Value efficiency in processes and operations, and display a preference for automation over manual processes (“allergic to ops work”) Be familiar with a major cloud provider (AWS, Azure, or GCP) and possess the ability to build and operate systems in a multi-cloud environment A strong understanding of how to run a large scale Linux environment, including low level fundamentals Firm grasp of at least one modern programming language, beyond basic scripting (Go, Ruby, Python) Solid understanding of web and network protocols and standards (HTTP, TLS, DNS, etc) Special Requirements: Be a US Citizen Expectations Participate in the development of a reliable and resilient multi-cloud platform that hosts business critic

pythonmongodbaws
View job →
🔔

Get new cloud operations system administrator jobs by email

Daily job updates · Unsubscribe anytime