Jobiba hiring network

Cloud Operations System Administrator Jobs

2,329 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. Technical Program Managers at OpenAI play a key leadership role in scaling these efforts, partnering deeply with product, engineering, design, and go-to-market teams to bring ambitious ideas to life and ensure clarity and discipline in execution. About the Role We are hiring a Technical Program Manager to support OpenAI's critical AI deployments across strategic cloud partners. This role is designed for a candidate who can operate as an end-to-end owner across internal engineering teams and external partner organizations. This role will drive the technical strategy and execution required to bring OpenAI models and platform capabilities into partner environments responsibly and at scale. The work spans engineering deliverables, shared roadmaps, model launch pipelines, technical integration, launch readiness, and post-launch follow-through. You will work closely with senior leaders across OpenAI engineering, infrastructure, product, safety, security, legal, finance, and go-to-market, as well as technical counterparts at our partners. The job is to turn broad partnership commitments into concrete execution plans, align both sides on what must land, and build repeatable mechanisms for launching OpenAI capabilities on third-party platforms. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end execution for major cloud partner programs spanning model deployment, product integration, operational readiness, launch follow-through, and partner-platform adoption. Own integrated technical roadma

awsrestai
View job →
N
13 days ago

NVIDIA is transforming how the world uses AI, cloud, and accelerated computing, and trust is at the center of that mission. Our Attestation and Trust Services team builds the secure cloud services that show customers their NVIDIA platforms are healthy, resilient, and ready for their most important workloads. In this role, you help design and run services that sit at the intersection of hardware, security, and large-scale distributed systems. We partner closely with security, silicon, platform, and cloud teams to bring new ideas into reliable production services that people rely on every day. We care about building systems that last, supporting each other, and creating space for learning and experimentation. If you enjoy solving complex problems, keeping services running smoothly, and collaborating with teammates from many disciplines, we would love to talk with you! What you’ll be doing: Your main focus will be on building and managing our core attestation cloud services. Day-to-day responsibilities include crafting APIs and integrations, boosting reliability, and working alongside NVIDIA teams to convert hardware trust mechanisms and standards into production-ready solutions. You will contribute significantly to shaping how customers verify that NVIDIA platforms are secure and prepared for their workloads. Crafting and evolving attestation cloud services, APIs, and SDK/CLI integration points that confirm the integrity of NVIDIA platforms across data center, AI, networking, and partner environments. Improving reliability and operational maturity through SLOs/SLIs, alerting, runbooks, incident response, and safe rollout practices. Crafting resilient service behavior that handles dependency failures, caching challenges, regional issues, customer-side resilience needs, and graceful degradation. Architecting trust-material distribution for certificate status, re

javaawsazure
View job →
R
Ramp
📍 United States• Full-time• From $10K/yr
1mo ago

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role Technical Consultants are on the frontlines working to establish partnerships with Ramp customers, and act as a liaison between customers and our internal technical teams. You'll be partnering with Customer Activation & Account Management teams to identify, prioritize and build financial solutions for new and existing customers so that Ramp continues to scale with their needs. You will act as the technical and financial subject matter expert for our customers and design solutions that address the customer's needs through Ramp and partner integrations. You will partner with Post-Sales teams to drive Product, Design and Engineering teams' roadmaps to evolve Ramp to better serve our customers at scale and offer increasingly higher value. You will define the customer's financial & technical success criteria from onboarding through activation and expansion. You will represent Product, Design, and Engineering externally and will operate with deep conviction and understanding of Ramp's technical capabilities. The role requires techni

restaigo
View job →
R
Ramp
📍 United States• Full-time• From $10K/yr
1mo ago

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role Technical Consultants are on the frontlines working to establish partnerships with Ramp customers, and act as a liaison between customers and our internal technical teams. You'll be partnering with Customer Activation & Account Management teams to identify, prioritize and build financial solutions for new and existing customers so that Ramp continues to scale with their needs. You will act as the technical and financial subject matter expert for our customers and design solutions that address the customer's needs through Ramp and partner integrations. You will partner with Post-Sales teams to drive Product, Design and Engineering teams' roadmaps to evolve Ramp to better serve our customers at scale and offer increasingly higher value. You will define the customer's financial & technical success criteria from onboarding through activation and expansion. You will represent Product, Design, and Engineering externally and will operate with deep conviction and understanding of Ramp's technical capabilities. The role requires techni

restaigo
View job →
J
Jamf
📍 Australia - Remote• Full-time• Remote
1mo ago

At Jamf, we believe in an open, flexible culture based on respect and trust. Our track record and thriving work environment all stem from the freedom we grant ourselves to get the job done right. We take pride in helping tens of thousands of customers around the globe succeed with Apple. The secret to our success lies in our connectivity, while operating with a high degree of flexibility. Work-life balance remains our priority while feeling connected is important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you’ll do at Jamf: At Jamf, we empower people to be their best selves and do their best work. As a Site Reliability Engineer II, you’ll help us balance development velocity with the reliability our customers depend on. You’ll work with engineering teams to implement how their services are measured, investigate production issues across the stack, and turn what you learn into automation, tooling, and documentation that improves reliability. You’ll use agentic development tools as part of your everyday practice, delegating well-bounded tasks, verifying the output, and contributing to the shared context that makes AI effective for the whole team. This is a hands-on individual contributor role at the intersection of Engineering, Product, Customer Success and Technical Support, where you’ll take ownership of part of your team’s systems, deliver reliability improvements end-to-end within your team and grow toward broader influence. This role if offered as remote in Australia. You may be required to work periodically at a Jamf office or collaborative work location with other Jamf employees in your area for certain events or moments that matter. We are only able to accept applications for those based in Australia and unrestricted work rights in Australia. At this time, we are unable to offer employer sponsorship or support employer-administered work authorisation for this position. # LI -Remote What you can expect to d

REMOTEpythonjavaaws
View job →

Senior Machine Learning Engineer Description - We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability. Key Responsibilities MLOps Platform and Architecture Design and implement a scalable MLOps platform using AWS and Databricks. Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models. Build standardized workflows that move models from development and validation into staging and production. Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure. Establish clear separation between development, testing, staging, and production environments. Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives. CI/CD and

pythonawsazure
View job →
M
Mongodb
📍 New York City; United States• Full-time• From $126K/yr
1mo ago

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. We are looking for a talented Senior Software Engineer to join our team as we execute on a multi-year roadmap and prepare to launch and rapidly scale services that handle petabytes of data. Come do some of the most interesting work of your career as we Think Big and Go Far for our customers! This role can be based out of our New York City office (hybrid working model) or remotely in the North America region. What you’ll do Design, build, and operate control plane services powering an elastic and multi-tenant storage layer for thousands of database instances. Solve problems around maintaining high availability and performance during load spikes, hardware failures, cloud provider outages, and other disruptions. Contribute to a culture of operational excellence through dashboards, playbooks, and on-call improvements. Lead complex technical projects from planning through successful deployment with clear stakeholder updates. Partner closely with peers across database, cloud, and infrastructure engineering teams as well as project management to investigate incidents and develop long-term roadmaps. Mentor junior engineers and foster a collaborative team environment. We’re looking for someone with 5+ years of professional software development experience building, deploying, and operating multi-tenant cloud services with a focus on operational excellence. Experience with large backend/compiled codebases, such as Rust or C/C++. Experience with containerization and orchestration platforms (e.g. Kubernetes). Experience with observability tooling (e.g. time series metrics, dashboards). Experience with distributed systems

mongodbawsazure
View job →
CI
Couchbase, Inc.
📍 Bengaluru• Full-time
16 days ago

Couchbase, the operational data platform for AI, empowers businesses to succeed by bringing data to life in new ways. Major market-leading companies rely on Couchbase for mission critical operational, analytical, mobile and AI workloads. Built to replace legacy infrastructure and fragmented data services, Couchbase empowers enterprises with a unified platform architected for performance, flexibility and global scale. With Couchbase, organizations bring their data to life, launching game‑changing customer experiences, exploring the limitless potential of AI, and seamlessly extending applications from the cloud to the edge and beyond. Couchbase’s AI‑ready technology and enterprise partnership model eliminate complexity and reduce total cost of ownership, enabling teams to stay agile, innovative and secure. Couchbase believes data should never slow you down, but act as the foundation for your next breakthrough. Discover why Couchbase is trusted to help the world’s biggest players scale, move fast and stay resilient, no matter what’s next on their roadmap. Visit couchbase.com and follow us on LinkedIn and X. Want to be part of our story? Apply today! AI Platform Engineering Location: Bangalore (Hybrid - in office at least 3 days/week) About the Role We are seeking an experienced and visionary technology leader to lead the development and scaling of our Operational AI platform capabilities. This role will own the strategy, architecture, delivery, and operational excellence of Couchbase AI Cloud Platform capabilities. This is a strategic leadership role at the intersection of distributed systems, cloud-native platforms, and AI. You will lead a large, multi-layered engineering organization responsible for delivering The Operational Data Platform for AI, while partnering closely with Product, Design, SRE, and Go-To-Market teams. Your leadership will directly influence company growth, customer adoption, platform reliability, and Couchbase’s competitive pos

pythonjavaaws
View job →
O
1mo ago

About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun

awsazuregcp
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical. The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly. About the Role We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity. In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity. This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions. This role is based in San Francisco, CA, with travel as needed. Key Responsibilities Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint Drive readiness to convert contracted compute capacity into schedulable production clusters Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives Build integrated schedules spanning procurement, logistics, installation, st

awsazurerest
View job →

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary The Executive Director, Digital Engineering- Aetna Member Care and Journey Services is a senior technology leader responsible for setting the technical vision, architectural direction, and engineering execution for member centric services. This role leads large-scale engineering teams that build high-performance backend APIs, microservices, and cloud-native systems that power member experiences across digital, agent, provider, and partner channels. The leader ensures exceptional service stability, resiliency, innovation velocity, and alignment with enterprise user experience and operational goals. Key Responsibilities 1. Backend API & Microservices Engineering Leadership • Lead the design, development, and delivery of scalable backend systems, APIs, and microservices powering member-facing capabilities. • Define API contract standards, and integration patterns used across Member Services platforms. • Drive service modernization by adopting cloud‑native architectures, containerization, and event-driven patterns. 2. Service Stability, Observability & Resiliency • Establish standards for availability, resiliency, performance, and disaster recovery across all services. • Implement SLO/SLI/error budget frameworks, health checks, and high‑availability architectures. <p

D
Datadog
📍 New York
12 days ago

About Datadog: Datadog is the essential monitoring and security platform for cloud applications. We bring together end-to-end traces, metrics, and logs to make your applications, infrastructure, and third-party services entirely observable. These capabilities help businesses secure their systems, avoid downtime, and ensure customers are getting the best user experience. The Team: The Financial Planning & Analysis (FP&A) team analyzes company financial data (revenue, customers, headcount, expenses, etc.) in order to support the business’ growth and success. Within FP&A, the R&D Finance team enables the financial strategy behind Datadog’s Engineering and Product organizations. We partner directly with technical leadership to drive decision-making around our most critical investments. Your work will be highly cross-functional and play a pivotal role in connecting the dots across the organization through a financial lens, ensuring operational alignment and informing decision making. The Opportunity: As a part of the R&D Finance team, this person will be key in supporting our product and engineering leadership team. Reporting to the Senior Manager, you will help create our annual budget and financial targets and work with operational leaders to support execution against our goals. Your work will be highly cross-functional and strategic, and you will play a pivotal role in connecting the dots across the organization through a financial lens, in order to help ensure operational alignment and inform decision making. What You’ll Do: Partner with senior business leaders to manage their departmental budgets on a regular basis, including but not limited to the company’s co-founder and CTO Manage financial forecasts and analytics, which include data across revenue, customers, product, company expenses, and headcount / workforce. Help manage significant portions of the company’s spend, including but not limited to AI and GPU spend across the compan

aiExcelTableau
View job →
CI
16 days ago

Couchbase, the operational data platform for AI, empowers businesses to succeed by bringing data to life in new ways. Major market-leading companies rely on Couchbase for mission critical operational, analytical, mobile and AI workloads. Built to replace legacy infrastructure and fragmented data services, Couchbase empowers enterprises with a unified platform architected for performance, flexibility and global scale. With Couchbase, organizations bring their data to life, launching game‑changing customer experiences, exploring the limitless potential of AI, and seamlessly extending applications from the cloud to the edge and beyond. Couchbase’s AI‑ready technology and enterprise partnership model eliminate complexity and reduce total cost of ownership, enabling teams to stay agile, innovative and secure. Couchbase believes data should never slow you down, but act as the foundation for your next breakthrough. Discover why Couchbase is trusted to help the world’s biggest players scale, move fast and stay resilient, no matter what’s next on their roadmap. Visit couchbase.com and follow us on LinkedIn and X. Want to be part of our story? Apply today! Job Description - Technical Support Engineer (Rotational Split Shift: Sat - Wed) We are looking for a Technical Support Engineer to assist our rapidly growing customer base. As part of our Technical Support team you will be the primary point of contact for Couchbase customers for all technical issues. NoSQL databases are the answer to the demand for high speed, highly available, extremely dynamic data storage and Couchbase is at the forefront of this technology. Working in Couchbase Technical Support you’ll acquire highly coveted skills, essential to this technology, by navigating the world of NoSQL databases. This includes working with Golang, learning about the principals and concepts of distributed systems, understanding what goes into good NoSQL database design, mobile data convergence, and becoming proficie

sqldockerkubernetes
View job →
DC
Diligent Corporation
📍 New York• Full-time• From $131K/yr
16 days ago

Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production

pythonawsazure
View job →
S
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We're hiring Software Engineers for our Data Platform team to build and evolve Snowflake's real-time stream processing and data transformation platform. If you've built or researched high-throughput streaming systems or scalable transformation engines - in industry or graduate work - we want to talk. AS A SOFTWARE ENGINEER, DATA PLATFORM AT SNOWFLAKE, YOU WILL: Design and implement low-latency stream processing and in-flight data transformation systems at global scale Own correctness and performance of streaming execution (watermarks, exactly-once delivery, out-of-order handling) Build transformation primitives and operators that run reliably at cloud-scale throughput Contribute to architectural decisions for our next-generation streaming and transformation platform Write production-quality systems code and collaborate across product and engineering teams OUR IDEAL CANDIDATE WILL HAVE: BS/MS/PhD in Computer Science or related field — graduate research in streaming, distributed systems, or query/transformation engines is a strong differentiator Deep distributed systems fundamentals: fault tolerance, consistency, state management Hands-on or experience with stream processing or transformation systems (Flink, Kafka Streams, Spark Structured Streaming, or similar) Proficiency i

pythonjavasql
View job →
🔔

Get new cloud operations system administrator jobs by email

Daily job updates · Unsubscribe anytime