Jobiba hiring network

Cloud Operations System Administrator Jobs

2,329 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

P
Peloton
📍 New York• Full-time• From $155.3K/yr
1mo ago

About the Role Join Peloton’s Global Network Services team as a Senior Engineer, sitting at the unique intersection of high-scale enterprise networking and high-stakes media production. In this role, you will architect and support the global infrastructure that powers our corporate offices, warehouses, and flagship New York Broadcast Studio. Acting as the key bridge between Global Cloud Engineering and Studio Operations, you will deploy highly resilient, scalable architectures that deliver live streaming content seamlessly to millions of members worldwide. Your Daily Impact Global Architecture & Deployment: Design, optimize, and secure Peloton's global network footprint, integrating Cisco, Meraki, Aruba, and Palo Alto Networks across hybrid on-premises and AWS environments. Studio & Broadcast Reliability: Lead deep-dive traffic analysis and troubleshooting for our NY studio, ensuring 24/7 uptime for live broadcasts, real-time video streaming, and OTT media delivery. Operational Leadership & Lifecycle Management: Manage the end-to-end network project lifecycle—from traffic shaping and SD-WAN optimization to establishing SOPs, security policies (with InfoSec), and handling Tier-3 disaster recovery. You Bring To Peloton Broad & Deep Network Expertise: 8+ years in Network Engineering, including 6+ years in complex SaaS environments and 4+ years designing public cloud networking (specifically AWS). Advanced Protocol & Security Mastery: Expert knowledge of L2/L3 protocols (BGP, OSPF, EIGRP), security protocols (IPsec, 802.1x, RADIUS), SD-WAN, and SDN/SDDC full-stack solutions. Media & Streaming Specialization: 3+ years optimizing IP networks specifically for live video delivery, utilizing multicast/unicast technologies and streaming protocols like HLS, RTMP, and SRT. Education & Elite Certifications: A Bachelor’s degree in Engineering or Computer Science, backed by active CCIE or JNCIE certifications (required). Collaborative Mindset: A curious

redisawsgit
View job →
S
Stripe
📍 New York• Full-time• $156.8K – $235.2K/yr
1mo ago

Who we are About Stripe #LI-DNI Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Work with a wide range of systems, processes and technologies to own and solve problems from end-to-end. Collaborate with engineers, designers, data scientists, and product managers to develop new features and products. Uphold our high engineering standards and bring consistency to the many codebases and processes you will encounter. Build elegant APIs and user experiences that enable merchants to run and scale their businesses on top of Stripe. Contribute to the design and architecture of the next generation of Stripe’s infrastructure, to meet the high growth needs of our company and customers for years to come. Who you are Minimum requirements Must have a Bachelor's degree or foreign equivalent in Computer Science, Software Engineering, or a related field, plus 2 years of experience in Software development. Must also have two years of experience in each of the following: Working across the stack and navigate codebases with different languages and tools; Developing user-facing experiences and writing queries for analyzing experiment results; Programming Language including Ruby, Java, Scala and SQL; Cloud based services including gRPC, GraphQL and AWS; Proficiency with MongoDB, Spark, Hadoop, Kafka, Grafana, Prometheus and AWS; and macOS and Unix-based operating systems. Salary: $156,800.00 - $235,200.00/yr. This salary range represents the base salary range for the rol

javasqlmongodb
View job →
G
Gitlab
📍 United States• Full-time• Remote• From $126.4K/yr
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Lead People Systems Engineer, you'll drive how GitLab's People systems evolve through an AI-first transformation mindset, putting AI to work across everything we build. You won't just support our people systems, you'll build them: owning the code, pipelines, and integrations that keep GitLab's People automation ecosystem running reliably at scale across Workato, Google Cloud, Claude, Workday, and the tools that power our global workforce. This is a strong fit if you think like a software engineer first, write clean and maintainable code, and thrive at the intersection of people systems, AI, innovatio

REMOTEpythonawsgcp
View job →
M
Mongodb
📍 Gurugram• Full-time
1mo ago

We are looking to speak to candidates who are based in Gurugram for our hybrid working Position Expectations As an Individual Contributor as part of the HR Shared Services, APAC team in India, you will play a vital role responsible for transitioning work and making sure relevant SLA’s are met Learn and assist in executing day-to-day HR data processes, such as employee record updates, org structure changes, and data validations Support the cleansing and maintenance of HR data in alignment with company policies and global standards Adhere to and demonstrate high proficiency in agreed critical metrics & SLAs Assist in preparing reports and trackers for internal HR teams and audits Deliver high quality and consistent service delivery to all internal customers and partners and follow Standard Operating Procedures Maintain confidentiality and accuracy when handling sensitive employee data Key Skills & Abilities 1+ years experience in HR Operations / Shared Services, HR Data Entry / Management in an HR Shared Services role (APAC/ India, EMEA, NAMER regions) Basic understanding of HR systems, data, or Excel/Google Spreadsheet; familiarity with ticketing tools (e.g., Zendesk) is a plus Excellent analytical, communication, and problem-solving skills Ability to learn quickly, ask questions, and follow through on assigned work Comfortable working in a structured, process-driven environment Good understanding of data governance frameworks and data quality About MongoDB MongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined the data platform for the AI era, enabling builders to create, transform, and disrupt industries with software. MongoDB’s unified data platform, the most widely available, globally distributed data platform on the market, helps organizations modernize legacy workloads, embrace innovation, and unleash AI. Our cloud-native platform, MongoDB Atlas, i

mongodbawsazure
View job →
D
Discord
📍 San Francisco Bay Area• Full-time
1mo ago

Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. We are seeking an experienced Senior Salesforce Technical Lead to join Discord's Business Systems team. In this role, you will lead the architecture, design, and development of Salesforce solutions across Ad Sales. You will partner closely with the Technical PM, business stakeholders, and cross-functional teams to deliver scalable, high-quality Salesforce solutions that power Discord's core business operations. What you'll be doing Lead the architecture, design, and development of Salesforce solutions across Ad Sales (Sales Cloud). Create and execute the Ad Sales Roadmap; drive delivery against 2026 priority items including platform enhancements, new feature builds, and ongoing maintenance. Lead the 3rd party development team by establishing technical direction, conducting code reviews, and collaborating with the MuleSoft team on developing and maintaining integrations. Architect Salesforce and middleware integrations via REST/SOAP APIs and MuleSoft; partner with the Technical PM and RevOps to translate business requirements into robust, scalable solutions. Manage AppExchange apps and HubSpot/Asana integrations; ensure ITGC and SOX audit readiness across all platform changes and deployments. Proactively identify technical debt and coordinate remediation efforts. Own end-to-end DevOps standards for the team: define coding standards, manage CI/CD pipelines and deployment trackers, and ensure SDLC/ITGC compliance across all phases of the sprint/dev cycle. Step in to troubleshoot and diagnose critical issues when they arise. Provide L2 (Incident and Problem Management) and L3 (Bug Fixes and Minor Enhancemen

ci/cdgitrest
View job →
SC
Sigma Computing
📍 San Francisco• Full-time• $140K – $180K/yr
15 days ago

Data Engineer, Data Platform About the Role We are building out our Data Platform team at Sigma, with a relentless focus on developing data models that fuel trusted insights across the company. As a Data Platform Engineer, you will be responsible for the underlying data architecture across Snowflake and Databricks, as well as building and optimizing various ETL pipelines to fuel both internal and external (demo) use cases. Reporting to the VP of Data & Revenue Engineering, this is a high visibility role with the opportunity to work on greenfield projects. If you’re a data engineer with a builder mindset who wants to leverage a best-in-class stack, and genuinely is invested in Sigma’s mission, let’s chat! What You Will Be Doing Architect and manage our production data pipelines in Snowflake and how they are consumed in Sigma ( Tech we use : Fivetran, dagster, dlt, terraform, dbt, Snowflake, Sigma, Hightouch, Metaplane) Build foundational processes for scaling our demo asset data across various Cloud Data Warehouses Scale our terraform deployment across all of our Snowflake assets Continue to advance our data governance policies Work cross functionally to accomplish all of the above! You’ll work across Product, Engineering, GTM—with users of all skills and levels Qualifications We Need Strong knowledge and experience of working with APIs and building data pipelines from various systems into Cloud Data Platforms (e.g., Snowflake, Databricks) Strong communication and collaboration skills. You will primarily partner with the Analytics and Infrastructure Engineering teams internally at Sigma; your ability to work and collaborate closely with them will be integral to your success. Experience deploying data governance frameworks with a scalable and repeatable process Ability to thrive in ambiguous environments and get stuff done. We move fast and iterate quickly, and we want you to feel empowered to do exactly that 3+ years of relevant experience wor

pythonsqlaws
View job →
N
1mo ago

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys

pythonjavakubernetes
View job →
N
Nvidia
📍 Remote, United States• Remote
12 days ago

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data

REMOTEpythonsqlaws
View job →
M
Mongodb
📍 United States• Full-time• From $1.3M/yr
1mo ago

The Opportunity MongoDB’s partner ecosystem — systems integrators, ISVs, and the major cloud providers — is a strategic growth engine for the business. We’re looking for a Senior Manager, Sales Plays and Offerings to design the joint go-to-market motions that turn partner relationships into pipeline and revenue. This is a highly cross-functional, strategic role: you’ll build the sales plays and joint offerings themselves, partner with Enablement to get the field and our partners ready to sell them, instrument how they perform, and work with the Programs lead to make sure incentives reward the behavior we want to see. You will report directly to the VP of Partner Strategic Operations and act as a connective layer between Partnerships, Sales, Enablement, and Programs — translating ecosystem strategy into repeatable, measurable, field-ready motions. We are looking to speak to candidates who are based anywhere in the US for our hybrid working model. What You'll Do Build sales plays and joint offerings Design and package partner sales plays and joint solution offerings with priority ISVs, SIs, and cloud partners — defining the joint value proposition, target segment, competitive positioning, and playbook for how field and partner sellers execute it Partner with Product Marketing, Solutions Engineering, and partner counterparts to validate technical integration stories and translate them into a compelling, sellable narrative Prioritize which plays to build and scale based on market opportunity, partner readiness, and alignment to MongoDB’s strategic pillars (e.g., AI, migrations, industry verticals) Own the lifecycle of each play from concept through launch, iteration, and eventual retirement or refresh Drive partner and field enablement Partner closely with the Enablement team to translate each sales play into field- and partner-facing assets: pitch decks, battlecards, demo scripts, certification conte

mongodbawsazure
View job →
N
Nvidia
📍 Santa Clara, United States
1mo ago

NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel

pythonkuberneteslinux
View job →
R
Roblox
📍 San Mateo• Full-time• From $243.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Infrastructure Compute Site Reliability Engineering mission is to own and manage the successful operation of our underlying cell infrastructure system, along with elements of service discovery, secrets management and related software layers. We’re looking for a skilled Senior Site Reliability Engineer with strong programming skills to help us build Roblox's private cloud, productionize our growing Kubernetes-based infrastructure, and institute reliability best practices across the Roblox Compute team. You will: Design and Develop systems & libraries that promote fault-tolerance and resilience, automate much of the management and lifecycle of our clusters, and ensure systems are observable. Promote and Institute reliability best practices across the Infra Compute group, drive common reliability initiatives. Provides collaborative technical reviews and operational guidance to strengthen system reliability. Build, Automate and Standardize process automation to create a "golden path" of tooling and platform support that powers the fundamental Roblox ecosystem. Create Tooling that provides production guardrails, by evaluating release candidate capacity with load testing tooling before de

javaawskubernetes
View job →
M
Mongodb
📍 New York City• Full-time• From $126K/yr
1mo ago

The MongoDB Cloud Services Team is a diverse group of contributors working together to help our users manage MongoDB at global scale. The Cloud Team is responsible for MongoDB Atlas: our database as a service offering, and fastest growing product, which allows users to deploy fault-tolerant, globally distributed MongoDB clusters in just minutes. The Backup Team delivers essential infrastructure to help our customers in their hour of need - providing the ability to quickly restore a massive, distributed database to any point in time at the click of a button. The Backup Team’s mission is to make MongoDB backup more reliable, faster, and also cheaper. This team is responsible for the Backup Agent (Go), the extensive server-side infrastructure (Java) which manages 100s of TB of data and processes billions of operations per day, and the user interface (Javascript) that customers use to manage their backups. Common project themes are performance, scaling, and ease of use. We are looking to speak to candidates who are based in New York for our hybrid working model. We're looking for someone who is Skilled at writing large-scale, distributed backend systems in a compiled language (Java, C#, Go, etc.) Fond of chasing down tough problems in a distributed systems environment Cool under pressure - has wrangled production crises, and secretly finds this a little fun Experienced with Linux, and able to correlate application performance problems with underlying hardware limits Comfortable working across the stack of a modern web application Always striving to expand their knowledge Curious, collaborative and intellectually honest Responsibilities Work closely with product teams, considering the user’s perspective while helping the team achieve success Collaborate with team members over best practices and core concepts Hold yourself accountable to your actions, maintaining the balance between accomplishing goals with research & development Own our

javascriptjavamongodb
View job →
P
Pagerduty
📍 Atlanta• Full-time• From $98K/yr
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office, you'll help build and operate the foundational infrastructure that powers PagerDuty's real-time digital operations platform. Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You'll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases. Key Responsibilities ● Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems. ● Contribute to the reliability and scalability of PagerDuty's core platform by hardening existing systems and supporting the rollout of new infrastructure

pythonawsazure
View job →
P
Pagerduty
📍 Atlanta• Full-time• $113K – $171.6K/yr
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. As a Site Reliability Engineer II on the Core Infrastructure team in our Atlanta office, you'll help build and operate the foundational infrastructure that powers PagerDuty's real-time digital operations platform. Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You'll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases. Key Responsibilities ● Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems. ● Contribute to the reliability and scalability of PagerDuty's core platform by hardening existing systems and supporting the rollout of new infrastructur

pythonawsazure
View job →
PE
Private Employer
📍 Washington• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Palantir's Offensive Security team probes our own products, infrastructure, and cloud environment the way a real attacker would. As an Offensive Security Engineer, you will test internal applications and infrastructure, chain findings into realistic attack paths, and work directly with the engineers responsible for the affected systems to get issues closed. A growing part of the team’s work is also developing agentic offensive security tooling: LLM-driven systems that encode operator tradecraft and apply it continuously. You will help in the design of that tooling and prove it out on live assessments. You will also help to coordinate third-party penetration testing engagements, scoping the work with specialized external firms, keeping assessments focused on the risks that matter, and turning their results into concrete remediation.

🔔

Get new cloud operations system administrator jobs by email

Daily job updates · Unsubscribe anytime