Jobiba hiring network

Cloud Operations System Administrator Jobs

2,329 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About the Team The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads. As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI’s rapidly expanding compute environment. About the Role We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI’s third-party infrastructure partners, including major cloud service providers and strategic compute vendors. In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership. Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure. Key Responsibilities Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations. Develop operationa

awsazuregcp
View job →
T
Toradex
📍 Bengaluru• Full-time
15 days ago

Toradex is a global company strongly focused on engineering & technology. We’re powered by a diverse & uniquely gifted workforce. We pursue the best people to propel our innovative vision of embedded computing and IoT. If you’re interested in being a driving force at an agile technology company, engineering clever computing solutions & helping other companies bring their products to life, we should talk. Description We are looking for a DevOps Engineer to strengthen our cloud operations and engineering practices, with a focus on reliable website delivery, secure AWS foundations, and fast but controlled delivery of new services. The position combines AWS operations, infrastructure as code, CI/CD, automation, and pragmatic software engineering. The person should be confident working with services for edge delivery, compute, storage, databases, DNS, security, and observability without relying on manual console changes as the default operating model. The role also supports on-premises to cloud migration, global service optimization, and practical responses to increasing AI-driven traffic. We value candidates who can use modern AI-assisted development effectively to spin up proof-of-concept projects quickly, while still applying disciplined Git, review, security, and deployment practices. About you You enjoy building stable, secure, and maintainable infrastructure that supports business-critical services. You can work independently and take ownership of cloud environments, deployments, and operational improvements. You are comfortable balancing speed, reliability, cost, and security when making technical decisions. You communicate clearly with technical and non-technical stakeholders and explain trade-offs in a practical way. You document your work well and create clear runbooks and support material for future maintenance. You are methodical when troubleshooting incidents and stay calm when systems are under pressure. You are curious about modern traffic patt

javascripttypescriptpython
View job →
A
Appspace
📍 Texas• Full-time• Remote
1mo ago

About Appspace: At Appspace, we’re passionate about creating better work experiences for people everywhere, and we’re looking for people that feel the same way. Our global office locations and flexible work culture help you work wherever and however you’re at your best. Plus, we take the time to help you enjoy your work, build lasting connections, and grow your role. Join the Appspace team and be a part of a culture that’s helping people everywhere love where they work. Your Role as a Cloud Security Engineer : We are seeking a highly skilled Cloud Security Engineer to join our dynamic team. This is a crucial customer-facing role where you will be instrumental in designing, implementing secure cloud configurations, being proactive by recommending security by design standards and, securing complex cloud environments for our clients across Google Cloud Platform (GCP), Microsoft Azure, and Amazon Web Services (AWS), with a strong emphasis on GCP. You will leverage your deep expertise in SaaS security, network security, and compliance to provide strategic guidance and hands-on support, ensuring our clients' cloud infrastructures are robust, resilient, and compliant with industry standards. What You'll Do: Cloud Security Operations: Design, implement, and optimize robust cloud security architectures to enhance, build, monitor and address all security alerts from our SIEM and other security systems. This is an operational role whereby you will be available M-F 8am-5pm EDT and, when needed, on-call shifts on evenings and weekends. Network Security Expertise: Your network security and cloud security expertise will be required to respond to customer questionnaires, customer calls and create artifacts including network diagrams, architecture diagram, data flow diagrams and other artifacts to support customer requests. Strong written skills will be required here and attention to detail. SIEM Integration & Optimization: As a Level-2 Security Operation

REMOTEpythonawsazure
View job →

About the Team At OpenAI, our User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We translate real-world user and operational signals into timely decisions, practical interventions, and improvements to our products and systems. This role will take on new, ambiguous, or underdeveloped operational risks and help mature them into scalable capabilities. We work across USRO and partner closely with Product, Engineering, Data Science, Product Policy, Legal, Safety, Support, and external vendors or partnership stakeholders. About the Role We are seeking a Senior Operations Analyst to take on complex, ambiguous safety and risk problems and turn them into practical operational solutions that can scale. This is a senior individual-contributor role for a versatile operator who is comfortable moving between queues, investigation, analysis, workflow design, hands-on execution, and cross-functional leadership. Depending on team needs, the role may focus on emerging-risk incubation, cloud deployment partnerships, or other new operational areas. You will be expected to move quickly, work hands-on, and create structure without waiting for perfect requirements or a large support team. The work starts with the problem, not a prescribed process. You may investigate unstructured user signals, stand up a lightweight workflow, build an AI-assisted tool, improve an existing operation, or help a new launch become operationally ready. The goal is to produce durable systems that other people can run, not simply complete a series of individual tasks. The portfolio will change with company priorities and may span established harm areas, emerging-risk incubation, cloud deployments and partnerships, device safety, or new product launches. Some hires may focus primarily on cloud deployment operations, including launch readiness, partner coordination, safety workflows, and operational monitoring. You will ty

sqlawsrest
View job →
E
10 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. The Role As part of the Core Products Business Unit (CPBU) Systems Software team, you will join a high-leverage engineering initiative dedicated to transforming upgrade workflows and lifecycle operations. Our goal is to shift hardware upgrades from manual, support-intensive workflows toward a safer, guided, automated self-service experience integrated directly into Pure1 and Skyline. In this role as a Systems Software Engineer (MTS3), located on-site in Santa Clara, CA, you will build dedicated systems infrastructure and controller-upgrade automation for our next-generation upgrade framework. You will bridge on-array software execution with cloud-guided operations, eliminating manual choreography and runbooks while maintaining uncompromising standards for safety and availability. What You'll Do Design & Build Upgrade Infrastructure: Architect, implement, and maintain high-reliability systems software components for controller-upgrade automation and the next-generation hardware upgrade framework. Automate Upgrade Workflows: Transition complex hardware replacement and upgrade tasks into repeatable, self-service workflows, significantly reducing support and field dependencies. Engineered Safety & Resilience: Build robust state machines, health checks, workflow orchestration, pre/post-upgrade validations, failure-handling mechanisms, and automated recovery paths. Cross-Functional Collaboration: Partner closely with int

pythonaic++
View job →
P
Pagerduty
📍 Lisbon• Full-time
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. About the role PagerDuty’s Operations Cloud runs on a platform that ingests billions of signals and turns them into real-time action for thousands of customers. We’re looking for a Senior AI/ML Engineer who lives at the intersection of two disciplines: large-scale distributed systems and applied AI. In this role you will design and ship AI systems that run in production at PagerDuty’s scale — powering Incident Management AI Agents, event intelligence, and the LLM-powered capabilities embedded across our platform. You’ll own the full lifecycle, from framing the problem to serving reliably at scale. We are looking for a candidate who is genuinely passionate about building with modern AI — LLMs, agents, and retrieval — but grounded in the realities of building resilient, high-throughput systems. What you’ll do Design and build AI-powered features — LLM agents, retrieval, and event intelligence — that operate on high-volume, real-time event streams, from problem framing through production deployment and monitoring. Architect and own the systems behind them: agent and prompt orchestration, retrieval pipelin

awsazuregcp
View job →
P
Pagerduty
📍 Lisbon• Full-time
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. About the role PagerDuty’s Operations Cloud runs on a platform that ingests billions of signals and turns them into real-time action for thousands of customers. We’re looking for an early-career AI/ML Engineer who is excited to grow at the intersection of two disciplines: large-scale distributed systems and machine learning. In this role you will help build and ship AI systems that run in production at PagerDuty’s scale — powering Incident Management AI Agents, event intelligence, and the LLM-powered capabilities embedded across our platform. You’ll work alongside senior engineers on real production problems, learning how AI features go from a prototype to something that serves reliably at scale. We are looking for a candidate who is genuinely excited about building with modern AI — LLMs, agents, and retrieval — eager to learn how resilient, high-throughput systems are built, and motivated to grow into an engineer who is strong in both. What you’ll do Contribute to AI-powered features — LLM agents, retrieval, and event intelligence — that operate on high-volume, real-time data, with support and guidanc

awsazuregcp
View job →
P
Pagerduty
📍 Washington Dc Baltimore Area• Full-time• Remote• $160K – $187K/yr
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. PagerDuty is seeking a Federal Systems Integrator (FSI) Account Executive (AE) to drive strategic growth across top-tier FSIs. In this role, you will partner with major integrators to embed PagerDuty's Operations Cloud into large-scale federal program bids and expand adoption across their portfolios. You will collaborate with senior FSI executives to deliver automated incident response and AIOps solutions that support critical government missions. This is a high-impact opportunity to scale PagerDuty’s public sector presence, close high-value deals, and lead complex cloud sales within the federal ecosystem. KEY RESPONSIBILITIES Own and grow a defined set of Federal Systems Integrator (FSI) accounts by driving upsell, cross-sell, and expansion opportunities. Build and maintain trusted senior relationships through regular in-person engagement and consultative selling. Develop and execute strategic account plans to identify growth areas, expansion pathways, and competitive positioning. Drive adoption of PagerDuty’s Operations Cloud by articulating clear business value an

REMOTEawsgitrest
View job →
H
Hyreo
📍 Bengaluru• Full-time
15 days ago

Own the architecture of Myntra’s new product platforms to drive business results Drive and own the architecture and design of some of the most advanced & complex software systems / products in the industry to create company wide impact Help build, mentor and coach a team of very talented Engineers, Architects, Quality engineers, System Operation Engineers and DevOps engineers in architectural and design best practices Experience in distributed systems, cloud service development, deployment and delivery Accountable for the design, for the ease of evolution, quality of the systems, performance, scaling, and availability characteristics and limitations of the systems Envision and develop the long-term architectural direction, with emphasis on platforms/ reusable components while adopting an agile delivery process. Establish structures and processes that ensure a high level of quality and reliability and extensibility of deliverables Drive the creation of next generation extensible web, mobile and fashion commerce platforms, security protocols, customisation and tools to support continuous scaling, internationalisation and platform extensions Drive code and design reviews of components / systems / products in scope and drives the architectural governance for them Set directional paths for the teams/department for adoption of new technology stacks for solving business problems Represent multiple technology domains and Myntra in external technical forums Work with product management, business stakeholders and other engineering leaders to help define mid-term, long-term roadmaps and shape business directions Initiate and deliver leadership training within the engineering organisation, including training new managers, and drive the growth of leaders to create a strong leadership bench. Qualifications & Experience 8+ years of experience in software product development Must have a degree in Computer Science o

javasqlagile
View job →
H
Hyreo
📍 Bengaluru• Full-time
15 days ago

Roles and Responsibilities Own the architecture of Myntra’s new product platforms to drive business results Drive and own the architecture and design of some of the most advanced & complex software systems / products in the industry to create company wide impact Help build, mentor and coach a team of very talented Engineers, Architects, Quality engineers, System Operation Engineers and DevOps engineers in architectural and design best practices Experience in distributed systems, cloud service development, deployment and delivery Accountable for the design, for the ease of evolution, quality of the systems, performance, scaling, and availability characteristics and limitations of the systems Envision and develop the long-term architectural direction, with emphasis on platforms/ reusable components while adopting an agile delivery process. Establish structures and processes that ensure a high level of quality and reliability and extensibility of deliverables Drive the creation of next generation extensible web, mobile and fashion commerce platforms, security protocols, customisation and tools to support continuous scaling, internationalisation and platform extensions Drive code and design reviews of components / systems / products in scope and drives the architectural governance for them Set directional paths for the teams/department for adoption of new technology stacks for solving business problems Represent multiple technology domains and Myntra in external technical forums Work with product management, business stakeholders and other engineering leaders to help define mid-term, long-term roadmaps and shape business directions Initiate and deliver leadership training within the engineering organisation, including training new managers, and drive the growth of leaders to create a strong leadership bench. Qualifications & Experience 8+ years of experience in software product development Must have a d

javasqlagile
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

awslinuxrest
View job →
O
Okta
📍 Bengaluru• Full-time
16 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Company Description: Okta - we are the World’s Identity Company. We don’t just protect logins; we secure the digital life of the Fortune 100. Built from the ground up in the cloud, Okta securely and simply connects people to their applications from any device, anywhere, at any time. Okta integrates with existing directories and identity systems, as well as thousands of on-premises, cloud and mobile applications, and runs on a secure, reliable and extensively audited cloud-based platform. Who are we looking for: We are looking for a Senior Software Engineer in Test. An individual who takes ownership and builds viable solutions. A team oriented individual who can demonstrate working independently, as an individual contributor, in a distributed working environment. Detail oriented and methodical in approaching tasks with excellent research and analytical skills. An ideal candidate will be someone who is passionate about automation & appling those automation skills in testing, cloud native, large-scale, mission-critical software in a fast-paced agile environment while partnering with cloud Infrastructure and operations teams. Okta engineering strongly believes in automated testing, and an iterative process to build high-quality next generation software. This role is mainly focused on automation of Cloud Infrastructure testing and supporting SREs, including bespoke solutions rolled out by Developer Productivity teams. The Quality Engineering team work

pythonsqlaws
View job →
N
1mo ago

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build cloud-native data and storage services for hybrid and multi-cloud infrastructure, including dataset discovery, ingestion, governance, checkpointing, observability, and low-latency access. Develop scalable cloud-native services and APIs that support exabyte-scale, high-performance GPU training and inference workflows. Work closely with product managers, internal AI teams, platform teams, and partner engineering teams to understand requirements and turn them into reliable production systems. Collaborate with SRE, operations, and support teams to improve service reliability, performance, observability, on-call readiness, and operational scale. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, and verification. What we need to see: BS in Computer Science, Information Systems, Computer Engineering, or equivalent experience, with 5+ years of software engineering experience. Strong foundation in algorithms, data structures, distributed systems, and practi

pythonjavaaws
View job →
O
1mo ago

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

pythonawsazure
View job →
M
Modal
📍 New York• Full-time
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for a People Operations Generalist to join our growing People team. You'll touch the employee lifecycle end-to-end — from offer acceptance through offboarding — while helping to build the processes and documentation that let our People function scale with the business. This is a great fit for a highly organized, systems-oriented people person who thrives in a fast-paced environment and wants to build operational foundations, not just maintain them. What you’ll do Own and continuously improve the new hire onboarding experience, ensuring employees are set up for success and internal tasks are tracked and completed on time. Serve as a first point of contact for employee questions across the full HR spectrum, triaging and routing more complex issues to the right People team member or external partner. Maintain and improve self-service resources (FAQs, Not

🔔

Get new cloud operations system administrator jobs by email

Daily job updates · Unsubscribe anytime