Jobiba hiring network

Cloud Operations System Administrator Jobs

2,329 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside leading cloud providers, engineering firms, construction partners, utilities, and equipment manufacturers, we are delivering hyperscale AI campuses that enable the next generation of frontier AI models. The Strategic Sourcing team develops and executes the commercial strategies that ensure our infrastructure programs have reliable access to the equipment, materials, and strategic partners needed to deliver at unprecedented scale. We partner closely with Infrastructure Delivery, Capacity Planning, Design Engineering, Hardware Operations, Finance, Legal, and our external suppliers to build a resilient global supply network capable of supporting Industrial Compute's long-term growth. As we continue expanding globally, strategic sourcing becomes a critical competitive advantage, ensuring our infrastructure programs remain cost-effective, resilient, and capable of executing against aggressive deployment timelines. About the Role We are seeking a Strategic Sourcing Manager, Data Center Infrastructure to lead sourcing strategy for the critical infrastructure systems that power Industrial Compute campuses. This role will develop commercial strategies, negotiate strategic supplier agreements, and manage relationships across engineering, construction, manufacturing, and infrastructure partners responsible for delivering mission-critical facilities. You will work closely with Infrastructure Delivery, Capacity Planning, Engineering, Finance, Construction, and external suppliers to ensure Industrial Compute has the capacity, supplier relationships, and commercial frameworks required to support rapid global expansion. The ideal candidate has experience sourcing major infrastructure systems for hyperscale data centers, mission-critical facilities, industrial construction, semiconductor manufacturing, energy infrastructure, or similarly comple

awsrestai
View job →

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside leading cloud providers, engineering firms, construction partners, utilities, and equipment manufacturers, we are delivering hyperscale AI campuses that enable the next generation of frontier AI models. The Strategic Sourcing team develops and executes the commercial strategies that ensure our infrastructure programs have reliable access to the equipment, materials, and strategic partners needed to deliver at unprecedented scale. We partner closely with Infrastructure Delivery, Capacity Planning, Design Engineering, Hardware Operations, Finance, Legal, and our external suppliers to build a resilient global supply network capable of supporting Industrial Compute's long-term growth. As we continue expanding globally, strategic sourcing becomes a critical competitive advantage, ensuring our infrastructure programs remain cost-effective, resilient, and capable of executing against aggressive deployment timelines. About the Role We are seeking a Strategic Sourcing Manager, Data Center Infrastructure to lead sourcing strategy for the critical infrastructure systems that power Industrial Compute campuses. This role will develop commercial strategies, negotiate strategic supplier agreements, and manage relationships across engineering, construction, manufacturing, and infrastructure partners responsible for delivering mission-critical facilities. You will work closely with Infrastructure Delivery, Capacity Planning, Engineering, Finance, Construction, and external suppliers to ensure Industrial Compute has the capacity, supplier relationships, and commercial frameworks required to support rapid global expansion. The ideal candidate has experience sourcing major infrastructure systems for hyperscale data centers, mission-critical facilities, industrial construction, semiconductor manufacturing, energy infrastructure, or similarly comple

awsrestai
View job →

Overview The Associate Director of Service Engineering leads the reliability, availability, and operational excellence of Natera’s lab-facing platforms. This role ensures that clinical systems, laboratory equipment workflows, and data pipelines operate with high reliability, scalability, and compliance in a regulated healthcare environment. You will lead a team responsible for production stability, incident response, and service health, partnering closely with Production Engineering, Lab Operations, Bioinformatics, Infrastructure, Facilities, and Compliance to support mission-critical genetic testing and diagnostics. Key Responsibilities Leadership & Team Development Lead, mentor, and scale a team of Service Engineers / SREs supporting clinical production systems Establish clear expectations around ownership, on-call readiness, and operational excellence Drive hiring, onboarding, performance management, and career growth Foster a blameless, learning-oriented culture focused on patient impact and reliability Service Reliability & Production Operations Own reliability and availability for production services supporting laboratory operations, reporting, and customer delivery Define and manage SLAs, SLOs, and operational KPIs aligned with clinical and business priorities Lead major incident response, ensuring rapid triage, clear communication, and thorough post-incident reviews Oversee on-call rotations, escalation paths, and operational playbooks Ensure operational readiness and go-live support for new assays, pipelines, and platform capabilities Technical Strategy & Execution Partner with Engineering and Development teams to design resilient, fault-tolerant systems Drive best practices for monitoring, alerting, logging, and observability across lab and cloud platforms Reduce operational toil through automation, tooling, and process improvements Advocate for reliability, performance, and scalability requirements early in t

D
12 days ago

We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices

REMOTEpythonnode.jsai
View job →
SC
Sigma Computing
📍 San Francisco• Full-time• $170K – $235K/yr
15 days ago

About the Role Sigma Computing is redefining business intelligence by making complex data analysis accessible through a high-performance platform built for the modern data stack. The Compiler Team plays a foundational role in this mission by transforming user-driven spreadsheet interactions into highly optimized SQL queries, enabling seamless exploratory analytics on cloud data warehouses. As a member of the Compiler Team, you will join a group of engineers dedicated to building the core systems and abstractions that power Sigma’s intuitive spreadsheet interface, ensuring speed, reliability, and scalability for all users. What You Will Be Doing Tackle core challenges at the intersection of data modeling, query compilation, and large-scale interactive analytics—making it possible for end-users to query data warehouses efficiently without deep technical knowledge Design, build, and maintain sophisticated compiler infrastructure and intermediate representations that translate spreadsheet operations into optimized query plans Apply advanced optimization strategies to improve performance and accuracy across a wide range of query workloads and data architectures Contribute to both backend (Rust) and key frontend foundations (TypeScript), evolving critical abstractions that enable end-to-end workflow optimizations and new features Debug, analyze, and resolve complex issues, ensuring robustness and maintainability in a rapidly evolving product Collaborate with engineers and product stakeholders to review designs and code, driving technical best practices and architectural decisions throughout the team and company Qualifications We Need 5+ years experience engineering high-quality software systems Demonstrated success building and maintaining complex infrastructure or core platform services Deep understanding of Computer Science fundamentals, particularly in compilers, algorithms, SQL Optimization Passion for teamwork, technical ownership, and continually

typescriptpythonsql
View job →
RS
Redwood Software
📍 Ontario• Full-time• C$110K – C$135K/yr
15 days ago

OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric—with AI embedded across the automation lifecycle—Redwood accelerates the path to autonomous operations. Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT We are seeking a highly skilled and passionate Full Stack Software Developer with a strong focus on Java to join our growing engineering team. In this role, you will be instrumental in designing, developing, and maintaining robust and scalable full-stack applications that power our automation and SaaS platforms. You will work across the entire software development lifecycle, from concept to deployment, collaborating closely with product managers, designers, and other engineers to deliver high-quality, impactful solutions. Design, develop, and implement highly performant and scalable full-stack applications using Java, Javascript, and related technologies. Build and maintain robust back-end services, APIs, and microservices. Develop responsive and intuitive front-end user interfaces. Collaborate with product management to understand requirements and translate them into technical specifications. Participate in all phases of the software development lifecycle, including planning, design, codin

javascripttypescriptjava
View job →
RS
15 days ago

OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric—with AI embedded across the automation lifecycle—Redwood accelerates the path to autonomous operations. Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT We are looking for a Software Engineer, Platform & Integrations . Working closely with senior and lead engineers, you will design, develop, and maintain high-quality features that power enterprise data exchange for more than 1,000 customers worldwide. This is an incredible opportunity to deepen your expertise in cloud-native architectures, enterprise security, and modern DevOps practices in a fast-growing product environment. Feature Development & Design: Write clean, maintainable, and well-tested code using Java and Spring Boot to deliver scalable backend services and microservices. Platform Reliability: Contribute to enhancing the monitoring, logging, and observability of our core platform to ensure high availability and performance. Security & Compliance: Implement secure coding practices to safeguard data exchange and maintain compliance across our cloud infrastructure. Collaborative Execution: Work within an agile team, collaborating closely with QA, Product, and fellow engineers to deliver high-qual

javasqlpostgresql
View job →
RS
Redwood Software
📍 Ontario• Full-time• C$100K – C$125K/yr
15 days ago

OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leader in full stack automation fabric solutions for mission-critical business processes. Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric—with AI embedded across the automation lifecycle—Redwood accelerates the path to autonomous operations. Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT We are hiring a high-impact Technical Program Manager (TPM) to own end-to-end delivery health within a product area. This is a senior role embedded within a core product engineering group. You will partner with the Engineering Directors as your peer enabling engineering leaders to focus on technical quality while the TPM will own planning, governance, and stakeholder visibility and be the critical link ensuring our roadmap commitments translate into high-quality, enterprise-ready releases. You will operate at the intersection of Product, Engineering, Platform, Support, and Go-To-Market teams, and be directly accountable for delivery predictability, release readiness, dependency management, and executive visibility. This role demands a combination of strategic systems thinking

awsrestagile
View job →
RS
15 days ago

OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric—with AI embedded across the automation lifecycle—Redwood accelerates the path to autonomous operations. Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT The Implementation Consultant role is instrumental in assisting Redwood customers ranging from Fortune global enterprises to midsized businesses to leverage our solutions to orchestrate automation across distributed and heterogeneous platforms and enterprise business applications including SAP, Oracle, PeopleSoft, data warehousing and OS tasks. The range of responsibilities includes product installation and configuration, training, conducting workshops and providing best practices guidance to advance our customer’s automation business needs. Assist customers with installation, configuration, training and use of Redwood solutions. Perform analysis, scoping, and effort estimation to meet customer requirements. Provide configuration, architecture & installation recommendations based on product best practices & customer needs. Coordinate with other Redwood resources including Engagement Managers to update project status, timeline, effort, and expected completion. Communicate effectively and p

javaawsazure
View job →
TI
TEGNA India
📍 Chennai• Full-time
15 days ago

TEGNA Inc. helps people thrive in their local communities by providing the trusted local news and services that matter most. With 64 television stations in 51 U.S. markets, TEGNA reaches more than 100 million people monthly across web, mobile apps, streaming, and linear television, while also maintaining a strong global presence in India with offices in Bangalore and Chennai that support technology, product, and business operations initiatives. Together, we are building a sustainable future for local news. Senior DevOps Engineer About TEGNA TEGNA Inc. (NYSE: TGNA) helps people thrive in their local communities by providing trusted local news and services. With 64 television stations across 51 U.S. markets, TEGNA reaches more than 100 million people monthly across digital, mobile, streaming, and television platforms. We are focused on innovation, technology excellence, and building scalable solutions that create meaningful impact. Position Overview TEGNA is looking for a highly skilled Senior DevOps Engineer with strong expertise in AWS, Kubernetes, and Infrastructure as Code to design, automate, and manage scalable cloud infrastructure. The ideal candidate will have hands-on experience operating Kubernetes workloads in production, building CI/CD pipelines, and implementing monitoring and security best practices. This role requires deep technical expertise, strong troubleshooting skills, and the ability to work in fast-paced, distributed environments. You will play a key role in ensuring platform reliability, automation maturity, and production stability across cloud-native microservices systems. What You’ll Do Design and manage cloud infrastructure using Infrastructure as Code (AWS CDK, CloudFormation, Terraform). Build and maintain CI/CD pipelines using GitHub Actions and Jenkins to enable automated and reliable deployments. Deploy, manage, and scale Kubernetes clust

pythonsqlmongodb
View job →
S
16 days ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Proactive Threat team is responsible for identifying vulnerabilities and security weaknesses across Stripe's systems, applications, networks, and cloud infrastructure — before adversaries do. We operate as a hybrid offensive function: conducting penetration testing, emulating real-world threat actors through red team operations, and partnering closely with our defensive security teams to validate detection capabilities and improve Stripe's overall security posture. We are builders first. Our team develops custom tooling, automation frameworks, and internal platforms that scale our offensive capabilities and enable repeatable, high-fidelity assessments. We believe the best offensive security engineers are equal parts hacker and engineer. The team is distributed across the United States, primarily operating in Eastern and Pacific time zones, and collaborates regularly with security, engineering, and product stakeholders across Stripe — including teams in Europe and Asia. What you'll do As an Offensive Security Engineer on the Proactive Threat team, you will simulate the tactics, techniques, and procedures (TTPs) of real-world adversaries to uncover security risks across Stripe's products and infrastructure. You'll conduct hands-on penetration testing, lead red team engagements, and collaborate with blue team counterparts to validate and improve detection and response capabilities. Your work will directly influence how Stripe builds, ships,

pythonawsazure
View job →
M
Mongodb
📍 Japan• Full-time
18 days ago

MongoDB Pre-Sales Solutions Architects are technical business advisors who help customers design, justify, and adopt reliable, scalable systems using MongoDB’s data platform. They own the technical strategy across complex opportunities—from discovery and qualification through architecture, proof of value, executive alignment, and successful adoption—and connect technical decisions to measurable business outcomes. You’ll partner closely with Account Executives, Sales Leadership, Customer Success, Professional Services, and ecosystem partners to shape multi-threaded account strategies, build champions, de-risk complex architectures, and drive expansion. You’ll serve as a trusted advisor to developers, architects, operations leaders, and business executives, helping organizations modernize legacy systems, build AI-powered applications, and realize measurable value from MongoDB. . We are looking to speak to candidates who are based in Tokyo for our hybrid working model. As an ideal candidate, you will have: Ideally 8 to 11 years of related experience in a customer facing role, with 5 to 7 years of experience in pre-sales with enterprise software Minimum of 3 years experience with modern scripting languages (e.g. Python, Node.js, SQL) and/or popular programming languages (e.g. C/C++, Java, C#) in a professional capacity Experience designing with scalable and highly available distributed systems in the cloud and on-prem Demonstrated ability to lead architecture reviews for complex, multi-component applications and platforms, identifying risks, evaluating trade-offs, and providing clear guidance to modernize, optimize, and de-risk the solution Excellent presentation, communication, and interpersonal skills, with the ability to convey complex technical and business concepts in a clear and compelling manner to technology and business leadership Ability to partner with Sales Leadership and Account Executives on multi-threaded account and territory strategies, prioritize oppor

pythonjavanode.js
View job →
D
Datadog
📍 Remote, France• Full-time• Remote
19 days ago

We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, http://incident.io , Rootly, Jira Cloud Platform, Cortex, or similar and know how to navigate operational challenges of different

REMOTEpythonnode.jsai
View job →
AG
Adani Group
📍 Ahmedabad• Full-time
1mo ago

We are seeking a detail-oriented and analytical Data Steward – Operational Technology (OT) to join our Enterprise Data team. The ideal candidate will have a minimum of 3 to 5 years of experience in data management, data quality, and operational technology datasets. In this role, you will be responsible for managing data quality, cataloguing, asset modeling, and data standardization across operational systems (PLC, DCS, SCADA, IoT) and modern cloud data platforms like Databricks. You will work closely with plant operations, IT, data engineering, and analytics teams to ensure high data reliability, consistency, and alignment with enterprise data governance standards. Source: Adani Group | Job ID: 56028

O
1mo ago

About the Team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. Safely delivering increasingly capable AI systems requires scalable technical safeguards, clear ownership of emerging risks, rigorous deployment readiness, and close coordination across research, engineering, product, operations, legal, policy, and external partners. Our Technical Program Managers lead complex, high-stakes initiatives that turn safety commitments into deployed systems and measurable outcomes. We work across model development, infrastructure, product, and operational response to help ensure our technology is deployed responsibly and cannot be used to cause serious real-world harm. About the Role We’re seeking Technical Program Managers to drive complex product, platform, and safety initiatives across ChatGPT, API, enterprise, and related deployment environments. These roles operate at the intersection of technical strategy and execution: you will turn safety and product priorities into actionable plans, influence architectural and operational decisions, and deliver durable capabilities across model, infrastructure, application, and platform layers. Depending on the role, you may enable sensitive or high-impact model deployments, integrate safeguards into cloud and API platforms, prevent violent misuse and other serious harms, improve detection and enforcement systems, create platform solutions for safety or establish new programs as risks evolve. You will partner deeply with engineers, researchers, product managers, and operational teams while communicating technical tradeoffs and program decisions to senior leadership. You bring technical fluency, product judgment, and a strong execution record. You’re comfortable navigating ambiguity, advocating for users and developers, balancing safety with model usefulness, and leading cross-functional work with urgency, rigor, and empathy. Specific focus areas and scope will vary by opening and level. Thi

awsrestmachine learning
View job →
🔔

Get new cloud operations system administrator jobs by email

Daily job updates · Unsubscribe anytime