Jobiba hiring network

Cloud Operations System Administrator Jobs

2,329 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Storage Infrastructure team builds and operates the storage foundation behind OpenAI’s most demanding workloads. We work directly with research to design storage systems for rapidly evolving experiments, while also powering production at scale. We own the platform end to end: backend systems, user-facing services and APIs, and the control planes that manage how data is placed, moved, and retained over time. Our stack spans cloud and in-house object stores across very different workload profiles, from GPU-attached systems to dedicated storage hardware. We also build the federation layer that unifies these backends behind a simple interface and routes each workload to the right storage solution. About the Role You will help build the storage platform that powers OpenAI’s research and production systems. This is a hands-on infrastructure role for engineers who want to work on deeply technical systems at scale and own them in production. You’ll work across object storage, cross-region data movement, lifecycle management, and the federation layer that provides a unified interface across multiple backends. Much of our stack runs on Kubernetes, and we primarily build services in Rust. In this role, you will: Build and operate storage services that underpin OpenAI’s research infrastructure Develop object storage systems across cloud and in-house environments Build systems for cross-region data movement, replication, and recovery Design lifecycle management capabilities that keep data durable, available, and cost-effective Evolve the federation layer that unifies multiple backend systems behind a simple interface Improve performance, reliability, and operational excellence across the platform Collaborate closely with researchers and infrastructure teams to support rapidly evolving workloads You might thrive in this role if you: Have experience building or operating distributed systems in production Have worked on storage infrastructure, object stores, dist

awskubernetesrest
View job →
S
Snowflake
📍 United Kingdom• Full-time• From $400K/yr
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake brings AI-native observability to the Snowflake AI Data Cloud helping engineering and data teams debug, optimize, and understand systems operating at massive scale. Modern systems don’t break at GB/day they break at TB/day. Traditional, index-based observability tools weren’t built for this world. Observe by Snowflake unifies telemetry and business data, applies AI to connect context automatically, and reduces time to resolution from hours to minutes. We’re entering the next phase of growth: scaling to $1B in the next 5 years. We’ve moved successfully upmarket into complex enterprise deals, larger ACVs, and strategic customer relationships. This is a high-impact role for a seller who wants to operate at the intersection of AI, data infrastructure, and modern engineering and help define a new category. WHAT YOU’LL DO Own full-cycle enterprise sales, from self-sourced pipeline through close Build and convert pipeline in a high-velocity, high-ACV environment (now ~$400K+ enterprise deals) Lead technical, consultative conversations around AI-driven observability and data strategy Sell into engineering, SRE, and data leaders solving problems at TB-scale Drive $500K plus land-and-expand motions with multi-million dollar growth potential Work cross-functionall

aigorust
View job →
I
Instacart
📍 Canada - Remote• Remote
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Instacarts Detection Engineering team sits at the core of our Security organization, building and operating the systems that identify, surface, and respond to threats across one of North America's largest grocery technology platforms. We own the full detection lifecycle, from telemetry collection and signal design to automated response, across a complex, cloud-native environment spanning endpoint, cloud, container, and SaaS. As a Senior Detection Engineer II, you'll be a technical anchor on the team: developing high-fidelity detection logic, hunting for novel attacker techniques, and raising the bar for how we think about coverage, quality, and scale. You'll work closely with Engineering, Red Team, Incident Response, Fraud, and Trust & Safety to ensure our detections reflect real-world adversary behavior; not just signatures. We operate with a detectio

REMOTEpythonawsazure
View job →
I
Instacart
📍 United States - Remote• Remote
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Instacarts Detection Engineering team sits at the core of our Security organization, building and operating the systems that identify, surface, and respond to threats across one of North America's largest grocery technology platforms. We own the full detection lifecycle, from telemetry collection and signal design to automated response, across a complex, cloud-native environment spanning endpoint, cloud, container, and SaaS. As a Senior Detection Engineer II, you'll be a technical anchor on the team: developing high-fidelity detection logic, hunting for novel attacker techniques, and raising the bar for how we think about coverage, quality, and scale. You'll work closely with Engineering, Red Team, Incident Response, Fraud, and Trust & Safety to ensure our detections reflect real-world adversary behavior; not just signatures. We operate with a detectio

REMOTEpythonawsazure
View job →
D
Datadog
📍 Massachusetts• Full-time• From $296K/yr
1mo ago

Datadog’s Cloud Observability group is one of the core data retrieval and processing groups powering our foundational product, Infrastructure Monitoring. The group’s scope includes integration with all major hyperscalers (AWS, Azure, GCP, OCI), as well as both regional and GPU-specific cloud providers. As Director, you will own engineering for all clouds, generating more than 10 million metric points per second, managing ~40 engineers through a team of Engineering Managers. You’ll partner with Senior Directors and product leadership to shape the roadmap, not just execute against it, managing the growth of one of Datadog’s foundational teams. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You'll Do: Own engineering for all of Cloud Observability Manage ~40 engineers through a layer of Engineering Managers; this is a manager-of-managers role Shape the roadmap alongside product leadership rather than simply executing against it — push back on, iterate on, and help author the strategy for your area Drive AI adoption across the engineering org, from tooling and workflows to product features and team practices Navigate cross-team dependencies across the Agent, Telemetry Onboarding, Integrations, Action Platform, and Infrastructure Monitoring. Build and retain engineering talent in NYC, Boston, and Paris, mentor Engineering Managers toward Director readiness, and participate in the on-call rotation Who You Are: You have directly managed Engineering Managers, not just individual contributors You have deep experience with one or more cloud providers, ideally with experience operating large-scale systems in the cloud. You have a solid understanding of cloud economics, as well as how to balance performance and cos

awsazuregcp
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal a dependable, scalable foundation for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between physical hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Security Engineer, Host Assurance to help build the trust foundations for bare-metal platforms across OpenAI’s global infrastructure. This is a deeply hands-on engineering role for a builder who can design, implement, and operate the core security infrastructure that establishes trust in hardware platforms before they are eligible to run workloads. Success in this role requires strong technical judgment, the ability to work comfortably at low levels of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—including novel hardware platforms and emerging deployment models– to make the secure path the easiest path. This role is well suited for engineers who enjo

awsrestagile
View job →
C
1mo ago

As a Staff Software Engineer on Coder’s Agentic Engineering team, you’ll shape the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on while setting the team's technical direction. You’ll lead complex work, make sound architectural decisions, and help other engineers do their best work. What you’ll do here Set technical direction across Coder’s agent harness, integrations, and workflows. Design and build production systems in Go, with work across React and TypeScript where needed. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Lead complex projects from early ambiguity through production. Raise the engineering bar through design reviews, code reviews, and technical mentorship. Partner with Product and Design on clear, useful agent experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Deep experience building and operating production software systems. Strong hands-on experience with Go. Experience with React and TypeScript. Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS. A track record of setting technical direction without formal authority. Strong architectural judgment and comfort working through ambiguity. Someone who makes the engineers around them better. Our tech stack Backend: Go, Postgres Frontend: TypeScript, React Infrastructure: AWS, Kubernetes Observability: Prometheus, Grafana CI/CD: GitHub Actions Bonus tacos if you have (Tacos? If you need an ice-breaker, ask how we say thanks by giving tacos!) Experience building coding agents, developer tools, or cloud development environm

typescriptreactaws
View job →
C
Coder
📍 United Kingdom• Full-time• Remote
1mo ago

As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute. Experience building integrations across multiple model providers. Experience with AW

REMOTEtypescriptreactaws
View job →
C
1mo ago

As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. To provide substantive overlap with the team, this position must be in Eastern Time. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute.

REMOTEtypescriptreactaws
View job →
C
13 days ago

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary: As a Senior Software Development Engineer at Aetna, you will play a critical leadership role in the design, development, and continuous enhancement of enterprise-scale Provider Applications. You will drive technical solutions for complex business problems, ensure application stability, and lead cross-functional initiatives as a Project Owner. This role requires a balance of hands-on engineering expertise, technical leadership, and delivery ownership, including overseeing vendor/contractor teams, ensuring alignment with enterprise architecture, and delivering high-impact solutions that improve provider data systems and operational efficiency. Required Qualifications: 5+ years of hands-on application development experience with Python and Google Cloud Platform (GCP) 2+ years of experience leading or contributing to large-scale application development initiatives Preferred Qualifications: Experience working in Agile/SCRUM environments Proven experience in project/program management, including planning, execution tracking, and delivery management Strong organizational, leadership, and planning skills with the ability to manage multiple priorities Experience working with distributed teams and cross-functional stakeholders Prior exposure to

REMOTEpythongcp
View job →

What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.

J
Jamf
📍 Us Remote• Full-time• Remote• From $113.3K/yr
1mo ago

At Jamf, we believe in an open, flexible culture based on respect and trust. Our track record and thriving work environment all stem from the freedom we grant ourselves to get the job done right. We take pride in helping tens of thousands of customers around the globe succeed with Apple. The secret to our success lies in our connectivity, while operating with a high degree of flexibility. Work-life balance remains our priority while feeling connected is important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: The Senior Software Engineer is responsible for building the tools required to help organizations succeed with Apple. Lead others on the agile team to break down problems and apply the appropriate designs and practices to build Jamf products. Subject matter expert in various Jamf components and product offerings. Mentor and coach others while delivering new components and features with high quality and reliability. You may be required to work periodically at a Jamf office or collaborative work location with other Jamf employees in your area for certain events or moments that matter. What you can expect to do in this role : Break down customer problems into work you and the team can execute on. Independently complete tasks from start to finish with high quality. Ability to communicate technical concepts to stakeholders. Use your knowledge of Engineering best practices to ask the right questions, solve problems and build great software with a high level of quality. Produce designs for new and existing features. Clearly communicate technical concepts with others in the organization (Technical Communication, Support, Product and Cloud). Performs all job responsibilities in alignment with the core values, mission and purpose of the organization. Adheres to the highest moral, ethical and legal standards to deliver and environment that promotes respect, innovation and creativity

REMOTErestagileai
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As the Engineering Manager for Baseten's Cloud Platform team, you will directly manage a team of cloud platform engineers responsible for building the systems and processes that keep our infrastructure scalable, reliable, and efficient — from automated deployments and monitoring to performance optimization and incident response. You are a people-first leader with a strong cloud infrastructure background. You set a high bar for reliability and operational excellence, engage credibly in technical discussions and code reviews, and know how to build a culture of ownership and accountability. You'll spend most of your time close to the work: unblocking your team, shaping technical direction on day-to-day decisions, and developing your engineers. At Baseten, we work closely with our users to understand their struggles operationalizing ML — you'll keep your team connected to that mission and translate user learnings into better infrastructure. RESPONSIBILITIES Recruit, hire, and grow a high-performing team of cloud platform engineers; provide ongoing coaching, feedback, and career development through regular 1:1s. Set clear performance expectations, hold a high bar, and create an environment where engineers do their best work. Foster a culture of ownership, accountability, and continuous improvement. Drive day-to-day technical decisions through design reviews, code reviews, and architectural discussions; translate th

kubernetesci/cdgit
View job →
R
Ramp
📍 New York City• Full-time• From $10K/yr
1mo ago

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role The Security Engineering team helps make Ramp the most secure place for our customers to collect, manage, and put to work their business’ financial information Our work centers in three areas: Ramp builds products with an eye for security Ramp detects and responds to threats before they cause harm Security powers Ramp’s growth Check out our Engineering Blog for more on our tech stack, mission and values! What You’ll Do Drive our cloud security roadmap: review our cloud deployments to identify opportunities for improvement Design and build security-focused infrastructure primitives and integrate them into our existing products and development processes Lead remediation of prioritized issues across our technology stack Partner with infrastructure, data, and devops teams to design and deploy solutions that are inherently secure What You Need Minimum 5 years of experience building software Minimum 3 years of experience building in AWS (with Terraform) A strong sense of ownership: you need to drive projects from inception to scaling it in

pythonawsazure
View job →
O
1mo ago

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. Technical Program Managers at OpenAI play a key leadership role in scaling these efforts, partnering deeply with product, engineering, design, and go-to-market teams to bring ambitious ideas to life and ensure clarity and discipline in execution. About the Role We are hiring a Technical Program Manager to support OpenAI's critical AI deployments across strategic cloud partners. This role is designed for a candidate who can operate as an end-to-end owner across internal engineering teams and external partner organizations. This role will drive the technical strategy and execution required to bring OpenAI models and platform capabilities into partner environments responsibly and at scale. The work spans engineering deliverables, shared roadmaps, model launch pipelines, technical integration, launch readiness, and post-launch follow-through. You will work closely with senior leaders across OpenAI engineering, infrastructure, product, safety, security, legal, finance, and go-to-market, as well as technical counterparts at our partners. The job is to turn broad partnership commitments into concrete execution plans, align both sides on what must land, and build repeatable mechanisms for launching OpenAI capabilities on third-party platforms. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end execution for major cloud partner programs spanning model deployment, product integration, operational readiness, launch follow-through, and partner-platform adoption. Own integrated technical roadma

awsrestai
View job →
🔔

Get new cloud operations system administrator jobs by email

Daily job updates · Unsubscribe anytime