Jobiba hiring network

Cloud Operations System Administrator Jobs

2,288 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

SA
Scale AI
📍 San Francisco• Full-time• From $252K/yr
17 days ago

About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. Role Overview As a Forward Deployed AI Engineering Manager on our Enterprise team, you'll be the technical bridge between Scale AI's cutting-edge AI capabilities and our most strategic customers. You'll work with enterprise clients to understand their unique challenges, lead a team that architects specific AI solutions, and ensure successful deployment and adoption of AI systems in production environments. This is a Management role that combines deep engineering and AI expertise, leading a team, and working on customer-facing problems. You'll work directly with customer engineering teams to integrate AI into their critical workflows. Key Responsibilities Customer Integration & Deployment Partner directly with enterprise customers to understand their technical infrastructure, data pipelines, and business requirements Design and implement custom integrations between Scale AI's platform and customer data environments (cloud platforms, data warehouses, internal APIs) Build robust data connectors and ETL pipelines to ingest, process, and prepare customer data for AI workflows Deploy and configure AI models and agents within customer security and compliance boundaries AI Agent Development Develop production-grade AI agents tailored to customer use cases across domains like customer support, data analysis, content generation, and workflow automation Architect multi-agent systems that orchestrate between different models, tools, and data sources Implement evaluation frameworks to measure agent performance and iterate toward business objectives Design human-in-the-loop workflows and feedback mechanisms for continuous agent improvement Prompt Engineeri

pythonawsazure
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $288K/yr
17 days ago

About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. Role Overview As a Senior Staff Frontier Agents Engineer on our Enterprise team, you'll be the technical bridge between Scale AI's cutting-edge AI capabilities and our most strategic customers. You'll work with enterprise clients to understand their unique challenges, architect custom AI solutions, and ensure successful deployment and adoption of AI systems in production environments. This is a hands-on technical role that combines deep engineering expertise with customer-facing problem solving. You'll work directly with customer engineering teams to integrate AI into their critical workflows. Key Responsibilities Customer Integration & Deployment Partner directly with enterprise customers to understand their technical infrastructure, data pipelines, and business requirements Design and implement custom integrations between Scale AI's platform and customer data environments (cloud platforms, data warehouses, internal APIs) Build robust data connectors and ETL pipelines to ingest, process, and prepare customer data for AI workflows Deploy and configure AI models and agents within customer security and compliance boundaries AI Agent Development Develop production-grade AI agents tailored to customer use cases across domains like customer support, data analysis, content generation, and workflow automation Architect multi-agent systems that orchestrate between different models, tools, and data sources Implement evaluation frameworks to measure agent performance and iterate toward business objectives Design human-in-the-loop workflows and feedback mechanisms for continuous agent improvement Prompt Engineering & Optimization Create sophisticate

pythonawsazure
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $216K/yr
17 days ago

Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Platform Engineering team. In this role, you will support the design and development of shared platforms used across Scale. This includes designing our foundational data platforms and lifecycle, architecting Scale’s core cloud infrastructure and orchestration stack, and redefining how engineers develop, build, test, and deploy software at Scale. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the design, and implementation of our foundational platforms and systems, working closely with stakeholders and internal customers to understand and refine requirements. Collaborating with cross-functional teams to define, design, and deliver new features. Proactively identifying opportunities for, and driving improvements to, current p

sqlmongodbaws
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $216K/yr
17 days ago

The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Senior Software Engineer, you will lead the development of a vertical feature or a horizontal capability to include defining requirements with stakeholders and implementation until it is accepted by the stakeholders. You will: Lead the design and implementation of scalable backend systems and distributed architectures for Federal customers. Manage the full lifecycle of feature development from requirement definition to deployment on classified networks. Direct the orchestration of asynchronous agent fleets to meet mission requirements. Lead customer engagements to translate mission needs into technical requirements. Own the communication with stakeholders to ensure implementation meets defined acceptance criteria. Conduct technical reviews and identify risks within machine learning infrastructure and model serving. Drive the platform roadmap by providing technical specifications for Federal product offerings. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in developing and deploying applications in a cloud-native environment. Understanding of containerization (e.g., Docker) and contai

awsazuregcp
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $216K/yr
17 days ago

Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Platform Engineering team. In this role, you will support the design and development of shared platforms used across Scale. This includes designing our foundational data platforms and lifecycle, architecting Scale’s core cloud infrastructure and orchestration stack, and redefining how engineers develop, build, test, and deploy software at Scale. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the design, and implementation of our foundational platforms and systems, working closely with stakeholders and internal customers to understand and refine requirements. Collaborating with cross-functional teams to define, design, and deliver new features. Proactively identifying opportunities for, and driving improvements to, current p

sqlmongodbaws
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
17 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As a Director, Strategy & Solutions in our GTM team, this role defines how Tenstorrent shows up in the market for AI/ML workloads—from positioning and messaging to real customer solutions. Sitting at the intersection of product, sales, engineering, and marketing, this role turns deep technical capability into clear, compelling value across industries and use cases. This role is hybrid OR remote, based out of The United States. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced at connecting complex AI/ML systems to real customer outcomes and business value. Comfortable acting as a technical storyteller for developers, architects, and executive audiences. Naturally curious and forward-looking about customer workloads, competitive dynamics, and emerging AI trends. What We Need Ownership of GTM positioning and messaging for Tenstorrent’s AI hardware, software stack, and solutions. Development of application - and vertical-specific value propositions across enterprise, cloud, and regulated markets. Value proposition & GTM strategy, including creation of technical collateral (whitepapers, sales decks, benchmarks, dem

awsairust
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
17 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As a Manager, Strategy & Solutions in our GTM team, this role defines how Tenstorrent shows up in the market for AI/ML workloads—from positioning and messaging to real customer solutions. Sitting at the intersection of product, sales, engineering, and marketing, this role turns deep technical capability into clear, compelling value across industries and use cases. This role is hybrid OR remote, based out of The United States. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced at connecting complex AI/ML systems to real customer outcomes and business value. Comfortable acting as a technical storyteller for developers, architects, and executive audiences. Naturally curious and forward-looking about customer workloads, competitive dynamics, and emerging AI trends. What We Need Ownership of GTM positioning and messaging for Tenstorrent’s AI hardware, software stack, and solutions. Development of application - and vertical-specific value propositions across enterprise, cloud, and regulated markets. Value proposition & GTM strategy, including creation of technical collateral (whitepapers, sales decks, benchmarks, demo

awsairust
View job →
G
17 days ago

About Graphcore At Graphcore, we’re building the future of AI compute. We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence . Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible. Responsibilities and Duties Develop, own, and maintain tools and services to support AI research and engineering teams Deploy and maintain services with Kubernetes and Docker Manage our Cloud Infrastructure using tools such as Terraform Candidate Profile Essential:

pythonjavaaws
View job →
C-
CLEAR - Corporate
📍 New York• Full-time• $180K – $220K/yr
17 days ago

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We’re looking for an experienced Software Engineer to help us build the next generation of products which will go beyond just ID and enable our members to leverage the power of a networked digital identity. As a Software Engineer at CLEAR, you will participate in the design, implementation, testing, and deployment of applications to build and enhance our platform- one that interconnects dozens of attributes and qualifications while keeping member privacy and security at the core. A brief highlight of our tech stack: Java / Kafka / Postgres AWS cloud What you'll do: Advance our capabilities across a wide array of industries and domains and gain hands-on experience with privacy, security, data modeling and architecture Develop and deliver code across the full stack, driving engineering excellence by defining to best practices in testing, documentation and observability Partner with product and other stakeholders to uncover requirements, to innovate, and to solve complex problems Have a strong sense of ownership, responsible for architectural decision-making and striving for continuous improvement in technology and processes at CLEAR What you’re great at: 3+ years of backend software development experience in Java Working with cloud-based application development, and being fluent in at least a few of: Cloud service providers like AWS Containerization technologies like Docker and Kubernetes Collaboration, integration, and deployment tools like Github, Argo, and Jenkins Arti

javaawsdocker
View job →
C-
CLEAR - Corporate
📍 New York• Full-time• $180K – $220K/yr
17 days ago

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We’re looking for a Data Engineer II to help us build the next generation of products which will go beyond just ID and enable our members to leverage the power of a networked digital identity. As a Data Engineer at CLEAR, you will participate in the design, implementation, testing, and deployment of applications to build and enhance our platform- one that interconnects dozens of attributes and qualifications while keeping member privacy and security at the core. A brief highlight of our tech stack: SQL / Python / Looker / Snowflake / dbt What you'll do: Build a scalable data system in which Analysts and Engineers can self-service changes in an automated, tested, secure, and high-quality manner Build processes supporting data transformation, data structures, metadata, dependency and workload management Develop and maintain data pipelines to collect, clean, and transform data (owning end to end data product from ingestion to visualization) Develop and implement data analytics models Partner with product and other stakeholders to uncover requirements, to innovate, and to solve complex problems Have a strong sense of ownership, responsible for architectural decision-making and striving for continuous improvement in technology and processes at CLEAR What you're great at: 4+ years of data engineering experience Working with cloud-based application development, and be fluent in at least a few of: Cloud services providers like AWS Data pipeline orchestration tools like Airflow, Dagster, Luigi, etc Big data tools like Spark, Ka

pythonsqlaws
View job →
C-
CLEAR - Corporate
📍 New York• Full-time• $225K – $300K/yr
17 days ago

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Fullstack Software Engineer on CLEAR’s Healthcare team, you will build and scale secure, interoperable identity and data solutions that connect patients, providers, and partners. You’ll operate at the intersection of modern web platforms, healthcare interoperability standards, and high-assurance identity systems powering frictionless, trusted healthcare experiences nationwide. A brief highlight of our tech stack: Python / React / Typescript AWS cloud What you’ll do: Design and deliver secure, scalable fullstack solutions that integrate with enterprise EHR systems and national health information exchange frameworks Build and maintain healthcare data integrations leveraging FHIR (RESTful APIs/JSON) and HL7 v2 messaging to enable compliant, real-time data exchange Develop identity resolution and patient matching capabilities using identifiers such as MRNs and NPIs to ensure integrity across disparate clinical systems Partner with Engineering, Security, Product, and Health Information Management teams to implement compliant, audit-ready workflows for regulated healthcare processes Collaborate with external vendors (e.g., Epic Technical Services) to troubleshoot integration issues, manage deployments across TST/PRD environments, and ensure production reliability How you’ll measure success: Successful delivery and stability of FHIR/HL7 integrations across healthcare partners Reduction in data integrity issues related to patient matching and identity resolution High system uptime and successful production deployments across tiered

typescriptpythonreact
View job →
C-
CLEAR - Corporate
📍 New York• Full-time• $175K – $215K/yr
17 days ago

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We’re looking for a Senior Software Engineer to join our Network Engineering team to accelerate building and scaling our innovative systems that support our growing identity platform. In this role, you will build the next-generation infrastructure that underpins all systems at CLEAR. The ideal candidate for this role will approach challenges with an eye toward reliability, simplicity, and scalability. What You'll Do: Develop and maintain a streamlined process for engineers to effortlessly build and deploy scalable and reliable software-defined networking solutions on AWS. Enhance our compute platform (Kubernetes) by integrating new functionalities and features, focusing on AWS networking services and concepts such as VPCs, Route Tables, Security Groups (SGs), ALBs/ELBs, and Route53, as well as implementing Kubernetes networking solutions like service mesh (Istio) to optimize service communication and management. Collaborate across engineering teams to advocate for and implement best practices in observability, utilizing tools like Splunk or Datadog to ensure robust network monitoring. Act as a product owner for our infrastructure, collecting feedback and requirements from engineering teams to address pain points and develop solutions, particularly in the realm of AWS networking and cloud-native design principles. What you're great at: 6+ years of extensive experience in infrastructure and platform development, particularly in software-defined networking and AWS cloud services. Proficient in writing production-grade softwar

pythonawskubernetes
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
17 days ago

Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is growing quickly. You will play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features, while also balancing scalability, reusability, and performance. Role Overview: We're looking for a Staff Platform Engineer to serve as the technical backbone of our Engineering organization. You'll own the technical strategy, and delivery of our platform — spanning architecture, DevOps, SRE, security, Dev-ex. This is a hands-on staff level role: you'll set technical direction, drive cross-team alignment, and be the senior escalation point for platform challenges. What you’ll do: Drive platform reliability, scalability, security, and cost efficiency across all environments. Technical Leadership: Provide technical leadership for platform components, Influence the technical strategy and architecture of our cloud platform, from CI/CD pipelines to observability and incident response. Design and implement platform components and reusable integration patterns that minimize custom development efforts, reduce the time spent on repetitive tasks, and ensure that integrations scale across multiple healthcare systems Partner closely with Architecture, DevOps, SRE, and Security teams to deliver cohesive platform solutions Cross-Functional Collaboration: Work closely with product teams, and solutions architects to understand integration needs and ensure the platform meets current and future business requirements. Serve as a senior escalation point for infrastructure and platform incidents Establish frameworks for: AI governance and compliance. Observability of systems. Traceability of decisions and outputs. Ensure enterprise readiness with security, auditability, and reliability in production environments. Security & Compliance : Ensure all p

awsci/cdgit
View job →
SL
Sauce Labs Inc.
📍 Gurugram• Full-time
17 days ago

About Us: Sauce Labs is the world’s largest full-lifecycle, test automation platform, and the company behind Selenium. Trusted by 80% of the world’s top ten largest financial institutions and over 300,000 enterprise users, Sauce Labs provides the only AI platform capable of turning business intent into autonomous testing and quality assurance. With a proprietary dataset of 8.7 billion test runs, Sauce Labs empowers the Fortune 2000 to bridge the gap between AI-driven code generation and enterprise-grade software quality. Learn more at saucelabs.com . The Role: We are seeking an innovative and experienced AI Architect to join our engineering leadership team. This is a strategic role that will be instrumental in designing and building the next generation of AI-powered features for our continuous testing platform. You will be responsible for architecting scalable and robust AI solutions that transform how our customers gain insights from their test data and production environments, and how they create tests. Responsibilities: Define AI Architecture: Lead the design and architecture of cutting-edge AI/ML solutions for new product offerings, ensuring scalability, performance, quality and reliability within a cloud-native environment. AI-Powered Insights (Test & Production): Architect AI systems to derive actionable insights from vast quantities of test run logs and analytics data. This includes identifying patterns, anomalies, and performance trends. Production Error Reporting Integration: Design AI solutions that integrate with our existing error reporting product to analyze production issues for mobile and web applications, providing deeper understanding and predictive capabilities. Unified Data Intelligence: Develop architectures for combining insights from both test runs and production data, creating a holistic view of application quality and user experience. Automated Failure Analysis & Remediation: Architect AI models and systems t

gcpmicroservicesmachine learning
View job →
DU
17 days ago

Role Overview We are seeking a Staff Simulation Engineer to build an end-to-end aerial autonomy simulation stack at DoorDash Labs. This is a highly technical, hands-on leadership role focused on defining and implementing the simulation architecture that underpins autonomy development, validation, CI/CD testing, and pilot training. You will operate as the technical authority for simulation: owning core architecture decisions, developing key components yourself, and setting engineering standards. You will build and mentor a small, high-caliber simulation team while remaining deeply involved in implementation and system design. This role is ideal for someone who has built simulation systems from first principles, understands simulator internals deeply, and is excited to create a world-class platform from scratch. Key Responsibilities Architect and implement an end-to-end simulation stack for aerial autonomy at DoorDash Labs.. Develop high-fidelity simulation capabilities, including: Flight dynamics modeling Contact modeling and constraint handling Sensor and perception simulation Autonomy software-in-the-loop (SITL) integration Design and implement scalable simulation infrastructure to support: Regression testing in CI/CD pipelines Continuous validation of flight autonomy and autopilot software stack Mission-level testing and scenario generation Build cloud-deployed simulation systems to enable large-scale parallel testing and pilot training. Partner closely with autonomy, controls, and aircraft teams to ensure simulation fidelity and validation alignment. Establish technical direction, architecture standards, and performance benchmarks for simulation. Mentor and grow a small team of simulation engineers while remaining deeply hands-on. Required Qualifications Master’s or PhD in Computer Science, Electrical Engineering, Mechanical Engineering, Robotics, Aerospace Engineering, or a related field. 10+ years of experience in robotics or physics-based simulation. Deep expe

awsci/cdgit
View job →
🔔

Get new cloud operations system administrator jobs by email

Daily job updates · Unsubscribe anytime