ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Observability team within the Infrastructure organization. As an early member of the Observability Team, you will be pivotal in building and shaping the observability experience for our internal and external customers. By joining this team, you’ll have a direct impact on the reliability and operational excellence of Basetens product systems. As Baseten scales its infrastructure across different cloud providers and diverse hardware, the volume and complexity of operational data is growing by orders of magnitude. This team is responsible for building high-throughput ingest pipelines, cost-efficient storage, and agentic diagnostic tools to ensure that we can detect, diagnose, and resolve issues in minutes rather than hours, even as the systems they operate become more complex. RESPONSIBILITIES Design and build scalable telemetry ingest and storage pipelines for metrics, logs, and traces across Baseten’s multi-cloud infrastructure Own and evolve core observability platforms, driving migrations and architectural improvements that improve reliability, reduce cost, and scale with organizational growth Build instrumentation libraries, SDKs, and integrations that make it easy for engineering teams to emit high-quality telemetry from their services Drive alerting and SLO infrastructure that enables teams to define, monitor, and respond to reliabi
Jobs in United States
Cloud Operations System Administrator in United States
698 active opportunities · Updated October 2026
Showing
15 jobs
Explore current cloud operations system administrator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. Remote (US East Coast preferred, for timezone coverage) About the team Cloud Infrastructure owns the platform every Synthesia product runs on — AWS, Kubernetes, MongoDB, Temporal, our observability stack, and the vendor and cost relationships underneath them. We're a small, high-leverage team scaling toward a domain-ownership model: small groups that both build and operate the systems they're accountable for. The role We're hiring a dedicated SRE to take real ownership of operational excellence across Cloud Infrastructure. Today, too much critical operational knowledge — vendor relationships, cost management, and incident response — lives with one or two people. Your mission is to take genuine ownership of those domains, make them resilient to any single person, and raise the bar on how reliably we run. This is not simply a ticket-queue or keep-the-lights-on role. You'll own domains end to end: understand them deeply, operate them well, and build the automation and tooling that make them boring . We deliberately pair operational and engineering work so the role grows rather than narrows. What you'll own Incident management & operational excellence — take custody of the incident process: on-call quality, resp
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: Join a team that builds robust, real-time distributed systems for a cutting-edge database. We care about performance, reliability, scalability, and most of all learning and having fun together. Whether you’re a seasoned coder or just getting started, if you’re passionate about technology and eager to learn, you’ll fit right in. Who we are: We show up to work, ready to collaborate and build technologies that make a difference, with people who genuinely care. We chase improvements such as tail latencies, bytes throughput, cache hit rate, and operational cost efficiency. We believe learning is ongoing and that even the most complex problems can have simple solutions. What You’ll Do: Collaborate with teammates to design and build database features that power AI applications. Learn how to tune performance and support reliability in distributed systems (don’t worry, we’ll guide you). Help Pinecone run smoothly on popular cloud providers. Take ownership of your work and grow your skills every day. Have fun. Who You Are: 5+ years of work experience - programming in Rust, Go, C++, or a comparable language. You’re genuinely curious about distributed systems and eager to dive deep into technical challenges. You approach problems with creativity and persistence, and you’re comfortable asking thoughtful questions or seeking feedback. You’re excited to learn, value constructive feedback, and appreciate mentorship. Bonus Points: You have hands-on experience with cloud platforms (AWS, GCP, Azure) or have demonstrated an ability to pick u
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake is a high-growth, cloud-native data platform company committed to empowering enterprises to achieve their full potential. With a culture built on impact, innovation, and collaboration, we offer an environment where you can build large-scale systems, move fast, and take your technology career to the next level. We are seeking a highly talented and experienced Staff Software Engineer to join the Unistore team in our Database Engineering org . The Unistore team is building a unified hybrid data platform that dissolves the boundaries between operational and analytical workloads. Our core product, Hybrid Tables, enables enterprises to manage both transactional and analytical data in a single place. By eliminating the need to move data between separate operational databases and analytical warehouses, we simplify architectures, reduce ETL complexity, and unlock high-performance, real-time application serving—all while maintaining Snowflake's industry-leading scale, performance, and ACID compliance. Joining this team means architecting the performance, scalability, and service infrastructure that will empower the world's largest enterprises to seamlessly govern their transactional and analytical data in one place. If you are passionate about building industry-leading tech
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake’s Release Engineering team builds and operates the systems that safely deliver infrastructure, platform, and product changes to production at global scale. We own the release platforms, rollout orchestration, and safety mechanisms that allow engineering teams across Snowflake to ship quickly while minimizing operational risk. Our mission is to make production deployments fast, safe, self-service, increasingly autonomous, and augmented by AI-driven intelligence and automation. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale multi-cloud infrastructure orchestration. At Snowflake, Release Engineering is a platform engineering function focused on building the systems, abstractions, and automation that make software delivery safe, scalable, and efficient across the company. In this role, you will Design and build continuous deployment and rollout infrastructure that safely ships changes across Snowflake’s large-scale, multi-cloud production environment. Build and evolve platform capabilities for progressive delivery, including staged rollouts, canarying, automated health checks, rollback controls, and guardrails that reduce blast radius during production change events. Improve engineering velocity by removi
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Vice President, Software Engineering Overview Decision Stream is Mastercard's next-generation AI-native decisioning platform, designed to power intelligent, real-time decisions across fraud, authentication, payments, and risk. Built with a startup mindset and enterprise-scale ambition, the platform combines innovation, speed, and engineering excellence to redefine decisioning across Mastercard's global ecosystem. We are seeking a visionary and hands-on VP of Software Engineering to help build and scale the platform. This leader will partner closely with Product, Architecture, AI, and Platform Engineering teams to drive technology strategy, architecture, engineering execution, and organizational growth. What You'll Do Lead Through Technical Excellence • Serve as a senior technology leader and role model for engineering teams. • Drive architecture, design, and technology decisions across distributed systems, streaming, AI, and cloud-native platforms. • Engage deeply with engineers, architects, and product leaders to solve complex technical challenges. • Influence engineering standards, software quality, and operational excellence. Build and Scale the Platform • Help shape and deliver a highly scalable, resilient, and secure decisioning platform operating at Mastercard scale. • Balan
From $128K/yr
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This position may be a hybrid or fully remote position, as decided by your manager. If designated as hybrid, you’ll divide your time between working remotely from your home and an office location, so you should live within commuting distance. If designated as remote, you’ll be working remotely from your home and may occasionally visit a GoDaddy office to meet with your team for events or meetings. Your hiring manager can share more about this role’s hybrid or remote designation. This position is not eligible to be performed in Alaska, Mississippi, North Dakota, or the Virgin Islands. GoDaddy is not currently considering candidates for this role in California, Seattle, or NYC. Join Our Team Join a team powering secure, scalable email services for millions of customers worldwide! As part of GoDaddy's Professional Email team, you'll solve complex challenges in distributed systems, cloud infrastructure, security, and AI while modernizing critical platforms that businesses rely on every day. If you enjoy owning impactful systems, working across a diverse technology stack, and building innovative solutions at scale, you'll feel right at home here. What you'll get to do... Design, build, and maintain highly available, scalable APIs and services used by millions of customers Deploy, manage, and optimize cloud infrastructure in AWS Architect and implement modern solutions that improve performance, reliability, and security Leverage AI technologies to enhance development workflows and create innovative customer experiences Monitor, troubleshoot, and resolve complex production issues using modern observability and monitoring tools Drive continuous improvement through automation, modernization, and operational excellence C
About the Team The Private Computing team works across product, engineering, security, and safety to build advanced privacy products and infrastructure at OpenAI. Our mission is to provide world-class security features to users so their private data remains private, even from OpenAI. We use technologies like confidential computing, trusted execution environments, and end-to-end encryption to ship product features across ChatGPT, the API, and our future consumer devices. About the Role We’re looking for software engineers to design, build, and scale novel privacy features and infrastructure across ChatGPT, API, and future consumer devices. In this role, you will: Ship fast while balancing difficult trade-offs in complex domains Build core abstractions for trusted execution environments and end-to-end-encryption Build product features for private inference and storage across ChatGPT, API, and future consumer devices Update build systems to increase trust and verifiability Integrate with safety and integrity infrastructure Operate systems at scale with high reliability, including an on-call rotation Collaborate with a diverse set of cross-functional teams across product, engineering, security, safety, policy, and legal You might thrive in this role if you: Care deeply about user privacy and security Have 5+ years of experience in professional software engineering Have experience building and scaling confidential computing or encryption technologies in production environments Have experience with Kubernetes and cloud orchestration systems Take pride in building and operating scalable, reliable, secure systems Can collaborate well and drive alignment in the face of difficult trade-offs Are comfortable with ambiguity and rapid change Workplace & Location This role is based in San Francisco, CA. We follow a hybrid model with 4 days a week in the office and offer relocation assistance to new employees. About OpenAI OpenAI is an AI research and deployment company dedicat
About the Team The Coding team is reimagining how software is built in the AI era. We build tools and workflows that help software engineers work faster, tackle more ambitious projects, and spend less time on repetitive tasks. AI has already transformed how code is written, but software engineering extends far beyond coding. Our mission is to apply AI across the entire software development lifecycle (SDLC) — from design and implementation to code review, testing, debugging, issue remediation, maintenance, documentation, and user support. The team is also responsible for developer-facing Codex experiences including the Codex IDE Extension and the terminal interface, which are used daily by developers ranging from individual open-source contributors to some of the world’s largest engineering organizations. The team also works closely with the open-source software community, building tools that help maintainers and contributors manage increasingly complex projects. We believe AI can make open-source development more sustainable by reducing the operational burden of reviewing contributions, triaging issues, maintaining quality, and supporting growing communities. By building the future of software development, we're helping advance OpenAI's mission of ensuring that the benefits of AI reach people around the world. About the Role We’re hiring a Full Stack Software Engineer to help invent the next generation of AI-powered software development workflows. “Full stack” in this role means much more than traditional frontend and backend development. You'll own complete product experiences, spanning user interfaces, workflow orchestration, agent and prompt design, backend systems, and cloud infrastructure. This is a highly product-oriented role. You'll work directly on the workflows developers use every day, identifying bottlenecks and rethinking how software gets built in a world where AI agents are active participants in the development process. The features you ship will inf
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake is redefining how enterprises bring data, applications, and AI together. As autonomous workflows begin taking actions on behalf of users, identity has officially become the new security perimeter. Every single interaction—whether initiated by a human, an application, a workload, or an AI agent—must be continuously authenticated, authorized, and governed. To unlock this next generation of enterprise software, Snowflake requires an identity platform that extends far beyond traditional workforce authentication to seamlessly support machine identities, fine-grained delegation, and policy-driven access at cloud scale. We are looking for a hands-on, high-impact Product Leader to define and build this foundational trust layer. Operating at a highly strategic intersection of product, engineering, partnerships, and executive-level customer engagement , you will own the core infrastructure that allows complex enterprise systems to securely interact, reason over sensitive data, and safely execute actions. AS A PRINCIPAL PRODUCT MANAGER AT SNOWFLAKE, YOU WILL : Set Portfolio Strategy: Own and define the long-term product strategy and roadmap for Snowflake’s IAM ecosystem, factoring in market-shifting competitive trends and technical evolutions. Build for AI era: Architect IAM
Cloud Analyst (Mid-Level or Senior) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations is currently seeking a Cloud Analyst (Mid-Level or Senior) **Sign on Bonus Potential** to join the team in Berkeley, MO, Dayton Beach, FL; or Seattle, WA . The Cloud Analyst plays a key role in supporting the testing and execution of automated scripts across a range of cloud technologies, including Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP). In this position, the selected candidate will help ensure smooth, reliable cloud operations while providing thoughtful recommendations to improve performance, efficiency, and automation. This role also requires strong customer relationship skills, with a focus on clear communication, responsiveness, and delivering high-quality support that builds trust and confidence. Join a team driving innovation in cloud technology and automation, and help shape secure, efficient, and reliable digital solutions. Position Responsibilities: Develop and maintain comprehensive documentation that may include detailed technical documents Create process flows, business requirements, functional specifications, and user guides, with a strong emphasis on cloud-based systems Collaborate with stakeholders to gather, analyze, and validate business and technical requirements related to cloud infrastructure and cost management Design and document process flows to support cloud service management, price tracking, and operational improvements Support testing activities by creating and executing test plans, test cases, and automa
From $118.4K/yr
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Sr Production Engineer to join our team. This role is available as a hybrid opportunity 3 days a week in San Jose, CA or Bellevue, WA reporting to Production Engineering in the Cloud Infrastructure & Operations department. Join Zscaler to be a force multiplier for the reliability of a global platform processing 200+ billion transactions daily across tens of millions of enterprise users. In this role, you will provide the technical vision and hands-on execution to drive an "automation-first" culture across the company. By maturing our observability and architectural standards, you will directly reduce our Mean Time to Mitigate (MTTM) and shape the scalability of our globally distributed, multi-cloud infrastructure. What you’ll do (Role Expectations) Implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments Drive an "automation-first" c
From $102.4K/yr
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Production Engineer to join our team. This role is available as a hybrid opportunity 3 days a week in San Jose, CA or Bellevue, WA reporting to Production Engineering in the Cloud Infrastructure & Operations department. Join Zscaler to be a force multiplier for the reliability of a global platform processing 200+ billion transactions daily across tens of millions of enterprise users. In this role, you will provide the technical vision and hands-on execution to drive an "automation-first" culture across the company. By maturing our observability and architectural standards, you will directly reduce our Mean Time to Mitigate (MTTM) and shape the scalability of our globally distributed, multi-cloud infrastructure. What you’ll do (Role Expectations) Implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments Drive an "automation-first" cult
$170K – $250K/yr
A Career with Point72’s Technology Team As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. What you’ll do Optimize cloud financial operations to maximize value from cloud investments, including rapidly growing artificial intelligence (AI) and machine learning workloads Provide actionable insights on cloud spend, SaaS license optimization, and emerging AI cost drivers, including model inference and usage-based consumption Implement tooling, tagging standards, and processes that improve cost visibility and optimization across cloud, SaaS, and AI workloads Monitor large language model API consumption and GPU-intensive infrastructure to identify cost trends, anomalies, and optimization opportunities Build financial models to forecast cloud, SaaS, and AI expenditures for budgeting cycles, commitment decisions, and vendor negotiations Design cost allocation, tagging, showback, and chargeback models that attribute spend to the teams, applications, and use cases driving it Educate engineering and business owners on cloud financial management practices th
Other cities to consider
More places hiring for this role
Get new cloud operations system administrator jobs in United States by email
Daily job updates · Unsubscribe anytime