At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake is now building a world class OLTP service on Postgres. We’re hiring Senior Postgres Engineers to help build the new Snowflake Postgres service. This is an exciting opportunity to help build a large scale, multi-cloud Postgres offering with access to an existing customer base. AS A SENIOR POSTGRES ENGINEER YOU WILL: Develop orchestration layer/ control-plane of large scale databases Work with AWS, Azure, and GCP APIs Build High Availability and Disaster Recovery solutions Tune Postgres to operate at scale for some of the largest datasets in the world Secure and ensure customer data is protected Work alongside some of the brightest minds in the industry and help redefine the space by inventing at various layers of the software stack — from broad distributed systems to low-level optimizations that fundamentally change the performance characteristics of the system to cater to our customer needs. OUR IDEAL SENIOR POSTGRES ENGINEER WILL HAVE: 7+ years industry experience designing, building and developing large scale systems in production Experience building and maintaining distributed, highly available, fault tolerant services Excellent understanding of low level operating systems concepts including multi-threading, memory management, networking, storage, performance,
Jobs in United States
Cloud Operations System Administrator in United States
698 active opportunities · Updated October 2026
Showing
15 jobs
Explore current cloud operations system administrator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform & node designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each bringing together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We're searching for a highly motivated, technical leader to drive the engineering roadmap and innovation for our rack system software architecture. From firmware, kernel drivers, operating systems, networking, fabrics and associated user mode drivers + manageability software. You will work with component leads internally and engage with industry leading hyperscalar / cloud service providers on taking these products to market. What you’ll be doing: Drive the software end-to-end architecture for NVIDIA's rack-scale products Maintain deep understanding of the product portfolio and roadmap; translate forward-looking plans into clear, formal software requirements that anchor execution across the organization. Ensure high quality & reliable software; serving as a trusted architectural partner to teams requiring
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Billing & Payments Platform team builds Snowflake's central data repository and infrastructure for customer resource consumption, revenue processing, invoicing, and reporting. Our systems power Snowflake's business and enable every other engineering team — and the architectures we ship double as reference patterns for our customers building on Snowflake. Computing Snowflake's bills, at its core, is a challenging distributed systems problem: real-time usage metering across every cloud and region, and supporting an ever-evolving catalog of pricing models — including the new commercial constructs we are inventing for Cortex AI, Snowflake Intelligence, and the broader agentic AI portfolio . Our applications must meet strict requirements for accuracy, auditability, and low-latency processing. This is a deeply cross-functional role. You will partner daily with Product, Finance, Legal, Growth, Go-to-Market Systems, Snowsight UI, Cortex AI, and product engineering teams across Snowflake to deliver experiences that customers and internal stakeholders depend on every day. What You'll Do As a Senior Software Engineer on Billing Platform, you will: Own medium-sized projects end-to-end — from design through launch and operation — and contribute as a key engineer on large, multi-
$170K – $250K/yr
At Playlist, life's richest moments happen when people step away from screens to move, connect, explore, and play. We're building the definitive platform for intentional living, connecting people with inspiring experiences in fitness, wellness, and beyond. With popular brands like Mindbody and ClassPass, Playlist empowers businesses and individuals, making it effortless for aspirations to become actions. Join us in reshaping technology's role to foster meaningful, real-world connections. Mindbody equips wellness entrepreneurs with technology to support thriving businesses and create exceptional experiences. Innovation and curiosity drive our culture, connecting businesses and individuals through cutting-edge solutions. Join us if you're passionate about enhancing wellness through technology. The Role You'll Play At Mindbody, Core Engineering builds and evolves the foundational systems that help our products run reliably at scale. In this staff role, you’ll bring clarity to complex technical problems, guide architecture, and strengthen how we design, deliver, and operate the backend services that power real-world experiences. Lead cross-team technical execution, aligning architecture and delivery across multiple squads and core domains Design and evolve microservices patterns that improve reliability, performance, and maintainability Drive cloud and deployment improvements across AWS, Mindbody’s cloud platform, and our containerized deployment environment Partner with engineering and product leaders to turn ambiguous problems into clear technical plans and milestones Establish and socialize standards for service design, APIs, and relational data modeling (SQL) Strengthen monitoring and operational visibility using New Relic and Kibana, turning insights into durable system improvements Mentor and unblock engineers through design reviews, pairing, and practical guidance Reduce technical risk and complexity while balancing
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally. What You Will Be Doing: Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software. Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can b
$293K – $405K/yr
About the team Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats to global security that could scale to an extreme level of severity. Our work involves: Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. Mitigation. Keeping misuse safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework , and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. About the role As AI agents become more capable at software engineering, and automate more of our internal work, they could become a dangerous cyber threat. People in this role will help OpenAI prepare for security threats from advanced AI agent insiders. In this role, you will: Identify paths by which capable future internal AI agents could compromise OpenAI. Design security controls - focusing on measures with long lead times that benefit from advanced preparation. Stress-test defenses with AI agent evaluations and penetration tests You might thrive in this role if you: Are deeply technical across security and modern infrastructure, and are comfortable digging into the details of operating systems, cloud, containers, CI/CD, or distributed systems. Have strong software engineering skills and enjoy building prototypes yourself. Are interested in engaging with stakeholders and can do so effectively. Bonus: have experience securing cloud infrastructure, and are deeply familiar with core components of the AI stack. Compensation Range: $293K - $405K USD About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Security Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments—GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Security Software Engineer, you will design and build critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Architect and implement production-grade security services (e.g., auth services, access brokers, secure proxies, key-management infrastructure) that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD. Partner with infrastructure and research engineers to embed security into high-performance compute clusters, enabling rapid model training and deployment without compromising protection. Develop automation and detection tooling to continuously identif
About the Team The Consumer Products team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. Within Consumer Products, the camera stack is a critical sensing component. The team partners closely with electrical engineering, silicon vendors, systems, and higher-level perception and product teams to bring up new hardware, stabilize capture pipelines, and ensure camera systems are robust, debuggable, and ready for real-world deployment. This work spans early prototypes through production, with a strong emphasis on correctness, repeatability, and long-term reliability. About the Role As a Camera Firmware Engineer, you will own low-level camera enablement on custom hardware—from early board bring-up through stable production capture. You will develop and maintain the firmware and software that makes camera sensors reliable, controllable, and debuggable, forming the foundation for higher-level camera pipelines and product features. This role is highly hands-on and systems-oriented. You will work close to the hardware, diagnose real-world timing and integration issues, and build tooling that accelerates iteration across the entire camera stack. This role is based in San Francisco, CA. We follow a hybrid work model with four days per week in the office and offer relocation assistance to new employees. In This Role, You Will Bring up new camera sensors and modules on prototype and production boards, including link stability, sensor control, and correct power, reset, and clock sequencing. Develop and maintain low-level camera software, including sensor drivers, board configuration, and camera subsystem integration across hardware revisions. Enable and validate core capture paths for development and production, including RAW capture for debugging, still capture, and hardware-
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: At Modal, we sell cloud services atop which our customers run their critical production systems. As a rapidly growing new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform, customer base, and our team. This role is for people who are deep systems thinkers, love stacking nines, and thrive from making others move faster at scale. Responsibilities include: Identifying architectural changes to improve reliability and performance. Fostering a culture of reliability across Modal’s engineering organization. Defining and implementing operational processes such as deployments, upgrades, etc. Operating systems like Kubernetes, Postgres, Redis, etc. Participating in on-call rotations, and responding to production incidents. Requirements: 5+ years of experience writing high-quality production code. 2+ years of
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: We're hiring a Product Partnerships Manager to own Replit's most important strategic partnerships end-to-end: from identifying the opportunity to shipping the outcome and measuring the impact. This is a product-focused role where you will work with our consumer technology partners, such as Stripe, Shopify, Google, and the broader connector ecosystem. You'll work at the intersection of product strategy, ecosystem thinking, and deal execution. You'll need to be a serious Replit power user. You'll speak credibly with engineers. You'll negotiate with senior partner stakeholders. And you'll be the person accountable for turning ambiguous ecosystem opportunities into shipped product and business outcomes. What You'll Do Develop a clear point of view on Replit's partner landscape and build a prioritized pipeline of high-leverage opportunities across consumer tooling, AI tooling, cloud and infrastructure, and payments technology partners. Lead partner conversations from early exploration through joint product thesis, business case, commercial terms, launch plan, and post-launch iteration. Partner with Product, Engineering, Partner Engineering, Legal, Finance, Marketing, and Sales to turn agreements into shipped outcomes. Work directly with Partner Engineers to scope integrations, demos, prototypes, reference apps, and partner enablement assets. Define success metrics before every launch activation: retained usage, apps created, deployments, revenue, partner-sourced users, and use them to decide when to scale, iterate, or sunset a partnership. Build lightweight operating systems: partner scorecards, launch checklists, partner roadmap tracking, and repeatable frameworks for evaluating new opportunities. Represent
About the Team The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. This role is based in San Francisco, CA. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change This role is exclusively based in our San Francisco HQ. We offer relocation assistance to new employees. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely depl
About the Team The Storage Infrastructure team builds and operates the storage foundation behind OpenAI’s most demanding workloads. We work directly with research to design storage systems for rapidly evolving experiments, while also powering production at scale. We own the platform end to end: backend systems, user-facing services and APIs, and the control planes that manage how data is placed, moved, and retained over time. Our stack spans cloud and in-house object stores across very different workload profiles, from GPU-attached systems to dedicated storage hardware. We also build the federation layer that unifies these backends behind a simple interface and routes each workload to the right storage solution. About the Role You will help build the storage platform that powers OpenAI’s research and production systems. This is a hands-on infrastructure role for engineers who want to work on deeply technical systems at scale and own them in production. You’ll work across object storage, cross-region data movement, lifecycle management, and the federation layer that provides a unified interface across multiple backends. Much of our stack runs on Kubernetes, and we primarily build services in Rust. In this role, you will: Build and operate storage services that underpin OpenAI’s research infrastructure Develop object storage systems across cloud and in-house environments Build systems for cross-region data movement, replication, and recovery Design lifecycle management capabilities that keep data durable, available, and cost-effective Evolve the federation layer that unifies multiple backend systems behind a simple interface Improve performance, reliability, and operational excellence across the platform Collaborate closely with researchers and infrastructure teams to support rapidly evolving workloads You might thrive in this role if you: Have experience building or operating distributed systems in production Have worked on storage infrastructure, object stores, dist
From $296K/yr
Datadog’s Cloud Observability group is one of the core data retrieval and processing groups powering our foundational product, Infrastructure Monitoring. The group’s scope includes integration with all major hyperscalers (AWS, Azure, GCP, OCI), as well as both regional and GPU-specific cloud providers. As Director, you will own engineering for all clouds, generating more than 10 million metric points per second, managing ~40 engineers through a team of Engineering Managers. You’ll partner with Senior Directors and product leadership to shape the roadmap, not just execute against it, managing the growth of one of Datadog’s foundational teams. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You'll Do: Own engineering for all of Cloud Observability Manage ~40 engineers through a layer of Engineering Managers; this is a manager-of-managers role Shape the roadmap alongside product leadership rather than simply executing against it — push back on, iterate on, and help author the strategy for your area Drive AI adoption across the engineering org, from tooling and workflows to product features and team practices Navigate cross-team dependencies across the Agent, Telemetry Onboarding, Integrations, Action Platform, and Infrastructure Monitoring. Build and retain engineering talent in NYC, Boston, and Paris, mentor Engineering Managers toward Director readiness, and participate in the on-call rotation Who You Are: You have directly managed Engineering Managers, not just individual contributors You have deep experience with one or more cloud providers, ideally with experience operating large-scale systems in the cloud. You have a solid understanding of cloud economics, as well as how to balance performance and cos
About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal a dependable, scalable foundation for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between physical hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Security Engineer, Host Assurance to help build the trust foundations for bare-metal platforms across OpenAI’s global infrastructure. This is a deeply hands-on engineering role for a builder who can design, implement, and operate the core security infrastructure that establishes trust in hardware platforms before they are eligible to run workloads. Success in this role requires strong technical judgment, the ability to work comfortably at low levels of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—including novel hardware platforms and emerging deployment models– to make the secure path the easiest path. This role is well suited for engineers who enjo
As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. To provide substantive overlap with the team, this position must be in Eastern Time. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute.
Other cities to consider
More places hiring for this role
Get new cloud operations system administrator jobs in United States by email
Daily job updates · Unsubscribe anytime