Jobs in United States

Cloud Operations Engineer in United States

698 active opportunities · Updated October 2026

Explore current cloud operations engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

SF
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -100%

From $88.1K/yr

Quick readStrong listing-quality and freshness signals

About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As a Platform Engineer, you will contribute to building and improving Stitch Fix’s cloud-native infrastructure and internal developer tooling. You’ll work on tools and automation that help product engineers deploy, operate, and debug services more easily, while learning modern platform engineering practices alongside experienced teammates. This role is ideal for engineers who enjoy improving developer experience and want to grow their skills in cloud infrastructure and CI/CD systems. Responsibilities: Contribute to the development and evolution of our internal platform-as-a-service used by application and service developers Build and maintain tooling that improves developer workflows, deployment reliability, and day-to-day productivity Collaborate with platform and application engineers to identify friction points and implement incremental improvements Learn and apply best practices around Infrastructure-as-Code, containerized workloads, and CI/CD pipelines Use, or are eager to adopt, AI-assisted development tools to improve productivity, and are excited to help explore and integrate LLM-powered solutions that automate internal support and operational workflows Have opportunities to propose ideas and improvements, with support and mentorship from the team Things you’ll get exposure to (and we don’t expect experience with everything): AWS Terraform, Pulumi CircleCI Docker, ECS, EKS Ruby, Golang, Python About You 2+ years of software development and infrastructure experience with significant contribut

PythonAWSDockerCI/CD
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Storage Infrastructure team builds and operates the storage foundation behind OpenAI’s most demanding workloads. We work directly with research to design storage systems for rapidly evolving experiments, while also powering production at scale. We own the platform end to end: backend systems, user-facing services and APIs, and the control planes that manage how data is placed, moved, and retained over time. Our stack spans cloud and in-house object stores across very different workload profiles, from GPU-attached systems to dedicated storage hardware. We also build the federation layer that unifies these backends behind a simple interface and routes each workload to the right storage solution. About the Role You will help build the storage platform that powers OpenAI’s research and production systems. This is a hands-on infrastructure role for engineers who want to work on deeply technical systems at scale and own them in production. You’ll work across object storage, cross-region data movement, lifecycle management, and the federation layer that provides a unified interface across multiple backends. Much of our stack runs on Kubernetes, and we primarily build services in Rust. In this role, you will: Build and operate storage services that underpin OpenAI’s research infrastructure Develop object storage systems across cloud and in-house environments Build systems for cross-region data movement, replication, and recovery Design lifecycle management capabilities that keep data durable, available, and cost-effective Evolve the federation layer that unifies multiple backend systems behind a simple interface Improve performance, reliability, and operational excellence across the platform Collaborate closely with researchers and infrastructure teams to support rapidly evolving workloads You might thrive in this role if you: Have experience building or operating distributed systems in production Have worked on storage infrastructure, object stores, dist

AWSKubernetesRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world. Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure. About the Role We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology. You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more. There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future. In this role, you will: Build, lead, and grow high-performing infrastructure engineering teams. Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search. Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach. Champion pragmatic use of agent technology to amplify execution velocity. Reduce operational toil and incident frequency through better abstractions, gua

AWSRestAIGo
O
📍 United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role We’re seeking an exceptional Staff - Principal level offensive security domain expert to build agents that continuously identify and coordinate remediation of vulnerabilities across OpenAI’s infrastructure and applications. You will be the technical owner of this effort, combining deep offensive security judgment with agent engineering to build a production system that can operate safely and reliably at scale. As OpenAI increasingly uses automation throughout the company, we believe our security testing must become increasingly automated as well. Advances in model capabilities create an opportunity to test more of our attack surface than would be possible through human effort alone and a need to ensure that we remain ahead of those same capabilities as they become available to attackers. In this role, you’ll build a portfolio of specialized agents that develop a deep understanding of OpenAI’s infrastructure, applications, processes, and security boundaries. These agents will combine internal context with feedback from running systems to explore our cloud environments, Kubernetes clusters, web applications, endpoints, external attack surface, and other high-value targets. The goal is for agents to not only discover vulnerabilities, but also to validate exploitability, document impact, drive remediation, and verify fixes. Success will be measured through outcomes like vulnerabilities fixed, attack surface covered, and performance on evals

AWSKubernetesLinuxRest
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for a talented and passionate Senior Software Engineer for our Snowpark Container Service , part of our Snowflake Compute Platform - to build our elastic, high-scale, high-performance, cloud native compute platform to enable bringing Compute to Data effortless and simple. Snowpark Container Services is a fully managed container offering that helps our customers easily deploy, manage, and scale containerized applications without having to move data out of Snowflake. As a fully managed service, it comes with Snowflake security, configuration, and operational best practices built in. You will be part of this highly productive, fast moving, and growing team that is critical to realizing Snowflake’s Data Cloud Mission. AS A SENIOR SOFTWARE ENGINEER, YOU WILL: ● Design and develop features, understand customer requirements and meet business goals. ● Lead a team of engineers, including mentoring and guiding them, and build technical direction and strategy for large and critical parts of the product surface area. ● Manage all aspects of the Project, including Design, Coding, Reviews, Testing, Observability, Tooling and On-Call support. ● Build highly reliable and fault-tolerant software to meet the needs of the largest customers. ● Ensure operational readiness and ma

JavaVueAWSAzure
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake’s Release Engineering team builds and operates the systems that safely deliver infrastructure, platform, and product changes to production at global scale. We own the release platforms, rollout orchestration, and safety mechanisms that allow engineering teams across Snowflake to ship quickly while minimizing operational risk. Our mission is to make production deployments fast, safe, self-service, increasingly autonomous, and augmented by AI-driven intelligence and automation. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale multi-cloud infrastructure orchestration. At Snowflake, Release Engineering is a platform engineering function focused on building the systems, abstractions, and automation that make software delivery safe, scalable, and efficient across the company. In this role, you will Design and build continuous deployment and rollout infrastructure that safely ships changes across Snowflake’s large-scale, multi-cloud production environment. Build and evolve platform capabilities for progressive delivery, including staged rollouts, canarying, automated health checks, rollback controls, and guardrails that reduce blast radius during production change events. Improve engineering velocity by removi

PythonJavaKubernetesCI/CD
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. As a Senior Technical Support Engineer, you will be a trusted advisor and technical resource for our customers. This is a hands-on role for someone who thrives in dynamic environments, loves troubleshooting complex technical issues, and is passionate about delivering exceptional support experiences. You’ll be responsible for resolving high-impact technical issues, driving customer success, and collaborating closely with product, engineering, and customer su

PythonSQLAWSAzure
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lake using open formats like Apache Iceberg, delivering deep correlation and long-term analytics at dramatically lower cost. A dynamic Knowledge Graph and chat-based AI SRE provide rich context and guided workflows so teams can move from detection to root cause and resolution significantly faster. The Infrastructure team at Observe by Snowflake is responsible for building, scaling, and operating the development and production environments that power our observability platform. We are a small, highly collaborative team with a broad scope, focused on delivering reliable infrastructure while continuously improving the systems that support our engineers and customers. What You’ll Do Design, build, and operate scalable cloud infrastructure in AWS supporting a high-scale observability platform. Improve system reliability, performance, and operational visibility across development and production environments. Develop and maintain CI/CD pipelines and internal tooling to improve developer productivity and deployment safety. Identify and mitigate security risks, and help maintain intern

PythonAWSAzureGCP
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for a talented and passionate Senior Software Engineer for our Snowpark Container Service , part of our Snowflake Compute Platform - to build our elastic, high-scale, high-performance, cloud native compute platform to enable bringing Compute to Data effortless and simple. Snowpark Container Services is a fully managed container offering that helps our customers easily deploy, manage, and scale containerized applications without having to move data out of Snowflake. As a fully managed service, it comes with Snowflake security, configuration, and operational best practices built in. You will be part of this highly productive, fast moving, and growing team that is critical to realizing Snowflake’s Data Cloud Mission. AS A SENIOR SOFTWARE ENGINEER, YOU WILL: ● Design and develop features, understand customer requirements and meet business goals. ● Lead a team of engineers, including mentoring and guiding them, and build technical direction and strategy for large and critical parts of the product surface area. ● Manage all aspects of the Project, including Design, Coding, Reviews, Testing, Observability, Tooling and On-Call support. ● Build highly reliable and fault-tolerant software to meet the needs of the largest customers. ● Ensure operational readiness and ma

JavaVueAWSAzure
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - Container Platform (Menlo Park) About the Role We build the foundational container platform that runs Snowflake's production, AI/ML, and CI workloads across AWS, Azure, and GCP, including a rapidly growing AI/ML footprint. Hundreds of large Kubernetes clusters under management and growing. The work is to make that fleet reliable, automated, and invisible to the thousands of engineers building on top of it. This is a staff-level role on a senior, high-performing platform team. You'll own hard problems end to end, drive technical direction across teams, and build the automation and platform abstractions that make operating at this scale sustainable. There is significant unsolved work ahead: improving the developer experience for thousands of internal engineers and continuing to scale the platform to meet Snowflake's growth. What You'll Do Own the design and delivery of large, complex platform initiatives spanning cluster lifecycle management, multi-cloud automation, and internal developer tooling. Identify and drive cross-team technical improvements across the platform, from architecture through adoption. Make and defend architectural trade-offs grounded in reliability, scalability, and operational reality. Act as a technical anchor for the team, dev

AWSAzureGCPKubernetes
C
📍 Tampa Florida United States, United States
✓ High-confidence listingCompany trend +800%
Quick readStrong listing-quality and freshness signals

The Engineering Lead Analyst – SonarQube & Code Quality Engineering is a senior-level engineering role responsible for leading static code analysis, automated code quality governance, security vulnerability remediation, and AI-augmented developer enablement across enterprise software delivery pipelines. In this role, you will champion software reliability, maintainability, clean-coding standards, and automated quality gates. You will partner with development teams, system architects, and platform engineering to integrate and manage enterprise-scale code quality platforms (such as SonarQube) both on-premises and in cloud/SaaS environments. Additionally, you will drive modern engineering practices by embedding Behavior-Driven Development (BDD) within your own software delivery and leveraging Agentic AI workers and Model Context Protocol (MCP) architectures to optimize developer experience, streamline code governance, and boost engineering velocity. Key Responsibilities 1. Code Quality & Static Analysis Platform Ownership Lead the architecture, deployment, administration, and continuous enhancement of enterprise Static Application Security Testing (SAST) and Code Quality platforms (e.g., SonarQube , DeepSource, Codacy, Semgrep). Configure, calibrate, and enforce automated Quality Gates, code rulesets, technical debt calculation models, and code-coverage baselines across multi-language enterprise repositories. Oversee version upgrades, patching, high availability, and operational maintenance for on-premises and SaaS/cloud-hosted code quality infrastructure. 2. CI/CD & Pipeline Integration <li style=

JavaScriptTypeScriptPythonJava
C
📍 United States· Remote
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary: As a Senior Software Development Engineer at Aetna, you will play a critical leadership role in the design, development, and continuous enhancement of enterprise-scale Provider Applications. You will drive technical solutions for complex business problems, ensure application stability, and lead cross-functional initiatives as a Project Owner. This role requires a balance of hands-on engineering expertise, technical leadership, and delivery ownership, including overseeing vendor/contractor teams, ensuring alignment with enterprise architecture, and delivering high-impact solutions that improve provider data systems and operational efficiency. Required Qualifications: 5&#43; years of hands-on application development experience with Python and Google Cloud Platform (GCP) 2&#43; years of experience leading or contributing to large-scale application development initiatives Preferred Qualifications: Experience working in Agile/SCRUM environments Proven experience in project/program management, including planning, execution tracking, and delivery management Strong organizational, leadership, and planning skills with the ability to manage multiple priorities Experience working with distributed teams and cross-functional stakeholders Prior exposure to

N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform & node designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each bringing together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We're searching for a highly motivated, technical leader to drive the engineering roadmap and innovation for our rack system software architecture. From firmware, kernel drivers, operating systems, networking, fabrics and associated user mode drivers &#43; manageability software. You will work with component leads internally and engage with industry leading hyperscalar / cloud service providers on taking these products to market. What you’ll be doing: Drive the software end-to-end architecture for NVIDIA's rack-scale products Maintain deep understanding of the product portfolio and roadmap; translate forward-looking plans into clear, formal software requirements that anchor execution across the organization. Ensure high quality & reliable software; serving as a trusted architectural partner to teams requiring

C
📍 United States· Full-time
✓ Quality checkedCompany trend -94.7%

We are looking for a Senior Forward Deployed Engineer to join the Customer Solutions team. You will be the technical authority embedded with our most complex customers, guiding them through deployment, architecture, onboarding, and the adoption of agentic development workflows. You bring deep hands-on experience from prior roles and use that depth to advise, design repeatable patterns, and drive customer outcomes end to end. You operate autonomously, own the technical success of your customers, and bring their experience back to shape how Coder builds and delivers. This is not an execution-only role. You are equally comfortable doing deep technical work with a customer and stepping back to design the repeatable pattern behind it. You are energized by ambiguity, motivated by customer outcomes, and capable of influencing organizational change alongside the technical work. This position is required to sit in the Eastern Time Zone. What You'll Do Serve as the primary technical authority for post-sales customers, guiding deployment architecture, environment design, and adoption of Coder across both human and AI development workflows Own onboarding engagements end to end, ensuring customers move from contract to productive adoption with speed and confidence Lead Get Well engagements where architecture decisions, rollout patterns, or organizational dynamics are limiting customer health or growth Help customers implement the technical and organizational changes required to adopt agentic development practices at scale Design and document repeatable delivery patterns across onboarding, architecture, and adoption that can scale across customer segments Design and recommend reference architectures tailored to each customer's cloud environment, security posture, and organizational constraints Translate customer environment complexity into clear guidance on networking, ingress, identity, and infrastructure patterns Anticipate technical and operational risks, escalate to the right

AWSAzureKubernetesCI/CD
P
📍 New York City, New York, United States· Full-time
✓ Quality checkedCompany trend -85.7%

About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: Join a team that builds robust, real-time distributed systems for a cutting-edge database. We care about performance, reliability, scalability, and most of all learning and having fun together. Whether you’re a seasoned coder or just getting started, if you’re passionate about technology and eager to learn, you’ll fit right in. Who we are: We show up to work, ready to collaborate and build technologies that make a difference, with people who genuinely care. We chase improvements such as tail latencies, bytes throughput, cache hit rate, and operational cost efficiency. We believe learning is ongoing and that even the most complex problems can have simple solutions. What You’ll Do: Collaborate with teammates to design and build database features that power AI applications. Learn how to tune performance and support reliability in distributed systems (don’t worry, we’ll guide you). Help Pinecone run smoothly on popular cloud providers. Take ownership of your work and grow your skills every day. Have fun. Who You Are: 5+ years of work experience - programming in Rust, Go, C++, or a comparable language. You’re genuinely curious about distributed systems and eager to dive deep into technical challenges. You approach problems with creativity and persistence, and you’re comfortable asking thoughtful questions or seeking feedback. You’re excited to learn, value constructive feedback, and appreciate mentorship. Bonus Points: You have hands-on experience with cloud platforms (AWS, GCP, Azure) or have demonstrated an ability to pick u

AWSAzureGCPAI
🔔

Get new cloud operations engineer jobs in United States by email

Daily job updates · Unsubscribe anytime