We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary The Executive Director, Digital Engineering- Aetna Member Care and Journey Services is a senior technology leader responsible for setting the technical vision, architectural direction, and engineering execution for member centric services. This role leads large-scale engineering teams that build high-performance backend APIs, microservices, and cloud-native systems that power member experiences across digital, agent, provider, and partner channels. The leader ensures exceptional service stability, resiliency, innovation velocity, and alignment with enterprise user experience and operational goals. Key Responsibilities 1. Backend API & Microservices Engineering Leadership • Lead the design, development, and delivery of scalable backend systems, APIs, and microservices powering member-facing capabilities. • Define API contract standards, and integration patterns used across Member Services platforms. • Drive service modernization by adopting cloud‑native architectures, containerization, and event-driven patterns. 2. Service Stability, Observability & Resiliency • Establish standards for availability, resiliency, performance, and disaster recovery across all services. • Implement SLO/SLI/error budget frameworks, health checks, and high‑availability architectures. <p
Jobs in United States
Senior Cloud Operations Engineer in United States
1,941 active opportunities · Updated October 2026
Showing
15 jobs
Explore current senior cloud operations engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary Health100 is an AI‑native health technology platform that unifies pharmacies, providers, insurers, PBMs, and digital health solutions into a single, consumer‑focused ecosystem. Powered by Google Cloud AI, we’re reimagining personalized and connected health experiences. As a Senior Software Engineer for Health100, you will play a crucial role within a collaborative team — designing, developing, and maintaining backend services and APIs while ensuring releases are well-coordinated, fully prepared, and successfully deployed to production. The ideal candidate brings strong technical expertise in modern backend development, excellent problem-solving skills, and a proactive approach to production monitoring, issue triage, and cross-team coordination. This position is critical in maintaining high engineering standards, ensuring smooth release cycles, and driving operational excellence across the development lifecycle. *This role can be based anywhere in the US; hybrid or remote with preference for candidates to work out of our corporate headquarters in Woonsocket, RI. Responsibilities: Partner with technical leaders and the open-source community to contribute to technical designs, frameworks, roadmap definition, and requirements-gathering. Provide domain knowledge and engineering insight to guide early designs, ac
We are developing advanced multi-rack, multi-tenant AI/ML datacenters with NVIDIA GB200, and upcoming GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team to focus on the cloud-native stack for datacenter products like GB200. In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex scheduling challenges across racks, tenants, and clouds as part of the CSP engagements team. What you’ll be doing: Perform deep-dive debugging of multi-rack, multi-tenant clusters: scheduler behavior, container runtime issues, device-plugin crashes, RDMA/IB fabric anomalies, etc. Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom micro-services that expose new GPU capabilities. Drive joint architecture reviews and “whiteboard” sessions with CSP and internal platform teams; convert findings into RFCs and upstream pull requests. Create reproducible testbeds (Helm/Ansible/Terraform) that mirror customer environments; automate validation and benchmark suites. Deliver technical collateral-design docs, how-to guides, demo scripts-and present at customer on-sites, KubeCon, and SlurmUG. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Strong source-level expertise in Kubernetes internals (scheduler, CRI/CNI/CSI, operators) and Slurm (federation, power-save, plugins). Hands-on experience integrating next-gen GPUs (Blackwell/GB200/GB300) or comparable accelerators into containerized clusters. Proven track record debugging large-scale, cloud-native stacks across ne
We are looking for a Senior Software Engineer to become part of our storage management plane team. The management plane is a web-based application crafted to provide our storage customers the capabilities to handle and supervise our distributed storage infrastructure. Our team is continually dedicated to acquiring and implementing ground breaking technologies to overcome obstacles and innovate solutions for improving our ability to handle large clusters of machines efficiently. What You Will Be Doing: Maintain and develop Kubernetes operators and our Container Storage Interface (CSI) plugin. Develop a web-based solution that manages, operates and monitors our distributed storage. Work closely with other teams to define and implement new APIs. What We Need to See: B.Sc., M.Sc. or Ph.D. in Computer Science, or related discipline, or equivalent experience. 8+ years of experience in web development ( both client and server ) Proven experience with Kubernetes (K8s), including developing or maintaining operators and/or CSI plugins. Experience scripting with Python, Bash or similar. Experience with nodejs is a must At least 5 years of experience working in a Linux OS environment You’re smart and a quick learner You do what it takes to get the job done Passionate about coding and big challenges Ways to stand out from the crowd: NodeJS for the server side: dominant modules are async & express . Kafka, MongoDB, K8s JavaScript frameworks: React, jQuery, c3j
NVIDIA is transforming how the world uses AI, cloud, and accelerated computing, and trust is at the center of that mission. Our Attestation and Trust Services team builds the secure cloud services that show customers their NVIDIA platforms are healthy, resilient, and ready for their most important workloads. In this role, you help design and run services that sit at the intersection of hardware, security, and large-scale distributed systems. We partner closely with security, silicon, platform, and cloud teams to bring new ideas into reliable production services that people rely on every day. We care about building systems that last, supporting each other, and creating space for learning and experimentation. If you enjoy solving complex problems, keeping services running smoothly, and collaborating with teammates from many disciplines, we would love to talk with you! What you’ll be doing: Your main focus will be on building and managing our core attestation cloud services. Day-to-day responsibilities include crafting APIs and integrations, boosting reliability, and working alongside NVIDIA teams to convert hardware trust mechanisms and standards into production-ready solutions. You will contribute significantly to shaping how customers verify that NVIDIA platforms are secure and prepared for their workloads. Crafting and evolving attestation cloud services, APIs, and SDK/CLI integration points that confirm the integrity of NVIDIA platforms across data center, AI, networking, and partner environments. Improving reliability and operational maturity through SLOs/SLIs, alerting, runbooks, incident response, and safe rollout practices. Crafting resilient service behavior that handles dependency failures, caching challenges, regional issues, customer-side resilience needs, and graceful degradation. Architecting trust-material distribution for certificate status, re
We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role A strong and reliable platform is essential to scaling Sentry for the future. Our Platform organization is responsible for everything that powers Sentry—from cloud infrastructure and streaming systems to storage, deployment, and security. We own the core services and technical foundations that enable every product and engineering team at Sentry to move fast and build with confidence. We're looking for a passionate and pragmatic Senior Staff Software Engineer to help lead this evolution. In this role, you’ll report directly to the VP of Engineering and collaborate with teams across the company to shape the future of Sentry’s platform. What You’ll Do Architect the future of Sentry by translating business needs and product strategy into clear, scalable technical blueprints. Partner with product and engineering leaders to align technical roadmaps with company goals. Lead cross-cutting initiatives across the Platform org—owning them end-to-end and driving meaningful outcomes. Promote engineering excellence by mentoring platform engineers, sharing best practices, and setting high standards for system design, scalability, and operational quality. Review major architectural proposals and help ensure consistency, maintainability, and long-term technical health across the company. You’ll Love This Job If You... Enjoy designing and building platforms that help teams move faster and scale safely. Thrive on solving complex, multi-dimensional problems across product, infrastructure, and organizational layers. Want to make architectural decisions that shape Sentry’s long-term success. Bring new ideas, tools, and frameworks t
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. As a Senior Technical Support Engineer, you will be a trusted advisor and technical resource for our customers. This is a hands-on role for someone who thrives in dynamic environments, loves troubleshooting complex technical issues, and is passionate about delivering exceptional support experiences. You’ll be responsible for resolving high-impact technical issues, driving customer success, and collaborating closely with product, engineering, and customer su
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary: As a Senior Software Development Engineer at Aetna, you will play a critical leadership role in the design, development, and continuous enhancement of enterprise-scale Provider Applications. You will drive technical solutions for complex business problems, ensure application stability, and lead cross-functional initiatives as a Project Owner. This role requires a balance of hands-on engineering expertise, technical leadership, and delivery ownership, including overseeing vendor/contractor teams, ensuring alignment with enterprise architecture, and delivering high-impact solutions that improve provider data systems and operational efficiency. Required Qualifications: 5+ years of hands-on application development experience with Python and Google Cloud Platform (GCP) 2+ years of experience leading or contributing to large-scale application development initiatives Preferred Qualifications: Experience working in Agile/SCRUM environments Proven experience in project/program management, including planning, execution tracking, and delivery management Strong organizational, leadership, and planning skills with the ability to manage multiple priorities Experience working with distributed teams and cross-functional stakeholders Prior exposure to
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview Join a team focused on transforming how Mastercard's payment systems are built, scaled, and operated. As a Senior Software Engineer, you will lead the design and development of cloud-ready applications, microservices, and distributed systems that support large-scale payment processing platforms while helping advance modernization, automation, and engineering excellence across the organization. In this role, you will contribute to software architecture decisions, drive technical design discussions, and partner with engineers to deliver scalable, resilient, and maintainable software solutions. You'll have the opportunity to solve complex technical challenges, mentor other engineers, and influence how software is designed, developed, tested, and supported across critical technology platforms. What You Will Do •Design software solutions and contribute to software architecture decisions that support scalability, maintainability, and operational excellence. •Translate complex product requirements into technical designs and implementation plans. •Lead development of modular, extensible, high-performance applications. •Design and implement comprehensive unit, functional, and integration testing strategies. •Analyze, optimize, and improve application performance, scal
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse using open formats like Apache Iceberg — at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world's leading data platforms. We are hiring a Senior Software Engineer for Observe by Snowflake on the Data Management team. This team is responsible for the tables, views, and materialized views at the core of Observe's architecture. Observe's data lake approach lets customers correlate heterogeneous telemetry — logs, metrics, traces, events — across a unified data model. This role owns that data model: how customers define, shape, and query the semi-structured data that makes cross-si
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Senior Software Engineer to own key components of our AI native External Observability Platform . In this role, you will contribute to the technical road map for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex technical projects, and skilled at collaborating with the brightest technical minds in the industry. Key Responsibilities Develop and Scale Distributed Infrastructure: Design and implement key components of Snowf
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. We are hiring a Senior Frontend Engineer, AI Products team at Observe by Snowflake. As an AI Product engineer you'll always be thinking first about the user experience and how to create the best product, technical choices, and implementation decisions that stem from that product first thinking. This team builds the AI-powered products and developer tooling at the core of Observe's platform, including our flagship AI SRE product, real-tim
We are looking for a Senior Forward Deployed Engineer to join the Customer Solutions team. You will be the technical authority embedded with our most complex customers, guiding them through deployment, architecture, onboarding, and the adoption of agentic development workflows. You bring deep hands-on experience from prior roles and use that depth to advise, design repeatable patterns, and drive customer outcomes end to end. You operate autonomously, own the technical success of your customers, and bring their experience back to shape how Coder builds and delivers. This is not an execution-only role. You are equally comfortable doing deep technical work with a customer and stepping back to design the repeatable pattern behind it. You are energized by ambiguity, motivated by customer outcomes, and capable of influencing organizational change alongside the technical work. This position is required to sit in the Eastern Time Zone. What You'll Do Serve as the primary technical authority for post-sales customers, guiding deployment architecture, environment design, and adoption of Coder across both human and AI development workflows Own onboarding engagements end to end, ensuring customers move from contract to productive adoption with speed and confidence Lead Get Well engagements where architecture decisions, rollout patterns, or organizational dynamics are limiting customer health or growth Help customers implement the technical and organizational changes required to adopt agentic development practices at scale Design and document repeatable delivery patterns across onboarding, architecture, and adoption that can scale across customer segments Design and recommend reference architectures tailored to each customer's cloud environment, security posture, and organizational constraints Translate customer environment complexity into clear guidance on networking, ingress, identity, and infrastructure patterns Anticipate technical and operational risks, escalate to the right
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: Join a team that builds robust, real-time distributed systems for a cutting-edge database. We care about performance, reliability, scalability, and most of all learning and having fun together. Whether you’re a seasoned coder or just getting started, if you’re passionate about technology and eager to learn, you’ll fit right in. Who we are: We show up to work, ready to collaborate and build technologies that make a difference, with people who genuinely care. We chase improvements such as tail latencies, bytes throughput, cache hit rate, and operational cost efficiency. We believe learning is ongoing and that even the most complex problems can have simple solutions. What You’ll Do: Collaborate with teammates to design and build database features that power AI applications. Learn how to tune performance and support reliability in distributed systems (don’t worry, we’ll guide you). Help Pinecone run smoothly on popular cloud providers. Take ownership of your work and grow your skills every day. Have fun. Who You Are: 5+ years of work experience - programming in Rust, Go, C++, or a comparable language. You’re genuinely curious about distributed systems and eager to dive deep into technical challenges. You approach problems with creativity and persistence, and you’re comfortable asking thoughtful questions or seeking feedback. You’re excited to learn, value constructive feedback, and appreciate mentorship. Bonus Points: You have hands-on experience with cloud platforms (AWS, GCP, Azure) or have demonstrated an ability to pick u
Other cities to consider
More places hiring for this role
Get new senior cloud operations engineer jobs in United States by email
Daily job updates · Unsubscribe anytime