We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. The Release Engineer is responsible for implementing and supporting automated software delivery processes that enable reliable, secure, and repeatable deployments across development, test, non-production, and production environments. This role serves as a technical bridge between Software Engineering, Quality Engineering, Infrastructure, and Operations teams to improve deployment speed, stability, and operational efficiency. The Release Engineer is responsible for implementing and supporting automated software delivery processes that enable reliable, secure, and repeatable deployments across development, test, non-production, and production environments. This role serves as a technical bridge between Software Engineering, Quality Engineering, Infrastructure, and Operations teams to improve deployment speed, stability, and operational efficiency. Key Responsibilities Develop, support, and maintain automated CI/CD pipelines to streamline application delivery. Standardize build, release, and deployment processes across applications and platforms. Establish repeatable, auditable, and traceable release management practices to ensure deployment consistency and compliance. Manage and support Git-based source control, branching strategies, and release workflows. Integrate automated testing, quality gates, security controls, and compliance checks into deploym
Jobs in United States
Sre Operations Engineer in United States
44 active opportunities · Updated October 2026
Showing
15 jobs
Explore current sre operations engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. We are looking for a Security Operations Lead (SOC Lead) to build, mature, and operate our 24/7 detection and response capabilities across a modern cloud-native and AI-driven environment. This role leads the global SOC function—monitoring, SIEM ownership, detection engineering, alert triage, and operational readiness—while also evaluating and integrating emerging AI-based SOC products and autonomous response platforms . You will oversee monitoring across multi-cloud environments (GCP primary, AWS/Azure secondary), Kubernetes, SaaS services, endpoints, developer tools, and AI workloads . You’ll collaborate closely with Cloud Security, Compliance/GRC, SRE, Platform Engineering, IT/Endpoint teams, and AI Infrastructure to ensure our detection strategy scales and stays ahead of evolving threats. This is a hands-on leadership role perfect for someone who wants to shape the SOC of the future while solving complex challenges in a high-scale AI setting. What You’ll Do SOC Leadership & 24/7 Monitoring Lead, mentor, and scale a global SOC team responsible for 24/7 monitoring, alert intake, triage, correlation, and escalation. Build operational rigor: processes, runbooks, SLAs, metrics, and quality standards for high-scale environments. Cover monitoring across: Cloud infrastructure (GCP, AWS, Azure) Kubernetes/GKE/EKS/AKS clusters SaaS platforms (Google Workspace, GitHub, Slack, Okta, etc.) Endpoints (macOS, Linux, Windows) including EDR/XDR telemetry Developer platforms + CI/CD pipelines AI/ML systems and model-serving workflows AI-Based SOC Integration & Innovation Evaluate, adopt, and integrate AI-native SOC technologies for triaging, detection, and correlation Identify opportunities to automate triage, investigations,
From $127K/yr
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently. You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. EXAMPLE INITIATIVES You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services. Building AI-assisted tooling for incident triage and response. RESPONSIBILITIES Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking. Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code. Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution. Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations. Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, con
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. At CVS Health, Site Reliability Engineering (SRE) is fundamental to delivering the reliable, secure, and scalable technology experiences that support millions of patients, customers, pharmacists, and healthcare professionals every day. Our SRE organization drives operational excellence across critical healthcare and retail platforms through innovation, automation, observability, and engineering best practices. The Executive Director, Site Reliability Engineering serves as the strategic leader responsible for the reliability, resilience, and performance of CVS Health's retail and pharmacy technology ecosystem. This executive will define and execute a comprehensive reliability strategy, oversee large global engineering teams, and establish a long-term vision for observability, automation, and operational excellence across thousands of store locations. Working closely with senior business and technology leaders, the Executive Director will champion modern SRE practices, accelerate incident response capabilities, and deliver real-time operational visibility that enables proactive issue prevention and exceptional customer and patient experiences. Key Responsibilities Strategic Leadership & Vision Define and lead the enterprise-wide Site Reliability Engineering strategy supporting CVS Health's retail and pharmacy operations. Align reliability and operational
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build cloud-native data and storage services for hybrid and multi-cloud infrastructure, including dataset discovery, ingestion, governance, checkpointing, observability, and low-latency access. Develop scalable cloud-native services and APIs that support exabyte-scale, high-performance GPU training and inference workflows. Work closely with product managers, internal AI teams, platform teams, and partner engineering teams to understand requirements and turn them into reliable production systems. Collaborate with SRE, operations, and support teams to improve service reliability, performance, observability, on-call readiness, and operational scale. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, and verification. What we need to see: BS in Computer Science, Information Systems, Computer Engineering, or equivalent experience, with 5+ years of software engineering experience. Strong foundation in algorithms, data structures, distributed systems, and practi
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys
From $156K/yr
As a Product Manager – IaC Detection, you will define, build, and launch capabilities that proactively detect infrastructure issues in code (e.g. Terraform, Helm) before they can be deployed into production and escalate into production incidents. The Infrastructure Monitoring team has pioneered shift-left detection in the industry with Bits Infrastructure Operations , and we’re looking for a Product Manager to expand this capability to a broader set of use cases Customers (and thus developers) are increasingly standardizing on IaC tools to deploy and maintain ever-growing infrastructure in the cloud. At the same time, SREs and Infra teams struggle with an increasing number of production incidents. By shifting-left and identifying high-impact infra changes before they are deployed, we help reduce production incidents, reduce waste, and free up SRE time to focus on value-added tasks. You will own the roadmap to expand IaC detection to a broader set of use cases, including cost detection, blast radius impact, as well as configuration changes on infrastructure powering applications like nginx, postgres and more. You’ll partner closely with Engineering, Design, and customers to build and iterate on the roadmap, build product market fit, drive customer adoption (including internal usage), and focus on coverage and correctness of the AI system. This is an opportunity to lead an initiative at the intersection of AI, infrastructure operations, and autonomous observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the product roadmap for IaC Detection, enabling customers to proactively detect and catch high-impact infrastructure and configuration changes before they are deployed into production and escalate into incidents. Define the end-to
Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. Remote (US East Coast preferred, for timezone coverage) About the team Cloud Infrastructure owns the platform every Synthesia product runs on — AWS, Kubernetes, MongoDB, Temporal, our observability stack, and the vendor and cost relationships underneath them. We're a small, high-leverage team scaling toward a domain-ownership model: small groups that both build and operate the systems they're accountable for. The role We're hiring a dedicated SRE to take real ownership of operational excellence across Cloud Infrastructure. Today, too much critical operational knowledge — vendor relationships, cost management, and incident response — lives with one or two people. Your mission is to take genuine ownership of those domains, make them resilient to any single person, and raise the bar on how reliably we run. This is not simply a ticket-queue or keep-the-lights-on role. You'll own domains end to end: understand them deeply, operate them well, and build the automation and tooling that make them boring . We deliberately pair operational and engineering work so the role grows rather than narrows. What you'll own Incident management & operational excellence — take custody of the incident process: on-call quality, resp
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking a strategic and results-driven engineering leader who is passionate about cloud agnostic infrastructure, operational excellence, and enabling engineering teams to operate autonomously and build with confidence. As Head of Infrastructure, you'll lead a talented and geographically distributed team of engineers across the SF Bay Area, India, and Europe, fostering a culture of collaboration, ownership, and continuous improvement. You'll own the infrastructure that underpins one of the world's most widely used API platforms, an environment handling ~80,000 requests per second at the front door, and be responsible for its reliability, scalability, and evolution. In addition to infrastructure, you'll own the Site Reliability Engineering (SRE) function at Postman, setting the standards and practices that keep the platform reliable at scale. You'll work closely with engineering managers, product managers, and platform teams to drive the technical roadmap for our cloud agnostic infrastructure and reliability practices, ensuring we can support a large and rapidly growing engineering organization. If you're p
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . The Production Engineering organization at Pinterest is accountable for ensuring overall Pinterest availability as well as enhancing Engineering teams' capability to design, build and operate robust systems at scale. Pinterest's applications and infrastructure handle billions of monthly page views and petabytes of data as Pinterest continues to grow and scale. As a Senior Production Engineer on Solutions Engineering, you will design and build AI agents, platforms, tools, frameworks and methodologies to assure the reliability of our large-scale distributed systems serving hundreds of millions of monthly active users, handling hundreds of thousands of requests per second, and managing tens of petabytes of data. You'll lead infrastructure modernization initiatives, build intelligent automation that eliminates operational toil and amplifies engineer
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are looking for a highly skilled PSIRT Engineer to lead the vulnerability response program for Replit’s cloud-native AI platform. You will own the lifecycle of security vulnerabilities affecting our products and services—from intake to validation, remediation coordination, and public disclosure. This role requires strong technical ability to reproduce vulnerabilities , deep understanding of web/app/cloud exploit classes, and experience operating bug bounty and coordinated disclosure programs. You will work closely with Engineering, Cloud Security, SecOps, SRE, and IT teams to ensure vulnerabilities are fixed quickly and communicated responsibly. What You’ll Do Vulnerability Intake, Triage & Validation Manage intake from bug bounty platforms (HackerOne preferred), customer reports, automated scanners, pentest reports, and coordinated disclosure channels. Independently validate, reproduce, severity-score, and document findings. Identify duplicates and maintain a clean vulnerability records pipeline. Assess relevance and exploitability using OWASP, cloud misconfiguration patterns, and identity/authentication/authorization risks (Oauth, OIDC). Remediation Coordination & SLA Management Work with Engineering, SecOps, IT, SRE, and Cloud Security to confirm product impact and drive remediation. Provide detailed reproduction steps, proof-of-concepts, and technical analyses. Track SLAs, remediation progress, regression testing, and systemic improvements. Support SOC 2, ISO 27001, and pentest evidence needs as part of vulnerability lifecycle governance. Bug Bounty & Vulnerability Disclosure Program Management Design and evolve the bug bounty program, including scope, rules, and reward structures. Man
From $126K/yr
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An Overview of This Role Staff Systems Engineer, IT is a senior, hands-on engineering role for a generalist who is comfortable owning a broad set of platforms. You'll own the systems and integrations behind the employee lifecycle, onboarding, role changes, and offboarding, and you'll build them the way we build software: as infrastructure-as-code, with SRE practices behind them so they are versioned, observable, and reliable. You'll be the technical owner of GitLab's ITSM platform and the AI capabilities layered on top of it, designing virtual agents, agentic workflows, and knowledge experiences that resolve requests before they
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lake using open formats like Apache Iceberg, delivering deep correlation and long-term analytics at dramatically lower cost. A dynamic Knowledge Graph and chat-based AI SRE provide rich context and guided workflows so teams can move from detection to root cause and resolution significantly faster. The Infrastructure team at Observe by Snowflake is responsible for building, scaling, and operating the development and production environments that power our observability platform. We are a small, highly collaborative team with a broad scope, focused on delivering reliable infrastructure while continuously improving the systems that support our engineers and customers. What You’ll Do Design, build, and operate scalable cloud infrastructure in AWS supporting a high-scale observability platform. Improve system reliability, performance, and operational visibility across development and production environments. Develop and maintain CI/CD pipelines and internal tooling to improve developer productivity and deployment safety. Identify and mitigate security risks, and help maintain intern
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse using open formats like Apache Iceberg — at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world's leading data platforms. We are hiring a Senior Software Engineer for Observe by Snowflake on the Data Management team. This team is responsible for the tables, views, and materialized views at the core of Observe's architecture. Observe's data lake approach lets customers correlate heterogeneous telemetry — logs, metrics, traces, events — across a unified data model. This role owns that data model: how customers define, shape, and query the semi-structured data that makes cross-si
Other cities to consider
More places hiring for this role
Get new sre operations engineer jobs in United States by email
Daily job updates · Unsubscribe anytime