At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. We are hiring a Senior Frontend Engineer, AI Products team at Observe by Snowflake. As an AI Product engineer you'll always be thinking first about the user experience and how to create the best product, technical choices, and implementation decisions that stem from that product first thinking. This team builds the AI-powered products and developer tooling at the core of Observe's platform, including our flagship AI SRE product, real-tim
Jobs in United States
Sre Operations Engineer in United States
44 active opportunities · Updated October 2026
Showing
15 jobs
Explore current sre operations engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer on the Internal Tooling team, you will own the internal operating system that sits at the heart of how Baseten operates. Capacity helps unlock revenue by carefully balancing supply and demand. The operating system manages all aspects of the customer lifecycle: from onboarding to managing complex customer SLA requirements. This role is for engineers who want to own a product end to end, not just implement tickets. You will work directly with the Capacity, Sales, and Engineering teams to understand requirements, define solutions, and ship software that removes friction from some of the most high-stakes workflows in the company. If something is slow, manual, or error-prone in the capacity fulfillment lifecycle, you will be the one to fix it. You are a strong fit if you have strong product intuition, move fast without sacrificing quality, and take satisfaction in building tools that make the people around you measurably more effective. RESPONSIBILITIES Own the Capacity product end to end: scoping, design, implementation, and iteration based on feedback from internal stakeholders Translate complex operational requirements from Capacity, Sales, and SRE teams into clean, ergonomic product experiences Build and maintain full-stack features across the Capacity toolchain, including UI surfaces, APIs, and backend services Identify workflow bottlenecks and manual processes across the capacity lifecy
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse using open formats like Apache Iceberg — at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world's leading data platforms. We are hiring a Senior Software Engineer for the Observe Data Management team. This team owns the core pipelines that ingest and process over 1 petabyte of telemetry data per day — the foundational infrastructure powering Observe's entire observability stack. You'll be working at the intersection of massive scale, open-source innovation, and real-world reliability challenges for enterprise customers around the globe. AS A SENIOR SOFTWARE ENGINEER - OBSERVE
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. In this role you will: Develop interactive, data-rich user interfaces using React, TypeScript, and Vega, with a focus on integrating LLM-driven features (e.g., natural language querying, generative UI, and AI-assisted data storytelling). Lead the end-to-end delivery of substantial product features, ensuring AI outputs are presented with high reliability and low latency. Work closely with PMs, UX designers, and AI/ML engineers to bridge t
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler’s high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You’ll Do (Role Expectations) Maintain h
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. As a Senior Technical Support Engineer, you will be a trusted advisor and technical resource for our customers. This is a hands-on role for someone who thrives in dynamic environments, loves troubleshooting complex technical issues, and is passionate about delivering exceptional support experiences. You’ll be responsible for resolving high-impact technical issues, driving customer success, and collaborating closely with product, engineering, and customer su
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. The Team On the AI Backend team at Observe by Snowflake you'll be at the forefront of how AI is reshaping the way engineering teams operate; building the platform that powers intelligent, automated workflows across observability and beyond. Our team is small, and moves fast, with real ownership over hard problems that span agentic APIs, real-time pipelines, and AI quality. You'll work alongside talented engineers across ML, product, and
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. About the Role As a Senior Product Manager at Observe by Snowflake you will lead a cross-functional team to define and execute the product roadmap across core product areas -- Logs, Metrics, Traces, Data Management, AI and more. You will set strategy grounded in deep customer empathy, drive prioritization across competing enterprise demands, and shape a product processing thousands of TB/day of telemetry. This is a high visibility role where you will intera
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is a high-growth SaaS observability platform built on the Snowflake AI Data Cloud, enabling businesses to troubleshoot modern distributed applications 10x faster. Now, as a core part of Snowflake, we’ve reached a major milestone in the evolution of the Snowflake platform. By bringing AI-powered observability directly into the Snowflake ecosystem, we’ve created the first truly unified platform for telemetry and business data. We’re looking for a Technical Account Manager to partner with our most strategic enterprise customers and ensure they derive sustained operational value from Observe. This is a hands-on, post-sales technical role focused on long-term platform adoption, optimization, and technical partnership. You will work directly with SRE, DevOps, platform, and engineering teams to embed Observe into daily workflows, evolve telemetry strategy over time, and continuously improve reliability, performance, and cost efficiency. This role is ideal for an experienced observability practitioner who enjoys being deeply embedded with customer teams, solving real production challenges, and acting as a trusted technical advisor in complex enterprise environments. What You’ll Do Serve as the primary technical owner and trusted advisor for assigned strategic a
From $56K/yr
Locations: South Jordan, UT Salary: $56,000 Launch Your Career in Technology Every app, website, payment, and digital service relies on technology running smoothly behind the scenes. When something goes wrong, Production Support Engineers are the people who investigate the issue, restore service, and help prevent it from happening again. If you're curious, analytical, and enjoy solving problems, this is an opportunity to build hands-on experience with cloud platforms, Linux, automation, databases, and large-scale enterprise systems from day one. What Is Production Support? Production Support Engineers keep business-critical applications running reliably in live environments. Think of it this way: Software Engineers build the platform. QA Engineers test the platform. Production Support Engineers keep the platform running when it matters most. Working at the intersection of technology and business, you'll troubleshoot issues, automate processes, and help improve the reliability and performance of systems used by thousands, or even millions, of people every day. If you enjoy solving puzzles, working under pressure, and understanding how large-scale systems work, this could be the perfect place to start your career. What You'll Do As part of a global production engineering team, you'll: Help support large-scale applications and platforms used by leading organizations around the world. Monitor business-critical applications and services to ensure high availability and performance. Investigate and resolve production incidents across applications, infrastructure, databases, and cloud environments. Analyse logs, alerts, and system metrics to identify root causes and prevent recurring issues. Partner with software engineers, infrastructure teams, and business stakeholders to improve system reliabilit
From $139.8K/yr
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest brings millions of people the inspiration to create a life they love. Behind that experience is a complex infrastructure ecosystem that powers reliability, performance, measurement, and efficiency across the platform. As Pinterest grows, it’s increasingly important that we understand these systems clearly so we can make smarter decisions for both Pinners and the business. We’re looking for a Data Scientist to join our Infrastructure Data Science team. In this role, you’ll partner with engineering and cross-functional teams to make Pinterest’s infrastructure more measurable, intelligible, and actionable. Depending on the area, your work may span app performance, shopping infrastructure, metrics quality, infrastructure governance, or site reliability. You’ll help build the data foundations, measurement systems, and analytical fram
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Senior Manager, Platform Engineering / DevOps Who Are You You are an experienced Senior Manager / emerging Staff-level leader in DevOps and Platform Engineering with strong technical depth and demonstrated leadership in delivering enterprise-scale cloud platforms. You bring a balanced mix of hands-on engineering expertise, team leadership, and execution rigor. You excel in driving outcomes in complex, multi-stakeholder environments, guiding teams to deliver secure, scalable, and high-quality platform solutions. You are comfortable leading engineers, managing stakeholders, and owning delivery across multiple workstreams. You demonstrate: A strong ownership mindset with accountability for delivery and outcomes Ability to translate business needs into actionable engineering roadmaps Solid expertise in cloud-native platforms, DevOps practices, and SRE principles Capability to lead teams and influence without requiring extensive tenure Role Responsibilities Development & Enforcement Own and execute the H100 platform engineering roadmap, aligned to enterprise priorities and program milestones Drive delivery of GCP-based platform capabilities (GKE, networking, IAM, CI/CD, observability) Establish and enforce engineering standards, best practices, and ADR compliance <li
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Senior Manager, Platform Engineering / DevOps Who Are You You are an experienced Senior Manager / emerging Staff-level leader in DevOps and Platform Engineering with strong technical depth and demonstrated leadership in delivering enterprise-scale cloud platforms. You bring a balanced mix of hands-on engineering expertise, team leadership, and execution rigor. You excel in driving outcomes in complex, multi-stakeholder environments, guiding teams to deliver secure, scalable, and high-quality platform solutions. You are comfortable leading engineers, managing stakeholders, and owning delivery across multiple workstreams. You demonstrate: A strong ownership mindset with accountability for delivery and outcomes Ability to translate business needs into actionable engineering roadmaps Solid expertise in cloud-native platforms, DevOps practices, and SRE principles Capability to lead teams and influence without requiring extensive tenure Role Responsibilities Development & Enforcement Own and execute the H100 platform engineering roadmap, aligned to enterprise priorities and program milestones Drive delivery of GCP-based platform capabilities (GKE, networking, IAM, CI/CD, observability) Establish and enforce engineering standards, best practices, and ADR compliance</li
$200K – $240K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Platform org at Sentry is the engine that everything else runs on, spanning developer platform, core infrastructure underpinning the product, SRE, security, and IT. It's a broad, technically complex, and deeply consequential organization. We're looking for a Staff Technical Program Manager who can operate at the intersection of technical depth and strategic execution: someone who thrives in complexity, builds trust with senior engineering leaders, and has a gift for turning ambiguity into clarity and momentum. You'll report to the Head of Technical Program Management and work closely with the VP of Platform Engineering, their staff, and partner teams across Engineering, Product, and Design (EPD). This is a high-visibility role with real influence. You'll work directly with the CTO and senior leaders, shape how the Platform org operates, and help Sentry scale through one of its most important chapters. In this role, you will: Drive strategic execution. Partner with the VP of Platform Engineering and senior EPD leaders to translate priorities into clear, measurable plans. Own sequencing, milestones, and key decisions, and make sure leadership always has reliable visibility into progress and tradeoffs. Build data-driven delivery health. Establish the metrics and dashboards that reflect delivery confidence, risk, and engineering health across the Platform org. Keep planning and reporting high-signal and lightweight, focused on outcomes rather than activity. Manage capacity, dependencies, and risk. Create visibility into resourcing, cross-team dependencies, and constraints so leaders can align investment to the
Other cities to consider
More places hiring for this role
Get new sre operations engineer jobs in United States by email
Daily job updates · Unsubscribe anytime