At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is a high-growth SaaS observability platform built on the Snowflake AI Data Cloud, enabling businesses to troubleshoot modern distributed applications 10x faster. Now, as a core part of Snowflake, we’ve reached a major milestone in the evolution of the Snowflake platform. By bringing AI-powered observability directly into the Snowflake ecosystem, we’ve created the first truly unified platform for telemetry and business data. We’re looking for a Technical Account Manager to partner with our most strategic enterprise customers and ensure they derive sustained operational value from Observe. This is a hands-on, post-sales technical role focused on long-term platform adoption, optimization, and technical partnership. You will work directly with SRE, DevOps, platform, and engineering teams to embed Observe into daily workflows, evolve telemetry strategy over time, and continuously improve reliability, performance, and cost efficiency. This role is ideal for an experienced observability practitioner who enjoys being deeply embedded with customer teams, solving real production challenges, and acting as a trusted technical advisor in complex enterprise environments. What You’ll Do Serve as the primary technical owner and trusted advisor for assigned strategic a
Jobiba hiring network
Sre Operations Engineer Jobs
197 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current sre operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
**English version below** Doit être local à Montréal Vous souhaitez travailler dans le domaine de la technologie au sein d'une banque d'investissement? Nous recherchons une personne pour rejoindre une équipe dynamique en tant qu’ Ingénieure Fiabilité de Site (Site Reliability Engineer) pour l’un de nos clients. Le Site Reliability Engineering (SRE) est une discipline orientée production, axée sur l’amélioration de la disponibilité des services systèmes, de l’observabilité, de l’évolutivité, de la performance et de la fiabilité des produits technologiques, en appliquant de solides principes d’ingénierie logicielle et en adoptant les technologies et outils les plus récents. Nous serions ravis de vous rencontrer si vous : Vous intéressez aux systèmes distribués et au travail sur des services hautement évolutifs, fiables et à grande échelle. Aimez évoluer dans un environnement dynamique et n’avez pas peur de changer les choses pour les améliorer. Appréciez les nouveaux défis technologiques et la résolution de problèmes complexes. Croyez qu’une équipe qui collabore efficacement est véritablement plus intelligente que la personne la plus brillante qui la compose. Aspirez à évoluer en tant que personne, coéquipier·e et ingénieur·e. Faites preuve de détermination, de motivation et d’un profond sens des responsabilités. À propos de mtrois : Depuis 2010, mtrois aide ses clients à résoudre leurs défis commerciaux et technologiques. Nous sommes une société de conseil en technologie et en affaires avec une main-d'œuvre mondiale qui réalise des projets commerciaux et informatiques significatifs dans certaines des plus grandes organisations de services financiers du monde. Services principaux Consulting et Conseil Services gérés Programme de diplômés Alumni Programme Alumni Pro Nous avons une présence mondiale et sommes experts dans la fourniture d'une qualité exceptionnelle à notre base de clients, offrant des services de conseil dans les domaines du risque, de la réglementa
Want to work in technology at an investment bank? Graduate training, ongoing support, opportunities at leading global employers – the Alumni graduate program gives you everything you need. (And don’t worry, there’s no training bond. No exit fees, no hidden catches). Here at mthree, we pair great graduates with brilliant global businesses. Our clients include tier one investment banks and other organizations across a range of industries, from insurance to healthcare to travel. mthree has an exclusive partnership with Columbia Univ. School of Engineering. All mthree Alumni are eligible to receive two Executive Education certificates from Columbia Engineering as part of their Academy and industry placement experience at no cost. Further, all participating Alumni will have access to the Columbia Engineering network and ongoing training. What you'll do: Production support plays a vital role in enterprise technology, from algorithmic trading engines to regulatory reporting. Think of it as healthcare for technology. As a production support analyst with mthree, you’ll be on a shared mission to look after the technical systems and processes other teams rely on. How the Alumni program works: Apply via this job advert. Complete our assessment process. Get trained at mthree Academy in an online class for 4-8 weeks with other graduates. Join a mthree client for 12-24 months while receiving support and salary increases every 12 months. The vast majority then convert to permanent employees with the client at the end of the program. What you’ll learn at the mthree Academy: How to discuss production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity trading, derivatives, currency, treasury, regulation, and risk. How to write a basic computer program in Python, including user input
Want to work in technology at an investment bank? Paid graduate training, ongoing support, opportunities at leading global employers – the Alumni graduate program gives you everything you need. (And don’t worry, there’s no training bond. No exit fees, no hidden catches). Here at mthree, we pair great graduates with brilliant global businesses. Our clients include tier one investment banks and other organizations across a range of industries, from insurance to healthcare to travel. mthree has an exclusive partnership with Columbia Univ. School of Engineering. All mthree Alumni are eligible to receive two Executive Education certificates from Columbia Engineering as part of their Academy and industry placement experience at no cost. Further, all participating Alumni will have access to the Columbia Engineering network and ongoing training. What you'll do: Production support plays a vital role in enterprise technology, from algorithmic trading engines to regulatory reporting. Think of it as healthcare for technology. As a production support analyst with mthree, you’ll be on a shared mission to look after the technical systems and processes other teams rely on. How the Alumni program works: Apply via this job advert. Complete our assessment process. Get trained at mthree Academy in an online class for 4-8 weeks with other graduates. Join a mthree client for 12-24 months while receiving support and salary increases every 12 months. The vast majority then convert to permanent employees with the client at the end of the program. What you’ll learn at the mthree Academy: How to discuss production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity trading, derivatives, currency, treasury, regulation, and risk. How to write a basic computer program in Python, including user
Locations: South Jordan, UT Salary: $56,000 Launch Your Career in Technology Every app, website, payment, and digital service relies on technology running smoothly behind the scenes. When something goes wrong, Production Support Engineers are the people who investigate the issue, restore service, and help prevent it from happening again. If you're curious, analytical, and enjoy solving problems, this is an opportunity to build hands-on experience with cloud platforms, Linux, automation, databases, and large-scale enterprise systems from day one. What Is Production Support? Production Support Engineers keep business-critical applications running reliably in live environments. Think of it this way: Software Engineers build the platform. QA Engineers test the platform. Production Support Engineers keep the platform running when it matters most. Working at the intersection of technology and business, you'll troubleshoot issues, automate processes, and help improve the reliability and performance of systems used by thousands, or even millions, of people every day. If you enjoy solving puzzles, working under pressure, and understanding how large-scale systems work, this could be the perfect place to start your career. What You'll Do As part of a global production engineering team, you'll: Help support large-scale applications and platforms used by leading organizations around the world. Monitor business-critical applications and services to ensure high availability and performance. Investigate and resolve production incidents across applications, infrastructure, databases, and cloud environments. Analyse logs, alerts, and system metrics to identify root causes and prevent recurring issues. Partner with software engineers, infrastructure teams, and business stakeholders to improve system reliabilit
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest brings millions of people the inspiration to create a life they love. Behind that experience is a complex infrastructure ecosystem that powers reliability, performance, measurement, and efficiency across the platform. As Pinterest grows, it’s increasingly important that we understand these systems clearly so we can make smarter decisions for both Pinners and the business. We’re looking for a Data Scientist to join our Infrastructure Data Science team. In this role, you’ll partner with engineering and cross-functional teams to make Pinterest’s infrastructure more measurable, intelligible, and actionable. Depending on the area, your work may span app performance, shopping infrastructure, metrics quality, infrastructure governance, or site reliability. You’ll help build the data foundations, measurement systems, and analytical fram
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Position Overview: We are seeking a highly technical Staff Observability Site Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to delivering a world class, comprehensive, scalable Observability Platform that enables our SRE teams and business partners. You will treat infrastructure as code —utilizing Terraform and strong coding proficiency in Go, Python, or Ruby —to automate the deployment of agents and collectors across complex distributed systems. Key Responsibilities Automated Infrastructure: Design, build, and maintain scalable observability infrastructure using tools like Terraform. Splunk Engineering: Optimize the collection, processing, and storage of log data to ensure high reliability and low latency of our Splunk services Incident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "observability-driven development." Automation: Eliminate "toil" by automating the deployment and scaling of observability agents and collectors. Required Skills & Experience (The Essentials) Log Management: Minimum 5+ Experience scaling and managing Splunk Cloud at scale (1000+ SVCs), including Workload Management (WLM) and HEC optimization. Visualization: Expertise in creating intuitive, actionable Splunk dashboards that correlate data across multiple sources. SRE Mindset: Minimum 5+ years of experience in an SRE, Dev
As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams. Success in this role means smoother launches, clearer roadmaps, stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale. This role can be based out of our Dublin or Cork office or remotely in Ireland. What You'll Do Drive Program Planning & Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep work clearly tracked in Jira, and deliver on time Strengthen Production Reliability – Lead change management and launch readiness programs. Partner with SREs and product teams to define and operationalize SLOs/SLIs, and use incident data, metrics, and capacity signals to drive prioritization and continuous improvement Lead Cross-Functional Coordination – Align SRE with Security, Compliance, Cloud platform, and other engineering teams. Coordinate cross-team incident response, ensure clear follow-through, and build trust as the go-to driver of complex, multi-team efforts Build Scalable Systems & Processes – Design lightweight frameworks and communication patterns that help SRE deliver reliably at scale. Work yourself out of the "hero" role by leaving teams better-equipped to execute independently Requirements 8+ years in technical program management, engineering management, or a comparable technical role partnering with software engineering teams Proven track record leading large-scale, cross-team platform initiatives through ambiguity and change Strong knowledge of production change management, software development lifecycle, and reliability metrics (SLOs, SLIs) Skilled at shaping roadmaps and managing dependencies Able to query and interpret
As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams. Success in this role means smoother launches, clearer roadmaps, stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale. This role can be based remotely on the East Coast What You'll Do Drive Program Planning & Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep work clearly tracked in Jira, and deliver on time Strengthen Production Reliability – Lead change management and launch readiness programs. Partner with SREs and product teams to define and operationalize SLOs/SLIs, and use incident data, metrics, and capacity signals to drive prioritization and continuous improvement Lead Cross-Functional Coordination – Align SRE with Security, Compliance, Cloud platform, and other engineering teams. Coordinate cross-team incident response, ensure clear follow-through, and build trust as the go-to driver of complex, multi-team efforts Build Scalable Systems & Processes – Design lightweight frameworks and communication patterns that help SRE deliver reliably at scale. Work yourself out of the "hero" role by leaving teams better-equipped to execute independently Requirements 8+ years in technical program management, engineering management, or a comparable technical role partnering with software engineering teams Proven track record leading large-scale, cross-team platform initiatives through ambiguity and change Strong knowledge of production change management, software development lifecycle, and reliability metrics (SLOs, SLIs) Skilled at shaping roadmaps and managing dependencies Able to query and interpret metrics, logs, or other data s
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs. The InfraSec team collaborates closely with other engineering teams to ensure that our infrastructure adheres to the highest security standards. They build essential security infrastructure and implement controls that reinforce the platform’s security posture. This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions.This team is deeply involved in the technical aspects of security and the nuances of its actual implementation. This role can sit in our New York City, Austin, Seattle or San Francisco offices on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones. Responsibilities: Cloud Security Design and Implementation: Help lead the design and deployment of security solutions for cloud platforms (AWS, Azure, GCP), including network and compute security, identity management, and cloud security posture management (CSPM) Automation and Monitoring: Build automated solutions for real-time security monitoring, logging, and alerting in cloud environments. Leverage native cloud services and third-party tools for runtime security monitoring and anomaly detection Security Tooling: Evaluate, implement, and manage cloud-native security tools and platforms for endpoint security, identity management (IAM), and CSPM Qualifications: Experience: 6+ years of experience in SRE, infrastructure engineering or similar role, with a strong focus on security work, with ideally 2+ years in a senior or staff engineering role Security Mindset: A comprehensive understanding of all facets of cloud environment security, spanning from foundational OS networking laye
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Senior Manager, Platform Engineering / DevOps Who Are You You are an experienced Senior Manager / emerging Staff-level leader in DevOps and Platform Engineering with strong technical depth and demonstrated leadership in delivering enterprise-scale cloud platforms. You bring a balanced mix of hands-on engineering expertise, team leadership, and execution rigor. You excel in driving outcomes in complex, multi-stakeholder environments, guiding teams to deliver secure, scalable, and high-quality platform solutions. You are comfortable leading engineers, managing stakeholders, and owning delivery across multiple workstreams. You demonstrate: A strong ownership mindset with accountability for delivery and outcomes Ability to translate business needs into actionable engineering roadmaps Solid expertise in cloud-native platforms, DevOps practices, and SRE principles Capability to lead teams and influence without requiring extensive tenure Role Responsibilities Development & Enforcement Own and execute the H100 platform engineering roadmap, aligned to enterprise priorities and program milestones Drive delivery of GCP-based platform capabilities (GKE, networking, IAM, CI/CD, observability) Establish and enforce engineering standards, best practices, and ADR compliance <li
About AlphaSense: The world’s most sophisticated companies rely on AlphaSense to remove uncertainty from decision-making. With market intelligence and search built on proven AI, AlphaSense delivers insights that matter from content you can trust. Our universe of public and private content includes equity research, company filings, event transcripts, expert calls, news, trade journals, and clients’ own research content. The acquisition of Tegus by AlphaSense in 2024 advances our shared mission to empower professionals to make smarter decisions through AI-driven market intelligence. Together, AlphaSense and Tegus will accelerate growth, innovation, and content expansion, with complementary product and content capabilities that enable users to unearth even more comprehensive insights from thousands of content sets. Our platform is trusted by over 6,000 enterprise customers, including a majority of the S&P 500. Founded in 2011, AlphaSense is headquartered in New York City with more than 2,000 employees across the globe and offices in the U.S., U.K., Finland, India, Singapore, Canada, and Ireland. Come join us! About The Role: Our Site Reliability Engineering team is growing, and we are looking for a highly experienced Staff Site Reliability Engineer to help shape the future of reliability, scalability, and performance at AlphaSense. This is a hands-on, high-impact role where you will architect core reliability platforms, lead by example in incident response, and drive cultural adoption of SRE best practices across our global engineering organization. Our mission is to engineer our platform to the reliability standards of mission-critical systems, targeting 99.99% uptime, while continuously enhancing our systems and processes. This role is key to that mission and goes beyond traditional system maintenance; it’s about pioneering the platforms, practices, and culture that enable engineering to scale effectively. You will act as a force multiplier, mentoring fello
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Senior Manager, Platform Engineering / DevOps Who Are You You are an experienced Senior Manager / emerging Staff-level leader in DevOps and Platform Engineering with strong technical depth and demonstrated leadership in delivering enterprise-scale cloud platforms. You bring a balanced mix of hands-on engineering expertise, team leadership, and execution rigor. You excel in driving outcomes in complex, multi-stakeholder environments, guiding teams to deliver secure, scalable, and high-quality platform solutions. You are comfortable leading engineers, managing stakeholders, and owning delivery across multiple workstreams. You demonstrate: A strong ownership mindset with accountability for delivery and outcomes Ability to translate business needs into actionable engineering roadmaps Solid expertise in cloud-native platforms, DevOps practices, and SRE principles Capability to lead teams and influence without requiring extensive tenure Role Responsibilities Development & Enforcement Own and execute the H100 platform engineering roadmap, aligned to enterprise priorities and program milestones Drive delivery of GCP-based platform capabilities (GKE, networking, IAM, CI/CD, observability) Establish and enforce engineering standards, best practices, and ADR compliance</li
Opportunity Overview: We’re looking for a senior-level automation engineer who will help raise the bar on release quality, environment reliability, and change safety across Cohere’s platform. You’ll partner closely with Product, Engineering, Platform, and SRE to build scalable automation, guardrails, and validation systems that reduce production risk while increasing delivery velocity. This is not a “test scripts only” role. You’ll shape automation strategy, embed quality into the SDLC, and help define how changes move safely from dev → staging → UAT → prod in a fast-moving healthcare platform. You’ll help define how quality scales as Cohere grows. This role has real influence over release safety, platform reliability, and how engineering teams ship software in a regulated, high-impact domain. You won’t just test features — you’ll shape how Cohere delivers them safely to production. What you’ll do: Own and evolve Cohere’s end-to-end test automation strategy across UI, API, config changes, and critical workflows Design and maintain scalable E2E automation frameworks for multi-tenant, payer-specific workflows Build automated validation for deployment guardrails, release readiness, and production change safety Partner with Platform/DevOps to integrate automation into CI/CD pipelines and deployment workflows Create automated coverage for high-risk paths (authorization flows, partner integrations, file pipelines, feature flags, config changes) Drive test reliability, flake reduction, and actionable failure signals Define and enforce quality gates for prod releases, blue/green and canary deployments, and config changes Collaborate with Product and Engineering to ensure business outcomes are testable, measurable, and observable Improve test data management and environment stability to enable reliable automation at scale Mentor engineers on testability, automation best practices, and quality-first development Partner with SRE and Security to ensure production readines
Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is growing quickly. You will play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features, while also balancing scalability, reusability, and performance. Role Overview: We're looking for a Staff Platform Engineer to serve as the technical backbone of our Engineering organization. You'll own the technical strategy, and delivery of our platform — spanning architecture, DevOps, SRE, security, Dev-ex. This is a hands-on staff level role: you'll set technical direction, drive cross-team alignment, and be the senior escalation point for platform challenges. What you’ll do: Drive platform reliability, scalability, security, and cost efficiency across all environments. Technical Leadership: Provide technical leadership for platform components, Influence the technical strategy and architecture of our cloud platform, from CI/CD pipelines to observability and incident response. Design and implement platform components and reusable integration patterns that minimize custom development efforts, reduce the time spent on repetitive tasks, and ensure that integrations scale across multiple healthcare systems Partner closely with Architecture, DevOps, SRE, and Security teams to deliver cohesive platform solutions Cross-Functional Collaboration: Work closely with product teams, and solutions architects to understand integration needs and ensure the platform meets current and future business requirements. Serve as a senior escalation point for infrastructure and platform incidents Establish frameworks for: AI governance and compliance. Observability of systems. Traceability of decisions and outputs. Ensure enterprise readiness with security, auditability, and reliability in production environments. Security & Compliance : Ensure all p
Get new sre operations engineer jobs by email
Daily job updates · Unsubscribe anytime