Jobiba hiring network

Sre Operations Engineer Jobs

197 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current sre operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently. You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. EXAMPLE INITIATIVES You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services. Building AI-assisted tooling for incident triage and response. RESPONSIBILITIES Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking. Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code. Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution. Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations. Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, con

kubernetesgitmachine learning
View job →
S
Supabase
📍 Remote• Full-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We're looking for a Release Engineer (SRE) to join our Release Engineering team (part of EngOps) — a production-operations expert who brings an SRE mindset to how Supabase ships and runs, making deploys safe, observable, and recoverable at scale. Release Engineering's scope has grown well beyond build-and-ship: we increasingly own the operational reliability of the systems that deploy and run Supabase. In this role you'll treat our deployment pipelines, pre-production signal, and the control plane itself as production systems — with SLOs, error budgets, and on-call ownership — and you'll be the person teams lean on when reliability is on the line. This is not a "gatekeeper" role. You'll make the reliable path the easy path: standardising how we deploy, instrumenting what we ship, and ensuring that when something breaks, we detect it quickly and recover quickly. What You'll Be Responsible For In this role, you'll: Own the reliability of Supabase's deployment and release systems, and the control plane they run on, against clear SLOs and error budgets Turn pre-production into a trustworthy signal — standardizing and instrumenting today's fragmented, ad-hoc deployment workflows Drive disaster-recovery readiness, including making environments reproducibly deployable from scratch (untangling undocumented secrets, unclear configuration ownership, and circular service dependencies) Build and operate health and SLO monitoring for critical user flows, using synthetic testing to catch regressions before customers do Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents — which account for a large share of our incident load Participate in on-call, lead blameless po

awskubernetesai
View job →
S
1mo ago

Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. We're looking for an engineer to own the deployment and operational infrastructure of Multigres, our distributed Postgres platform. You'll be responsible for building and maintaining the Multigres Operator, ensuring reliable cloud deployments, and creating the tooling that powers our Kubernetes-based infrastructure. What You’ll Be Responsible for: Build and maintain the Multigres Operator - Maintain our Go-based Kubernetes operator that orchestrates distributed Postgres deployments Architect cloud deployment infrastructure - Design and implement robust deployment patterns for EKS and other Kubernetes platforms Manage storage and networking layers - Work with CSI drivers, persistent volumes, and cross-cloud networking to ensure data reliability and connectivity Develop deployment tooling - Create internal tools and automation for provisioning, scaling, and managing Multigres clusters Ensure operational excellence - Build monitoring, alerting, and diagnostic capabilities into the deployment layer Collaborate across teams - Work with database engineers, SRE, and product teams to deliver seamless deployment experiences You Might Be a Good Fit If You have: Strong systems programming skills - Proficiency in Go and experience building production-grade operators or controllers Deep Kubernetes expertise - Hands-on experience with Kubernetes internals, custom resources, and cloud-managed Kubernetes services (EKS, GKE, AKS) Database operations knowledge - Understanding of database deployment patterns, backup/restore, replication, and high availability Distributed systems experience - Familiarity with consensus protocols, failure scenarios, and designing for resilience Cloud infrastructure background - Experience with cl

kubernetesrestai
View job →

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workflows Pr

mongodbawsazure
View job →
O
Okta
📍 Washington• Full-time• From $165K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Technology, Data and Intelligence Team Message Okta’s Technology, Data and Intelligence (TDI) team delivers the systems, tools, and services that power internal operations across the company. From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering , this role will help build, improve, and maintain our cloud platform services by designing and implementing complex cloud-based engineering enablement systems. With a strong focus on automation, testing, and operational excellence, you will deliver foundational infrastructure capabilities that enable corporate engineering teams to operate securely, reliably, and at scale. What you'll be doing Secure Cloud Infrastructure & Pipelines: Design, build, and modernize scalable cloud environments and development tools while strictly enforcing security policies and standards for regulated environments. Cross-Functional Collaboration & Advocacy: Partner with software engineering teams to champion DevOps and SRE best practices, deliver excellent internal customer service, and actively contribute to Agile workflows (e.g., demos, architecture sessions). Technical Documentation & Operations: Create and maintain comprehensive technical documentation, including network diagrams, runbooks, and disaster recovery procedures to en

pythonawskubernetes
View job →
M
Mongodb
📍 New York City• Full-time• From $157K/yr
1mo ago

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in New York City for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workf

mongodbawsazure
View job →

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in Cork for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workflows Pref

mongodbawsazure
View job →
O
Okta
📍 Washington• Full-time• From $174K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. The Staff SRE, Classified Opportunity This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a technical leader within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. Security Clearance: Active U.S. TS/SCI clearance with Full Scope Poly Compliance Expertise: Proven experience navigating Federal and DoD compliance frameworks, specifically FedRAMP and Impact Level 6 (IL6) What you’ll be doing: Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting hi

pythonsqlpostgresql
View job →
I
Instacart
📍 United States - Remote• Full-time• Remote• From $160K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Are you passionate about technology and ready to dive into the world of Site Reliability Engineering? We are looking for enthusiastic individuals to join our team as Site Reliability Engineer II, where you'll play a crucial role in ensuring the reliability and performance of our platform. This is a top-notch opportunity to learn from skilled engineers and contribute to solving complex technical challenges. You will be involved in monitoring systems, responding to incidents, and developing automation to streamline operations. We are looking for someone eager to grow their skills, learn new technologies, and contribute to a culture of reliability. The Site Reliability Engineering (SRE) team integrates software and systems engineering to design and manage large-scale, distributed, and fault-tolerant systems. The team is responsible for ensuring high reliabili

REMOTEawsazuregcp
View job →
PE
1mo ago

Position: Engineering Manager - Database Job Location: Noida Role Overview We are seeking a Database Engineering Manager (Individual Contributor) with deep expertise in MySQL and strong working knowledge of MongoDB, PostgreSQL, and Cassandra. This role combines hands-on database administration and optimization with strategic ownership of database reliability, automation, and cloud adoption. The candidate will lead by example—driving technical excellence, influencing best practices, and partnering cross-functionally with DevOps, SRE, and product engineering teams to deliver highly available, secure, and scalable database platforms. Key Responsibilities 1. End-to-End Ownership of MySQL databases in production & staging—availability, performance, and reliability. 2. Architect, manage, and support MongoDB, PostgreSQL, and Cassandra clusters for scale and resilience. 3. Define and enforce backup, recovery, HA, and DR strategies across all critical database platforms. 4. Drive database performance engineering—tuning queries, optimizing schemas, indexing, and partitioning for high-volume workloads. 5. Own replication, clustering, and failover architectures ensuring business continuity. 6. Champion automation & AI-driven operations—design self-healing scripts, predictive scaling, and proactive monitoring solutions. Collaborate with Cloud/DevOps teams on AWS database services (RDS, Aurora, DynamoDB, EC2, S3) to optimize cost, security, and performance. 7. Establish monitoring dashboards & alerting mechanisms for slow queries, replication lag, deadlocks, and capacity planning. Ensure compliance & security standards—encryption, auditing, and regulatory requirements. 8. Lead incident management & on-call rotations, ensuring rapid response and minimal MTTR. 9. Act as a strategic technical partner, contributing to database roadmaps, automation strategy, and adoption of AI-driven DBA practices. Required Skills & Experience 1. 6–10 years of p

sqlpostgresqlmysql
View job →

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. At CVS Health, Site Reliability Engineering (SRE) is fundamental to delivering the reliable, secure, and scalable technology experiences that support millions of patients, customers, pharmacists, and healthcare professionals every day. Our SRE organization drives operational excellence across critical healthcare and retail platforms through innovation, automation, observability, and engineering best practices. The Executive Director, Site Reliability Engineering serves as the strategic leader responsible for the reliability, resilience, and performance of CVS Health's retail and pharmacy technology ecosystem. This executive will define and execute a comprehensive reliability strategy, oversee large global engineering teams, and establish a long-term vision for observability, automation, and operational excellence across thousands of store locations. Working closely with senior business and technology leaders, the Executive Director will champion modern SRE practices, accelerate incident response capabilities, and deliver real-time operational visibility that enables proactive issue prevention and exceptional customer and patient experiences. Key Responsibilities Strategic Leadership & Vision Define and lead the enterprise-wide Site Reliability Engineering strategy supporting CVS Health's retail and pharmacy operations. Align reliability and operational

awsazuregcp
View job →

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. As a Staff Site Reliability Engineer (SRE) at GitLab, you’ll help keep all user-facing services and production systems reliable, scalable, and efficient. Our SREs combine a pragmatic operations mindset with strong software engineering practices to drive automation, reduce toil, and improve resilience across our platform. In the Environment Automation specialization, your focus is on operating and automating hundreds of GitLab environments—from initial provisioning to day-to-day maintenance tasks. Unlike other SRE roles, this position centers on automating the lifecycle of many tenant environments, ensuring they remain secur

awsgcpkubernetes
View job →
MR
14 days ago

Want to work in technology at an investment bank? Graduate training, ongoing support, opportunities at leading global employers – the Alumni graduate program gives you everything you need. (And don’t worry, there’s no training bond. No exit fees, no hidden catches).Here at mthree, we pair great graduates with brilliant global businesses. Our clients include tier one investment banks and other organizations across a range of industries, from insurance to healthcare to travel.This is an exciting opportunity to join an FX Front Office support team in a major North American bank working on the Toronto Trading floor, supporting both front office users and a progressive eFX programme. What you'll do: Support IT solutions for various business lines globally including FX, Money Market, STIR and Options Offer technical expertise and support for the systems used Manage and resolving incidents and outages Manage requests for changes, system releases and capacity planning Disaster recovery, planning and execution How the Alumni program works: Apply via this job advert. Complete our assessment process. Get trained at mthree Academy in an online class for 4-8 weeks with other graduates. Join a mthree client for 12-24 months while receiving support and salary increases every 9 months. The vast majority then convert to permanent employees with the client at the end of the program. What you’ll learn at the mthree Academy: How to discuss production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity trading, derivatives, currency, treasury, regulation, and risk. How to write a basic computer program in Python, including user input, common data structures, and flow of control. How to use MySQL to perform CRUD (create, read, update and delete) operations on a relational database stored in

pythonsqlmysql
View job →
MR
14 days ago

Want to work in technology at an investment bank? Graduate training, ongoing support, opportunities at leading global employers – the Alumni graduate program gives you everything you need. (And don’t worry, there’s no training bond. No exit fees, no hidden catches).Here at mthree, we pair great graduates with brilliant global businesses. Our clients include tier one investment banks and other organizations across a range of industries, from insurance to healthcare to travel.This is an exciting opportunity to join an FX Front Office support team in a major North American bank working on the Toronto Trading floor, supporting both front office users and a progressive eFX programme. What you'll do: Support IT solutions for various business lines globally including FX, Money Market, STIR and Options Offer technical expertise and support for the systems used Manage and resolving incidents and outages Manage requests for changes, system releases and capacity planning Disaster recovery, planning and execution How the Alumni program works: Apply via this job advert. Complete our assessment process. Get trained at mthree Academy in an online class for 4-8 weeks with other graduates. Join a mthree client for 12-24 months while receiving support and salary increases every 9 months. The vast majority then convert to permanent employees with the client at the end of the program. What you’ll learn at the mthree Academy: How to discuss production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity trading, derivatives, currency, treasury, regulation, and risk. How to write a basic computer program in Python, including user input, common data structures, and flow of control. How to use MySQL to perform CRUD (create, read, update and delete) operations on a relational database stored in

pythonsqlmysql
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng

pythonsqlpostgresql
View job →
🔔

Get new sre operations engineer jobs by email

Daily job updates · Unsubscribe anytime