ABOUT THE ROLE The Reliability Engineer will be working with the Peloton Hardware Design and Operations teams to lead our new product qualification testing and on-going reliability testing in a fast-growing global fitness business. You will lead Projects, troubleshoot, innovate, and make decisions to help Peloton ensure quality and safety as we continue to innovate our product offerings. You are an individual who is independent, highly motivated, super organized, hyper-flexible and enjoys working in a collaborative team environment to improve product quality. YOUR DAILY IMPACT AT PELOTON Build and maintain test plans for new product development and sustaining engineering activities, including sample size estimation and reliability goal Perform product validation and verification with off-the-shelf testing equipment as well as custom built machinery. Lead the execution to keep the schedule on track Manage ODM/OEM partners’ lab and third party lab to complete test task with their own test resource Lead test development projects.Design, build, and automate custom test rigs/fixtures to interact with various pieces of exercise equipment Deliver Validation status and results regularly to facilitate tracking the progress with cross functional team Summarize Validation result with field failure rate estimation to help crossfunctional team understand the margin impact Correspond with NY and Seattle based team with test reports and status updates, which will require morning/evening calls Provide technical support to technicians for various testing activities Actively perform preliminary root cause/failure analysis to collaborate with & drive cross functional team to resolution Utilize statistical tools to compile the result into impact to real world application Work multi-functionally with other departments and prioritize work queues through scoping the request and capturing business impact Standardize test documentation, including fixtures, test procedures, schematics an
Jobiba hiring network
Reliability Engineer Jobs
2,049 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office, you'll help build and operate the foundational infrastructure that powers PagerDuty's real-time digital operations platform. Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You'll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases. Key Responsibilities ● Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems. ● Contribute to the reliability and scalability of PagerDuty's core platform by hardening existing systems and supporting the rollout of new infrastructure
PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. As a Site Reliability Engineer II on the Core Infrastructure team in our Atlanta office, you'll help build and operate the foundational infrastructure that powers PagerDuty's real-time digital operations platform. Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You'll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases. Key Responsibilities ● Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems. ● Contribute to the reliability and scalability of PagerDuty's core platform by hardening existing systems and supporting the rollout of new infrastructur
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve our production infrastructure. This is a single application for Site Reliability Engineering opportunities across our Infrastructure Platforms department. Rather than asking you to choose the right team or level upfront, we evaluate your skills holistically and match you to the opportunity that best aligns with your experience
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams. We are looking to speak to candidates who can work East Coast hours. The ideal candidate should Have 6+ years of experience in software development and operating distributed systems Proficiency in Python, Go, or a similar language Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Expectations Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products This includes, but is not limited to, contributing to open-source projects, or engineering software-based
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams. We are looking to speak to candidates who can work East Coast hours. The ideal candidate should Have 6+ years of experience in software development and operating distributed systems Proficiency in Python, Go, or a similar language Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Expectations Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products This includes, but is not limited to, contributing to open-source projects, or engineering software-based
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta’s Workforce Identity Cloud Security Engineering group is looking for an experienced and passionate Staff Site Reliability Engineer to join a team focused on designing and developing Security solutions to harden our cloud infrastructure. We embrace innovation and pave the way to transform bright ideas into excellent security solutions that help run large-scale, critical infrastructure. We encourage you to prescribe defense-in-depth measures, industry security standards and enforce the principle of least privilege to help take our Security posture to the next level. Our Infrastructure Security team has a niche skill-set that balances Security domain expertise with the ability to design, implement, rollout infrastructure across multiple cloud environments without adding friction to product functionality or performance. We are responsible for the ever-growing need to improve our customer safety and privacy by providing security services that are coupled with the core Okta product. This is a high-impact role in a security-centric, fast-paced organization that is poised for massive growth and success. You will act as a liaison between the Security org and the Engineering org to build technical leverage and influence the security roadmap. You will focus on engineering security aspects of the systems used across our services. Join us and be part of a company that is about to change the cloud computing landscape forever. Bring all the passion and dedicat
We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Are you passionate about technology and ready to dive into the world of Site Reliability Engineering? We are looking for enthusiastic individuals to join our team as Site Reliability Engineer II, where you'll play a crucial role in ensuring the reliability and performance of our platform. This is a top-notch opportunity to learn from skilled engineers and contribute to solving complex technical challenges. You will be involved in monitoring systems, responding to incidents, and developing automation to streamline operations. We are looking for someone eager to grow their skills, learn new technologies, and contribute to a culture of reliability. The Site Reliability Engineering (SRE) team integrates software and systems engineering to design and manage large-scale, distributed, and fault-tolerant systems. The team is responsible for ensuring high reliabili
Who we are About Stripe Stripe is a technology company focused on improving the conditions for economic growth and prosperity. We build programmable financial infrastructure, rethinking from first principles how financial services should work, to make it easier and cheaper for any business to start and scale. More than 10 million businesses build on Stripe, spanning the economic frontier—from solo founders to established enterprises—united by a practical focus on growth. The most ambitious companies in the world use Stripe as core infrastructure to grow faster. They process trillions of dollars a year on Stripe, equivalent to around 1.6% of global GDP. While economic growth makes everyone better off, open markets also enable greater variety. When any business can easily serve a global customer base, the quality and diversity of products in the world increase, and craft and creativity are unleashed into the smallest niches. Our own growth is wholly contingent on the success of the businesses building on Stripe. We therefore invest back into our technology at an unusual rate. We make upgrades to our products every single day to deliver compounding gains to our customers. We maintain some of the most reliable APIs on the internet. We build entirely new pieces of financial infrastructure to enable new ideas. And our significant advances in risk and fraud infrastructure over many years are making the internet economy safer and more accessible. Though people at Stripe don’t tend to take themselves seriously, Stripe is a fairly serious place: our customers are depending on us for their livelihoods. We admire ambition, intensity, curiosity, humility, and rigor. The most effective people become knowledgeable about many domains besides their own. Any company is an applied exercise in understanding some aspect of society or the market. In working with so many (especially the new and innovative ones), we think that Stripe is one of the very best places to learn about how
Who we are About Stripe Stripe is a technology company focused on improving the conditions for economic growth and prosperity. We build programmable financial infrastructure, rethinking from first principles how financial services should work, to make it easier and cheaper for any business to start and scale. More than 10 million businesses build on Stripe, spanning the economic frontier—from solo founders to established enterprises—united by a practical focus on growth. The most ambitious companies in the world use Stripe as core infrastructure to grow faster. They process trillions of dollars a year on Stripe, equivalent to around 1.6% of global GDP. While economic growth makes everyone better off, open markets also enable greater variety. When any business can easily serve a global customer base, the quality and diversity of products in the world increase, and craft and creativity are unleashed into the smallest niches. Our own growth is wholly contingent on the success of the businesses building on Stripe. We therefore invest back into our technology at an unusual rate. We make upgrades to our products every single day to deliver compounding gains to our customers. We maintain some of the most reliable APIs on the internet. We build entirely new pieces of financial infrastructure to enable new ideas. And our significant advances in risk and fraud infrastructure over many years are making the internet economy safer and more accessible. Though people at Stripe don’t tend to take themselves seriously, Stripe is a fairly serious place: our customers are depending on us for their livelihoods. We admire ambition, intensity, curiosity, humility, and rigor. The most effective people become knowledgeable about many domains besides their own. Any company is an applied exercise in understanding some aspect of society or the market. In working with so many (especially the new and innovative ones), we think that Stripe is one of the very best places to learn about how
DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You'll Do: Deploy, configure, and maintain Kubernetes clusters for our microservices architecture. Utilize Git and Helm for version control and deployment management. Implement and manage monitoring solutions using Prometheus and Grafana. Work on continuous integration and continuous deployment (CI/CD) pipelines. Containerize applications using Docker and manage orchestration. Manage and optimize AWS services, including but not limited to EC2, S3, RDS, and AWS CDN. Maintain and optimize MySQL databases, Airflow, and Redis instances. Write automation scripts in Bash or Python for system administration tasks. Perform Linux administration tasks and troubleshoot system issues. Utilize Ansible and Terraform for configuration management and infrastructure as code. Demonstrate knowledge of networking and load-balancing principles. Collaborate with development teams to ensure applications meet reliability and performance standards. Who you are: Bachelor’s degree in engineering (CS / IT) or equivalent degree from a well-known Institute / University. 2+ years of experience in a Site Reliability Engineer role or similar. Proven experience with Kubernetes, Git, Helm, Prometheus, Grafana, CI/CD, Docker, and microservices architecture. Strong knowledge of AWS services, MySQL, Airflow, Redis, AWS CDN. Proficient in scripting languages such as Bash or Python. Hands-on experience with Linux administration. Familiarity with Ansible and Terraform fo
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. As a Staff Site Reliability Engineer (SRE) at GitLab, you’ll help keep all user-facing services and production systems reliable, scalable, and efficient. Our SREs combine a pragmatic operations mindset with strong software engineering practices to drive automation, reduce toil, and improve resilience across our platform. In the Environment Automation specialization, your focus is on operating and automating hundreds of GitLab environments—from initial provisioning to day-to-day maintenance tasks. Unlike other SRE roles, this position centers on automating the lifecycle of many tenant environments, ensuring they remain secur
About Backblaze Backblaze is the object storage leader in the open cloud movement, fueling customer success with cloud storage built purposefully to unlock budgets, unburden administrators, and unleash innovators. Together with our partners, we’re helping customers break free from the restrictive, overpriced legacy solutions that hold them back, and blaze forward with the full power of the open cloud in their hands. Founded in 2007, we scaled the business with less than $3 million in outside funding until 2021, when we did a traditional IPO on the Nasdaq stock exchange. Today, Backblaze generates over $136M ARR and is the leading specialized storage cloud, managing over three billion gigabytes of data storage for 500K+ customers in 175+ countries, including businesses, developers, IT professionals, and individuals. But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us. We’re seeking a Sr. Reliability Engineer ll (DBA) to join our team! About the Role We are seeking a Site Reliability Engineer (SRE) with a DBA (Database Administration) focus to help ensure the stability, scalability, and reliability of our production database systems - primarily Vitess (distributed MySQL) and Cassandra - alongside the rest of our services and infrastructure. This role operates within procedures and runbooks established by our senior DBA SREs, and focuses on building automation, maintaining observability, and supporting incident response to keep customer-facing systems performing at their best. The SRE will collaborate with engineering, product, and operations teams to embed reliability practices into day-to-day development and operations while contributing to tools and processes that improve efficiency and reduce manual effort Key Responsibilities Database Administration Operating and maintaining high-availability database systems — primarily Vitess (distributed MySQL) and Cassandra — against established architecture and runbooks. Op
Reolink , a leader in intelligent visual technology for homes and businesses, was founded in 2009 by a group of engineers with a strong commitment to and passion for smarter security solutions. Our products are now trusted by millions of users across more than 110 countries and regions worldwide. Building on this trust, we continue expanding our presence and bringing our innovations to more markets around the globe. Reolink remains committed to delivering advanced, reliable, and user‑centric solutions that empower people to protect what matters most. 5 Work Days Per Week Office at Tai Seng Exchange Tower B Near Tai Seng MRT, Singapore Insurance Coverage Entitled to Yearly Bonus & Performance Bonus Responsibilities (Site Reliability Engineer - Senior / Lead ) High Availability and Stability Maintenance of Application Systems: Includes daily monitoring, alert response, emergency handling, on-call duties, regular system health checks, and performance optimization. Compliance and Secure Access Construction for Application Systems: Ensure operational design, processes, and data management comply with relevant privacy and data protection laws. Ensure compliance with full auditing and regulatory checks and provide auditing materials as required. Change and Release Management: Best practices for application system changes, including change control, version management, and rollback strategies, while ensuring operational duties during release windows. Automation and Infrastructure Optimization: Drive operational automation by designing and implementing automated tools and processes, ensuring resource allocation is optimized and supporting business scalability. Other Operational Practices and Work Arrangements: Providefeedback and suggestions for business architecture design and continuously produce operational technical documentation. Qualifications Bachelor's Degree or above; a degree in computer science or a related field is preferred. Experiences as Senior SRE or
What We Do At GoGuardian, we’re helping build a future where all learners are ready and inspired to solve the world’s greatest challenges. Our award-winning system of learning solutions is purpose-built for K-12 and trusted by school leaders to promote effective teaching and equitable engagement while helping empower educators to keep students safe. What It’s Like to Work at GoGuardian We are an outcomes-focused learning company with a steadfast focus on improving learning environments, one classroom at a time. Working with us means joining a remote team of diverse, committed, mission-driven employees who are inspired by our vision, dedicated to our customers, and ready to roll up their sleeves. Guardians put their heads together to solve problems, learn together from experiments that fail, and stand together by their work with full accountability. We balance our diligence with an inclusive culture that invites everyone to bring their whole self to work. Join us and learn why “I love the people here” is one of the most frequent comments we hear from Guardians. The Role We’re looking for a Site Reliability Engineer (SRE) II to help build, maintain, and scale the infrastructure that powers our core products and services. In this role, you’ll work alongside engineering teams to support operational excellence, optimize system performance, and ensure high availability across production environments. This position sits on Tech Foundation, a team that manages core cloud infrastructure, shared data services, and developer tooling to empower our product teams to deliver software efficiently and securely. The ideal candidate has practical experience with cloud infrastructure, automation, and core reliability practices, with a strong desire to solve complex operational challenges in a collaborative environment. _________________________________________________________________________________________ What You'll Do Build, maintain, and support scalable cloud
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime