Jobs in Canada

Reliability Engineer in Canada

138 active opportunities · Updated October 2026

Explore current reliability engineer jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.

L
📍 Toronto, Canada
✓ Quality checkedCompany trend -72.4%

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Rider Loyalty team is where riders become members. We build the membership, rewards, and benefits products that give people a reason to choose Lyft on every trip, and we make sure the value a rider has earned shows up at the moment it matters. Loyalty sits inside the Rider Loyalty, Partnerships, and Rider Pay (PLP) group. You will lead a team of engineers across iOS, Android, and Server. You will own the membership and rewards platform end to end and work daily with Product, Design, Data Science, and Partnerships. Responsibilities: Own the Loyalty roadmap from strategy through delivery. Turn goals like member growth and retention into an engineering plan, and manage the dependencies that run through Partnerships and Rider Pay. Build and scale the systems behind membership, rewards earning and redemption, and benefit delivery. Hold a high technical bar through architecture reviews, tech debt management, observability, reliability, and on-call. Grow engineers by matching people to the right opportunities, setting clear expectations, and giving feedback early. Experience: 5+ years building software professionally, including 2+ years directly managing engineers. You have managed a team that shipped both mobile and backend work, and you can still read and review code in at least one of those areas. You have owned a consumer product used by millions of people each month. You use AI tools in your own work and have a clear view of where they help and where they do not. BS/MS in Computer Science, Computer Engineering, or a related field, or equivalent practical experience. Benefits: Extended health and dental coverage options, along with life insurance and disability benefits Mental health benefits Family building benefits Child care and pet benefits Access to a Lyft funded Health Care Savings Accou

Artificial IntelligenceAI
R
📍 Menlo Park, CA· Full-time
✓ Quality checkedCompany trend -76.2%

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Load and Fault team sits within Robinhood's Developer Infrastructure organization, with a mission to give every engineering team the tools they need to test their services under real-world conditions before those conditions test them in production. We build the platforms and frameworks that enable load testing, fault injection, and resilience validation at scale — treating reliability as a developer productivity problem, not just an operations one. Our work directly raises the quality bar for every service Robinhood ships, and we partner closely with engineering teams across the organization to make resilience testing a seamless part of the development workflow. As a Senior Software Engineer on the Load and Fault Environments team, you will design and build the infrastructure that lets Robinhood's engineers simulate load, inject faults, and validate system behavior under stress — at the scale of a fast-growing financial platform. You'll own meaningful components of the load testing and fault injection platform, write production-quality code, and collaborate with engineers across infrastructure and product teams to ensure the tooling you build gets adopted and drives real

PythonVueAWSKubernetes
C
📍 Canada· Full-time
✓ Quality checkedCompany trend -100%

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 ClickUp is looking for an experienced Engineering Manager to lead our fullstack team responsible for building and scaling our flagship products. As the leader of the team that owns the APIs and core experiences powering ClickUp, you will play a pivotal role in shaping the future of our platform. You will guide engineers working across the stack, from frontend experiences to backend infrastructure. Your focus will be on driving the development of new features, addressing performance and reliability challenges, and ensuring operational excellence as we continue to grow. This is an opportunity to make a significant impact on our core product while fostering a culture of technical excellence and collaboration. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our produ

Node.jsAWSMachine LearningAI
S
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -100%

C$162K – C$420K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Events Analytics Platform (EAP) team is responsible for the infrastructure that powers all of Sentry's time-series data and searching capabilities across billions of events with sub-second latency. We started this initiative by building Snuba, the primary storage and query service for Sentry's event data powered by ClickHouse, and we are now focused on unlocking deeper visibility and reporting across the terabytes of event data our users generate. As a Senior Software Engineer, you will lead efforts to push the boundaries of data visibility at Sentry. You will do this by expanding the capabilities of our search infrastructure, building new capabilities on top of our state-of-the-art storage layer and increasing the performance and integrity of Sentry’s core data services. You will also help shape Infrastructure's technical direction at Sentry and collaborate with Product and other Engineering teams to turn that vision into a reality. If you want to solve the hard problems that come with scaling event data into the petabyte range, this could be the job for you. In this role you will: Expand EAP's ability to deliver data at world-class speed and reliability. Architect and automate services and systems to scale reliably under growing demand. Make architectural trade-offs that balance product requirements with engineering constraints. Maintain and grow the team's code quality initiatives by regularly reviewing code and contributing to design decisions. Lead design and discussions around deliverables the team is working towards. Improve the maintainability and developer experience of the codebases EAP owns. Exa

PythonSQLPostgreSQLRedis
S
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -100%

C$162K – C$420K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. About the role As a Senior Software Engineer, you will architect and scale the core systems that power Sentry's billing infrastructure, ensuring accuracy and reliability at massive scale. You will collaborate on building the next generation of Sentry’s usage tracking pipeline, processing hundreds of billions of events daily with low latency and financial-grade accuracy. You will help design flexible pricing primitives that support everything from per-event usage billing to complex enterprise contracts, enabling product and sales teams to experiment rapidly while maintaining revenue accuracy and reduced time-to-market for new products. You will contribute to technical decisions on data consistency challenges unique to billing—like handling event delays, retroactive pricing changes, and distributed count reconciliation across our infrastructure. You'll love this job if you Want to solve the "easy to explain, hard to build" problems—like ensuring a customer's bill matches their usage perfectly, even when processing hundreds of billions of events daily across distributed

O
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -63.6%

From C$184K/yr

Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know the team The Developer Platform team at Auth0 (an Okta company) owns the platform that developers build identity on. Increasingly they build alongside AI agents, and that raises the bar on everything underneath: interfaces have to hold up whether a person or an agent is calling them, and the systems behind them have to stay reliable and coherent as usage grows. We move fast, we own problems end to end, and we care deeply about the platform we put in front of the developers and agents who depend on it. The opportunity We're hiring a Principal Engineer (P5) to serve as the technical leader and compass for the Developer Platform team. You'll work across the breadth of the platform, tackling the highly complex, vaguely specified problems that span it and turning them into clear technical direction the team can execute against, without day-to-day oversight. Above all, you'll own how the platform is architected to scale: the distributed systems behind it, the reliability and consistency guarantees developers depend on, and the coherence that keeps it easy to build on as usage grows. You'll champion the team's technical execution, raise the engineering bar, mentor the people around you, and partner with tech leads across teams to keep the wider platform aligned. You'll have real influence over how our platform holds up in a world where developers and agents are both first-class consumers. What you'll be doing Own the platform architecture: Set the tech

Node.jsAWSRestMachine Learning
O
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -63.6%

From C$160K/yr

Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Streaming Foundations team builds services and operates data pipeline infrastructure to support event streaming, messaging, and analytics use cases. We are looking for a Software Engineer who is passionate about distributed systems, platform engineering, and solving data-intensive problems at scale. In this high-impact role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Auth0 to scale for years to come. What you’ll be doing Help set the technical direction for the team and influence the engineering roadmap for the Platform’s streaming capabilities Design and lead the implementation of our most complex and critical systems for data-intensive use cases. Research and champion new technologies and architectural patterns to solve strategic challenges and scale the platform. Lead and influence cross-functional initiatives, ensuring technical alignment and successful execution across multiple teams. Improve the operational posture of our systems by designing for observability, reliability, and scalability, and by mentoring others in operational best practices. Coach and mentor senior engineers and act as a technical leader across the engineering organization. Collaborate with different stakeholders like product teams whenever needed. What you’ll bring to our teams 7+ years of software development experience in a fast-paced, agile environment Experience working with Golang or Java is preferred H

TypeScriptJavaReactAWS
R
📍 Menlo Park, CA· Full-time
✓ Quality checkedCompany trend -76.2%

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Observability team's mission is to build and own Robinhood's full-stack observability platform — the foundation that keeps every product, service, and customer experience running reliably at scale. We design and operate the systems that give engineers deep visibility into how Robinhood's infrastructure behaves, ensuring that when something goes wrong, the right people know immediately and can act fast. Our work spans metrics, logs, distributed tracing, and alerting pipelines, and we partner closely with the Robinhood Command Center to ensure our observability systems meet or exceed 99.9% uptime. We believe in monitoring our own monitors — the observability infrastructure is a product, not just a tool. As a Staff Software Engineer on the Observability team, you will be the technical lead shaping the roadmap for how Robinhood observes itself at scale. You will own the observability control plane end-to-end, making architectural decisions that directly impact the reliability and operational health of Robinhood's entire product surface. You'll lead a team of six engineers, drive the strategy for cost-efficient telemetry ingestion, build in-house solutions where off-the-shelf

PythonVueAWSKubernetes
R
📍 Menlo Park, CA· Full-time
✓ Quality checkedCompany trend -76.2%

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team & role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Robinhood Command Center (RCC) is a newly formed reliability team that serves as the front line for detecting, coordinating, and mitigating production incidents across Robinhood. As part of Robinhood’s broader reliability initiative, RCC works closely with product engineering, reliability, observability, infrastructure, and business teams to reduce customer impact and shorten incident duration. As a Senior Engineer, you will be part of the founding RCC team, helping define how Robinhood responds to and learns from incidents at scale. This is a highly visible role focused on incident leadership, operational excellence, and reliability tooling. You will not own product services or core infrastructure, but you will own the processes and tools that enable fast, high-quality incident response. This role is based in our Menlo Park, California office, with in-person attendance expected at least 3 days per week. What you'll do: Serve as a senior technical leader driving the long-term reliability and observability strategy across Robinhood’s infrastructure Partner closely across many different types of engineers to raise the bar for operational excellence and incident r

VueAWSAIGo
R
📍 Bellevue, WA; Menlo Park, CA· Full-time
✓ Quality checkedCompany trend -76.2%

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Security Platform team is responsible for the secure lifecycle, governance, and protection of Robinhood’s most sensitive user data. This team builds foundational systems such as secure data pipelines, tokenization services, and privacy compliance infrastructure to ensure customer data is handled responsibly and in line with regulations. The team works closely with security, infrastructure, and product engineering partners to ensure data is both usable and protected. You will contribute to systems that support authentication, third-party integrations, and emerging AI-driven use cases. As a Senior Software Engineer, you will design and build backend systems that securely process and manage customer data across Robinhood’s platform. You will own the systems that handle authentication, authorization, and privacy-preserving data operations. This role involves close collaboration with engineers focused on access management, infrastructure, and data systems. Your work will directly support efforts to improve system reliability, strengthen data protections, and enable new product capabilities using secure data. This role is based in our Bellevue, WA, and Menlo Park, CA off

PythonJavaVueSQL
R
📍 Menlo Park, CA· Full-time
✓ Quality checkedCompany trend -76.2%

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Data Engineering team builds and maintains the foundational datasets that power decision-making across Robinhood. We design reliable, scalable data systems that support product analytics, growth strategy, financial reporting, experimentation, and machine learning. The team partners closely with Product, Engineering, Data Science, and Finance to ensure accurate, well-modeled data is available to teams across the company. Our work directly influences how Robinhood measures performance, improves customer experience, and scales its products. As a Senior Data Engineer, you will design, build, and evolve core datasets that track product performance and company-wide metrics. You will develop scalable data pipelines that ingest application events and database snapshots into our data lake, ensuring high data quality and reliability. You’ll collaborate with application engineers to improve data generation patterns and with analytics teams to design intuitive, well-documented data models. This is an opportunity to shape the technical foundation that supports data-informed decisions across the organization! This role is based in our Menlo Park, CA office, with in-person attend

PythonVueSQLAWS
L
📍 San Francisco, CA· Full-time
✓ High-confidence listingCompany trend -72.4%
Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft is looking for an Engineering Manager from across multiple disciplines. We are growing our team with people who want to build, improve, and incorporate technologies that make the lives of our community more enriched. As an Engineering Manager at Lyft, you'll collaborate with other teams and orgs, product managers, designers, data scientists, analysts, and operations on technology that empowers us to iterate quickly, while focusing on delighting our passengers and drivers. The Rider organization is focused on building a seamless, best-in-class rideshare experience for riders. From the foundational functionality of requesting a ride to the tailored interactions with specific verticals like your flight, we sweat the small stuff to help make Lyft the best transportation solution. As an Engineering Manager on the Rider vertical team, you will act as a critical technical leader, taking holistic ownership of supporting current projects and delivering new products from 0 to 1. This includes developing complex systems, defining strategic roadmaps, driving cross-functional alignment, and ensuring engineering excellence to improve the rideshare experience. Responsibilities: Drive team execution, proactively resolve bottlenecks and make decisive trade-offs. Partner with cross-functional teams (Product, Design, Marketing, Science, and Analytics) to define the team's strategic direction. Translate high-level business goals into actionable projects. Own a team roadmap from conception to delivery, managing cross-team dependencies and mitigating risks. Maintain operational excellence through contributing to best practices for observability, reliability, and on-call processes. Ensure technical excellence through architecture reviews, tech debt management, and engineering guidance. Maintain team health

L
📍 San Francisco, CA· Full-time
✓ High-confidence listingCompany trend -72.4%
Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. As an Engineering Manager on the Core Rider team, you will act as a critical technical leader in a highly visible area of the Rider Organization. Taking holistic ownership of our foundational systems and core rider functionality, you will be directly responsible for moving top-line business metrics and delivering a seamless rideshare experience. You will partner with multiple product managers, data scientists, and cross-functional teams to develop complex systems, define strategic roadmaps, scale the product and infrastructure that powers how millions of riders request and experience their rides every day. Responsibilities: Drive team execution, proactively resolve bottlenecks and make decisive trade-offs. Partner with cross-functional teams (Product, Design, Marketing, Science, and Analytics) to define the team's strategic direction. Translate high-level business goals into actionable projects. Own a team roadmap from conception to delivery, managing cross-team dependencies and mitigating risks. Maintain operational excellence through contributing to best practices for observability, reliability, and on-call processes. Ensure technical excellence through architecture reviews, tech debt management, and engineering guidance. Maintain team health by creating and refining team processes, fostering a supportive team culture, and managing resourcing. Develop team member careers by matching them with opportunities, setting clear expectations, and providing timely feedback. Build relationships across engineering teams to share learnings, align on standards, and identify collaboration opportunities. Serve as a dependable, high-bar interviewer and active participant in Lyft’s broader community. Experience: BS/MS or equivalent in Computer Engineering, Computer Science, or a related field, or equivalent practic

L
📍 San Francisco, CA· Full-time
✓ High-confidence listingCompany trend -72.4%
Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Lyft Business Product Platform team builds the systems and experiences that power Lyft's B2B products — enabling companies, organizations, and their employees to seamlessly access Lyft's transportation network. We sit at the intersection of product and platform, owning both the customer-facing features and the underlying infrastructure that makes them reliable at scale. Our work directly impacts how businesses integrate with Lyft, how admins manage their programs, and how millions of riders get where they need to go. Responsibilities: Drive architecture and technical design for systems that are highly available, scalable, and built to last — not just for today's requirements but for where the product is heading Own features end-to-end: from shaping the technical spec and design through to production rollout and operational health Think critically about how AI capabilities can be incorporated into Lyft Business products to improve the experience for business admins and riders — and bring that perspective into roadmap and architecture conversations Make well-reasoned trade-off decisions and communicate them clearly to peers, leads, and cross-functional partners Write clean, well-tested, maintainable code and hold a high bar for the same in code reviews Partner across engineering, product, and design to align on direction and get buy-in on technical approaches Proactively engage in incident response, contributing both to resolution and to long-term reliability improvements Grow the team's technical culture through design reviews, tech talks, and mentorship Experience: 5+ years of software engineering experience, with a track record of designing and shipping production systems at scale Strong system design instincts — you can reason through distributed systems trade-offs, identify failure modes, and

SQLAWSAzureGCP
S
📍 South San Francisco, CA· Full-time
✓ High-confidence listingCompany trend -91.4%

$212K – $318K/yr

Quick readStrong listing-quality and freshness signals

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Design state-of-the-art ML models and large-scale ML systems for underwriting and portfolio management for Stripe Capital based on ML principles, domain knowledge, risk, regulatory and engineering constraints. Design systems to speed up the time from idea to deployment of new models. Experiment and iterate on ML models (using tools including PyTorch and TensorFlow) to achieve key business goals and drive efficiency. Develop pipelines and automated processes to train and evaluate models in offline and online environments. Integrate ML models into production systems and ensure their scalability and reliability. Collaborate with product and strategy partners to propose, prioritize, and implement new product features. Engage with the latest developments in ML/AI and take calculated risks in transforming innovative ML ideas into productionized solutions. Who you are Minimum requirements Must have a Bachelor's degree or foreign equivalent in Computer Science, Machine Learning, Mathematics, Physics, Statistics, or a related field, plus two (2) years of experience in Building and shipping ML systems in production. Must have two (2) years of experience in each of the following: ML algorithms and model architectures; Designing, training and evaluating machine learning models; Productionizing and deploying machine learning models at scale; Orchestrating data pipelines and leveraging large-s

Machine LearningAIGo
🔔

Get new reliability engineer jobs in Canada by email

Daily job updates · Unsubscribe anytime