At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. As a Senior Software Engineer, Data on the Mapping team, you will collaborate with our world-class team of engineers, product managers, and scientists to grow and improve the quality of recommended routes and accuracy of our travel time estimations. You will lead the architecture and long-term technical direction of our offline experimentation tooling and route simulation services — the systems that let Lyft test routing changes safely before they reach production. You'll also build scalable data pipelines for experimentation, analytics, and machine learning models, along with the data governance and observability systems that keep them trustworthy. Your work will enable integration with partner teams and allow stakeholders across Engineering, Data Science, and Product to make data-informed decisions that directly impact Lyft’s growth and profitability. Our technology stack is based on the latest technologies such as AWS, Databricks, Kubernetes and Airflow. You will work with incredibly passionate and talented colleagues from software engineering, machine learning and data science on projects that directly impact millions of riders and drivers. Responsibilities Own core data pipelines end-to-end, building deep subject matter expertise in the systems you manage and defining/managing SLAs for pipelines, services, and datasets to ensure reliability at scale Serve as the technical owner and architectural lead for our offline experimentation platform and route simulation services, setting technical direction, evaluating trade-offs, and ensuring the systems scale with Lyft's routing and mapping ambitions Continuously evolve data models and schemas to meet business and engineering requirements Develop AI tools that support self-service management of data pipelines (ETL) and schema evolution, and perform han
Jobs in Canada
Software Reliability Engineer in Canada
493 active opportunities · Updated October 2026
Showing
15 jobs
Explore current software reliability engineer jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Load and Fault team sits within Robinhood's Developer Infrastructure organization, with a mission to give every engineering team the tools they need to test their services under real-world conditions before those conditions test them in production. We build the platforms and frameworks that enable load testing, fault injection, and resilience validation at scale — treating reliability as a developer productivity problem, not just an operations one. Our work directly raises the quality bar for every service Robinhood ships, and we partner closely with engineering teams across the organization to make resilience testing a seamless part of the development workflow. As a Senior Software Engineer on the Load and Fault Environments team, you will design and build the infrastructure that lets Robinhood's engineers simulate load, inject faults, and validate system behavior under stress — at the scale of a fast-growing financial platform. You'll own meaningful components of the load testing and fault injection platform, write production-quality code, and collaborate with engineers across infrastructure and product teams to ensure the tooling you build gets adopted and drives real
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The internal infrastructure team is responsible for building world-class infrastructure and tools used to train, evaluate and serve Cohere's foundational models. By joining our team, you will work in close collaboration with AI researchers to support their AI workload needs on the cutting edge, with a strong focus on stability, scalability, and observability. You will be responsible for building and operating superclusters across multiple clouds. Your work will directly accelerate the development of industry-leading AI models that power Cohere's platform North. Please Note: All of our infrastructure roles require participating in a 24x7 on-call rotation, where you are compensated for your on-call schedule. As a Staff Software Engineer, you will: Build and scale ML-optimized HPC infrastructure : Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads. Optimize for AI/ML training : Collaborate with cloud providers to fine-tune infrastructure for cost efficiency, reliability, and performance , leveraging technologies like R
C$162K – C$420K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Events Analytics Platform (EAP) team is responsible for the infrastructure that powers all of Sentry's time-series data and searching capabilities across billions of events with sub-second latency. We started this initiative by building Snuba, the primary storage and query service for Sentry's event data powered by ClickHouse, and we are now focused on unlocking deeper visibility and reporting across the terabytes of event data our users generate. As a Senior Software Engineer, you will lead efforts to push the boundaries of data visibility at Sentry. You will do this by expanding the capabilities of our search infrastructure, building new capabilities on top of our state-of-the-art storage layer and increasing the performance and integrity of Sentry’s core data services. You will also help shape Infrastructure's technical direction at Sentry and collaborate with Product and other Engineering teams to turn that vision into a reality. If you want to solve the hard problems that come with scaling event data into the petabyte range, this could be the job for you. In this role you will: Expand EAP's ability to deliver data at world-class speed and reliability. Architect and automate services and systems to scale reliably under growing demand. Make architectural trade-offs that balance product requirements with engineering constraints. Maintain and grow the team's code quality initiatives by regularly reviewing code and contributing to design decisions. Lead design and discussions around deliverables the team is working towards. Improve the maintainability and developer experience of the codebases EAP owns. Exa
C$162K – C$420K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. About the role As a Senior Software Engineer, you will architect and scale the core systems that power Sentry's billing infrastructure, ensuring accuracy and reliability at massive scale. You will collaborate on building the next generation of Sentry’s usage tracking pipeline, processing hundreds of billions of events daily with low latency and financial-grade accuracy. You will help design flexible pricing primitives that support everything from per-event usage billing to complex enterprise contracts, enabling product and sales teams to experiment rapidly while maintaining revenue accuracy and reduced time-to-market for new products. You will contribute to technical decisions on data consistency challenges unique to billing—like handling event delays, retroactive pricing changes, and distributed count reconciliation across our infrastructure. You'll love this job if you Want to solve the "easy to explain, hard to build" problems—like ensuring a customer's bill matches their usage perfectly, even when processing hundreds of billions of events daily across distributed
From C$184K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know the team The Developer Platform team at Auth0 (an Okta company) owns the platform that developers build identity on. Increasingly they build alongside AI agents, and that raises the bar on everything underneath: interfaces have to hold up whether a person or an agent is calling them, and the systems behind them have to stay reliable and coherent as usage grows. We move fast, we own problems end to end, and we care deeply about the platform we put in front of the developers and agents who depend on it. The opportunity We're hiring a Principal Engineer (P5) to serve as the technical leader and compass for the Developer Platform team. You'll work across the breadth of the platform, tackling the highly complex, vaguely specified problems that span it and turning them into clear technical direction the team can execute against, without day-to-day oversight. Above all, you'll own how the platform is architected to scale: the distributed systems behind it, the reliability and consistency guarantees developers depend on, and the coherence that keeps it easy to build on as usage grows. You'll champion the team's technical execution, raise the engineering bar, mentor the people around you, and partner with tech leads across teams to keep the wider platform aligned. You'll have real influence over how our platform holds up in a world where developers and agents are both first-class consumers. What you'll be doing Own the platform architecture: Set the tech
From C$160K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Streaming Foundations team builds services and operates data pipeline infrastructure to support event streaming, messaging, and analytics use cases. We are looking for a Software Engineer who is passionate about distributed systems, platform engineering, and solving data-intensive problems at scale. In this high-impact role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Auth0 to scale for years to come. What you’ll be doing Help set the technical direction for the team and influence the engineering roadmap for the Platform’s streaming capabilities Design and lead the implementation of our most complex and critical systems for data-intensive use cases. Research and champion new technologies and architectural patterns to solve strategic challenges and scale the platform. Lead and influence cross-functional initiatives, ensuring technical alignment and successful execution across multiple teams. Improve the operational posture of our systems by designing for observability, reliability, and scalability, and by mentoring others in operational best practices. Coach and mentor senior engineers and act as a technical leader across the engineering organization. Collaborate with different stakeholders like product teams whenever needed. What you’ll bring to our teams 7+ years of software development experience in a fast-paced, agile environment Experience working with Golang or Java is preferred H
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Observability team's mission is to build and own Robinhood's full-stack observability platform — the foundation that keeps every product, service, and customer experience running reliably at scale. We design and operate the systems that give engineers deep visibility into how Robinhood's infrastructure behaves, ensuring that when something goes wrong, the right people know immediately and can act fast. Our work spans metrics, logs, distributed tracing, and alerting pipelines, and we partner closely with the Robinhood Command Center to ensure our observability systems meet or exceed 99.9% uptime. We believe in monitoring our own monitors — the observability infrastructure is a product, not just a tool. As a Staff Software Engineer on the Observability team, you will be the technical lead shaping the roadmap for how Robinhood observes itself at scale. You will own the observability control plane end-to-end, making architectural decisions that directly impact the reliability and operational health of Robinhood's entire product surface. You'll lead a team of six engineers, drive the strategy for cost-efficient telemetry ingestion, build in-house solutions where off-the-shelf
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team & role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Robinhood Command Center (RCC) is a newly formed reliability team that serves as the front line for detecting, coordinating, and mitigating production incidents across Robinhood. As part of Robinhood’s broader reliability initiative, RCC works closely with product engineering, reliability, observability, infrastructure, and business teams to reduce customer impact and shorten incident duration. As a Senior Engineer, you will be part of the founding RCC team, helping define how Robinhood responds to and learns from incidents at scale. This is a highly visible role focused on incident leadership, operational excellence, and reliability tooling. You will not own product services or core infrastructure, but you will own the processes and tools that enable fast, high-quality incident response. This role is based in our Menlo Park, California office, with in-person attendance expected at least 3 days per week. What you'll do: Serve as a senior technical leader driving the long-term reliability and observability strategy across Robinhood’s infrastructure Partner closely across many different types of engineers to raise the bar for operational excellence and incident r
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Security Platform team is responsible for the secure lifecycle, governance, and protection of Robinhood’s most sensitive user data. This team builds foundational systems such as secure data pipelines, tokenization services, and privacy compliance infrastructure to ensure customer data is handled responsibly and in line with regulations. The team works closely with security, infrastructure, and product engineering partners to ensure data is both usable and protected. You will contribute to systems that support authentication, third-party integrations, and emerging AI-driven use cases. As a Senior Software Engineer, you will design and build backend systems that securely process and manage customer data across Robinhood’s platform. You will own the systems that handle authentication, authorization, and privacy-preserving data operations. This role involves close collaboration with engineers focused on access management, infrastructure, and data systems. Your work will directly support efforts to improve system reliability, strengthen data protections, and enable new product capabilities using secure data. This role is based in our Bellevue, WA, and Menlo Park, CA off
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Data Engineering team builds and maintains the foundational datasets that power decision-making across Robinhood. We design reliable, scalable data systems that support product analytics, growth strategy, financial reporting, experimentation, and machine learning. The team partners closely with Product, Engineering, Data Science, and Finance to ensure accurate, well-modeled data is available to teams across the company. Our work directly influences how Robinhood measures performance, improves customer experience, and scales its products. As a Senior Data Engineer, you will design, build, and evolve core datasets that track product performance and company-wide metrics. You will develop scalable data pipelines that ingest application events and database snapshots into our data lake, ensuring high data quality and reliability. You’ll collaborate with application engineers to improve data generation patterns and with analytics teams to design intuitive, well-documented data models. This is an opportunity to shape the technical foundation that supports data-informed decisions across the organization! This role is based in our Menlo Park, CA office, with in-person attend
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Lyft Business Product Platform team builds the systems and experiences that power Lyft's B2B products — enabling companies, organizations, and their employees to seamlessly access Lyft's transportation network. We sit at the intersection of product and platform, owning both the customer-facing features and the underlying infrastructure that makes them reliable at scale. Our work directly impacts how businesses integrate with Lyft, how admins manage their programs, and how millions of riders get where they need to go. Responsibilities: Drive architecture and technical design for systems that are highly available, scalable, and built to last — not just for today's requirements but for where the product is heading Own features end-to-end: from shaping the technical spec and design through to production rollout and operational health Think critically about how AI capabilities can be incorporated into Lyft Business products to improve the experience for business admins and riders — and bring that perspective into roadmap and architecture conversations Make well-reasoned trade-off decisions and communicate them clearly to peers, leads, and cross-functional partners Write clean, well-tested, maintainable code and hold a high bar for the same in code reviews Partner across engineering, product, and design to align on direction and get buy-in on technical approaches Proactively engage in incident response, contributing both to resolution and to long-term reliability improvements Grow the team's technical culture through design reviews, tech talks, and mentorship Experience: 5+ years of software engineering experience, with a track record of designing and shipping production systems at scale Strong system design instincts — you can reason through distributed systems trade-offs, identify failure modes, and
From C$108K/yr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. We are building and maintaining a highly scalable asynchronous platform that empowers our organization to handle critical business cases. As a software engineering team, our mission is to create robust and innovative solutions that drive the success of our business and deliver unparalleled value to our customers. We adopt Infrastructure as Code practice to automate the provisioning and configuration of our resources, which helps reduce manual configuration and improve consistency. Our team culture is built on collaboration, open communication, and a supportive environment where each member's ideas are valued and contributions are recognized. We believe in the importance of fostering a positive workplace culture that inspires innovation and creativity. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and production readiness reviews, platform management and capacity planning ceremonies with cross-functional teams Document Infrastructure operations process and insights, identify repeatable actions and ruthlessly automate repetitive tasks Participate in our teams on-call rotations, respond to incidents and support other teams mitigate customer impacting events Experience: 5+ years experience working on teams responsible for software development, automation and systems engineering Experience building large-scale infrastructure, distributed systems or networks. Knowledge with SQS,
From C$108K/yr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As an Observability team member, you are responsible for the operation and maintenance of our logging and metrics infrastructure. You ensure all teams at Lyft are aware of the operational health of their products by monitoring system availability and take a holistic view of our platform performance. You build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you find opportunities to improve our systems in order to push our platform forward. You provide our partners with the support they need to help them build robust large scale distributed systems. We count on the reliability of our infrastructure to empower Lyft teams to provide our customers rich experiences that are highly available with rock solid performance to ensure our transportation platform continues to connect people and places. As we grow our team, we are seeking experienced Infrastructure Engineer to ensure that as our Infrastructure continues to scale, our platform continues to provide an essential and dependable service that transports millions of people every day. Specifically we are searching for someone who brings fresh perspectives, enjoys collaborating with cross-functional teams in order to continually improve our products and services for our customers. Responsibilities: Maintain, improve, and develop tooling and systems that enhance the reliability, scalability, and efficiency of our platform. Assist engineering teams in defining service-level objectives (SLOs) and provide the necessary toolin
About ElevenLabs ElevenLabs is an AI research and product company transforming how we interact with technology. We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including Andreessen Horowitz, ICONIQ Growth and Sequoia. We've raised $781M in funding and our last valuation was $11B - multiples of 11, always. We have expanded from voice into three main platforms: ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale. ElevenCreative empowers creators and marketers to generate and edit speech, music, image, and video across 70+ languages. ElevenAPI gives developers access to our leading AI audio foundational models. Everything we do is the result of the creativity and commitment of our team - builders doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to work hard and create lasting positive impact, we want to hear from you. How we work High-velocity: Rapid experimentation, lean autonomous teams, and minimal bureaucracy. Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you. AI first: We use AI to move faster with higher-quality results. We do this across the whole company—from engineering to growth to operations. Excellence everywhere: Everything we do should match the quality of our AI models. Global team: We prioritize your talent, not your location. What we offer Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible. Growth paths: Joining ElevenLabs means joining a
Other cities to consider
More places hiring for this role
Get new software reliability engineer jobs in Canada by email
Daily job updates · Unsubscribe anytime