Jobs in Canada

Software Reliability Engineer in Toronto

155 active opportunities · Updated October 2026

Explore current software reliability engineer jobs in Toronto. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -63.6%

From C$160K/yr

Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Staff Software Reliability Engineer - Data Platform About the Team The Data Platform team is responsible for the foundational data services, systems, and data products for Okta that benefit our users. Today, the Data Platform team solves challenges and enables: Streaming analytics Interactive end-user reporting Data and ML platform for Okta to scale Telemetry of our products and data Our elite team is fast, creative and flexible. We encourage ownership. We expect great things from our engineers and reward them with stimulating new projects, new technologies and the chance to have significant equity in a company. Okta is about to change the cloud computing landscape forever. About the Position This is an opportunity for experienced Software Reliability Engineers to join our fast growing Data Platform organization that is passionate about scaling high volume, low-latency, distributed data-platform services & data products. In this role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Okta to scale for years to come. As a member of the Data Platform team, you will be responsible for designing, building, and deploying the systems that power our data analytics and ML. Our analytics infrastructure stack sits on top of many modern technologies, including Kinesis, Flink, ElasticSearch, and Snowflake. We are looking for experienced Software Engineers who can help desi

JavaAWSKubernetesRest
T-
📍 Toronto, Canada· Full-time
✓ High-confidence listing

From C$1.4M/yr

Quick readStrong listing-quality and freshness signals

About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission is to engineer resilience from the ground up, enabling our product teams to innovate rapidly while ensuring our users have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through a culture of data-driven decision-making, blameless learning, and relentless automation. As a Senior Site Reliability Engineer, you are a hands-on engineer who blends deep software development expertise with a passion for operational excellence. You will be responsible for designing, building, and running the resilient, scalable, and increasingly self-healing systems that power our products. You will apply sound engineering principles to solve our most complex reliability challenges, with a mandate to automate everything, eliminate toil, and write robust, maintainable code. You will be a force multiplier, mentoring other engineers and elevating the site reliability bar for the entire organization. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: System Architecture & Design: Design, build, and maintain scalable, highly available, and fault-tolerant distributed systems. Partner with development teams as a reliability consultant, reviewing designs and influencing architectural decisions to ensure new services are built with reliability, observability, and performance as core principles, not afterthoughts. Automation & Software Development: Write robust, performant, and maintainable code to automate operational tasks, and CI/CD pipelines. Build the internal tools, libraries, and frameworks that enable engineering teams to self-service their

TypeScriptPythonAWSKubernetes
T
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation of AI computing systems. In this highly visible technical leadership role, you'll drive reliability from architecture through production, partnering across hardware, software, and manufacturing teams to build high-performance AI platforms that set the standard for uptime, durability, and quality. If you're passionate about solving complex engineering challenges and influencing products at scale, you'll have the opportunity to shape technology powering the future of AI. This role is hybrid, based out of Toronto, Canada. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You've spent 8+ years in reliability engineering, ideally in high-performance computing, AI hardware, or data center systems. You're comfortable with the statistical side of the job, HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA are all familiar territory. You can work through a technical problem in a thermal lab and then explain the risks and trade-offs clearly to leadership. You're good at bringing people together, mechanical, electrical, thermal, softw

AWSAIGoSEM
O
📍 Toronto, Ontario, Canada
✓ High-confidence listingCompany trend -63.6%
Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Okta Privileged Access Management (PAM) is an identity-centric approach to a common and critical privileged access use case. Our elegant Zero Trust architecture is purpose-built for the modern cloud and helps customers solve challenging security and operations pain points at scale. We are looking for a software engineer to join our fast-growing team with a focus on scalability, reliability, and enhancing the core building blocks of the product. In this role you will: Be deeply involved in evolving the core architecture of PAM. Work in our product development teams to build scalable, composable components of our platform. Be responsible for designing and implementing scalable architecture patterns. Delight our customers by providing world class UX using our React-based design system Design and build APIs that customers rely on for access to production infrastructure. Work on backend components written in Go and frontend components written in React. You might be a good fit if you: Have 3-5 years of software development experience with a background in Golang or similar programming languages. Proficient in React or similar front-end UI stacks. Experienced working with relational databases like PostgreSQL or similar RDBMS technologies. have the ability to complete a feature end to end from designing database models to backend APIs and frontend UI components. Experienced working with any cloud provider such as AWS, GCP or Azure. Thrive in a collaborativ

ReactPostgreSQLAWSAzure
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$40/hr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Interns work side-by-side with top engineers in the industry while having autonomy from the get-go. They contribute to user-facing products and are able to see their work go live quickly. Lyft fosters a collaborative environment in the office, so there's always a sharp mind eager to hear about your next idea. So what's yours? Responsibilities: Own your project, while checking in with other team members throughout the day with questions and updates You leave the code in a better state than when you found it (progressive refactor) You value reliability, ensured by testing (unit, integration and load tests) Participate in code reviews to ensure code quality and distribute knowledge Continuous integration and deployment Go home knowing that your work today is meaningfully improving the lives of every Lyft driver and every Lyft passenger! Experience: Currently pursuing a Bachelor's or Master's degree in Computer Science from a university in Canada (required) , with a graduation date between December 2027 and Summer 2028 (required). For any candidates who are master's students who worked between their bachelor's and master's programs: candidates should also have less than 2 years of relevant full-time work experience Available during Summer 2027 for the internship in Toronto Strong knowledge of CS fundamentals Excellent communication skills Passion for community, sustainability, and/or transportation Ability to thrive in a startup environment Experience with real-time technology problems Contributions to open source projects Experience working with databases Experience solving real-time technology problems Experience with mobile development Benefits: Mental health benefits In addition to holidays, interns receive 2 days paid time off and 3 days sick time off Subsidized commuter benefits and Lyft ride credi

AIGoExcelHR
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$40/hr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Interns work side-by-side with top engineers in the industry while having autonomy from the get-go. They contribute to user-facing products and are able to see their work go live quickly. Lyft fosters a collaborative environment in the office, so there's always a sharp mind eager to hear about your next idea. So what's yours? Responsibilities: Own your project, while checking in with other team members throughout the day with questions and updates You leave the code in a better state than when you found it (progressive refactor) You value reliability, ensured by testing (unit, integration and load tests) Participate in code reviews to ensure code quality and distribute knowledge Continuous integration and deployment Go home knowing that your work today is meaningfully improving the lives of every Lyft driver and every Lyft passenger! Experience: Currently pursuing a Bachelor's or Master's degree in Computer Science from a university in Canada (required) , with a graduation date between December 2027 and Summer 2028 (required). For any candidates who are master's students who worked between their bachelor's and master's programs: candidates should also have less than 2 years of relevant full-time work experience Available during Summer 2027 for an internship in Toronto Strong knowledge of CS fundamentals Knowledge of Python, JavaScript, CSS, and HTML Experience working with leading JavaScript frameworks, like React Experience with modern frontend testing tools, such as Webpack, Babel, Jest, Jasmine, Protractor, and WebDriver Understanding of how browsers and DOM work Experience with Git or other distributed version control systems Experience with browser developer tools Experience with the Unix command line interface Solid understanding of web performance Experience with TypeScript Experience with CSS

JavaScriptTypeScriptPythonJava
O
📍 Toronto, Ontario, Canada
✓ High-confidence listingCompany trend -63.6%
Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are looking for an experienced Staff Software Engineer in Test to join our Identity Management Engineering (IDM) team serving the Privileged Access Team (PAM). This team is passionate about delivering large-scale, mission-critical software in a fast-paced Agile environment. In this role you'll be working with a team of highly-skilled and talented engineers, responsible for delivering sophisticated backend solutions that help Okta reliably operate at large scale and be highly available. As part of the team, you’ll be ensuring projects are completed with the highest quality and reliability using automation at every level for fast, robust and secure releases. Job Duties and Responsibilities: Review requirements and design specs to develop relative test plans and test cases Automate API tests, end-to-end tests, reliability/scale tests Work with engineering management to scope and plan engineering efforts Communicate and document QE plans for scrum teams to review Review application code, identify bug and other areas of weakness, architect tools for future coverage Automate all critical features to maintain zero-debt cadence Release features with solid quality Respond to production issues/alerts and customer issues during on-call rotation Be a strong customer advocate with a strong quality DNA. Requirements: 5+ years of QE experience preferably in an enterprise SaaS company 3+ years experience in quality engineering for enterprise level software. 5+ yea

PythonJavaAWSKubernetes
L
📍 Toronto, Canada
✓ High-confidence listingCompany trend -72.4%
Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Driver team is dedicated to fostering a platform of high-quality service by empowering drivers to perform their best. We are looking for a product-minded engineer who wants to build and improve products that sit at the center of the core driver experience. Products you drive will solve pain points that matter most to drivers by streamlining key interactions, reducing friction, and creating systems that feel intuitive, fair, and supportive. As an engineer at Lyft, you'll collaborate with teams like product, data science, analytics, and operations on code that empower us to iterate quickly, while focusing on delighting our passengers and drivers. Responsibilities: Design and implement backend features end-to-end with clear ownership, delivering well-scoped work from technical design through to production with moderate guidance from senior engineers Write clean, reliable, well-tested code that meets team standards and holds up in code review Participate actively in code reviews, giving specific and constructive feedback while continuing to develop your own review instincts Debug and resolve issues across backend services including performance bottlenecks, reliability problems, and data integrity issues Collaborate with product managers, designers, and partner engineering teams to clarify requirements and surface technical constraints early Contribute to technical discussions and help evaluate implementation approaches for new features Write unit and integration tests for your own code and develop familiarity with the team's broader testing and observability practices Address technical debt and make incremental improvements to existing services as part of regular development work Participate in on-call rotations, respond to production incidents, and support teammates in mitigating custome

PythonJavaAWSGCP
F
📍 Toronto, Canada· Full-time
✓ High-confidence listing

From C$190K/yr

Quick readStrong listing-quality and freshness signals

About Forma.ai: Forma.ai is a Series B startup that's revolutionizing how sales compensation is designed, managed and optimized. We handle billions in annual managed commissions for market leaders like Edmentum, Stryker, and Autodesk. Our growth has been fuelled by our passion for fundamentally changing and shaping how companies use sales intelligence to drive business strategy. We’re welcoming equally driven individuals who are excited about creating something big! About the Team We build enterprise software that helps organizations optimize sales performance and improve go-to-market agility. Our engineering organization includes multiple product application teams responsible for delivering core customer-facing capabilities. We are seeking Senior Backend Engineers to join our application teams. You’ll work alongside staff, senior, and early-career engineers to design, build, and scale backend systems that power enterprise-grade product workflows. This is an opportunity to work on complex product and data problems while contributing meaningfully to technical decisions, system quality, and team delivery. We are low on meetings and high on accountability. Most of the team is in the EST time zone, with a few located in AST, PST, and Central as well. What you’ll be doing You will play an important role in the continued evolution of our application stack. You will design and build backend capabilities for complex product workflows, contribute to system design discussions, and help ensure our systems remain maintainable, reliable, and scalable as we grow. As a Senior Backend Engineer, you are expected to operate with strong ownership and sound technical judgment. This includes identifying risks in the work you own, surfacing edge cases, asking thoughtful questions, and proposing improvements that strengthen the quality and reliability of the system. You will: Design and build backend services that power complex product workflows. Contribute to d

JavaScriptTypeScriptPythonJava
T
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a skilled Software Engineer with a passion for building high-performance, low-level systems software. In this role, you’ll contribute to the development and optimization of the infrastructure that powers our cutting-edge processors, with a primary focus on C/C++ development and low-level programming. You'll work closely with large inference and training model development to further drive Scale Out software and hardware performance. This role is hybrid, based out of Toronto, ON. Who You Are Strong C or C++ systems engineer with a deep understanding of memory, threading, I/O, and low-level execution models. Experienced building low-level software, drivers, embedded systems, or performance-critical infrastructure. Comfortable working close to hardware and curious about how systems behave under the hood. Proficient with Linux systems programming and debugging tools such as gdb, strace, and perf. Structured problem solver who thrives in fast-paced, highly technical environments. What We Need Design, develop, and maintain core infrastructure software that interfaces directly with Tenstorrent hardware. Build low-level libraries and APIs for communication and synchronization across compute nodes. Optimize system-level software for performance, scalability, and reliability in distributed environments. Support hardware

AWSLinuxAIC++
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$136K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As a Infrastructure Engineer at Lyft, you will run our Production Infrastructure by monitoring system availability and take a holistic view of our platform health. You will build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you will seek opportunities to optimize our systems in order to push our platform forward, anticipating our customers' needs in order to continually improve the platform. You will provide Lyft partner teams with operational support to help them build robust large scale distributed systems. About the Team Data Pipelines is at the heart of all critical data flowing through Lyft supporting hundreds of services that impact millions of drivers and passengers every day. Our team’s mission is to empower Lyft engineers to self-serve in building and maintaining data pipelines as needed to support products that deliver the world’s best transportation experience. We leverage a variety of technologies to store, stream and manage data making it available to our internal customers. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and

PythonAWSDockerKubernetes
L
📍 Toronto, Canada
✓ Quality checkedCompany trend -72.4%

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. As a Senior Software Engineer, Data on the Mapping team, you will collaborate with our world-class team of engineers, product managers, and scientists to grow and improve the quality of recommended routes and accuracy of our travel time estimations. You will lead the architecture and long-term technical direction of our offline experimentation tooling and route simulation services — the systems that let Lyft test routing changes safely before they reach production. You'll also build scalable data pipelines for experimentation, analytics, and machine learning models, along with the data governance and observability systems that keep them trustworthy. Your work will enable integration with partner teams and allow stakeholders across Engineering, Data Science, and Product to make data-informed decisions that directly impact Lyft’s growth and profitability. Our technology stack is based on the latest technologies such as AWS, Databricks, Kubernetes and Airflow. You will work with incredibly passionate and talented colleagues from software engineering, machine learning and data science on projects that directly impact millions of riders and drivers. Responsibilities Own core data pipelines end-to-end, building deep subject matter expertise in the systems you manage and defining/managing SLAs for pipelines, services, and datasets to ensure reliability at scale Serve as the technical owner and architectural lead for our offline experimentation platform and route simulation services, setting technical direction, evaluating trade-offs, and ensuring the systems scale with Lyft's routing and mapping ambitions Continuously evolve data models and schemas to meet business and engineering requirements Develop AI tools that support self-service management of data pipelines (ETL) and schema evolution, and perform han

PythonSQLPostgreSQLMySQL
S
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -100%

C$162K – C$420K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Events Analytics Platform (EAP) team is responsible for the infrastructure that powers all of Sentry's time-series data and searching capabilities across billions of events with sub-second latency. We started this initiative by building Snuba, the primary storage and query service for Sentry's event data powered by ClickHouse, and we are now focused on unlocking deeper visibility and reporting across the terabytes of event data our users generate. As a Senior Software Engineer, you will lead efforts to push the boundaries of data visibility at Sentry. You will do this by expanding the capabilities of our search infrastructure, building new capabilities on top of our state-of-the-art storage layer and increasing the performance and integrity of Sentry’s core data services. You will also help shape Infrastructure's technical direction at Sentry and collaborate with Product and other Engineering teams to turn that vision into a reality. If you want to solve the hard problems that come with scaling event data into the petabyte range, this could be the job for you. In this role you will: Expand EAP's ability to deliver data at world-class speed and reliability. Architect and automate services and systems to scale reliably under growing demand. Make architectural trade-offs that balance product requirements with engineering constraints. Maintain and grow the team's code quality initiatives by regularly reviewing code and contributing to design decisions. Lead design and discussions around deliverables the team is working towards. Improve the maintainability and developer experience of the codebases EAP owns. Exa

PythonSQLPostgreSQLRedis
S
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -100%

C$162K – C$420K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. About the role As a Senior Software Engineer, you will architect and scale the core systems that power Sentry's billing infrastructure, ensuring accuracy and reliability at massive scale. You will collaborate on building the next generation of Sentry’s usage tracking pipeline, processing hundreds of billions of events daily with low latency and financial-grade accuracy. You will help design flexible pricing primitives that support everything from per-event usage billing to complex enterprise contracts, enabling product and sales teams to experiment rapidly while maintaining revenue accuracy and reduced time-to-market for new products. You will contribute to technical decisions on data consistency challenges unique to billing—like handling event delays, retroactive pricing changes, and distributed count reconciliation across our infrastructure. You'll love this job if you Want to solve the "easy to explain, hard to build" problems—like ensuring a customer's bill matches their usage perfectly, even when processing hundreds of billions of events daily across distributed

O
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -63.6%

From C$184K/yr

Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know the team The Developer Platform team at Auth0 (an Okta company) owns the platform that developers build identity on. Increasingly they build alongside AI agents, and that raises the bar on everything underneath: interfaces have to hold up whether a person or an agent is calling them, and the systems behind them have to stay reliable and coherent as usage grows. We move fast, we own problems end to end, and we care deeply about the platform we put in front of the developers and agents who depend on it. The opportunity We're hiring a Principal Engineer (P5) to serve as the technical leader and compass for the Developer Platform team. You'll work across the breadth of the platform, tackling the highly complex, vaguely specified problems that span it and turning them into clear technical direction the team can execute against, without day-to-day oversight. Above all, you'll own how the platform is architected to scale: the distributed systems behind it, the reliability and consistency guarantees developers depend on, and the coherence that keeps it easy to build on as usage grows. You'll champion the team's technical execution, raise the engineering bar, mentor the people around you, and partner with tech leads across teams to keep the wider platform aligned. You'll have real influence over how our platform holds up in a world where developers and agents are both first-class consumers. What you'll be doing Own the platform architecture: Set the tech

Node.jsAWSRestMachine Learning
🔔

Get new software reliability engineer jobs in Toronto, Canada by email

Daily job updates · Unsubscribe anytime