About the Team DoorDash Labs, established in 2018, serves as the innovation hub for DoorDash, focusing on developing automation and robotics solutions to enhance last-mile logistics. The team's mission is to create technologies that support and augment human networks, aiming to improve efficiency for Dashers, merchants, and consumers alike. We’re ruthlessly focused on business impact. We are a highly senior team composed of former pioneers from a variety of different robotics industries. As of 2025, DoorDash has completed 10B lifetime deliveries. We’re focused on how to do the next 10B even better. About the Role We are seeking a highly motivated Senior Reliability & Test Engineer to join our team. This individual will play a key role in the development and validation of our unmanned platforms at the system and component levels. You will partner closely with EE, ME, and Autonomy teams to translate mission needs into robust, reliable hardware. The ideal candidate thrives in a fast-moving, cross-functional environment where reliability and test rigor determine program success. You will be hands-on in developing test methods and equipment to uncover failures before they happen in the field. You will partner closely with EE, ME, and Autonomy teams to translate mission needs into robust, reliable hardware. The ideal candidate thrives in a fast-moving, cross-functional environment where reliability and test rigor determine program success. You’re excited about this opportunity because you will… Architect and implement rigorous validation strategies, utilizing Python scripts for automation while leveraging CAD and shop tools to engineer bespoke test fixtures and hardware rigs. Oversee experimental execution across internal facilities and external laboratories, maintaining technical mastery over vibration tables, environmental chambers, DAQ systems, and ingress protection testing. Translate high-level vehicle reliability requirements into granula
Jobs in Canada
Senior Reliability Engineer in Canada
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current senior reliability engineer jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.
Typical salary
C$1.7M – C$1.7M/yr
Based on 1 salary observations
From C$1.4M/yr
About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission is to engineer resilience from the ground up, enabling our product teams to innovate rapidly while ensuring our users have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through a culture of data-driven decision-making, blameless learning, and relentless automation. As a Senior Site Reliability Engineer, you are a hands-on engineer who blends deep software development expertise with a passion for operational excellence. You will be responsible for designing, building, and running the resilient, scalable, and increasingly self-healing systems that power our products. You will apply sound engineering principles to solve our most complex reliability challenges, with a mandate to automate everything, eliminate toil, and write robust, maintainable code. You will be a force multiplier, mentoring other engineers and elevating the site reliability bar for the entire organization. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: System Architecture & Design: Design, build, and maintain scalable, highly available, and fault-tolerant distributed systems. Partner with development teams as a reliability consultant, reviewing designs and influencing architectural decisions to ensure new services are built with reliability, observability, and performance as core principles, not afterthoughts. Automation & Software Development: Write robust, performant, and maintainable code to automate operational tasks, and CI/CD pipelines. Build the internal tools, libraries, and frameworks that enable engineering teams to self-service their
From C$107K/yr
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operati
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron's Global Supplier Quality organization is seeking a Controller Quality Principal Engineer to lead the quality strategy, qualification, and continuous improvement of storage and memory controller (ASIC/SoC) manufacturers supporting Micron's SSD, embedded, and storage solutions portfolio. This is a senior technical leadership role responsible for driving controller supplier quality performance from design qualification through mass production and field support The successful candidate will collaborate across functions with ASIC Development Team, Compose Engineering, Product Engineering, Dependability, Manufacturing, and Commodity Management, as well as directly with controller IC vendors, foundries, third party reliability labs and OSAT (outsourced assembly and test) partners, to ensure controller quality, reliability, and supply continuity meet Micron's standards. This role can be based in Taiwan, Hyderabad, or San Jose and will work extensively across time zones with global partners and suppliers! Develops, evaluates, revises, and applies technical quality assurance protocols/methods to inspect and test in-process raw materials, production equipment, and finished products. Ensures activities and items are in compliance with both company quality assurance standards and applicable government regulations. Performs analysis and identifies trends in the inspection of finished products, in-process materials and bulk raw materials, and recommends corrective actions when vital. Ensures that established manufacturing inspection, sampling and statisti
About the Team DoorDash Labs, established in 2018, serves as the innovation hub for DoorDash, focusing on developing automation and robotics solutions to enhance last-mile logistics. The team's mission is to create technologies that support and augment human networks, aiming to improve efficiency for Dashers, merchants, and consumers alike. We’re ruthlessly focused on business impact. We are a highly senior team composed of former pioneers from a variety of different robotics industries. As of 2025, DoorDash has completed 10B lifetime deliveries. We’re focused on how to do the next 10B even better. About the Role We are seeking a highly motivated Senior/Staff Test Engineer to join our team. This individual will play a key role in the development and validation of our unmanned platforms at the system and component levels. The ideal candidate has a strong background in test development, test execution, and root cause analysis with a proven track record of collaboratively managing risk throughout a fast paced development process. You’re excited about this opportunity because you will… Run and monitor tests within our facility as well as at outside test labs. Collaborate with a tight knit team to identify and understand test failures. Be hands-on in developing test methods and equipment to uncover failures before they happen in the field. Find clarity through root cause analysis of lab and field failures and suggest design changes to prevent them. Use your creativity to create novel and scaled tests for autonomous systems. We’re excited about you because you have… A bachelors or advanced degree in a relevant engineering discipline. Mastery of test equipment such as environmental chambers, vibration tables, water testers, DAQs, etc.. Ability to bring order to complex test and development programs via clear technical communication and documentation. Experience designing and building testers and equipment. Ability to write Python scripts to automate
$170K – $240K/yr
Senior Software Engineer - Observability and Reliability About the Role We are growing the engineering team and looking for engineers who have the chops to build and deliver world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible. What You Will Be Doing Build observability tools and platforms, including: metrics, logging, distributed tracing, dashboarding, alerting, application performance management Build with modern tools and languages like Go, Open Telemetry and Kubernetes Participate in on-call rotation and ensure uptime of services Create runtime tools/processes that optimize cloud triaging and limit downtime Define best practices around making our systems and services measurable Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies. We expect successful candidates to be coding a majority of their time Qualifications We Need Strong Computer Science fundamentals 5+ years industry experience building and maintaining high-quality software, especially software other engineers use You apply a product mindset to infrastructure systems and feel accomplished enabling others Desire to be a great teammate and have fun at work Strong sense of craftsmanship, and a healthy academic curiosity Qualifications We Want (also, skills you’ll learn!) Experience building systems for data analytics Distributed systems monitoring and profiling skills Knowledge of cloud application security models Administered cloud service infrastructure (GCP, AWS, Azure) Startup experience Additional Job details Additional Job details The base salary range for this position is $170k - $240k annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience. Base pay is one part of the Total Package that is provided to compensate and recognize e
About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the technical direction for our in-house Spark deployment and shape the architecture that will run DoorDash's data, analytics, and ML compute for the next five years and beyond. You will own the deep, cross-cutting problems that span the runtime, the shuffle service, the scheduler, and the overall service reliability — making the architectural calls that compound across the platform's lifetime. You will partner with the Engineering Manager on technical roadmap, hiring, and team shape, and act as the senior technical voice in cross-team partnerships with Data Engineering, ML Platform, and product engineering teams that depend on the platform. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Set the multi-year technical direction for an in-house Spark-on-Kubernetes platform — runtime, shuffle, scheduler, reliability — and make the architectural calls that compound for years. Own the deepest distributed-systems problems on the team: shuffle architecture, multi-tenant scheduling, runtime performance, and the failure modes that only show up at scale. Partner with the Engineering Manager on technical roadmap, hiring, inte
From C$1.2M/yr
About the Role: We are looking for a talented Automation Engineer to join our Automation Engineering team in Toronto. In this role, you will be responsible for designing and implementing automated tests for Mobile development. You will collaborate closely with QA engineers and developers to build scalable test frameworks, improve automation coverage, and contribute to the efficiency of our multi-platform release process. You will also design data-driven end-to-end checks around playback and ad insertion , integrate them into CI/CD pipelines as quality gates, and operate a reliable device lab to prevent regressions from shipping. Your work will directly accelerate testing and release velocity while improving revenue-critical reliability across Tubi’s Android and IOS apps. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: Design, implement, and maintain automated tests for mobile development (Android & iOS) Contribute to the development and optimization of cross-platform automation frameworks. Write and maintain test scripts in JavaScript/TypeScript , using frameworks such as Puppeteer, Appium, WebDriverIO, Selenium. Ensure test cases are integrated into CI/CD pipelines and provide reliable feedback on product quality. Help identify flaky tests, investigate root causes, and improve test stability. Collaborate with developers and QA engineers to clarify requirements and improve test strategies. Participate in code reviews and follow best practices for test automation . Your Background: Bachelor’s degree or above in a technical field (e.g., Computer Science, Engineering, Mathematics), or equivalent industry experience. 3+ years of hands-on experience in automation testing for mobile devices Strong programming skills in JavaScript/TypeScript (preferred), or Python/Java. Experience with automation frameworks (e.g. Puppeteer, Appium,, WebDriverIO, Selenium, Playwright, T
From $180K/yr
Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform providing APIs for knowledge retrieval, inference, evaluation, and more. We are seeking a strong Senior Full-Stack Engineer to help us build, scale, and refine our rapidly growing product. The ideal candidate is deeply grounded in software engineering best practices and experienced in developing and scaling modern web applications end-to-end. You will work across the stack—from React/TypeScript frontends to Python-based backends—while integrating with LLMs and machine learning systems. You will solve complex challenges in scalability, reliability, and product experience while owning significant product areas in a fast-paced environment. What You’ll Do Own major full-stack product areas , driving features from design through production deployment. Build modern frontend experiences using React and TypeScript, ensuring performance, usability, and responsiveness. Develop reliable backend services in Python, working with distributed systems, data pipelines, and ML/LLM components. Integrate with LLMs, vector databases, and AI infrastructure to power intelligent product experiences. Deliver experiments and new features quickly , maintaining high quality and tight feedback loops with customers. Collaborate across product, ML, and infrastructure teams to shape the direction of Scale GP. Adapt quickly —learning new technologies, frameworks, and tools as needed across the stack. Ideal Experience 5+ years of full-time engineering experience , post-graduation. Strong experience developing full-stack applications using React, TypeScript, and Python . Experience scaling or shipping products at high-growth startups . Familiarity with LLMs, vector databases, embeddings, or other modern AI tooling (tinkering or production experience welcome). Proficiency with SQL and modern API development. Experience with Kubernetes , containerization, and microservice architectures. Experience working with at leas
#TeamNextdoor Nextdoor (NYSE: NXDR) is the essential neighborhood network. Neighbors, public agencies, and businesses use Nextdoor to connect around local information that matters in more than 350,000 neighborhoods across 11 countries. Nextdoor builds innovative technology to foster local community, share important news, and create neighborhood connections at scale. Download the app and join the neighborhood at nextdoor.com . Meet Your Future Neighbors As a Software Engineer at Nextdoor, you’ll join a focused team of developers, product managers, and designers who are passionate about using technology to cultivate a kinder world where everyone has a neighbor they can rely on. We are a small, high performing team of engineers that wear multiple hats and prioritize impact for our customers. We care about moving fast and delivering impact, without compromising on quality and reliability as guided by our Engineering Principles . You will have the opportunity to learn from your co-workers and teach them. As a team, we will make each other better and build great software. At Nextdoor, we operate in an AI-first environment and expect every team member to actively use AI tools as part of their workflow. We aren't looking for prompt engineers; we’re looking for people who use tools like Claude, Gemini, ChatGPT, and Glean to challenge their own thinking and take full ownership of AI-assisted outputs. We also offer a warm and inclusive work environment that embraces a hybrid employment model, blending an in office presence and work from home experience for our valued employees. The hiring team will go over these expectations with you if you are being considered for a role near one of our offices in San Francisco, Los Angeles, Chicago, Dallas, New York, and London. The Impact You’ll Make We believe in empowering our teams to own all aspects of bringing Nextdoor to life. As such, you’ll get the opportunity to make key contributions across our engineering stack - thi
$162K – $225K/yr
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Senior Software Engineer, Cloud Services, you will create scalable, reliable backend systems that support HP IQ's mission to transform the way people work. We value how we work as much as what we deliver. Our journey of continuous learning and evolution requires a flexible and resilient services platform that fosters innovation, experimentation, and adaptation. To enable this, we focus on designs and tools rooted in strong engineering principles like abstraction, composition, virtualization, automation, and iterative development cycles. What You Might Do Design, develop, and maintain backend services and RESTful APIs using Java and Spring Boot. Write clean, efficient, and well-tested code following established best practices. Collaborate with frontend developers, product managers, and other engineers to deliver end-to-end features. Integrate with databases and external services, ensuring performance, security, and reliability. Participate in code reviews, debugging, and performance optimization efforts. Contribute to CI/CD pipelines and support application deployment in cloud and
$140K – $225K/yr
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As the Senior Software Engineer, Tooling and Development Infrastructure, you will play a critical role in shaping the developer productivity tools and automated testing strategy. You’ll collaborate closely with design, development, and quality teams to plan, design, and implement robust automated tools and services that ensure the quality and reliability of our AI software stack. You will be highly hands-on in your work and collaborate closely with stakeholders. This position offers a unique opportunity to influence the development of cutting-edge automation frameworks, foster a culture of quality, and contribute to the long-term success of the organization. What You Might Do Develop and implement automation frameworks and testing strategies that cover the entire software stack, from backend systems to user-facing features. Identify, evaluate, and integrate new tools that streamline development. This includes everything from code quality tools and to Infrastructure-as-Code (IaC) solutions. Lead continuous improvement efforts for our build, release, and test systems, ensuring a robust
From C$190K/yr
About Forma.ai: Forma.ai is a Series B startup that's revolutionizing how sales compensation is designed, managed and optimized. We handle billions in annual managed commissions for market leaders like Edmentum, Stryker, and Autodesk. Our growth has been fuelled by our passion for fundamentally changing and shaping how companies use sales intelligence to drive business strategy. We’re welcoming equally driven individuals who are excited about creating something big! About the Team We build enterprise software that helps organizations optimize sales performance and improve go-to-market agility. Our engineering organization includes multiple product application teams responsible for delivering core customer-facing capabilities. We are seeking Senior Backend Engineers to join our application teams. You’ll work alongside staff, senior, and early-career engineers to design, build, and scale backend systems that power enterprise-grade product workflows. This is an opportunity to work on complex product and data problems while contributing meaningfully to technical decisions, system quality, and team delivery. We are low on meetings and high on accountability. Most of the team is in the EST time zone, with a few located in AST, PST, and Central as well. What you’ll be doing You will play an important role in the continued evolution of our application stack. You will design and build backend capabilities for complex product workflows, contribute to system design discussions, and help ensure our systems remain maintainable, reliable, and scalable as we grow. As a Senior Backend Engineer, you are expected to operate with strong ownership and sound technical judgment. This includes identifying risks in the work you own, surfacing edge cases, asking thoughtful questions, and proposing improvements that strengthen the quality and reliability of the system. You will: Design and build backend services that power complex product workflows. Contribute to d
From $264.8K/yr
Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. About the General Agents Team The General Agents team, part of Scale’s Enterprise organization, builds robust general agents for customer use cases and applications. The team sits at the intersection of frontier agent development and real-world deployment, translating state-of-the-art reasoning and agentic capabilities into reliable, production-grade systems that drive real economic value. Our agents are scalable systems built around recurring enterprise problem domains, with a strong emphasis on generalization, extensibility, and deployment across many customers. About the Role As a Senior/Staff Machine Learning Engineer (MLE) on the General Agents team, you’ll play a critical role in designing, building, and deploying production-ready AI agents that solve high-impact enterprise problems. You will work across the full agent lifecycle—from model and system design to evaluation, deployment, and iteration—bridging cutting-edge agentic techniques with the constraints and requirements of real customer environments. You will: Design and implement end-to-end agent systems that combine LLM reasoning, tool use, memory, and control logic to solve recurring enterprise use cases. Build scalable, reliable agent architectures that can be deployed across many customers with varying data, tools, and constraints. Develop evaluation frameworks, datasets, environments, and metrics to measure agent performance, reliability, and business impact in production settings. Collaborate closely with product managers, customers, data annotators, and other engineering teams to translate enterprise requirements into robust agent designs. Productionize frontier agent techniques (e.g.,
About the Team Data is at the foundation of DoorDash success. The Data Engineering team builds database solutions for various use cases including reporting, product analytics, marketing optimization and financial reporting. By implementing pipelines, data structures, and data warehouse architectures; this team serves as the foundation for decision-making at DoorDash. About the Role DoorDash is looking for a Senior Data Engineer to be a technical powerhouse to help us scale our data infrastructure, automation and tools to meet growing business needs. This is a hybrid position and you must be located in Sunnyvale, San Francisco, or Seattle. You're excited about this opportunity because you will... Work with business partners and stakeholders to understand data requirements Work with engineering, product teams and 3rd parties to collect required data Design, develop and implement large scale, high volume, high performance data models and pipelines for Data Lake and Data Warehouse Develop and implement data quality checks, conduct QA and implement monitoring routines Improve the reliability and scalability of our ETL processes Manage a portfolio of data products that deliver high-quality, trustworthy data Help onboard and support other engineers as they join the team We're excited about you because... 5+ years of professional experience 3+ years experience working in data engineering, business intelligence, or a similar role Proficiency in programming languages such as Python/Java 3+ years of experience in ETL orchestration and workflow management tools like Airflow, Flink, Oozie and Azkaban using AWS/GCP Expert in Database fundamentals, SQL and distributed computing 3+ years of experience with the Distributed data/similar ecosystem (Spark, Hive, Druid, Presto) and streaming technologies such as Kafka/Flink. Experience working with Snowflake, Redshift, PostgreSQL and/or other DBMS platforms Excellent communication skills and experience working with technical and non-tec
Higher-paying openings
Jobs with higher listed pay
Other cities to consider
More places hiring for this role
Get new senior reliability engineer jobs in Canada by email
Daily job updates · Unsubscribe anytime