About the Team The Storage teams build and operate online stateful systems and abstractions that are reliable, efficient, secure and easy to use for DoorDash Engineering. The teams are responsible for understanding Product Engineering’s evolving needs and developing platform and infrastructure capabilities to serve them. The team currently supports CockroachDB, Cassandra, Kafka and Redis as well as data abstraction services to reduce the complexity of interacting with storage systems for Product Engineers. About the Role The Storage team is building and operating a high-performance, scalable, and reliable data abstraction layer that optimizes both efficiency and reliability. Our goal is to create a platform that manages itself and fades into the background—empowering engineers to focus on delivering product experiences our customers love. This role is available across two teams within Storage, each solving unique and high-impact challenges: One team is building the orchestration layer for DoorDash’s storage platform—unifying lifecycle management, operations, and self-serve APIs for databases and streaming systems, turning complex, stateful infrastructure into reliable, developer-friendly services used across the company. One team builds and operates the distributed data platform powering DoorDash's largest stateful workloads -- including Cassandra, which backs critical product surfaces across DoorDash, Wolt, and Roo. You'll design high-throughput data abstractions, smart clients, and platform services that make distributed data reliable and easy to work with at multi-petabyte, multi-million-QPS scale, with opportunities to go deep on distributed systems internals and contribute to the open-source Cassandra ecosystem. If you're passionate about distributed systems, developer experience, and building foundational infrastructure at scale, we'd love to hear from you. You must be located in San Francisco, Sunnyvale, Seattle, or the New York Metro Area for this hybrid pos
Jobs in Canada
Engineering Manager Platform Reliability in Canada
681 active opportunities · Updated October 2026
Showing
15 jobs
Explore current engineering manager platform reliability jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.
From C$107K/yr
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operati
About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team The Payments organization owns some of Stripe’s most critical payment flows and a platform that processes hundreds of billions of dollars in payments a year. Our team builds and scales the infrastructure and financial partner integrations that enables Stripe to accept, manage, and payout money across many countries, currencies, and payment methods. Our work is core to Stripe’s business, and thousands of developers use our platform and infrastructure to create valuable products and services that billions of people use. Our goal is to increase the GDP of the internet by making it easy to build global products, services, and platforms that handle money. Technical Operations roles in Payments are a dynamic and key component of Stripe's success. Focused on financial partner integrations and funds flow expansion, we sit at the intersection of product/platform engineers and financial partners, connecting them to ensure that everyone thrives and nothing is lost in translation. Our team partners closely with various finance and infrastructure engineering teams to ensure timely delivery of accurate data between financial partners, internal stakeholders, and Stripe leaders. We report and trace all of Stripe’s money movement transactions, including payments in more than 30 currencies and dozens of countries. You’ll own building and scaling Stripe’s manual and programmatic financial reporting and reconciliation processes for intra-company and outgoing money m
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are seeking an experienced Staff Software Engineer to join Okta's Universal Directory Platform team within the Product Platform Pillar. The team serves as the intelligent core of the enterprise security fabric, maintaining the source of truth for all identity assets and their associated relationships. Opportunity This position will be involved into development, design, and maintenance of our highly performant, reliable, and scalable platform, which is critical for managing user lifecycles, groups, and memberships. The successful candidate will possess experience in building and deploying scalable, reliable software and infrastructure within a cloud environment. What you’ll be doing Understand our identity management group codebase and development process: Jira, Technical Designs, Code Review, Testing, and Deployment. Develop and implement frameworks and toolings for our Universal Directory Service platform. Design and implement high-performance distributed scalable and fault-tolerant software components. Quickly deliver high-quality bug fixes and handle customer-reported issues. Conduct quality code reviews and automated testings. Partner with our Product Development, QA, and Site Reliability Engineering teams for scoping the development and deployment work. What you’ll bring to the role The ideal candidate is someone who is experienced building software systems to manage and deploy reliable and performant infrastructure and prod
About the Team Data is at the foundation of DoorDash success. The Data Engineering team builds database solutions for various use cases including reporting, product analytics, marketing optimization and fi nancial reporting. Team serves as the foundation for decision-making at DoorDash. About the Role DoorDash is looking for a Sta ff Software Engineer,Data to be a technical lead and help architect and scale our data reliability, data infrastructure, automation and tools to meet growing business needs. You’re excited about this opportunity because you will... Own critical data systems that support multiple products/teams Develop, implement and enforce best practices for data infrastructure and automation Design, develop and implement large scale, high volume, high performance data models and pipelines for Data Lake and Data Warehouse Improve the reliability and scalability of our Ingestion, data processing, ETLs, Reporting tools and data ecosystem services Manage a portfolio of data products that deliver high-quality, trustworthy data Help onboard and support other engineers as they join the team We’re excited about you because... 8+ years of professional experience as a hands-on engineer and technical leader leading multiple projects 6+ years experience working in data platform and data engineering or a similar role You have proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software Pro fi ciency in programming languages such as Python/Kotlin/Scala 4+ years of experience in ETL orchestration and work fl ow management tools like Air fl ow Expert in database fundamentals, SQL, data reliability practices and distributed computing 4+ years of experience with the Distributed data/similar ecosystem (Spark, Presto) and streaming technologies such as Kaa/Flink/Spark Streaming Excellent communication skills and experience working
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Security Platform team is responsible for the secure lifecycle, governance, and protection of Robinhood’s most sensitive user data. This team builds foundational systems such as secure data pipelines, tokenization services, and privacy compliance infrastructure to ensure customer data is handled responsibly and in line with regulations. The team works closely with security, infrastructure, and product engineering partners to ensure data is both usable and protected. You will contribute to systems that support authentication, third-party integrations, and emerging AI-driven use cases. As a Senior Software Engineer, you will design and build backend systems that securely process and manage customer data across Robinhood’s platform. You will own the systems that handle authentication, authorization, and privacy-preserving data operations. This role involves close collaboration with engineers focused on access management, infrastructure, and data systems. Your work will directly support efforts to improve system reliability, strengthen data protections, and enable new product capabilities using secure data. This role is based in our Bellevue, WA, and Menlo Park, CA off
From $264.8K/yr
Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. About the General Agents Team The General Agents team, part of Scale’s Enterprise organization, builds robust general agents for customer use cases and applications. The team sits at the intersection of frontier agent development and real-world deployment, translating state-of-the-art reasoning and agentic capabilities into reliable, production-grade systems that drive real economic value. Our agents are scalable systems built around recurring enterprise problem domains, with a strong emphasis on generalization, extensibility, and deployment across many customers. About the Role As a Senior/Staff Machine Learning Engineer (MLE) on the General Agents team, you’ll play a critical role in designing, building, and deploying production-ready AI agents that solve high-impact enterprise problems. You will work across the full agent lifecycle—from model and system design to evaluation, deployment, and iteration—bridging cutting-edge agentic techniques with the constraints and requirements of real customer environments. You will: Design and implement end-to-end agent systems that combine LLM reasoning, tool use, memory, and control logic to solve recurring enterprise use cases. Build scalable, reliable agent architectures that can be deployed across many customers with varying data, tools, and constraints. Develop evaluation frameworks, datasets, environments, and metrics to measure agent performance, reliability, and business impact in production settings. Collaborate closely with product managers, customers, data annotators, and other engineering teams to translate enterprise requirements into robust agent designs. Productionize frontier agent techniques (e.g.,
From $158K/yr
Lithic is the modern card issuing and processing platform empowering ambitious financial companies to build the future of payments. Our infrastructure powers card programs for 100+ innovative clients, from fintechs reimagining credit and digital banking to platforms transforming disbursements and spend management. Companies like Mercury, Flex, and Novo rely on Lithic's developer-friendly APIs, direct network connections, and flawless reconciliation to launch and scale card programs in weeks, not years. We're building a future where access to better financial products materially improves people's lives, free from the constraints of 30-year-old mainframes and legacy processors. We're proud to be backed by world-class investors who share that vision, including Bessemer Venture Partners, Index Ventures, Spark Capital, Stripes, and Mastercard, along with many others. We're a team of 170+ across 26 states and 7 countries, headquartered in New York City. We are hiring for our Treasury team Software Engineers at various levels (II and Senior) who are curious and willing to dive deep and understand our technology and domain in order to solve interesting and hard problems.The Treasury team maintains and builds the backend services that manage the flow of funds between Lithic and third parties. This includes our ledger, ACH and wire infrastructure, and associated reconciliation. The systems we maintain have high standards of reliability and correctness. You will become an expert in the card payments space. The Treasury team primarily uses Python for their tech stack. What You'll Do: Ensure high reliability and correctness for Lithic’s ledger and orchestrated funds flows Develop new features to better serve Lithic customers Ensure that the team is delivering reliable, secure, and scalable code with minimal tech debt Own initiatives from planning to launch, keeping stakeholders informed and aligned along the way Lead efforts to improve systems and processes wit
C$132K – C$165K/yr
OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric—with AI embedded across the automation lifecycle—Redwood accelerates the path to autonomous operations. Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT As a Senior Full Stack Software Developer, you will be responsible for leading the design, development, and delivery of scalable full-stack applications, shaping system architecture, and driving engineering excellence across Redwood’s automation and SaaS platforms. Design, develop, and implement scalable, secure, and high-performance full-stack applications using Java, JavaScript, and related technologies Architect and build backend services, APIs, and microservices with a focus on scalability, reliability, and maintainability Develop responsive, accessible, and high-quality front-end user experiences Partner with product managers and stakeholders to define technical strategy and translate business requirements into system designs Own and contribute across the full software development lifecycle, from architecture and design to deployment and optimization Establish and promote best practices in coding, testing, observability, performance optimization, and AI usage Lead architectural discussions a
#TeamNextdoor Nextdoor is where you connect to the neighborhoods that matter to you so you can belong. Our purpose is to cultivate a kinder world where everyone has a neighborhood they can rely on. Neighbors around the world turn to Nextdoor daily to receive trusted information, give and get help, get things done, and build real-world connections with those nearby — neighbors, businesses, and public services. Today, neighbors rely on Nextdoor in more than 300,000 neighborhoods across 11 countries. Meet Your Future Neighbors As a Software Engineer at Nextdoor, you’ll join a focused team of developers, product managers, and designers who are passionate about using technology to cultivate a kinder world where everyone has a neighbor they can rely on. We are a small team of engineers that wear multiple hats and work across different languages and services to deliver value to our members. We care about moving fast and delivering impact, without compromising on quality and reliability. You will have the opportunity to learn from your co-workers and teach them. As a team, we will make each other better and build great software. What You’ll Bring To The Team If you didn't see an opportunity listed that looked right for you, we would still love to hear from you and consider you for other opportunities! At Nextdoor, we empower our employees to build stronger local communities. To create a platform where all feel welcome, we want our workforce to reflect the diversity of the neighbors we seek to serve. We encourage everyone interested in our purpose to apply. We do not discriminate on the basis of race, gender, religion, sexual orientation, age, or any other trait that unfairly targets a group of people. In accordance with the San Francisco Fair Chance Ordinance, we always consider qualified applicants with arrest and conviction records. #LI-DNI
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. AS A SENIOR SOFTWARE ENGINEER YOU WILL: Drive high-impact initiatives that span our product areas and full tech stack, including golang and Python on the backend and TypeScript/React on the frontend. Own and deliver features across the notebook service, container runtimes, and UI — designing and shipping medium-to-large projects independently, from ambiguous problem statements through production and post-launch. Advance core platform initiatives such as runtime management and patching, environment reproducibility and replication, security and compliance, and observability for notebooks. Extend the product to operate reliably in regulated and air-gapped environments, where security, compliance, and operational rigor are paramount. Promote strong collaboration within a cross-functional team and partner closely with embedded product managers and designers, as well as platform organizations across Snowflake Be a strong contributor to the product vision and drive team planning. Build for scale, reliability, and high performance, and participate in the on-call rotation to keep a Tier-1 production service healthy. Mentor, coach, and empower more junior team members, and raise the engineering bar through high-quality design and code review. OUR IDEAL CANDIDATE WILL HAVE: 7+ years o
From C$1.3M/yr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Safety and Customer Care (SCC) team at Lyft manages over 1.7 million monthly human and AI interactions and serves as Lyft's primary direct touchpoint with riders and drivers. We handle critical infrastructure that powers both human associates and AI agents to make riders and drivers feel safe and comfortable while riding or driving with Lyft, transforming every support interaction into a moment of genuine connection. As a Data Engineer on the SCC team, you will have ownership over the data modeling and pipelines that power SCC’s Associate and AI Agent Platform . Your efforts will be critical to the reliability of our pipelines, execution of third party data integrations, accurate reporting of agents performance, and efficiency improvements that can save millions of dollars / year. You will work cross-functionally to bridge Lyft's business goals with data engineering. Your efforts will allow access to business and user behavior insights, using huge amounts of Lyft data to fuel several teams such as Analytics, Data Science, Engineering, and many others. Responsibilities: Owner of the core data pipeline, responsible for scaling up data processing flow to meet the rapid data growth at Lyft Evolve data model and data schema based on business and engineering needs Implement systems tracking data quality and consistency Develop tools supporting self-service data pipeline management (ETL) SQL and MapReduce job tuning to improve data processing performance Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and distribute knowledge Collaborate cross-functionally with product, engineering, data science, and marketing teams to understand business problems and align on prioritization and solutions Experience: Bachelor's degree in Compute
C$125K – C$200K/yr
We deliver foundational systems that shape the future of how technology is used in a top performing quantitative equity fund that manages over $78+ billion USD in financial assets. This is a fantastic opportunity in the exciting intersection of finance and technology where investment decisions are made using technology. Quantitative equity funds use programmed investment strategies and as a result, our technology team is crucial to its success. The team is headquartered and deeply rooted in West Coast Vancouver. We place high value on maintaining an entrepreneurial spirit and creating a culture where each of us has opportunities to succeed. What You Will Do The technology infrastructure team plays an essential role through innovative technologies on our hybrid (on-premise and cloud based) platform: distributed computing, petabyte-scale data storage, containerization, non-traditional high-performance databases, process orchestration, monitoring, data visualization and DevOps. You own the entire technology infrastructure life cycle: Engineer and support software and systems infrastructure. Introduce new foundational technologies that advance our software engineering capabilities to the next level. Collaborate with our software development teams on support issues and improvements to our infrastructure tools, processes, and software. Act as a conduit between our application development teams, and IT, network security, and other stakeholders to align priorities and translate business requirements into technical designs. Improve systems infrastructure reliability. Gather and analyze metrics from operating systems and applications to assist in performance tuning, fault finding and business continuity planning. Design, plan and implement solutions in an entrepreneurial spirit. What You Bring Programming Knowledge – you have an undergraduate, graduate, or post-graduate degree in a computer-related field OR exceptional programming skills gain
Overview: Qsight is a high-growth division of Guidepoint focused on building data intelligence solutions for the healthcare sector. Qsight leverages proprietary datasets and rigorous analysis of alternative data sources to generate actionable insights for top-tier institutional investors, medical device manufacturers, and pharmaceutical companies. The Qsight team develops market intelligence products designed to be highly relevant, accurate, and scalable – delivering superior insights to a diverse, global client base. We are seeking an experienced, motivated Tehnical Operations Engineer to join our growing team. This is a multiple-hats role focused on SaaS/platform operations and tier-2 support for client-facing systems. You will own the administration and reliability of key tools, troubleshoot and resolve escalations with clear documentation, and build lightweight automation and reporting to reduce manual work as we scale. You will partner closely with Customer Success, Product, and Engineering to proactively monitor, support, and improve critical systems. Through practical, creative problem-solving, you will strengthen reliability, accelerate time to resolution, and increase operational visibility. Day to day, you will triage and resolve client technical questions, manage vendor license administration and renewals, and produce reporting that informs operational decisions. This role is a launchpad toward an SRE/Platform Engineering track as you grow into deeper automation, reliability engineering, and systems design work. This is a hybrid position based out of our Toronto office. What You’ll Do: Platform Support Own routine ops and configuration changes for critical SaaS platforms – Including Tableau, Freshdesk, Datadog, and our own client facing and internal portals Configure and maintain Freshdesk portals, routing, SLAs, permissions, integrations, etc. based on business requirements. Automate manual operations with Python, PowerAutomate, and shell scri
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Lyft Business Product Platform team builds the systems and experiences that power Lyft's B2B products — enabling companies, organizations, and their employees to seamlessly access Lyft's transportation network. We sit at the intersection of product and platform, owning both the customer-facing features and the underlying infrastructure that makes them reliable at scale. Our work directly impacts how businesses integrate with Lyft, how admins manage their programs, and how millions of riders get where they need to go. Responsibilities: Drive architecture and technical design for systems that are highly available, scalable, and built to last — not just for today's requirements but for where the product is heading Own features end-to-end: from shaping the technical spec and design through to production rollout and operational health Think critically about how AI capabilities can be incorporated into Lyft Business products to improve the experience for business admins and riders — and bring that perspective into roadmap and architecture conversations Make well-reasoned trade-off decisions and communicate them clearly to peers, leads, and cross-functional partners Write clean, well-tested, maintainable code and hold a high bar for the same in code reviews Partner across engineering, product, and design to align on direction and get buy-in on technical approaches Proactively engage in incident response, contributing both to resolution and to long-term reliability improvements Grow the team's technical culture through design reviews, tech talks, and mentorship Experience: 5+ years of software engineering experience, with a track record of designing and shipping production systems at scale Strong system design instincts — you can reason through distributed systems trade-offs, identify failure modes, and
Related career options
Similar roles with stronger pay
Demand 50/100 · 8 jobs
$1.3M – $1.3M/yr
Salary →Demand 67/100 · 15 jobs
$972.9K – $972.9K/yr
Salary →Other cities to consider
More places hiring for this role
Get new engineering manager platform reliability jobs in Canada by email
Daily job updates · Unsubscribe anytime