Jobs in Canada

Distributed Systems Engineer in Toronto

44 active opportunities · Updated October 2026

Explore current distributed systems engineer jobs in Toronto. Filter by work mode, employment type, experience, department, date posted and distance.

L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$108K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As an Observability team member, you are responsible for the operation and maintenance of our logging and metrics infrastructure. You ensure all teams at Lyft are aware of the operational health of their products by monitoring system availability and take a holistic view of our platform performance. You build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you find opportunities to improve our systems in order to push our platform forward. You provide our partners with the support they need to help them build robust large scale distributed systems. We count on the reliability of our infrastructure to empower Lyft teams to provide our customers rich experiences that are highly available with rock solid performance to ensure our transportation platform continues to connect people and places. As we grow our team, we are seeking experienced Infrastructure Engineer to ensure that as our Infrastructure continues to scale, our platform continues to provide an essential and dependable service that transports millions of people every day. Specifically we are searching for someone who brings fresh perspectives, enjoys collaborating with cross-functional teams in order to continually improve our products and services for our customers. Responsibilities: Maintain, improve, and develop tooling and systems that enhance the reliability, scalability, and efficiency of our platform. Assist engineering teams in defining service-level objectives (SLOs) and provide the necessary toolin

PythonAWSKubernetesAI
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$108K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. We are building and maintaining a highly scalable asynchronous platform that empowers our organization to handle critical business cases. As a software engineering team, our mission is to create robust and innovative solutions that drive the success of our business and deliver unparalleled value to our customers. We adopt Infrastructure as Code practice to automate the provisioning and configuration of our resources, which helps reduce manual configuration and improve consistency. Our team culture is built on collaboration, open communication, and a supportive environment where each member's ideas are valued and contributions are recognized. We believe in the importance of fostering a positive workplace culture that inspires innovation and creativity. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and production readiness reviews, platform management and capacity planning ceremonies with cross-functional teams Document Infrastructure operations process and insights, identify repeatable actions and ruthlessly automate repetitive tasks Participate in our teams on-call rotations, respond to incidents and support other teams mitigate customer impacting events Experience: 5+ years experience working on teams responsible for software development, automation and systems engineering Experience building large-scale infrastructure, distributed systems or networks. Knowledge with SQS,

PythonAWSAzureGCP
S
📍 Toronto, Ontario, Canada· Full-time
✓ Quality checkedCompany trend -100%

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Issue Workflow is Sentry's primary product surface. Our issue platform processes billions events daily and turns them into actionable insights that help millions of developers fix bugs faster. As a Staff Software Engineer on the Issue Workflow team, you'll architect the systems that power this experience. You'll work at the intersection of high-scale distributed systems and product engineering, building real-time data pipelines, search backends, and analysis systems that surface signal from noise. This is product engineering at massive scale—where every architectural decision impacts millions of debugging sessions. You'll be the technical leader who shapes how Sentry groups issues, how we make search lightning-fast, how we enable sophisticated agentic workflows, and how we ensure that the product is performant even at billions-of-events scale. Your work will define what's possible for the most trafficked part of Sentry's platform. In this role you will Drive technical strategy and roadmap. Partner with engineering leadership, product, and design to shape the multi-quarter technical vision for Issue Workflow platform. Make strategic calls about architectural direction, technology choices, and technical debt. Ensure the team is building a strong foundation to scale with Sentry's growth. Solve complex performance and scalability challenges. Champion product quality and user experience. Build features that don't just work—they delight. You understand that milliseconds matter in the developer experience. You sweat the details of interfaces, error messages, loading states, and edge cases. You instrument everything s

TypeScriptPythonSQLPostgreSQL
T-
📍 Toronto, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Role: As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams to build world-class machine learning inference platforms. These platforms power essential services like personalized recommendations, search, and content understanding across Tubi. A core responsibility of this team is developing and maintaining low-latency ML model serving systems that support Deep Learning, LLM, and Search models. This involves building self-service infrastructure and critical components such as the inference engine, feature store, vector store, and experimentation engine. You will improve the way we deploy and operate our services and even contribute to open-source projects. This role grants the architectural freedom to explore new frameworks, lead critical cross-functional projects, and transform the capabilities of our ML and Product teams. Responsibilities: Design and build scalable, high throughput, and low latency distributed systems using Scala Build reusable components and services that serve various ML applications like Personalization, Search, Ads and Exploration Partner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art. Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessary Mentor other engineers on the team on system design, effective incident management, interviewing, leveraging LLMs for work, etc. Collaborate with ML, Product, and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi. Your Background: Experience designing and building scalable, distributed systems in any modern backend language (e.g., Scala, Java, Python, Go, C++); experience with Scala or JVM b

PythonJavaSQLRedis
PE
📍 Toronto, Canada· Full-time· Hybrid
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Role We are a small team of AI builders in Paytm Labs. As a Staff AI Platform Engineer, you will work across inference and agentic systems. You will contribute to Paytm's AI inference platform (Pi), serving internal teams and enterprise customers - running our own coding and domain-specific models (voice, vision, risk, fintech workflows) as well as third-party models. You will also architect and build the platform that enables autonomous AI agents to operate safely and reliably in production - the runtime, orchestration, and developer tooling for agents to reason, plan, use tools, and execute complex multi-step workflows, automating both software development and business processes. You will work at the intersection of LLMs, distributed systems, and production fintech infrastructure, helping define how inference and agentic AI are built and deployed across payments, risk, fraud, collections, support, and developer experience.

T-
📍 Toronto, Canada· Full-time
✓ High-confidence listing

From C$1.4M/yr

Quick readStrong listing-quality and freshness signals

About the Role: Tubi's content platform is the engine behind one of the largest free streaming services in the world. Every play, every deal, every creator, every frame of video flows through systems CPE owns, and the surface area is enormous. Distributed services running on the hottest path of Tubi's traffic. Video pipelines processing one of the largest workloads in streaming. Workflow engines automating the operations that used to consume entire teams. Creator-facing products turning a back-office process into a real platform. And on top of all of it, an AI-native rebuild of the CMS that most companies aren't willing to attempt. This isn't a single-domain role. It's a platform where backend, frontend, video, infrastructure, and applied AI all collide at the scale where decisions actually matter, where an architectural choice ripples across millions of titles and billions of requests, and where the difference between "good enough" and "great" shows up in revenue. We're looking for builders who want to range across domains — backend one quarter, frontend the next, applied AI the one after that — and who want their work to be felt: by viewers when a title plays instantly, by creators when they go live the same day, by Content Ops when a workflow runs itself, and by the business when the platform stops being a cost center and starts being a force multiplier. The infrastructure is already there. The mandate is already there. What's missing is the people who want to build the thing, not talk about it. Come build it. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: You'll work on systems that sit at the heart of Tubi's business, where the content pipeline meets the viewer, the creator, and increasingly, the AI agent. The work spans the full stack of a modern content platform: distributed services, video infrastructure, workflow automation, and applied AI, all running at

TypeScriptPythonReactKubernetes
T-
📍 Toronto, Canada· Full-time
✓ High-confidence listing

From C$1.4M/yr

Quick readStrong listing-quality and freshness signals

About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission is to engineer resilience from the ground up, enabling our product teams to innovate rapidly while ensuring our users have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through a culture of data-driven decision-making, blameless learning, and relentless automation. As a Senior Site Reliability Engineer, you are a hands-on engineer who blends deep software development expertise with a passion for operational excellence. You will be responsible for designing, building, and running the resilient, scalable, and increasingly self-healing systems that power our products. You will apply sound engineering principles to solve our most complex reliability challenges, with a mandate to automate everything, eliminate toil, and write robust, maintainable code. You will be a force multiplier, mentoring other engineers and elevating the site reliability bar for the entire organization. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: System Architecture & Design: Design, build, and maintain scalable, highly available, and fault-tolerant distributed systems. Partner with development teams as a reliability consultant, reviewing designs and influencing architectural decisions to ensure new services are built with reliability, observability, and performance as core principles, not afterthoughts. Automation & Software Development: Write robust, performant, and maintainable code to automate operational tasks, and CI/CD pipelines. Build the internal tools, libraries, and frameworks that enable engineering teams to self-service their

TypeScriptPythonAWSKubernetes
S
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -100%

$100K – $125K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role At Sentry, Support is an engineering discipline. Our customers are the greatest technical minds in the world—developers at elite enterprises building the future of software—and they deserve answers that go deeper than a knowledge base link. We are architecting the Technical Support engine . We’re looking for a veteran engineer to help us redefine the standard of technical support by combining deep human expertise with autonomous agentic systems. You are a debugger of both code and systems. You will treat support volume as a data signal to build automated resolution paths, ensuring our human engineers only touch the most complex, high-impact architectural puzzles. Sentry Support Engineers aren't just clearing queues; they are Orchestrators . You will engage with our users across GitHub, Discord, and our internal systems, while acting as the Technical Lead for our Agentic Ops. You ensure that when a developer asks a complex question, our systems have the right context and a seamless "Human-in-the-Loop" path to you when deep, nuanced expertise is required. In this role you will Master the Sentry Ecosystem & Support Elite Developers Deep-Dive Debugging: Perform root-cause analysis on complex issues and distributed tracing gaps across polyglot environments. Support the Great Minds: Act as a strategic consultant for senior engineers at our largest enterprise customers, solving high-stakes architectural challenges that push the boundaries of observability. Troubleshoot SDK Implementations: Go deep into the source code of Sentry’s SDKs to help developers instrument complex frameworks and custom environments. Engin

JavaScriptPythonJavaReact
T
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a skilled Software Engineer with a passion for building high-performance, low-level systems software. In this role, you’ll contribute to the development and optimization of the infrastructure that powers our cutting-edge processors, with a primary focus on C/C++ development and low-level programming. You'll work closely with large inference and training model development to further drive Scale Out software and hardware performance. This role is hybrid, based out of Toronto, ON. Who You Are Strong C or C++ systems engineer with a deep understanding of memory, threading, I/O, and low-level execution models. Experienced building low-level software, drivers, embedded systems, or performance-critical infrastructure. Comfortable working close to hardware and curious about how systems behave under the hood. Proficient with Linux systems programming and debugging tools such as gdb, strace, and perf. Structured problem solver who thrives in fast-paced, highly technical environments. What We Need Design, develop, and maintain core infrastructure software that interfaces directly with Tenstorrent hardware. Build low-level libraries and APIs for communication and synchronization across compute nodes. Optimize system-level software for performance, scalability, and reliability in distributed environments. Support hardware

AWSLinuxAIC++
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$136K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As a Infrastructure Engineer at Lyft, you will run our Production Infrastructure by monitoring system availability and take a holistic view of our platform health. You will build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you will seek opportunities to optimize our systems in order to push our platform forward, anticipating our customers' needs in order to continually improve the platform. You will provide Lyft partner teams with operational support to help them build robust large scale distributed systems. About the Team Data Pipelines is at the heart of all critical data flowing through Lyft supporting hundreds of services that impact millions of drivers and passengers every day. Our team’s mission is to empower Lyft engineers to self-serve in building and maintaining data pipelines as needed to support products that deliver the world’s best transportation experience. We leverage a variety of technologies to store, stream and manage data making it available to our internal customers. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and

PythonAWSDockerKubernetes
S
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -100%

C$162K – C$420K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. About the role As a Senior Software Engineer, you will architect and scale the core systems that power Sentry's billing infrastructure, ensuring accuracy and reliability at massive scale. You will collaborate on building the next generation of Sentry’s usage tracking pipeline, processing hundreds of billions of events daily with low latency and financial-grade accuracy. You will help design flexible pricing primitives that support everything from per-event usage billing to complex enterprise contracts, enabling product and sales teams to experiment rapidly while maintaining revenue accuracy and reduced time-to-market for new products. You will contribute to technical decisions on data consistency challenges unique to billing—like handling event delays, retroactive pricing changes, and distributed count reconciliation across our infrastructure. You'll love this job if you Want to solve the "easy to explain, hard to build" problems—like ensuring a customer's bill matches their usage perfectly, even when processing hundreds of billions of events daily across distributed

O
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -63.6%

From C$160K/yr

Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Streaming Foundations team builds services and operates data pipeline infrastructure to support event streaming, messaging, and analytics use cases. We are looking for a Software Engineer who is passionate about distributed systems, platform engineering, and solving data-intensive problems at scale. In this high-impact role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Auth0 to scale for years to come. What you’ll be doing Help set the technical direction for the team and influence the engineering roadmap for the Platform’s streaming capabilities Design and lead the implementation of our most complex and critical systems for data-intensive use cases. Research and champion new technologies and architectural patterns to solve strategic challenges and scale the platform. Lead and influence cross-functional initiatives, ensuring technical alignment and successful execution across multiple teams. Improve the operational posture of our systems by designing for observability, reliability, and scalability, and by mentoring others in operational best practices. Coach and mentor senior engineers and act as a technical leader across the engineering organization. Collaborate with different stakeholders like product teams whenever needed. What you’ll bring to our teams 7+ years of software development experience in a fast-paced, agile environment Experience working with Golang or Java is preferred H

TypeScriptJavaReactAWS
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

$136K – $170K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Marketplace teams are at the heart of our products and decision-making, owning everything from rider pricing to driver earnings, incentives, and efficient matching. We’re looking for passionate, driven engineers to build systems that empower our riders and drivers to have the best transportation experience possible through prediction, adaptivity, and personalization. We’re looking for someone who is excited about working in a fast-paced, innovative, and impactful environment to create reliable solutions to distributed computing, ML, and data problems. The Pricing team is a centerpiece of Lyft’s Marketplace org, determining prices for all rideshare products and supporting new initiatives. We work with Product & Science to solve and implement complex pricing requirements, balancing the needs of riders, drivers, and the business goals. As an owner of one of the most critical flows in the company, you will work on a wide array of challenges such as latency-sensitive concurrency problems, large scale distributed systems, and experimentation. If you’re interested in playing a large part in demand / supply management and improving the Lyft customer experience, this could be a great fit for you. Responsibilities: Help define the roadmap and architecture based on technology and business needs Unblock, support, effectively communicate, and obtain buy-in across teams to achieve results Lead projects of multiple people from idea to positive execution Write clear, scalable and clear design documentation Write well-crafted, well-tested, readable, maintainable code Utilize your expertise in Python, Golang, AWS to deliver robust and scalable solutions Participate in code reviews to ensure code quality and distribute knowledge Proactively participate in resolving ongoing incidents Share your kno

PythonAWSAzureGCP
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$108K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Marketplace teams are at the heart of our products and decision-making, owning everything from rider pricing to driver earnings, incentives, and efficient matching. We’re looking for passionate, driven engineers to build systems that empower our riders and drivers to have the best transportation experience possible through prediction, adaptivity, and personalization. We’re looking for someone who is excited about working in a fast-paced, innovative, and impactful environment to create reliable solutions to distributed computing, ML, and data problems. The Pricing team is a centerpiece of Lyft’s Marketplace org, determining prices for all rideshare products and supporting new initiatives. Rider Engagement develops rider-facing engagement levers and optimizes user pricing experience to drive both short term and long term business outcomes. We work with Product & Science to solve and implement complex pricing requirements, balancing the needs of riders, drivers, and the business goals. As an owner of one of the most critical flows in the company, you will work on a wide array of challenges such as latency-sensitive concurrency problems, large scale distributed systems, and experimentation. If you’re interested in playing a large part in demand / supply management and improving the Lyft customer experience, this could be a great fit for you. Responsibilities: Drive high-impact projects and innovate new solutions to provide the best user experience. Work closely with cross-functional teams and partner teams to develop solutions based on technology and business needs, and advance team’s goals and priorities Independently lead features from idea to positive execution and launch Unblock, support and communicate with internal partners to achieve results Write well-crafted, well-tested, readable, maintaina

PythonAWSRestAI
C
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -91.5%

From £215K/yr

Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the role. We’re building the next generation of agentic AI infrastructure at Cohere. This team sits at the intersection of ML systems, distributed infrastructure, and developer experience, creating the platform that powers autonomous AI agents at scale. You’ll work on hard, forward-looking problems with few established patterns, including secure code execution, agent state management, model routing, identity and authentication, and resource management for long-running agent workflows. This role is a strong fit for someone who combines systems depth with ML intuition. You should be comfortable building reliable infrastructure, thinking through distributed systems tradeoffs, and understanding how emerging agentic capabilities shape platform design. What you’ll work on. Secure execution environments for agent-generated code Identity, authentication, and trust boundaries for agents Model routing and orchestration across different model types and environments Rate limiting, quotas, and resource management for agent workflows State management, memory, and filesystem abstractions for agents. In this role you will: Turn emerging M

KubernetesGitRestAI
🔔

Get new distributed systems engineer jobs in Toronto, Canada by email

Daily job updates · Unsubscribe anytime