Jobs in Canada

Distributed Systems Engineer in Canada

108 active opportunities · Updated October 2026

Explore current distributed systems engineer jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.

L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$108K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As an Observability team member, you are responsible for the operation and maintenance of our logging and metrics infrastructure. You ensure all teams at Lyft are aware of the operational health of their products by monitoring system availability and take a holistic view of our platform performance. You build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you find opportunities to improve our systems in order to push our platform forward. You provide our partners with the support they need to help them build robust large scale distributed systems. We count on the reliability of our infrastructure to empower Lyft teams to provide our customers rich experiences that are highly available with rock solid performance to ensure our transportation platform continues to connect people and places. As we grow our team, we are seeking experienced Infrastructure Engineer to ensure that as our Infrastructure continues to scale, our platform continues to provide an essential and dependable service that transports millions of people every day. Specifically we are searching for someone who brings fresh perspectives, enjoys collaborating with cross-functional teams in order to continually improve our products and services for our customers. Responsibilities: Maintain, improve, and develop tooling and systems that enhance the reliability, scalability, and efficiency of our platform. Assist engineering teams in defining service-level objectives (SLOs) and provide the necessary toolin

PythonAWSKubernetesAI
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$108K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. We are building and maintaining a highly scalable asynchronous platform that empowers our organization to handle critical business cases. As a software engineering team, our mission is to create robust and innovative solutions that drive the success of our business and deliver unparalleled value to our customers. We adopt Infrastructure as Code practice to automate the provisioning and configuration of our resources, which helps reduce manual configuration and improve consistency. Our team culture is built on collaboration, open communication, and a supportive environment where each member's ideas are valued and contributions are recognized. We believe in the importance of fostering a positive workplace culture that inspires innovation and creativity. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and production readiness reviews, platform management and capacity planning ceremonies with cross-functional teams Document Infrastructure operations process and insights, identify repeatable actions and ruthlessly automate repetitive tasks Participate in our teams on-call rotations, respond to incidents and support other teams mitigate customer impacting events Experience: 5+ years experience working on teams responsible for software development, automation and systems engineering Experience building large-scale infrastructure, distributed systems or networks. Knowledge with SQS,

PythonAWSAzureGCP
P
📍 Toronto, ON, CA· Full-time
✓ Quality checkedCompany trend -100%

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . People use Pinterest to find ideas and brands that they love. We aspire to help our advertisers and partners reach their audiences with inspiring content. The API team is responsible for ensuring our first party clients (Android, iOS, and Web) have a stable API to develop on top of. Our customers are Pinterest developers who want to build product features, and who rely on a highly available system to do so. As the Engineering Manager for API, you will lead a talented and growing engineering team responsible for growing the existing portfolio. The ideal candidate should have experience building backend web applications, be driven to become an expert in their domain, have some knowledge of distributed systems engineering, and have a passion for leadership. What you’ll do: Collaborate with stakeholders across the organization to architect so

AWSRestAIGo
L
📍 San Francisco, CA· Full-time
✓ High-confidence listingCompany trend -72.4%
Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft Infrastructure builds the systems engineers depend on to ship stable, scalable, and efficient services. We're hiring a Senior Technical Program Manager to run cross-functional programs across our infrastructure and data platform teams. This role blends program delivery with product sense: you'll own the roadmap for your area, set priorities, and act as the voice of the customer back into how we build. Responsibilities Run infrastructure programs end to end, from kickoff through delivery Own the roadmap for your platform area: shape the strategy, sequence the work, and make the prioritization calls Drive data platform migration and modernization work, coordinating across engineering, data, and platform teams to keep dependencies and timelines under control Be the voice of the customer: partner with engineering teams across Lyft, surface their pain points, and feed that back into priorities and roadmaps Define success metrics and adoption goals, gather and document customer requirements, and make sure what ships actually solves the problem Build feedback loops with customer teams and turn what you hear into concrete improvements Partner with engineering and infrastructure leads to build plans, call out risks early, and keep stakeholders aligned Own program health: track milestones, surface blockers before they slip, and keep decision-makers in the loop Use your technical background in distributed systems and data infrastructure to ask sharp questions and build plans the team believes in Share in the team's release oncall rotation Experience 5+ years in Technical Program Management or a TPM/PM hybrid role A background in software, data, or systems engineering, enough to go deep with engineers Experience owning a roadmap: setting strategy, prioritizing across competing demands, and defining what suc

E
📍 New York, NY or Los Angeles, Canada· Full-time
✓ High-confidence listing

$180K – $295K/yr

Quick readStrong listing-quality and freshness signals

The Opportunity This is a critical and exciting time at Enigma. Our customers consistently tell us that our data products create tremendous value and are deeply aligned with their most important workflows. As demand grows, we have an urgent opportunity to improve both the intelligence of our data and the systems through which customers access it. We are looking for an experienced Senior/Staff Machine Learning Engineer to join our Match Team and help shape the next generation of Enigma’s customer-facing data products. In this role, you will combine advanced statistical and machine learning research with the engineering systems required to power fast, relevant, and reliable search experiences at scale. This is a uniquely high-impact role sitting at the intersection of information retrieval, ranking systems, semantic search, distributed systems, and customer data delivery. The Role At the core of Enigma’s product is our data, which makes both data science and delivery systems central to what we build. As a Senior/Staff ML Engineer on the Match Team, you will lead efforts that improve the relevance, latency, and scalability of our customer-facing data products. You’ll work across the full lifecycle: framing retrieval and ranking problems, developing models and experimentation strategies, evaluating results using real-world signals, and implementing high-throughput search and retrieval systems. This role is ideal for someone who is excited by both hard ranking/search problems and the systems challenges of turning those solutions into low-latency, production-grade retrieval systems. What You'll Do Develop innovative solutions to complex problems in information retrieval, ranking, semantic search, query understanding, and recommendation systems Build and optimize low-latency, high-throughput search APIs, indexing pipelines, and retrieval systems using Python, Typesense, and AWS Evaluate and evolve our search technology stack, driving technical design decisions across index

PythonAWSMachine LearningAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team The Storage teams build and operate online stateful systems and abstractions that are reliable, efficient, secure and easy to use for DoorDash Engineering. The teams are responsible for understanding Product Engineering’s evolving needs and developing platform and infrastructure capabilities to serve them. The team currently supports CockroachDB, Cassandra, Kafka and Redis as well as data abstraction services to reduce the complexity of interacting with storage systems for Product Engineers. About the Role The Storage team is building and operating a high-performance, scalable, and reliable data abstraction layer that optimizes both efficiency and reliability. Our goal is to create a platform that manages itself and fades into the background—empowering engineers to focus on delivering product experiences our customers love. This role is available across two teams within Storage, each solving unique and high-impact challenges: One team is building the orchestration layer for DoorDash’s storage platform—unifying lifecycle management, operations, and self-serve APIs for databases and streaming systems, turning complex, stateful infrastructure into reliable, developer-friendly services used across the company. One team builds and operates the distributed data platform powering DoorDash's largest stateful workloads -- including Cassandra, which backs critical product surfaces across DoorDash, Wolt, and Roo. You'll design high-throughput data abstractions, smart clients, and platform services that make distributed data reliable and easy to work with at multi-petabyte, multi-million-QPS scale, with opportunities to go deep on distributed systems internals and contribute to the open-source Cassandra ecosystem. If you're passionate about distributed systems, developer experience, and building foundational infrastructure at scale, we'd love to hear from you. You must be located in San Francisco, Sunnyvale, Seattle, or the New York Metro Area for this hybrid pos

JavaSQLRedisAWS
S
📍 Toronto, Ontario, Canada· Full-time
✓ Quality checkedCompany trend -100%

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Issue Workflow is Sentry's primary product surface. Our issue platform processes billions events daily and turns them into actionable insights that help millions of developers fix bugs faster. As a Staff Software Engineer on the Issue Workflow team, you'll architect the systems that power this experience. You'll work at the intersection of high-scale distributed systems and product engineering, building real-time data pipelines, search backends, and analysis systems that surface signal from noise. This is product engineering at massive scale—where every architectural decision impacts millions of debugging sessions. You'll be the technical leader who shapes how Sentry groups issues, how we make search lightning-fast, how we enable sophisticated agentic workflows, and how we ensure that the product is performant even at billions-of-events scale. Your work will define what's possible for the most trafficked part of Sentry's platform. In this role you will Drive technical strategy and roadmap. Partner with engineering leadership, product, and design to shape the multi-quarter technical vision for Issue Workflow platform. Make strategic calls about architectural direction, technology choices, and technical debt. Ensure the team is building a strong foundation to scale with Sentry's growth. Solve complex performance and scalability challenges. Champion product quality and user experience. Build features that don't just work—they delight. You understand that milliseconds matter in the developer experience. You sweat the details of interfaces, error messages, loading states, and edge cases. You instrument everything s

TypeScriptPythonSQLPostgreSQL
V
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$150K – $240K/yr

Quick readStrong listing-quality and freshness signals

Location: San Francisco, CA (Remote/Hybrid Available) What is Verse? The race to AI has become the race to power. Every breakthrough in artificial intelligence depends on one thing: access to electricity. But across the country, aging grid infrastructure and years-long interconnection queues are slowing the deployment of the data centers that will power the next generation of innovation. Solving this challenge isn't just about energy—it's about unlocking the future of AI. At Verse, we're building the energy intelligence platform for the AI economy. Our software helps the world's largest energy consumers achieve faster, cheaper, and cleaner power by combining real-time control of energy assets with complete visibility into their energy portfolio. Backed by Bessemer Venture Partners, GV, Coatue, and NVIDIA, and built by pioneers in grid-scale batteries, energy markets, and enterprise software, we're redefining how the world's most ambitious organizations access and manage energy. The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in collaboration with some of the brightest industry experts in the field building cloud-native applications that scale to trillions of data points collected from electricity markets globally. You will be a part of a dynamic, robust team primarily supporting the backend needs of our Aria software product spanning hundreds of data sources, sinks, services, and jobs. Your expertise will not only have a direct impact on product decisions, but you also be well-positioned to drive the development and trajectory of our entire platform and infrastructure and influence important architectural decisions that affect the whole organization. Key Responsibilities Foster a culture and mindset of well-designed systems, test-driven software, and transparent communication with a high caliber of mutual respect and consideration for stakeholders Read and write a lot of Go, Python, and Protobuf Build, test, debug, maint

PythonJavaKubernetesMicroservices
T-
📍 Toronto, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Role: As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams to build world-class machine learning inference platforms. These platforms power essential services like personalized recommendations, search, and content understanding across Tubi. A core responsibility of this team is developing and maintaining low-latency ML model serving systems that support Deep Learning, LLM, and Search models. This involves building self-service infrastructure and critical components such as the inference engine, feature store, vector store, and experimentation engine. You will improve the way we deploy and operate our services and even contribute to open-source projects. This role grants the architectural freedom to explore new frameworks, lead critical cross-functional projects, and transform the capabilities of our ML and Product teams. Responsibilities: Design and build scalable, high throughput, and low latency distributed systems using Scala Build reusable components and services that serve various ML applications like Personalization, Search, Ads and Exploration Partner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art. Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessary Mentor other engineers on the team on system design, effective incident management, interviewing, leveraging LLMs for work, etc. Collaborate with ML, Product, and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi. Your Background: Experience designing and building scalable, distributed systems in any modern backend language (e.g., Scala, Java, Python, Go, C++); experience with Scala or JVM b

PythonJavaSQLRedis
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $189.6K/yr

Quick readStrong listing-quality and freshness signals

Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering MLEs, researchers, data scientists and operators for fast and automatic training and evaluation of LLM's, as well as evaluation of data quality. Scale is uniquely positioned at the heart of the field of AI as an indispensable provider of training and evaluation data and end-to-end solutions for the ML lifecycle. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development. You will be building and optimizing the platform to enable our next generation of LLM training, inference and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you! You will: Build, profile and optimize our training and inference framework Collaborate with ML teams to accelerate their research and development and enable them to develop the next generation of models and data curation Research and integrate state-of-the-art technologies to optimize our ML system Ideally you’d have: Strong excitement about system optimization Experience with multi-node LLM training and inference Experience with developing large-scale distributed ML systems Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc. Strong written and verbal communication skills and the ability to operate in a cross functional team environment Nice to haves: Demonstrated expertise in post-training methods &/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the positi

AWSRestAIGo
PE
📍 Toronto, Canada· Full-time· Hybrid
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Role We are a small team of AI builders in Paytm Labs. As a Staff AI Platform Engineer, you will work across inference and agentic systems. You will contribute to Paytm's AI inference platform (Pi), serving internal teams and enterprise customers - running our own coding and domain-specific models (voice, vision, risk, fintech workflows) as well as third-party models. You will also architect and build the platform that enables autonomous AI agents to operate safely and reliably in production - the runtime, orchestration, and developer tooling for agents to reason, plan, use tools, and execute complex multi-step workflows, automating both software development and business processes. You will work at the intersection of LLMs, distributed systems, and production fintech infrastructure, helping define how inference and agentic AI are built and deployed across payments, risk, fraud, collections, support, and developer experience.

E
📍 Ontario, Canada· Full-time· Remote
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE: Are you passionate about driving innovation in the data storage industry? Are you skilled in crafting technical solutions that exceed customer expectations? If so, we have an exciting opportunity for you! Everpure (formerly Pure Storage), a leader in the data storage and flash technology space, is seeking a talented and motivated Senior Pre-Sales Systems Engineer to join our dynamic team. As a Senior Pre-Sales Systems Engineer, you will play a crucial role in understanding our customers' unique challenges and tailoring Everpure solutions to meet their specific needs. Collaborating closely with the sales team, you will act as a technical expert during the sales process, helping to showcase the value of our products and services. WHAT YOU’LL DO: Develop an exhaustive understanding of what drives a customer’s business and what motivates their decision making Connect the dots from technology solutions, inclusive of the Everpure portfolio and others from the ecosystem, to measurable customer business outcomes Partner closely with account managers, specialists and channel partners to create a seamless and holistic customer experience and strategy to drive revenue growth and net new business Delight customers and teammates with your technical leadership and domain expertise on storage products, distributed storage architectures, file systems, and competitive storage offerings in the DAS, NAS and SAN product spaces

SQLAWSAzureGCP
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

Scale GP is Scale's enterprise Generative AI platform—APIs and infrastructure for knowledge retrieval, inference, evaluation, and intelligent automation. We power mission-critical workflows for leading enterprises, helping teams turn complex data and models into reliable, production-ready AI systems. We're building a new AI Enablement team to create the next generation of agent-powered tools that ground AI in real operational workflows. Our goal: help internal teams demystify their own workflows, then deploy agentic systems that reason over data, take action, and deliver measurable outcomes. We don't build in a vacuum. You'll use our own platform to solve real business problems internally—then selectively commercialize that same stack for customers. What we run on is what we sell. This is a 0→1 team. We're looking for a sharp, product-minded engineer who thrives in ambiguity, moves fast, and loves building systems from scratch alongside customers and cross-functional partners. You'll work closely with product, forward-deployed engineers, data scientists, and applied AI teams to turn real-world problems into scalable production solutions. If you like shipping fast, owning outcomes, and working across the stack—from polished frontends to distributed backends to LLM integrations—this role is for you. What You’ll Do Own full-stack features and projects end-to-end — from design through production deployment — within a larger product area Sample surfaces - Accounting Agents, Finance Copilots, GTM Agents, Agentic Experimentation Platforms Develop reliable backend services in Typescript/Python, work with distributed systems, data pipelines, and AI/ML infrastructure Integrate LLMs, vector databases, and agentic frameworks to power intelligent workflows Ship quickly through tight experimentation loops while maintaining high quality and reliability Adapt across the stack and learn new tools as needed to solve real problems end-to-end Ideal Experience 3+ years of full-tim

TypeScriptPythonAWSRest
T-
📍 Toronto, Canada· Full-time
✓ High-confidence listing

From C$1.4M/yr

Quick readStrong listing-quality and freshness signals

About the Role: Tubi's content platform is the engine behind one of the largest free streaming services in the world. Every play, every deal, every creator, every frame of video flows through systems CPE owns, and the surface area is enormous. Distributed services running on the hottest path of Tubi's traffic. Video pipelines processing one of the largest workloads in streaming. Workflow engines automating the operations that used to consume entire teams. Creator-facing products turning a back-office process into a real platform. And on top of all of it, an AI-native rebuild of the CMS that most companies aren't willing to attempt. This isn't a single-domain role. It's a platform where backend, frontend, video, infrastructure, and applied AI all collide at the scale where decisions actually matter, where an architectural choice ripples across millions of titles and billions of requests, and where the difference between "good enough" and "great" shows up in revenue. We're looking for builders who want to range across domains — backend one quarter, frontend the next, applied AI the one after that — and who want their work to be felt: by viewers when a title plays instantly, by creators when they go live the same day, by Content Ops when a workflow runs itself, and by the business when the platform stops being a cost center and starts being a force multiplier. The infrastructure is already there. The mandate is already there. What's missing is the people who want to build the thing, not talk about it. Come build it. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: You'll work on systems that sit at the heart of Tubi's business, where the content pipeline meets the viewer, the creator, and increasingly, the AI agent. The work spans the full stack of a modern content platform: distributed services, video infrastructure, workflow automation, and applied AI, all running at

TypeScriptPythonReactKubernetes
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform providing APIs for knowledge retrieval, inference, evaluation, and more. We are seeking a strong Senior Full-Stack Engineer to help us build, scale, and refine our rapidly growing product. The ideal candidate is deeply grounded in software engineering best practices and experienced in developing and scaling modern web applications end-to-end. You will work across the stack—from React/TypeScript frontends to Python-based backends—while integrating with LLMs and machine learning systems. You will solve complex challenges in scalability, reliability, and product experience while owning significant product areas in a fast-paced environment. What You’ll Do Own major full-stack product areas , driving features from design through production deployment. Build modern frontend experiences using React and TypeScript, ensuring performance, usability, and responsiveness. Develop reliable backend services in Python, working with distributed systems, data pipelines, and ML/LLM components. Integrate with LLMs, vector databases, and AI infrastructure to power intelligent product experiences. Deliver experiments and new features quickly , maintaining high quality and tight feedback loops with customers. Collaborate across product, ML, and infrastructure teams to shape the direction of Scale GP. Adapt quickly —learning new technologies, frameworks, and tools as needed across the stack. Ideal Experience 5+ years of full-time engineering experience , post-graduation. Strong experience developing full-stack applications using React, TypeScript, and Python . Experience scaling or shipping products at high-growth startups . Familiarity with LLMs, vector databases, embeddings, or other modern AI tooling (tinkering or production experience welcome). Proficiency with SQL and modern API development. Experience with Kubernetes , containerization, and microservice architectures. Experience working with at leas

TypeScriptPythonReactSQL
🔔

Get new distributed systems engineer jobs in Canada by email

Daily job updates · Unsubscribe anytime