About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus
Jobs in Canada
Aws And Tooling Platform Lead in Canada
545 active opportunities · Updated October 2026
Showing
15 jobs
Explore current aws and tooling platform lead jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the technical direction for our in-house Spark deployment and shape the architecture that will run DoorDash's data, analytics, and ML compute for the next five years and beyond. You will own the deep, cross-cutting problems that span the runtime, the shuffle service, the scheduler, and the overall service reliability — making the architectural calls that compound across the platform's lifetime. You will partner with the Engineering Manager on technical roadmap, hiring, and team shape, and act as the senior technical voice in cross-team partnerships with Data Engineering, ML Platform, and product engineering teams that depend on the platform. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Set the multi-year technical direction for an in-house Spark-on-Kubernetes platform — runtime, shuffle, scheduler, reliability — and make the architectural calls that compound for years. Own the deepest distributed-systems problems on the team: shuffle architecture, multi-tenant scheduling, runtime performance, and the failure modes that only show up at scale. Partner with the Engineering Manager on technical roadmap, hiring, inte
From C$108K/yr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. We are building and maintaining a highly scalable asynchronous platform that empowers our organization to handle critical business cases. As a software engineering team, our mission is to create robust and innovative solutions that drive the success of our business and deliver unparalleled value to our customers. We adopt Infrastructure as Code practice to automate the provisioning and configuration of our resources, which helps reduce manual configuration and improve consistency. Our team culture is built on collaboration, open communication, and a supportive environment where each member's ideas are valued and contributions are recognized. We believe in the importance of fostering a positive workplace culture that inspires innovation and creativity. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and production readiness reviews, platform management and capacity planning ceremonies with cross-functional teams Document Infrastructure operations process and insights, identify repeatable actions and ruthlessly automate repetitive tasks Participate in our teams on-call rotations, respond to incidents and support other teams mitigate customer impacting events Experience: 5+ years experience working on teams responsible for software development, automation and systems engineering Experience building large-scale infrastructure, distributed systems or networks. Knowledge with SQS,
From C$160K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know the team The Developer Platform team at Auth0 (an Okta company) owns the surfaces developers use to build with Auth0 - APIs, SDKs, CLIs, and the tooling that shapes how developers experience identity. Software is increasingly built by developers working alongside AI agents, and that shift raises the bar on every interface we ship: APIs have to be consumable by agents, tools have to be discoverable, and our platform has to hold up when an agent - not just a person, is the one calling it. We're looking for an engineer who already builds with those skills. We move fast, we own problems end-to-end, and we care deeply about the developer experience we put in front of people. The opportunity We're hiring a Staff Software Engineer to build the platform surfaces that developers and AI agents use to configure, extend, and interact with Auth0. You'll work at the intersection of API platform design, developer tooling, and emerging standards like MCP (Model Context Protocol). You'll take on ambiguous, high-leverage problems alongside a group of strong senior and staff engineers, and you'll have real influence over how our platform is built for a world where agents are first-class consumers. What you'll be doing Design and build developer-facing platform surfaces - APIs, MCP tools, CLI integrations, that hold up whether the caller is a developer or an agent Own technical direction for key platform surfaces that make Auth0 consumable by agents Drive cross-cut
From C$136K/yr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As a Infrastructure Engineer at Lyft, you will run our Production Infrastructure by monitoring system availability and take a holistic view of our platform health. You will build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you will seek opportunities to optimize our systems in order to push our platform forward, anticipating our customers' needs in order to continually improve the platform. You will provide Lyft partner teams with operational support to help them build robust large scale distributed systems. About the Team Data Pipelines is at the heart of all critical data flowing through Lyft supporting hundreds of services that impact millions of drivers and passengers every day. Our team’s mission is to empower Lyft engineers to self-serve in building and maintaining data pipelines as needed to support products that deliver the world’s best transportation experience. We leverage a variety of technologies to store, stream and manage data making it available to our internal customers. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and
From C$108K/yr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As an Observability team member, you are responsible for the operation and maintenance of our logging and metrics infrastructure. You ensure all teams at Lyft are aware of the operational health of their products by monitoring system availability and take a holistic view of our platform performance. You build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you find opportunities to improve our systems in order to push our platform forward. You provide our partners with the support they need to help them build robust large scale distributed systems. We count on the reliability of our infrastructure to empower Lyft teams to provide our customers rich experiences that are highly available with rock solid performance to ensure our transportation platform continues to connect people and places. As we grow our team, we are seeking experienced Infrastructure Engineer to ensure that as our Infrastructure continues to scale, our platform continues to provide an essential and dependable service that transports millions of people every day. Specifically we are searching for someone who brings fresh perspectives, enjoys collaborating with cross-functional teams in order to continually improve our products and services for our customers. Responsibilities: Maintain, improve, and develop tooling and systems that enhance the reliability, scalability, and efficiency of our platform. Assist engineering teams in defining service-level objectives (SLOs) and provide the necessary toolin
About the Team The Storage organization builds and operates the online stateful systems and abstractions that DoorDash Engineering depends on: reliable, efficient, secure, and easy to use. Within Storage, the Distributed Caching team owns every caching offering at DoorDash end to end, including ElastiCache (Redis/Valkey), Boulder (our KVRocks-based key-value store for high-QPS feature serving), Entity Cache (read Bill Shen’s engineering blog post, “ High-Performance Proxy Cache for DoorDash Services ”), and the Distributed Lock Service, plus the smart clients (asgard-redis, valkey-go) that sit in front of them. These systems back critical product surfaces across DoorDash, Wolt, and Deliveroo: the team runs roughly 400 ElastiCache clusters serving hundreds of millions of GET requests per second in aggregate, and Boulder, our offline-to-online feature store, serves billions of feature lookups per second at peak. About the Role The team owns provisioning of clusters and the smart clients that sit in front of them, baking in sensible defaults so that other engineering teams get a turnkey caching solution instead of having to run their own. You'll help drive Boulder's evolution to scale further, improve cost efficiency, enhance performance, and support real-time updates; re-platform the Distributed Lock Service onto a strongly consistent backend; and build the self-serve tooling and recommendation engine that let customers describe a workload (QPS, TTL, payload size, latency profile) and get the right backend without talking to a human. You'll go deep on cache invalidation, replication, sharding, compaction, and failover, while shipping the guardrails, automation, and observability that keep this scale operable by a small team. You must be located in San Francisco, Seattle, or the New York Metro Area for this hybrid position. You will report to the Engineering Manager on the Distributed Caching team within the Storage organization. You’re excited about this opportunity b
$240K – $270K/yr
About the Role At Sigma, we’re not just adding AI—we’re building the future of how people work with data. Our platform already lets users explore billions of rows of data in seconds with a spreadsheet-like interface, analyze and present their data in workbooks, and build data apps and workflows. Now we’re pushing further, applying AI to reshape how people build in Sigma, discover insights, and make smarter decisions—fast. That’s where you come in. As an AI/ML Engineer, you’ll join a growing team focused on building the AI foundation that will power Sigma for the future. Your work will become an integral part of the workflow for the thousands of enterprises that run on Sigma. What You’ll Do Partner with product, design, and engineering teams to identify high-impact AI/ML opportunities Prototype and productionize AI systems that feel intuitive but do a lot under the hood—recommendations, natural language interfaces, agentic workflows, and more Develop and scale AI/ML infrastructure that powers both internal tooling and customer-facing features Tackle novel UX problems at the intersection of AI, BI, and apps What You Bring Bachelor’s degree in Computer Science, Engineering, Mathematics, or a related field (required) 10+ years of experience building and deploying production-grade AI/ML systems Deep knowledge of machine learning, deep learning, and applied AI Experience across the full ML lifecycle: data curation, training, deployment, monitoring A track record of building things that ship—whether it’s recommendations, search, machine translation, or something equally complex Experience adapting or training foundation models (language or multimodal) for novel domains Bonus Points (or skills you’ll build here) You've built agents that can plan, reason, and use tools You know your way around cloud infrastructure (AWS, GCP, Azure) You’ve worked in a fast-moving startup or high-growth environment Additional Job details The base salary range for this posit
$240K – $270K/yr
About the Role At Sigma, we’re not just adding AI—we’re building the future of how people work with data. Our platform already lets users explore billions of rows of data in seconds with a spreadsheet-like interface, analyze and present their data in workbooks, and build data apps and workflows. Now we’re pushing further, applying AI to reshape how people build in Sigma, discover insights, and make smarter decisions—fast. That’s where you come in. As an AI/ML Engineer, you’ll join a growing team focused on building the AI foundation that will power Sigma for the future. Your work will become an integral part of the workflow for the thousands of enterprises that run on Sigma. What You’ll Do Partner with product, design, and engineering teams to identify high-impact AI/ML opportunities Prototype and productionize AI systems that feel intuitive but do a lot under the hood—recommendations, natural language interfaces, agentic workflows, and more Develop and scale AI/ML infrastructure that powers both internal tooling and customer-facing features Tackle novel UX problems at the intersection of AI, BI, and apps What You Bring Bachelor’s degree in Computer Science, Engineering, Mathematics, or a related field (required) 10+ years of experience building and deploying production-grade AI/ML systems Deep knowledge of machine learning, deep learning, and applied AI Experience across the full ML lifecycle: data curation, training, deployment, monitoring A track record of building things that ship—whether it’s recommendations, search, machine translation, or something equally complex Experience adapting or training foundation models (language or multimodal) for novel domains Bonus Points (or skills you’ll build here) You've built agents that can plan, reason, and use tools You know your way around cloud infrastructure (AWS, GCP, Azure) You’ve worked in a fast-moving startup or high-growth environment Additional Job details The base salary range for this posit
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. Robinhood's Enterprise Security team is at the forefront of AI security, designing safety guardrails for cutting-edge AI and driving modern enterprise-wide security tooling standards. As a Senior Corporate Security Engineer on our Corporate Systems team, you will operate at the center of our enterprise security infrastructure, maintaining the control plane for user-facing security tooling and building innovative solutions that keep our employees and platform secure! In this role, you will combine security engineering, systems administration, and hands-on AI agent development to drive key automation and safety initiatives. Joining a lean, high-impact team, you will have high visibility and direct empowerment to take strategic initiative, solve complex technical challenges, and shape Robinhood's AI security posture. This role is based in our Menlo Park or Bellevue office(s), with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams.
About Us Twitch is the world’s biggest live streaming service, with global communities built around gaming, entertainment, music, sports, cooking, and more. It is where thousands of communities come together for whatever, every day. We’re about community, inside and out. You’ll find coworkers who are eager to team up, collaborate, and smash (or elegantly solve) problems together. We’re on a quest to empower live communities, so if this sounds good to you, see what we’re up to on LinkedIn and X , and discover the projects we’re solving on our Blog . Be sure to explore our Interviewing Guide to learn how to ace our interview process. About the Team Twitch's Enterprise Platform & Technology (EPT) organization is looking for a Software Development Engineer II to build the systems where SAP meets AWS — extending our enterprise ERP with cloud-native services and agentic AI to power Finance and enterprise functions across Twitch and Amazon. This is a hands-on engineering role where you'll work on SAP-side development (ABAP, CDS, RAP, OData) and AWS-native engineering (Lambda, Step Functions, MCP servers, Bedrock, CDK) — owning integrations end-to-end. About the Role You'll work at the intersection of SAP enterprise systems and modern cloud engineering — building AWS-hosted APIs and pipelines that talk to SAP, standing up agentic AI tooling (MCP servers, Bedrock agents, Claude-powered workflows) that automates Finance operations, and shipping production AWS services alongside SAP customizations. The business problems are well understood; how you solve them with SAP + AWS + AI is up to you. We need someone equally comfortable writing ABAP inside SAP GUI and Python/TypeScript infrastructure code against AWS — a full-stack enterprise engineer who ships across both worlds and raises the bar on automation. You Will SAP Engineering (~50%) Design, build, and maintain integrations between SAP (S/4HANA, BTP) and downstream systems — own
From C$102K/yr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. We're looking for a Workato/Boomi Integration Engineer to build and support enterprise automation, with a focus on Legal systems integration (CLM, e-signature, matter management, compliance). You'll work across iPaaS, APIs, and AI/agentic tooling (Workato Genie, AI Skills, MCP servers) to connect Legal and other business systems. Responsibilities: Design and build Recipes and process automations across multiple systems, including AI-powered workflows using Workato Genie and MCP servers. Build and support AI Agents (Genies) for Legal and business-process automation, including prompt design and IDP for unstructured Legal content (contracts, filings). Design and manage secure API endpoints (REST/SOAP) and SFTP connections; build ETL/data-sync pipelines across formats (JSON, XML, CSV, EDI). Deploy and manage On-Premises Agents (OPA) for secure, behind-firewall connectivity. Maintain platform environments and security (RBAC, SSO, versioning); monitor jobs and troubleshoot failures. Build employee-facing tools (Workbot/ChatOps, Workflow/AgentX Apps, Slack integrations) with human-in-the-loop steps. Integrate business applications (CRM, HR, Finance, LMS, Legal, Budget & Forecast) via Boomi, with emphasis on Legal, Financial, and Supply Chain integrations. Write and optimize SQL/PL-SQL and custom scripts (Java/Groovy/JavaScript/Python); manage code via Git/GitHub; provide production support and on-call as needed. Partner with Legal, Compliance, and IT stakeholders to gather requirements and translate them into working integrations. Experience: 6–8 years of integration/automation experience, with hands-on Workato and Boomi. 4–6 years building and managing APIs (API Gateway, OAuth/SAML/JWT) and SFTP-based file transfers. 2+ years supporting Legal systems integrations (CLM, e-signature, matter management, c
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team of bold thinkers and sharp problem-solvers who are wired to make an impact. The Ops Platform organization develops internal platforms that replace repetitive manual processes with AI-driven systems. These tools support key areas such as Fraud Operations, Account Operations, Financial Crimes Operations, and Retirement Services. The team works closely with product, data science, and operations partners to deliver reliable systems that improve decision-making and efficiency! As a Software Developer, you will design and build platforms that enable operational teams to investigate and resolve issues more quickly and accurately. You will work with large datasets and signals to create tooling that supports fraud investigation and other operational workflows. You will collaborate with data scientists and machine learning engineers to translate manual processes into automated systems. Your work will focus on improving system reliability, reducing operational effort, and increasing the speed at which new products and features can be supported across Robinhood’s offerings. This role is based in our Toronto, ON office, with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. What you’ll do You will define technical direction and make architectural decisions for systems that support operational workflows across multiple product lines You will build tools that proces
We take play seriously. We’re looking for curious adventurers ready to find their party, fueled by imagination and drive to build what’s never been built before. At Hasbro and Wizards of the Coast, you’ll collaborate with passionate teams to reimagine our iconic brands and create experiences that spark joy, connection, and community through the magic of play. This is your chance to shape legendary play that lasts a lifetime. We’re building something new inside our AI Studio, and we’re looking for a Full Stack Engineer to help build it with us. This role sits at the center of a new platform that is being actively designed and shipped. The architecture is still taking shape, and the integration surface is constantly expanding. What excites us is the opportunity to build the infrastructure that makes never-before-imagined character experiences possible. We work where storytelling, imagination, and technology converge. Our focus is on building the core platform: the APIs, the admin tooling, the service integrations, and the delivery infrastructure that makes AI-powered character interactions real. This work connects what brand and creative teams envision with what audiences ultimately experience. This role is foundational to everything the platform does. Working closely with the Engineering Manager and Principal AI Engineer, this Full Stack Engineer will own significant portions of the system end to end — from the tooling that configures and versions character behavior to the API layer that external partners and products consume. Success here means shipping features that hold up under real-world conditions and integrating reliably with the AI services that bring characters to life. What You’ll Do Design and build features for our AI character platform. This includes: the admin tooling, as well as the API layer consumed by external integrations and par
From C$110K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Team The Auth0 Platform Tools team owns the incident management tooling, Slack-based tooling, StatusPage, and local development environments that Auth0 engineers rely on every day. That includes incident.io and the services we have built around it, Statuspage, custom Slack bot applications that automate our incident response and engineering operations workflows, the customer-facing web application behind status.auth0.com, Vivaldi, and Tilt - the tools engineers use to run Auth0 locally. We are seeking an engineer to help build new features across all of these tools. Our stack is primarily TypeScript and Node.js, with a React and Next.js front end, backed by Postgres and Redis, and deployed on Kubernetes on AWS. A significant portion of our incident and engineering operations automation is built on Tines, a no-code automation platform. Prior no-code experience is welcome, but we expect you to learn Tines here and become effective with it. Current initiatives include extending our incident tooling to meet FedRAMP requirements, taking full ownership of the status page, and improving how we communicate incident status to customers. There is real room to improve along the way, from test coverage to resilience to inherited technical debt. We build for two audiences: Auth0 engineers, who depend on our tooling every day, and Auth0's customers, who rely on the status page during incidents. We are looking for an engineer who cares about both and enjoys working wi
Other cities to consider
More places hiring for this role
Get new aws and tooling platform lead jobs in Canada by email
Daily job updates · Unsubscribe anytime