About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, leading the design and architecture of our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly. You’re excited about this opportunity because you will… Lead the design of infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Own and evolve our open-weights serving stack — real-time GPU endpoints, high-thr
Jobs in Canada
Lead Software Engineer Machine Learning Infrastructure Consultant in Canada
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current lead software engineer machine learning infrastructure consultant jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.
Hiring demand
76/100
rising · 185 related jobs
Hiring trend
+28.4%
Job postings compared with the previous 30 days
Remote options
2.7%
Share of matching jobs listed as remote
Typical salary
$232.5K – $232.5K/yr
Based on 82 salary observations
From $216K/yr
The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Senior Software Engineer, you will lead the development of a vertical feature or a horizontal capability to include defining requirements with stakeholders and implementation until it is accepted by the stakeholders. You will: Lead the design and implementation of scalable backend systems and distributed architectures for Federal customers. Manage the full lifecycle of feature development from requirement definition to deployment on classified networks. Direct the orchestration of asynchronous agent fleets to meet mission requirements. Lead customer engagements to translate mission needs into technical requirements. Own the communication with stakeholders to ensure implementation meets defined acceptance criteria. Conduct technical reviews and identify risks within machine learning infrastructure and model serving. Drive the platform roadmap by providing technical specifications for Federal product offerings. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in developing and deploying applications in a cloud-native environment. Understanding of containerization (e.g., Docker) and contai
$140K – $225K/yr
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As the Senior Software Engineer, Tooling and Development Infrastructure, you will play a critical role in shaping the developer productivity tools and automated testing strategy. You’ll collaborate closely with design, development, and quality teams to plan, design, and implement robust automated tools and services that ensure the quality and reliability of our AI software stack. You will be highly hands-on in your work and collaborate closely with stakeholders. This position offers a unique opportunity to influence the development of cutting-edge automation frameworks, foster a culture of quality, and contribute to the long-term success of the organization. What You Might Do Develop and implement automation frameworks and testing strategies that cover the entire software stack, from backend systems to user-facing features. Identify, evaluate, and integrate new tools that streamline development. This includes everything from code quality tools and to Infrastructure-as-Code (IaC) solutions. Lead continuous improvement efforts for our build, release, and test systems, ensuring a robust
From C$160K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Streaming Foundations team builds services and operates data pipeline infrastructure to support event streaming, messaging, and analytics use cases. We are looking for a Software Engineer who is passionate about distributed systems, platform engineering, and solving data-intensive problems at scale. In this high-impact role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Auth0 to scale for years to come. What you’ll be doing Help set the technical direction for the team and influence the engineering roadmap for the Platform’s streaming capabilities Design and lead the implementation of our most complex and critical systems for data-intensive use cases. Research and champion new technologies and architectural patterns to solve strategic challenges and scale the platform. Lead and influence cross-functional initiatives, ensuring technical alignment and successful execution across multiple teams. Improve the operational posture of our systems by designing for observability, reliability, and scalability, and by mentoring others in operational best practices. Coach and mentor senior engineers and act as a technical leader across the engineering organization. Collaborate with different stakeholders like product teams whenever needed. What you’ll bring to our teams 7+ years of software development experience in a fast-paced, agile environment Experience working with Golang or Java is preferred H
$184K – $1.2M/yr · Jobiba est.
About the Role: As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams to build world-class machine learning inference platforms. These platforms power essential services like personalized recommendations, search, and content understanding across Tubi. A core responsibility of this team is developing and maintaining low-latency ML model serving systems that support Deep Learning, LLM, and Search models. This involves building self-service infrastructure and critical components such as the inference engine, feature store, vector store, and experimentation engine. You will improve the way we deploy and operate our services and even contribute to open-source projects. This role grants the architectural freedom to explore new frameworks, lead critical cross-functional projects, and transform the capabilities of our ML and Product teams. Responsibilities: Design and build scalable, high throughput, and low latency distributed systems using Scala Build reusable components and services that serve various ML applications like Personalization, Search, Ads and Exploration Partner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art. Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessary Mentor other engineers on the team on system design, effective incident management, interviewing, leveraging LLMs for work, etc. Collaborate with ML, Product, and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi. Your Background: Experience designing and building scalable, distributed systems in any modern backend language (e.g., Scala, Java, Python, Go, C++); experience with Scala or JVM b
From C$136K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. At Okta, we’re building the next generation of authentication for the Agentic AI era. We’re looking for a Senior Engineer to join our AI Authentication team – a group of product-minded, deeply technical engineers delivering features to define what identity and access mean in a world increasingly powered by artificial intelligence. This is a great opportunity for an Senior engineer who thrives in collaborative environments, is eager to grow their technical depth, and wants to build scalable systems that solve real-world security problems. Auth0 Emerging Tech is the Engineering organization where we take care of the hottest technology out there: we ship fast, we don't break things. We are a dynamic and collaborative distributed and diverse team. We value ownership, learning and innovation. We launched our Auth for GenAI offering and we are looking for an Senior Engineer to join us to lead the team taking care of the Identity Protocols parts of the game. Auth0 works with NodeJS ( Javascript or Typescript), a hint of Go and MongoDB or PostgreSQL databases. What will you do: Bring expertise in identity and security while building innovative features and standards that will secure the Agentic AI world Build and maintain scalable services using TypeScript, NodeJS PostgreSQL/MongoDB. Collaborate with other product managers, designers, and senior engineers to deliver features that improve identity and access for GenA
From $269K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About Okta for AI Agents Okta secures access for 20,000 organizations and billions of users. Okta for AI Agents extends that work to the agentic shift. Deploying an AI agent is not like deploying traditional software. You are putting professional work output into production, and it needs deep integration, continuous tuning, and change management. Every agent needs an identity, a scope, an audit trail, and a way to be shut down when it goes wrong. Most enterprises have not built this yet. We are. We hire builders who see the cracks in enterprise agent identity that everyone else has learned to live with. The Role You are the most senior technical field authority for agent identity at Okta. Where a Senior FDE owns the outcome inside one account, you own the patterns that every account and every FDE inherits. You take the hardest and most strategic deployments yourself, set the reference architecture the team builds from, and turn what the field learns into the direction the product takes. You still write code. You also multiply the people around you, and you are the person product and engineering leadership call when an agent identity problem has no precedent. Responsibilities Own the reference architecture. Define the canonical agent identity, delegation, audit, and kill-switch patterns that Senior FDEs deploy across the portfolio, and keep them current as the standards and the product move. Lead the hardest accounts. Personally own the most strategic, regul
From $302.4K/yr
Director of Engineering, Physical AI Role Overview The Director of Engineering will report to the General Manager of Physical AI, and will be responsible for leading a multi-disciplinary engineering organization. In this senior leadership role, you will own the execution of the Physical AI Data Engine — the platform powering the next generation of Physical AI/Embodied AI. You will collaborate closely with Operations and GTM to guide product direction and help solve the data bottleneck that stands between today's robotics research and real-world deployment. This role requires significant ownership in a fast-paced environment and you will motivate internal teams to set the pace for business growth. Travel will come into play. Key Responsibilities: Set and drive the technical vision across data collection infrastructure, teleoperation systems, ML training pipelines, model evaluation frameworks, annotation tooling, and research Lead a multidisciplinary engineering organization—spanning engineering managers, software engineers, ML engineers, and ML research scientists—while designing the organizational structure, talent strategy, and culture required to scale rapidly without compromising on quality or strategic alignment Maintain exceptional technical and operational excellence by deeply understanding team deliverables, asking incisive questions, identifying slipping standards early, and knowing precisely when to step in Drive cross-functional alignment across Engineering, Operations, and GTM on platform architecture, release processes, and shared priorities Collaborate with researchers and clients to architect and deliver scalable, production-grade data infrastructure tailored for complex robotics workloads Required Qualifications: Bachelor's degree in Engineering, Robotics, Computer Science, or a related technical field 8+ years of engineering experience in fast-paced environments, including 4+ years direct people management demonstrated history of recruiting, mentorin
From $24K/yr
Amplitude is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Winter 2026 Report, Amplitude is the best-in-class solution for product, data, and marketing teams. Learn more at amplitude.com . As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do. Amplitude’s Commitment to Diversity Equity & Inclusion (DEI): Amplitude believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive. About the Role & Team We’re looking for an Engineering Manager to lead the Data Infrastructure team within Statsig Experiment at Amplitude. You will lead a multidisciplinary team of software engineers, data engineers, and data scientists responsible for the systems that power experimentation at scale. The team owns three critical areas: Data ingestion: Collecting and importing experiment exposures, custom events, OpenTelemetry data, and real user monitoring data across SDKs, streaming systems, cloud storage, and customer data warehouses. Data computation: Building distributed computation systems that transform raw data into accurate, timely experiment results. Stats engine: Developing and productionizing the statistical methods that help customers make trustworthy decisions from their experiments. This is not a traditional data engineering management ro
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 ClickUp is looking for an experienced Engineering Manager to lead our fullstack team responsible for building and scaling our flagship products. As the leader of the team that owns the APIs and core experiences powering ClickUp, you will play a pivotal role in shaping the future of our platform. You will guide engineers working across the stack, from frontend experiences to backend infrastructure. Your focus will be on driving the development of new features, addressing performance and reliability challenges, and ensuring operational excellence as we continue to grow. This is an opportunity to make a significant impact on our core product while fostering a culture of technical excellence and collaboration. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our produ
From C$146K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity Reporting to the Director of Quality & Performance, this role as Engineering Manager of Performance & Resilience will drive performance and resiliency improvements for the Auth0 product at Okta. Here you'll be working with some of the most advanced technology in the space, helping to streamline and secure billions of access requests a year. In this role, you will work closely with architects, platform team members, and product engineers to build the testing infrastructure that keeps Auth0 performant at scale — including the frameworks, tooling, and realistic datasets that make that testing meaningful. The ideal candidate is passionate about software quality and architecture, a self-starter, intellectually curious, and brings deep experience with performance testing, load testing frameworks, dataset generation, performance analysis, monitoring tooling, and chaos engineering. What you’ll be doing Collaborate with architects, tech lead, product owners, security and operations engineers to implement best practices related to performance and resiliency Communicate and organize cross-team projects with high business
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. As a Senior Software Engineer, Data on the Mapping team, you will collaborate with our world-class team of engineers, product managers, and scientists to grow and improve the quality of recommended routes and accuracy of our travel time estimations. You will lead the architecture and long-term technical direction of our offline experimentation tooling and route simulation services — the systems that let Lyft test routing changes safely before they reach production. You'll also build scalable data pipelines for experimentation, analytics, and machine learning models, along with the data governance and observability systems that keep them trustworthy. Your work will enable integration with partner teams and allow stakeholders across Engineering, Data Science, and Product to make data-informed decisions that directly impact Lyft’s growth and profitability. Our technology stack is based on the latest technologies such as AWS, Databricks, Kubernetes and Airflow. You will work with incredibly passionate and talented colleagues from software engineering, machine learning and data science on projects that directly impact millions of riders and drivers. Responsibilities Own core data pipelines end-to-end, building deep subject matter expertise in the systems you manage and defining/managing SLAs for pipelines, services, and datasets to ensure reliability at scale Serve as the technical owner and architectural lead for our offline experimentation platform and route simulation services, setting technical direction, evaluating trade-offs, and ensuring the systems scale with Lyft's routing and mapping ambitions Continuously evolve data models and schemas to meet business and engineering requirements Develop AI tools that support self-service management of data pipelines (ETL) and schema evolution, and perform han
From C$1.2M/yr
About the Role: Tubi is seeking a highly skilled and experienced Senior QA Automation Engineer to lead quality assurance initiatives for our cutting-edge streaming and AI-driven product features. This pivotal role involves ensuring exceptional end-to-end user experiences, robust streaming playback, and the accuracy and integrity of our AI/ML features across web, mobile, and OTT platforms. We're looking for a candidate with a strong background in streaming QA and deep technical knowledge of media workflows. You'll be instrumental in collaborating with engineering, product, and data science teams to define comprehensive QA strategies that guarantee both functional excellence and data-level quality. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office three days/week. What You'll Do: Design and lead test strategies for streaming workflows, playback systems, and AI-powered features. Test across platforms (web, mobile, and connected TV) to ensure functional parity and playback stability. Validate streaming performance—including ABR logic, encoding pipelines, and DRM integrations—under diverse real-world conditions. Debug with precision using tools like Charles Proxy, Chrome DevTools, ADB, and Xcode. Collaborate with data and ML teams to validate AI model updates, recommendations, and personalization accuracy. Leverage AI-assisted QA tools to enhance regression coverage, UI validation, and anomaly detection. Contribute to automation and CI/CD frameworks, driving faster, more reliable releases. Help drive a shift-left testing approach by engaging early in the software development lifecycle, partnering with product managers, engineers, and data scientists to identify quality risks, define test strategies, and ensure testability during requirements and design phases. Oversee QA deliverables for multiple concurrent releases and ensure seamless sign-off for production launches. Monitor live environments for playback or reco
From $184.9K/yr
About Us Twitch is the world’s biggest live streaming service, with global communities built around gaming, entertainment, music, sports, cooking, and more. It is where thousands of communities come together for whatever, every day. We’re about community, inside and out. You’ll find coworkers who are eager to team up, collaborate, and smash (or elegantly solve) problems together. We’re on a quest to empower live communities, so if this sounds good to you, see what we’re up to on LinkedIn and X , and discover the projects we’re solving on our Blog . Be sure to explore our Interviewing Guide to learn how to ace our interview process. About the Role Creators are the backbone of Twitch, and they rely on their ability to earn a living doing what they love. As an Engineering Leader focused on growing creators' revenue, you will lead people to build products, features and experiences that enable such a living. You will be empowered to work across the stack and across functions. You will take high ownership of your services - architecting, building and operating them. You will partner with other engineering leaders, product managers, designers, data specialists and program managers to deliver solutions. Our teams own features like Subscriptions, Gifting, Turbo, Bits and Hype Train. This role can be based in San Francisco, CA or Seattle, WA. You Will: Lead a team of engineers to design, build and operate full stack applications. Establish a clear vision for your team and generate urgency to deliver for our customers. Have a laser focus on delivering results - even if that means disrupting process or protocol. Articulate behaviors that can grow careers. Your people development radar is always on. Quickly dissect the nuances of trade-offs, making conscious choices and overcoming challenges. Constructively parse, ingest, challenge and reframe proposals / plans for the better. Adjust
From C$100K/yr
Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. Appian Customer Success is obsessed with delivering exceptional customer outcomes and driving mission-critical business impact. Grounded in our core values of Excellence and Intensity, we act as elite technical advisors to our commercial clients. By joining this high-performance team, you will champion a culture of candid communication and excellence while accelerating global adoption of our AI-Powered Process Automation platform. As a Principal Consultant, you will occupy a leadership position at the intersection of enterprise business strategy and complex technical delivery. This role matters now more than ever because you won't just execute projects - you will define the client experience and how they are delivered. You will partner directly with Appian Architects and Technical Delivery Managers, lead large-scale consulting initiatives, and serve as a vital mentor to the next generation of engineers, ensuring our clients achieve world-class Enterprise-Grade Orchestration. What You’ll Do Lead Enterprise Delivery: Command the entire project lifecycle to define, design, and implement custom automation solutions using the Appian platform for major commercial clients. Direct High-Performing Teams: Lead and mentor consultants through fast-paced software implementations, instilling a dedication to going "beyond completion." Partner with Leadership: Collaborate directly with Appian Architects and Technical Delivery Managers to design resilient, scalable, and secure system architectures. Architect Complex Integrations: Build secure, high-throughput APIs
Higher-paying openings
Jobs with higher listed pay
Staff Software Engineer, Lyft Business
Lyft · San Francisco, CA
$2.1M – $2.6M/yr
Senior ML Software Engineer, Mapping
Lyft · San Francisco, CA
$2M – $2.4M/yr
Staff Software Engineer
Amplitude · San Francisco, Canada
From $2.4M/yr
Staff Software Engineer - DevX
Amplitude · San Francisco, Canada
From $2.4M/yr
Senior Staff Software Engineer - Pricing and Packaging
Gusto, Inc. · San Francisco, CA
From $2.3M/yr
ML Software Engineer, ETA
Lyft · San Francisco, CA
$1.7M – $2.1M/yr
Related career options
Similar roles with stronger pay
Demand 67/100 · 15 jobs
$1.8M – $1.8M/yr
Salary →Demand 50/100 · 8 jobs
$1.3M – $1.3M/yr
Salary →Demand 67/100 · 38 jobs
$315K – $315K/yr
Salary →Demand 51/100 · 10 jobs
$274.5K – $274.5K/yr
Salary →Demand 59/100 · 45 jobs
$251K – $251K/yr
Salary →Demand 54/100 · 24 jobs
$246.3K – $246.3K/yr
Salary →Other cities to consider
More places hiring for this role
Get new lead software engineer machine learning infrastructure consultant jobs in Canada by email
Daily job updates · Unsubscribe anytime