Jobs in Canada

Inference Technical Lead in Canada

125 active opportunities · Updated October 2026

Explore current inference technical lead jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.

MT
📍 San Jose, California, Canada
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Job Summary We are seeking a motivated engineer to join the DRAM Systems Engineering team, focusing on the development, evaluation, and optimization of next-generation memory systems for AI accelerators. This role emphasizes research and development across hardware architecture, operating systems, and performance analysis to support Agentic AI inference workloads. If you are ambitious and eager to make an impact in the exciting world of AI and memory systems, this is the perfect opportunity for you! Responsibilities Characterize AI inference workloads and examine memory behavior Build and evaluate tiered memory hierarchies for AI accelerators Study KV cache lifecycles, MoE models, and data placement strategies Compare and optimize explicit versus hardware-assisted data movement Develop, test, debug, and detail system-level and OS components Prototype and evaluate agentic AI systems by building agents and multi-agent workflows using modern frameworks and orchestration patterns (planning, tool use, memory, and context management). Apply these technologies both as workloads under study and as accelerators for internal engineering workflows <h2 style="color:!importan

PythonLinuxAIRecruitment
T
📍 Toronto, Ontario, Canada
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As a Datacenter Liquid Cooling Architect, you will define, design, and architect next-generation liquid cooling infrastructure for Tenstorrent’s large-scale AI training and inference clusters. You will partner with systems engineering, mechanical engineering, software, and cross-functional design teams to develop chassis-, rack-, and cluster-scale cooling solutions, including CDU integration, telemetry and control, leak detection, and resilient operating strategies. This role will help shape reliable AI datacenter architectures and deployments for both internal and external customers. This role is on-site, based out of Toronto, Canada, Austin, Texas or Santa Clara, California. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A datacenter and system thermal design professional with 10+ years of experience architecting cooling infrastructure for complex computing environments. An experienced liquid cooling architect who can design chassis- and rack-scale solutions for large AI training and inference clusters. A systems thinker who understands how mechanical, electrical, software, facility, and systems engineering decisions come toge

DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $102K/yr

Quick readStrong listing-quality and freshness signals

About the Team The DoorDash Research Fellowship is a 3-month program (extendable to 6 months) looking for Summer and Fall 2026 cohorts, for researchers and engineers who want to work on the hardest applied ML and AI problems in local commerce. Fellows are given the resources, autonomy, and access to real-world operational data needed to pursue ambitious research directions — with the goal of producing work that influences both the field and how DoorDash operates at scale. This program is modeled on the best external research fellowships: fellows are treated as independent researchers, not as junior employees on a product team. You pick the problem (within a set of priority areas), you own the direction, and you publish or ship the outcome. You’re excited about this opportunity because you will receive… Dedicated compute allocation sized to the research agenda — GPU clusters for training and inference budgets for experimentation Full access to DoorDash's research infrastructure — our internal RL stack, training and evaluation pipelines, RL environments built on real operational systems, agent evaluation harnesses, and the tooling our own research teams use day-to-day. Fellows are first-class users, not sandboxed visitors. Access to DoorDash operational data — real-world datasets spanning logistics, merchant operations, consumer behavior, and marketplace dynamics, under appropriate data governance Research mentorship from senior researchers and engineering leaders at DoorDash, plus a named research sponsor for each fellow who meets with you weekly and is accountable for unblocking your work Speaker series featuring leading researchers and practitioners from academia and industry — faculty from top ML programs, research leads from frontier AI labs, and senior operators from across tech. Fellows get dedicated 1:1 time with speakers when possible. A cohort of fellows working alongside you — a small, tight-knit group of researchers tackling different problems but sharing

GitRestAIGo
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$45/hr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft’s Data Science Team builds mathematical models underpinning the platform’s core services. Compared to other technology companies of a similar size, the set of problems that we tackle is incredibly diverse. They cut across optimization, prediction, modeling, inference, transportation, and mapping. We're looking for Masters or PhD students who are passionate about solving mathematical problems with data and are excited about working in a fast-paced, innovative and collegial environment. We are hiring for a variety of Data Science interns, focusing on the following specialties: Optimization: Construct and fit statistical or optimization models that facilitate automated decision making in the app. Machine Learning: Design, build, tune, and deploy machine learning models with a special emphasis on feature engineering and deployment. Inference: Design and analyze tests in our dynamic marketplace, estimating statistical and ML models to enable better decisions, and developing and evaluating algorithmic policies in our pricing, dispatch, and incentives systems. You will report into a Science Manager. Responsibilities: Partner with Engineers, Product Managers, and other cross-functional partners to frame problems, both mathematically and within the business context Perform exploratory data analysis to gain a deeper understanding of the problem Write production modeling code; collaborate with software engineers to implement algorithms in production Design and run both simulated and live traffic experiments Analyze experimental and observational data; communicate findings including working with partner teams and presentations; facilitate launch decisions Experience: Currently pursuing a Masters or PhD degree at a university in Canada (required) in mathematical sciences ( Opera

PythonSQLMachine LearningAI
L
📍 San Francisco, CA· Full-time
✓ High-confidence listingCompany trend -72.4%
Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. As an Applied Scientist specializing in Machine Learning and Operations Research on this team, you will develop mathematical models and launch algorithms that power these key pricing and ETA decisions. You will leverage your skills to build ML and optimization models and productionalize pipelines that can scale to millions of calls per day while solving critical business problems that have a big impact on the marketplace and rider experience. You will get exposure to a diverse set of real-world problems across optimization, prediction, machine learning, and inference and collaborate closely with teammates and stakeholders across Pricing, from Product Managers to Engineers and Analysts. We are looking for someone who is excited about working in a fast-paced, innovative, and impactful environment, and is adept at balancing complexity and efficiency to translate real world business problems into reliable solutions, systems and decision frameworks. Responsibilities Partner with Data Scientists, Engineers, Product Managers, and Business Partners to frame problems mathematically and within the business context Write production quality code. Design, build and deploy production-grade ML and Optimization models. Able to build custom methods and tooling beyond off-the-shelf libraries. Perform data analysis and build proof-of-concepts to explore and propose ML and Optimization solutions to both new and existing problems. Evaluate machine learning systems against business goals. Collaborate with Engineers to implement algorithms in live systems and ensure the robustness of the systems Establish metrics and development measurement methodologies to monitor the health of our products, as well as the impacts on user and marketplace outcomes Drive collaboration and coordination with cross-functional teams

PythonMachine LearningAIGo
PE
📍 Toronto, Canada· Full-time· Hybrid
✓ Quality checked

About the role There’s a wide gap between an agent that works in a demo and one that works across millions of live transactions. Closing it is the job. You’ll embed with the teams and customers who depend on AI — risk, fraud, collections, payments, support, developer experience — and design, build, and ship agentic systems into their production environments. You’ll also help build the platform underneath: Paytm’s AI inference platform (Pi) and the agentic runtime, orchestration, and tooling that lets agents reason, plan, use tools, and run multi-step workflows safely.

HI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

C$45 – C$51/hr

Quick readStrong listing-quality and freshness signals

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role HP IQ's AI Machine Learning (AML) team is building the foundational platform powering a new generation of agentic devices. This platform orchestrates the complete lifecycle of AI models: from creation and fine-tuning through optimized inference and intelligent orchestration. The team works to make complex AI capabilities run efficiently on-device, enabling locally-deployed agentic experiences that reduce token costs and improve privacy. In this internship role, you will contribute to one or more core pillars of the AML platform: model inference optimization, orchestration and agent workflows, or model creation and fine-tuning. You will partner closely with experienced engineers to ship features that directly impact HP's next-generation devices and gain visibility into how each component of an end-to-end AI system integrates and scales. What You Might Do Work on model hosting and inference optimization, learning how to profile, benchmark, and accelerate model execution on resource-constrained devices; experiment with quantization, distillation, or other optimization techniques to reduc

PythonRedisMachine LearningAI
HI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$190K – $270K/yr

Quick readStrong listing-quality and freshness signals

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role The AI team is building cutting-edge solutions that bring the power of AI directly to edge devices while seamlessly integrating with cloud infrastructure. We are looking for a Lead Software Engineer to design and develop high-performance, scalable services to support AI workloads across edge and cloud environments. What You Might Do Design, build, and maintain services that power AI-driven applications, ensuring scalability and performance. Develop APIs and microservices that facilitate seamless integration between cloud-based AI models and edge devices. Optimize data pipelines and storage solutions for real-time AI inference and processing. Implement security and privacy best practices for distributed AI systems. Work closely with AI researchers, infrastructure engineers, and frontend developers to deliver end-to-end AI-driven solutions. Build and optimize an agent orchestration runtime that enables tool use, memory management, and multi-step reasoning across LLMs, APIs, and edge-connected systems. Develop robust logging, monitoring, and alerting systems to ensure system reliability.

PythonJavaSQLRedis
HI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$140K – $225K/yr

Quick readStrong listing-quality and freshness signals

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role HP IQ's Connectivity team is seeking an Embedded Firmware Engineer with strong hands-on experience across RTOS and embedded Linux platforms. You'll bring up new hardware, develop and support firmware across core device subsystems, and help scale products from prototype to fleet deployment. A key part of this role is enabling on-device intelligence — bringing AI models, sensing algorithms, and local processing to lightweight, power-constrained devices at the edge. What You Might Do Design, develop, and debug firmware across RTOS and embedded Linux platforms. Lead hardware bring-up — board bring-up, driver integration, and firmware support through to production. Develop and maintain firmware for subsystems such as connectivity, power and other sensors. Integrate lightweight model inference and sensing/proximity algorithms within tight compute, memory, and power budgets. Support fleet-scale deployment: OTA updates, field diagnostics, and post-launch sustainment. Collaborate with hardware, systems, and QA teams to troubleshoot system-level issues and drive root-cause fixes to closure. Ess

RedisLinuxAIC++
HI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$149K – $240K/yr

Quick readStrong listing-quality and freshness signals

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role The AI team is building cutting-edge solutions that bring the power of AI directly to edge devices while seamlessly integrating with cloud infrastructure. We are looking for a Senior Software Engineer to design and develop high-performance, scalable services to support AI workloads across edge and cloud environments. What You Might Do Design, build, and maintain services that power AI-driven applications, ensuring scalability and performance. Develop APIs and microservices that facilitate seamless integration between cloud-based AI models and edge devices. Optimize data pipelines and storage solutions for real-time AI inference and processing. Implement security and privacy best practices for distributed AI systems. Work closely with AI researchers, infrastructure engineers, and frontend developers to deliver end-to-end AI-driven solutions. Build and optimize an agent orchestration runtime that enables tool use, memory management, and multi-step reasoning across LLMs, APIs, and edge-connected systems. Develop robust logging, monitoring, and alerting systems to ensure system reliabilit

PythonJavaSQLRedis
L
📍 Toronto, Canada
✓ Quality checkedCompany trend -72.4%

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Data Science is at the heart of Lyft’s products and decision-making. Data Scientists at Lyft operate in dynamic environments, moving quickly to build the world’s best transportation solutions. We tackle a wide range of challenges - from shaping long-term business strategy with data, to making critical short-term decisions, to developing algorithms and models that power both internal systems and customer-facing products. Driver Incentives Science owns the algorithms and systems behind incentive design, influencing driver engagement and marketplace efficiency — from real-time supply positioning to longer-horizon earnings and engagement programs. The team is responsible for designing pay and incentive mechanisms that are efficient and good for driver experience over the long run. As a Data Scientist specializing in Algorithms, you'll partner closely with product, engineering, and operations leaders to build and scale incentive systems, shape long-term mechanism design strategy, and deliver on critical business goals tied to marketplace efficiency and driver earnings. Candidates with strong optimization backgrounds — think mathematical programming, control theory, or operations research — are a great fit, though we welcome strong candidates from machine learning or causal inference as well. The ideal candidate thrives in a fast-paced environment and brings a hands-on, entrepreneurial mindset to drive results. Responsibilities: Collaborate with engineering and product teams to design, implement, and iterate on new features and algorithmic improvements for driver incentives and pay mechanisms. Design, develop, and deploy optimization models, algorithms, and systems for problems such as budget allocation, multidimensional cost-curve development, and incentive targeting. Write production model code; collabor

PythonMachine LearningArtificial Intelligence
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$108K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Data Science is at the heart of Lyft’s products and decision-making. Data Scientists at Lyft operate in dynamic environments, moving quickly to build the world’s best transportation solutions. We tackle a wide range of challenges - from shaping long-term business strategy with data, to making critical short-term decisions, to developing algorithms and models that power both internal systems and customer-facing products. Driver Incentives Science owns the algorithms and systems behind incentive design, influencing driver engagement and marketplace efficiency — from real-time supply positioning to longer-horizon earnings and engagement programs. The team is responsible for designing pay and incentive mechanisms that are efficient and good for driver experience over the long run. As a Data Scientist specializing in Algorithms, you'll partner closely with product, engineering, and operations leaders to build and scale incentive systems, shape long-term mechanism design strategy, and deliver on critical business goals tied to marketplace efficiency and driver earnings. Candidates with strong optimization backgrounds — think mathematical programming, control theory, or operations research — are a great fit, though we welcome strong candidates from machine learning or causal inference as well. The ideal candidate thrives in a fast-paced environment and brings a hands-on, entrepreneurial mindset to drive results. Responsibilities: Collaborate with engineering and product teams to design, implement, and iterate on new features and algorithmic improvements for driver incentives and pay mechanisms. Design, develop, and deploy optimization models, algorithms, and systems for problems such as budget allocation, multidimensional cost-curve development, and incentive targeting. Write production model code; collabor

PythonMachine LearningAIGo
C
📍 Toronto, Ontario, Canada· Full-time
✓ Quality checkedCompany trend -91.5%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Security Clearance: Active Secret+ clearance strongly preferred; candidates eligible and willing to obtain clearance will also be considered. More information about Canadian Security Clearance is available here . As an Infrastructure Security Engineer, your key responsibilities include: Deploy, and manage infrastructure for Protected B classified environments, ensuring compliance with ITSG-33 and Canadian government standards Design and implement security controls for cloud (AWS, GCP, Azure) and hybrid/multi-cloud deployments Evaluate, implement, and manage security tools and technologies for training cluster and inference infrastructure hardening Implement security best practices including IAM, encryption, logging, and monitoring Participate in security incident response activities, including detection, analysis, containment, and remediation Conduct regular vulnerability assessments and penetration testing of infrastructure components Maintain comprehensive security documentation, procedures, and configurations for classified environments Maintain active Secret+ security clearance and adhere to all Canadian government security

AWSAzureGCPKubernetes
PE
📍 Palo Alto, CA· Full-time· Hybrid
✓ Quality checked

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We are a software engineering team with expertise in enabling ML models in production. We deploy AI models to run in variety of environments: air-gapped government networks, forward-deployed defense environments, edge nodes, and enterprises with strict data sovereignty requirements. Our customers rely on us for frontier AI capabilities running on hardware they control, often with constrained GPU resources and limited direct access. Rising to that challenge and meeting those expectations is what Palantir's excels at. We treat models like any other software: continuously tested, continually delivered, packaged for reproducible deployment, and built for long-term maintainability. You will own services end-to-end, and work across the full stack, from inference engines, GPU scheduling to deployment pipelines, observability, and integration with Palantir's platform. The goal is to deliver new models and capabilities quickly and continuously. Join us if you want to solve problems at the intersection of infrastructure and machine learning that directly enable critical customers.

Machine LearningAIGo
🔔

Get new inference technical lead jobs in Canada by email

Daily job updates · Unsubscribe anytime