Jobs in Canada

Inference Technical Lead in San Francisco

48 active opportunities · Updated October 2026

Explore current inference technical lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $184K/yr

Quick readStrong listing-quality and freshness signals

Scale AI is seeking a highly skilled and motivated Software Engineer, Frontier AI Infrastructure to join our dynamic Public Sector Engineering team. As a part of this team, you will own the model inference layer - enabling state of the art models, debugging the latest AI tools, managing networking, debugging latency, and tracking pricing/usage metrics for AI models. You will lead technical discussions on the frontlines with cloud vendors and customers to deliver on critical contracts and to debug platform issues. You will also work upstream with Product to understand features before they break, moving us from "infra-only debugging" to proactive integration testing. You will: Design and implement secure scalable backend systems for Public Sector customers, leveraging Scale's modern and cloud-native AI infrastructure. Own services or systems and define their long-term health goals, while also improving the health of surrounding components Re-architect the stack to run in compliant or restrictive environments. This requires designing swappable components (auth, storage, logging) to meet government/security mandates without breaking the product. You will work with Product to build integration tests that catch issues early, shifting the focus from "infra-only debugging" to preventing failures upstream. Participate actively in customer engagements, working closely with stakeholders to understand requirements and deliver innovative solutions. Contribute to the platform roadmap and product strategy for Scale AI's Public Sector business, playing a key role in shaping the future direction of our offerings. Must have: At least an active secret clearance and the ability & willingness to up level to TS/SCI with CI Poly. This is a requirement and candidates will not be considered who do not hold at least a secret clearance Ideally you'd have: Full Stack Development: Proficiency in both front-end and back-end development, including experience with modern web develo

AWSAzureGCPDocker
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, leading the design and architecture of our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly. You’re excited about this opportunity because you will… Lead the design of infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Own and evolve our open-weights serving stack — real-time GPU endpoints, high-thr

PythonAWSGCPKubernetes
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team As DoorDash continues to expand rapidly, our Core Consumer team plays a pivotal role in shaping the consumer experience through enhancing personalization, search relevance, merchandising strategy, app quality, and the overall ordering experience. Our mission is to implement scalable data solutions and provide insights that directly influence our strategic product direction. We are looking for an Analytics Senior Manager to lead our Discovery team. About the Role As a Senior Manager of Data Science/Analytics, you’ll be leading a team of data scientists who are working on improving the consumer experience. You will develop strategic insights and work closely with the product, engineering, strategy and operations teams to actively build and measure the impact of new features. You will oversee our metrics and analytics strategy to inform strategic product direction, offer technical leadership and build processes to support velocity, and partner with data and product engineers to build robust data foundations. You're excited about this opportunity because you will… Lead and develop a team of Data Scientists in investigating complex issues and uncovering key drivers of our business, along with your own contributions as an Individual Contributor Influence the Product and Operations roadmap by making actionable recommendations based on data Interface frequently with senior leadership to showcase your team’s work and tackle complex business problems Drive measurement strategy for the area under scope, defining success metrics and implementing best practices around experiment design and statistical analysis Develop a strategic learning roadmap based on data observations, strategic questions, and hypotheses We're excited about you because you have… A degree in Math, Physics, Statistics, Economics, Computer Science, or a similar domain Experience managing a team of data scientists, and a track record of delivering impactful analyses 8+ years of experience i

AWSGitRestAI
A
📍 San Francisco, Canada· Full-time
✓ High-confidence listingCompany trend -86.2%

From $132K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join AirSupport helps Airbnb employees stay productive wherever and however they work. Within AirSupport, Executive Support provides a dedicated, high-touch technology experience for Airbnb’s senior leaders and their Executive Business Partners. The team combines deep technical expertise, proactive support, and strong cross-functional partnership across office, remote, travel, event, and other approved environments. Executive Support owns the executive technology experience, working closely with engineering, security, infrastructure, AV, workplace, and other technology partners who own the underlying technologies. The Difference You Will Make As a Senior Executive IT Support Engineer, you will be one of the most senior technical individual contributors within AirSupport and a trusted technology partner to Airbnb’s executive population. You will combine hands-on technical expertise with strong judgment, proactive planning, and cross-functional leadership. You will own complex and sensitive executive technology issues through resolution, anticipate risks before they become disruptions, and lead improvements that make the executive experience more reliable and supportable. Your impact will extend beyond the issues you personally resolve. You will identify patterns in what Executive Support is seeing, connect them to broader employee technology priorities, develop recommendations, and influence the teams responsible for the underlying technology. You will help shape Executive Support technical direction, represent the support perspective in major technology decisions, and identify opportuniti

AIRustExcelRecruitment
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team At DoorDash, design means making experiences for the people who order, the people who prepare, and the people who deliver. As a Design Manager at DoorDash, you want to build things that matter to real people. You're at your best when you can move from idea to shipped product quickly, bringing experiences to life that reach and influence users at massive scale. You'll care about whether the product you make solved a real problem for real people, or changed how someone experiences their day. You'll work shoulder-to-shoulder with Engineering and Product Management, dig into the data, and use LLM-powered tools alongside traditional design tools. If the LLM powered tools don’t exist yet, you build them. About the Role Trust is the foundation of every interaction on DoorDash. We're looking for a design leader who can define and drive a vision for designing for integrity at scale. In this role, you will join the Integrity organization. You’ll lead a team of designers working across fraud prevention, trust, safety, and compliance — protecting millions of consumers, Dashers, and merchants while keeping the platform seamless and trustworthy. You will report into the Head of Design for our Customer Experience & Integrity organization. This role is hybrid- 1–2 days per week in one of our Design Hubs. You’re excited about this opportunity because you will… Set the technical direction for Design across your product area — decide what gets built, in what order, and why; connect multiple teams around shared platforms so the work compounds instead of duplicating Work on ambiguous problems and turn them into architecture decisions and working systems that teams actually use; earn trust with leadership not through decks, but through prototypes and shipped code that make your point for you Stay close to the code and the craft across multiple projects at once — you're not just reviewing, you're building; the work you ship will move real metrics across the product area

TypeScriptReactAWSGit
DU
📍 San Francisco, Canada· Full-time· Remote
✓ High-confidence listing

From $1.1M/yr

Quick readStrong listing-quality and freshness signals

About the Team The Sales Analytics team at DoorDash is responsible for getting the right information and insights out to drive our business goals. As DoorDash grows both in scale and breadth of offering, the strength of our sales engine and organizational structure must grow with it. About the Role You will be focused on partnering with our sales & strategy teams, who help solve pain points for small business (SMB) merchants on the Doordash platform through various methods such as new product adoption. You will lead analyses to guide our sales team with data-driven insights to unlock growth opportunities within their accounts. This cross-functional role will partner with sales, operations, product, and strategy teams. You will report into a Senior Manager on our Sales Analytics team, which is part of our Sales Operations organization. You’re excited about this opportunity because you will… Strategize – Devise and execute initiatives against the overall sales org strategy for “winning the merchant” while managing stakeholders across multiple lines of business Experiment – Use data-driven decision-making and sound business judgment to run sales tests and lead market intelligence efforts Optimize – Build the best merchant acquisition engine so DoorDash continues to offer the highest quality selection for its customers Analyze – Build models to evaluate the economics, value, and opportunity costs of strategic initiatives to improve sales performance Influence – Manage cross-functional projects with our sales, partner management, operations, product, engineering, business operations and BD teams to improve the merchant experience and achieve targets We’re excited about you because… You have 2+ years of experience in Analytics, Consulting, or Strategy You have a bachelors degree or higher You are highly technical with advanced SQL, basic Python/R, AI, and data modeling experience. You are an excellent analytical thinker who can deliver actionable recommendations

PythonSQLAWSGit
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $179.4K/yr

Quick readStrong listing-quality and freshness signals

Scale GP (Scale Generative AI Platform) is an enterprise-grade AI platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale our core infrastructure in a fast-paced environment. The ideal candidate will have a strong understanding of software engineering principles and practices, as well as experience with large-scale distributed systems. You will implement solutions across multiple cloud providers (GCP, Azure, AWS) for customers in diverse, highly-regulated industries like healthcare, telecom, finance, and retail. What You’ll Do: Architect multi-cloud systems and abstractions to allow the SGP platform to run on top of existing Cloud providers Implement custom integrations between Scale AI's platform and customer data environments (cloud platforms, data warehouses, internal APIs) Collaborate with platform, product teams and our customers directly to develop and implement innovative infrastructure that scales to meet evolving needs. Deliver experiments at a high velocity and level of quality to engage our customers Work across the entire product lifecycle from conceptualization through production Be able, and willing, to multi-task and learn new technologies quickly What We’re Looking For: 4+ years of full-time engineering experience, post-graduation Experience scaling products at hyper growth startups Experience tinkering with or productizing LLMs, vector databases, and the other latest AI technologies Proficient in Python or Javascript/Typescript, and SQL Experience with Kubernetes Experience with major cloud providers (AWS, Azure, GCP) Excellent communication skills with the ability to explain technical concepts to both technical and non-technical audiences Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries fo

JavaScriptTypeScriptPythonJava
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is our evaluation platform — the unified evals backbone that lets teams measure, trace, and trust the quality of LLM and agent systems across the company, powering trace/score ingestion, LLM-as-judge workflows, agent simulations, and LLM observability for the tens of millions of daily requests flowing through our LLM Gateway. We also own core platform surfaces including the Agent Gateway, open-weights model serving and batch inference, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, with a primary focus on our evals and LLM observability platform: the systems that let teams evaluate, trace, and continuously improve the quality of LLM and agent products. You’ll work across evaluation frameworks and SDKs, OpenTelemetry-based trace/score ingestion, LLM-as-judge and offline/online eval pipelines, agent simulations, data pipelines, backend services, and observability. This role is ideal for an engineer who enjoys building reliable measurement and quality primitives in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and evaluation methodologies are evolving quickly. You’re excited about this opportunity because you will… Build the infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Work on our unified evals platform — evaluation SDKs, OpenTelemetry trace/score ingestion, LLM-as-judge, offline and online eval pipelines, and agent simulations — alongside the LLM Gatew

PythonSQLAWSGCP
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $252K/yr

Quick readStrong listing-quality and freshness signals

The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Staff Software Engineer, you will orchestrate the implementation of vertical features and horizontal capabilities to include mentoring other engineers on defining requirements with stakeholders and communication tradeoffs of technical implementations on feature and capabilities until they are accepted by the stakeholders. You will: Orchestrate feature implementation across the Federal engineering team to ensure architectural consistency. Define technical strategy for agentic guardrails, explainability, and fleet orchestration. Ensure system reliability and performance across multiple security classifications and network types. Mentor engineers in the process of defining requirements with stakeholders and gathering acceptance. Communicate high-level technical trade-offs and implementation strategies to senior government stakeholders and Scale C-Suite members. Influence the long-term product strategy and technical roadmap for the Federal business unit. Consult on the architecture of AI-powered solutions for large-scale federal contracts. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in

AWSAzureGCPDocker
TI
📍 San Francisco, Canada
✓ High-confidence listing

$84K – $120K/yr

Quick readStrong listing-quality and freshness signals

We believe communication belongs to everyone. We exist to democratize phone service. TextNow is evolving the way the world connects, and that's because we're made up of people with curious minds who bring an optimistic yet critical lens into the work we do. We're the largest provider of free phone service in the nation. And we're just getting started. Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. TextNow is looking for an experienced Data Developer with hands-on experience designing and developing data platforms. You will own the design, development, and maintenance of TextNow's data platform, enabling us to make effective data-informed decisions. You will be part of cross-functional efforts to build scalable and reliable frameworks that support allTextNow's business and data products. In this role, you can interact with different functional areas within the business and influence decision-making in a fast-growing mobile communications start-up. This role is about impact at scale. You’ll shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded into everyday engineering practice. Using AI is not optional, it’s expected. From design and architecture to implementation, testing, debugging, documentation, and operational analysis, you will actively leverage AI tools to increase velocity, improve code quality, and make better technical decisions. We provide a robust suite of AI-powered development tools and workflows to support you, and we expect you to continuously evolve how you use them to raise the bar for efficiency, clarity, and product excellence across the organization. What You'll Do Own TextNow's data warehouse, data pipelines, and integration points between various business systems. Design, develo

PythonSQLAWSArtificial Intelligence
TI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$88.9K – $127K/yr

Quick readStrong listing-quality and freshness signals

We believe communication belongs to everyone. We exist to democratize phone service. TextNow is evolving the way the world connects, and that's because we're made up of people with curious minds who bring an optimistic yet critical lens into the work we do. We're the largest provider of free phone service in the nation. And we're just getting started. Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. TextNow is looking for an experienced Data Developer with hands-on experience designing and developing data platforms. You will own the design, development, and maintenance of TextNow's data platform, enabling us to make effective data-informed decisions. You will be part of cross-functional efforts to build scalable and reliable frameworks that support allTextNow's business and data products. In this role, you can interact with different functional areas within the business and influence decision-making in a fast-growing mobile communications start-up. This role is about impact at scale. You’ll shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded into everyday engineering practice. Using AI is not optional, it’s expected. From design and architecture to implementation, testing, debugging, documentation, and operational analysis, you will actively leverage AI tools to increase velocity, improve code quality, and make better technical decisions. We provide a robust suite of AI-powered development tools and workflows to support you, and we expect you to continuously evolve how you use them to raise the bar for efficiency, clarity, and product excellence across the organization. What You'll Do Own TextNow's data warehouse, data pipelines, and integration points between various business systems. Design, develo

PythonSQLAWSAI
HI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$140K – $225K/yr

Quick readStrong listing-quality and freshness signals

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Wireless Systems Software Engineer , you will help build the core software libraries and system infrastructure that power the next generation of wireless experiences at HP IQ. You will work across hardware, firmware, embedded software, operating systems, and product teams to integrate and optimize multiple wireless technologies, including Bluetooth, Ultra-Wideband (UWB), NFC , and future connectivity solutions. You will play a key role in designing scalable software architectures that bridge hardware and software, enabling seamless communication across embedded devices and host platforms. This position offers the opportunity to work on challenging system-level problems, influence the architecture of a new wireless ecosystem, and help define technologies that will shape the future of work. What You Might Do Serve as a technic al expert across wireless technologies including Bluetooth, UWB, NFC , and related embedded communication interfaces. Architect, design, and develop reusable, scalable wireless software libraries with well-defined API

RedisLinuxAIC++
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $189.6K/yr

Quick readStrong listing-quality and freshness signals

Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering MLEs, researchers, data scientists and operators for fast and automatic training and evaluation of LLM's, as well as evaluation of data quality. Scale is uniquely positioned at the heart of the field of AI as an indispensable provider of training and evaluation data and end-to-end solutions for the ML lifecycle. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development. You will be building and optimizing the platform to enable our next generation of LLM training, inference and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you! You will: Build, profile and optimize our training and inference framework Collaborate with ML teams to accelerate their research and development and enable them to develop the next generation of models and data curation Research and integrate state-of-the-art technologies to optimize our ML system Ideally you’d have: Strong excitement about system optimization Experience with multi-node LLM training and inference Experience with developing large-scale distributed ML systems Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc. Strong written and verbal communication skills and the ability to operate in a cross functional team environment Nice to haves: Demonstrated expertise in post-training methods &/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the positi

AWSRestAIGo
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

Scale GP is Scale's enterprise Generative AI platform—APIs and infrastructure for knowledge retrieval, inference, evaluation, and intelligent automation. We power mission-critical workflows for leading enterprises, helping teams turn complex data and models into reliable, production-ready AI systems. We're building a new AI Enablement team to create the next generation of agent-powered tools that ground AI in real operational workflows. Our goal: help internal teams demystify their own workflows, then deploy agentic systems that reason over data, take action, and deliver measurable outcomes. We don't build in a vacuum. You'll use our own platform to solve real business problems internally—then selectively commercialize that same stack for customers. What we run on is what we sell. This is a 0→1 team. We're looking for a sharp, product-minded engineer who thrives in ambiguity, moves fast, and loves building systems from scratch alongside customers and cross-functional partners. You'll work closely with product, forward-deployed engineers, data scientists, and applied AI teams to turn real-world problems into scalable production solutions. If you like shipping fast, owning outcomes, and working across the stack—from polished frontends to distributed backends to LLM integrations—this role is for you. What You’ll Do Own full-stack features and projects end-to-end — from design through production deployment — within a larger product area Sample surfaces - Accounting Agents, Finance Copilots, GTM Agents, Agentic Experimentation Platforms Develop reliable backend services in Typescript/Python, work with distributed systems, data pipelines, and AI/ML infrastructure Integrate LLMs, vector databases, and agentic frameworks to power intelligent workflows Ship quickly through tight experimentation loops while maintaining high quality and reliability Adapt across the stack and learn new tools as needed to solve real problems end-to-end Ideal Experience 3+ years of full-tim

TypeScriptPythonAWSRest
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $290.4K/yr

Quick readStrong listing-quality and freshness signals

Scale's LLM post-training platform team builds our internal distributed framework for large language model training. The platform powers MLEs, researchers, data scientists, and operators for fast and automatic training and evaluation of LLMs. It also serves as the underlying training framework for the data quality evaluation pipeline. Scale is uniquely positioned at the heart of the field of AI as an indispensable provider of training and evaluation data and end-to-end solutions for the ML lifecycle. You will work closely with Scale’s ML teams and researchers to build the foundation platform which supports all our ML research and development works. You will be building and optimizing the platform to enable our next generation LLM training, inference and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you! You will: Build, profile and optimize our training and inference framework. Collaborate with ML and research teams to accelerate their research and development, and enable them to develop the next generation of models and data curation. Research and integrate state-of-the-art technologies to optimize our ML system. Ideally you’d have: Passionate about system optimization Experience with multi-node LLM training and inference Experience with developing large-scale distributed ML systems Experience with post-training methods like RLHF/RLVR and related algorithms like PPO/GRPO etc. Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc. Strong written and verbal communication skills to operate in a cross functional team environment. Nice to haves: Demonstrated expertise in post-training methods and/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc. Compensation packages at Scale for eligible roles include base salary, equity,

AWSRestAIGo
🔔

Get new inference technical lead jobs in San Francisco, Canada by email

Daily job updates · Unsubscribe anytime