About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus
Jobs in Canada
Ai Platform Engineer in San Francisco
330 active opportunities · Updated October 2026
Showing
15 jobs
Explore current ai platform engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About Us Twitch is the world’s biggest live streaming service, with global communities built around gaming, entertainment, music, sports, cooking, and more. It is where thousands of communities come together for whatever, every day. We’re about community, inside and out. You’ll find coworkers who are eager to team up, collaborate, and smash (or elegantly solve) problems together. We’re on a quest to empower live communities, so if this sounds good to you, see what we’re up to on LinkedIn and X , and discover the projects we’re solving on our Blog . Be sure to explore our Interviewing Guide to learn how to ace our interview process. About the Team Twitch Security Platform builds and operates the foundational software, data, and automation that enable security at scale across Twitch. As a Software Development Engineer II (SDE II) on the Security Platform team, you will design, build, and operate critical services, pipelines, and tooling that power Twitch's security, privacy, and compliance programs. About the Role Twitch Security Platform builds and operates the foundational software, data, and automation that enable security at scale across Twitch. As a Software Development Engineer II (SDE II) on the Security Platform team, you will design, build, and operate critical services, pipelines, and tooling that power Twitch's security, privacy, and compliance programs. In this role, you will work at the intersection of software, security, privacy, and data engineering, contributing production systems that handle large-scale security telemetry, automate security workflows, and provide reliable data and services to internal teams. You will partner closely with engineers and product teams to solve real-world security problems through well-designed software. You will own projects end-to-end from design and implementation through deployment and operational support for the systems that are business-critical and highly visible. The problems yo
About the Team DoorDash is a data driven organization and relies on timely, accurate and reliable data to drive many business and product decisions. The Core Data Platform organization owns all the infrastructure necessary to run an operationally efficient analytical data stack. About the Roles The Data Platform team spans data mobility frameworks, ingestion, infrastructure, tools, and governance. Together, they design and operate scalable compute and ingestion frameworks using technologies such as Spark, Flink, Kafka, Airflow, and modern lakehouse solutions, while also building abstractions and tools that simplify data workflows for engineers, analysts, and ML practitioners. In parallel, these teams establish strong data quality, cataloging, privacy, and compliance standards to ensure trust in analytics and regulatory adherence. As relatively high-impact teams, they offer engineers the opportunity to shape the roadmap, influence core platform decisions, and directly enable DoorDash’s business-critical insights and real-time personalization capabilities. You must be located in San Francisco, CA, Sunnyvale, CA, Seattle, WA, or New York, NY. You're excited about this opportunity because you will… Drive vision & strategy for building the frameworks charter and position it to handle the challenges of a rapidly growing business. Scale the analytical platform for the increasing amounts of data and use cases. You will bring your expertise in building and operating high scale systems with a focus on reliability, scalability and cost efficiency. Collaborate with stakeholders building solutions on top of the platform Foster a positive and supportive work culture, upleveling others. We're excited about you because you have… B.S., M.S., or PhD. in Computer Science or equivalent. 2+ years of industry experience at our I4 level, 5+ years of industry experience at our I5 level Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in th
About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, leading the design and architecture of our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly. You’re excited about this opportunity because you will… Lead the design of infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Own and evolve our open-weights serving stack — real-time GPU endpoints, high-thr
About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is our evaluation platform — the unified evals backbone that lets teams measure, trace, and trust the quality of LLM and agent systems across the company, powering trace/score ingestion, LLM-as-judge workflows, agent simulations, and LLM observability for the tens of millions of daily requests flowing through our LLM Gateway. We also own core platform surfaces including the Agent Gateway, open-weights model serving and batch inference, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, with a primary focus on our evals and LLM observability platform: the systems that let teams evaluate, trace, and continuously improve the quality of LLM and agent products. You’ll work across evaluation frameworks and SDKs, OpenTelemetry-based trace/score ingestion, LLM-as-judge and offline/online eval pipelines, agent simulations, data pipelines, backend services, and observability. This role is ideal for an engineer who enjoys building reliable measurement and quality primitives in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and evaluation methodologies are evolving quickly. You’re excited about this opportunity because you will… Build the infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Work on our unified evals platform — evaluation SDKs, OpenTelemetry trace/score ingestion, LLM-as-judge, offline and online eval pipelines, and agent simulations — alongside the LLM Gatew
$240K – $270K/yr
About the Role At Sigma, we’re not just adding AI—we’re building the future of how people work with data. Our platform already lets users explore billions of rows of data in seconds with a spreadsheet-like interface, analyze and present their data in workbooks, and build data apps and workflows. Now we’re pushing further, applying AI to reshape how people build in Sigma, discover insights, and make smarter decisions—fast. That’s where you come in. As an AI/ML Engineer, you’ll join a growing team focused on building the AI foundation that will power Sigma for the future. Your work will become an integral part of the workflow for the thousands of enterprises that run on Sigma. What You’ll Do Partner with product, design, and engineering teams to identify high-impact AI/ML opportunities Prototype and productionize AI systems that feel intuitive but do a lot under the hood—recommendations, natural language interfaces, agentic workflows, and more Develop and scale AI/ML infrastructure that powers both internal tooling and customer-facing features Tackle novel UX problems at the intersection of AI, BI, and apps What You Bring Bachelor’s degree in Computer Science, Engineering, Mathematics, or a related field (required) 10+ years of experience building and deploying production-grade AI/ML systems Deep knowledge of machine learning, deep learning, and applied AI Experience across the full ML lifecycle: data curation, training, deployment, monitoring A track record of building things that ship—whether it’s recommendations, search, machine translation, or something equally complex Experience adapting or training foundation models (language or multimodal) for novel domains Bonus Points (or skills you’ll build here) You've built agents that can plan, reason, and use tools You know your way around cloud infrastructure (AWS, GCP, Azure) You’ve worked in a fast-moving startup or high-growth environment Additional Job details The base salary range for this posit
$240K – $270K/yr
About the Role At Sigma, we’re not just adding AI—we’re building the future of how people work with data. Our platform already lets users explore billions of rows of data in seconds with a spreadsheet-like interface, analyze and present their data in workbooks, and build data apps and workflows. Now we’re pushing further, applying AI to reshape how people build in Sigma, discover insights, and make smarter decisions—fast. That’s where you come in. As an AI/ML Engineer, you’ll join a growing team focused on building the AI foundation that will power Sigma for the future. Your work will become an integral part of the workflow for the thousands of enterprises that run on Sigma. What You’ll Do Partner with product, design, and engineering teams to identify high-impact AI/ML opportunities Prototype and productionize AI systems that feel intuitive but do a lot under the hood—recommendations, natural language interfaces, agentic workflows, and more Develop and scale AI/ML infrastructure that powers both internal tooling and customer-facing features Tackle novel UX problems at the intersection of AI, BI, and apps What You Bring Bachelor’s degree in Computer Science, Engineering, Mathematics, or a related field (required) 10+ years of experience building and deploying production-grade AI/ML systems Deep knowledge of machine learning, deep learning, and applied AI Experience across the full ML lifecycle: data curation, training, deployment, monitoring A track record of building things that ship—whether it’s recommendations, search, machine translation, or something equally complex Experience adapting or training foundation models (language or multimodal) for novel domains Bonus Points (or skills you’ll build here) You've built agents that can plan, reason, and use tools You know your way around cloud infrastructure (AWS, GCP, Azure) You’ve worked in a fast-moving startup or high-growth environment Additional Job details The base salary range for this posit
From $302.4K/yr
Director of Engineering, Physical AI Role Overview The Director of Engineering will report to the General Manager of Physical AI, and will be responsible for leading a multi-disciplinary engineering organization. In this senior leadership role, you will own the execution of the Physical AI Data Engine — the platform powering the next generation of Physical AI/Embodied AI. You will collaborate closely with Operations and GTM to guide product direction and help solve the data bottleneck that stands between today's robotics research and real-world deployment. This role requires significant ownership in a fast-paced environment and you will motivate internal teams to set the pace for business growth. Travel will come into play. Key Responsibilities: Set and drive the technical vision across data collection infrastructure, teleoperation systems, ML training pipelines, model evaluation frameworks, annotation tooling, and research Lead a multidisciplinary engineering organization—spanning engineering managers, software engineers, ML engineers, and ML research scientists—while designing the organizational structure, talent strategy, and culture required to scale rapidly without compromising on quality or strategic alignment Maintain exceptional technical and operational excellence by deeply understanding team deliverables, asking incisive questions, identifying slipping standards early, and knowing precisely when to step in Drive cross-functional alignment across Engineering, Operations, and GTM on platform architecture, release processes, and shared priorities Collaborate with researchers and clients to architect and deliver scalable, production-grade data infrastructure tailored for complex robotics workloads Required Qualifications: Bachelor's degree in Engineering, Robotics, Computer Science, or a related technical field 8+ years of engineering experience in fast-paced environments, including 4+ years direct people management demonstrated history of recruiting, mentorin
From $184K/yr
Scale AI is seeking a highly skilled and motivated Software Engineer, Frontier AI Infrastructure to join our dynamic Public Sector Engineering team. As a part of this team, you will own the model inference layer - enabling state of the art models, debugging the latest AI tools, managing networking, debugging latency, and tracking pricing/usage metrics for AI models. You will lead technical discussions on the frontlines with cloud vendors and customers to deliver on critical contracts and to debug platform issues. You will also work upstream with Product to understand features before they break, moving us from "infra-only debugging" to proactive integration testing. You will: Design and implement secure scalable backend systems for Public Sector customers, leveraging Scale's modern and cloud-native AI infrastructure. Own services or systems and define their long-term health goals, while also improving the health of surrounding components Re-architect the stack to run in compliant or restrictive environments. This requires designing swappable components (auth, storage, logging) to meet government/security mandates without breaking the product. You will work with Product to build integration tests that catch issues early, shifting the focus from "infra-only debugging" to preventing failures upstream. Participate actively in customer engagements, working closely with stakeholders to understand requirements and deliver innovative solutions. Contribute to the platform roadmap and product strategy for Scale AI's Public Sector business, playing a key role in shaping the future direction of our offerings. Must have: At least an active secret clearance and the ability & willingness to up level to TS/SCI with CI Poly. This is a requirement and candidates will not be considered who do not hold at least a secret clearance Ideally you'd have: Full Stack Development: Proficiency in both front-end and back-end development, including experience with modern web develo
From $180K/yr
About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with the world's leading enterprises and government organizations to accelerate their AI transformation through frontier AI systems that solve real business problems. Every day, we work with organizations across finance, healthcare, manufacturing, media and telecommunications to build production AI agents that automate complex workflows, help humans, reason over enterprise knowledge, and operate safely at scale. The Opportunity Applied AI is moving faster than ever. New foundation models, reasoning techniques, agent architectures, and research papers emerge every week. Yet building AI systems that reliably solve real-world problems remains one of the hardest engineering challenges. As a Frontier Agent Engineer (Applied AI) , you'll bridge the gap between cutting-edge AI research and production deployment. You'll work directly with enterprise customers to design, evaluate, and deploy intelligent systems that combine frontier models with structured knowledge, retrieval, traditional machine learning, and enterprise software. Unlike traditional ML roles that focus on a single model or product, you'll work across a diverse portfolio of AI challenges spanning multiple industries and use cases. You may build a multi-agent research system and then participate in designing a customer intelligence platform, a healthcare copilot, or an autonomous workflow for a Fortune 100 company. If you enjoy reading new AI papers, experimenting with the latest models, and shipping production systems that create measurable business impact, you'll fit right in. What You'll Build Frontier AI Systems Design and deploy production AI agents that leverage the latest advances in large language models, reasoning, retrieval, memory, and tool use. Architect intelligent systems that combine LLMs, traditional machine learning, structured knowledge, enterprise data
From $216K/yr
About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with the world's leading enterprises and government organizations to accelerate their AI transformation through frontier AI systems that solve real business problems. Every day, we work with organizations across finance, healthcare, manufacturing, media and telecommunications to build production AI agents that automate complex workflows, help humans, reason over enterprise knowledge, and operate safely at scale. The Opportunity Applied AI is moving faster than ever. New foundation models, reasoning techniques, agent architectures, and research papers emerge every week. Yet building AI systems that reliably solve real-world problems remains one of the hardest engineering challenges. As a Senior Frontier Agent Engineer (Applied AI) , you'll bridge the gap between cutting-edge AI research and production deployment. You'll work directly with enterprise customers to design, evaluate, and deploy intelligent systems that combine frontier models with structured knowledge, retrieval, traditional machine learning, and enterprise software. Unlike traditional ML roles that focus on a single model or product, you'll work across a diverse portfolio of AI challenges spanning multiple industries and use cases. You may build a multi-agent research system and then participate in designing a customer intelligence platform, a healthcare copilot, or an autonomous workflow for a Fortune 100 company. If you enjoy reading new AI papers, experimenting with the latest models, and shipping production systems that create measurable business impact, you'll fit right in. What You'll Build Frontier AI Systems Design and deploy production AI agents that leverage the latest advances in large language models, reasoning, retrieval, memory, and tool use. Architect intelligent systems that combine LLMs, traditional machine learning, structured knowledge, enterpri
From $189.6K/yr
Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering MLEs, researchers, data scientists and operators for fast and automatic training and evaluation of LLM's, as well as evaluation of data quality. Scale is uniquely positioned at the heart of the field of AI as an indispensable provider of training and evaluation data and end-to-end solutions for the ML lifecycle. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development. You will be building and optimizing the platform to enable our next generation of LLM training, inference and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you! You will: Build, profile and optimize our training and inference framework Collaborate with ML teams to accelerate their research and development and enable them to develop the next generation of models and data curation Research and integrate state-of-the-art technologies to optimize our ML system Ideally you’d have: Strong excitement about system optimization Experience with multi-node LLM training and inference Experience with developing large-scale distributed ML systems Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc. Strong written and verbal communication skills and the ability to operate in a cross functional team environment Nice to haves: Demonstrated expertise in post-training methods &/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the positi
From $216K/yr
Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Identity Engineering team. In this role, you will help support the design and development of core software systems specifically focused on identity, access management, authorization, and authentication. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the design, and implementation of our identity infrastructure to ensure secure authentication and authorization across enterprise systems. Build software for authentication mechanisms such as Single Sign-On (SSO), Multi-Factor Authentication (MFA), and federated identity solutions (SAML, OAuth, OpenID Connect). Build software for authorization mechanisms such as Relation-based access control (ReBAC), Attribute-based access control (ABAC), Role-based access cont
From $252K/yr
About Scale AI At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. Scale Frontier Data is the organization behind the training and evaluation data that frontier labs depend on. We build the systems, tooling, and expert workflows that turn hard human expertise into signals that models can learn from, across reasoning, coding, agentic tool use, and domain expertise. Reinforcement learning environments are now the center of gravity for that work: the difference between a model that demos well and a model that reliably completes long-horizon work is almost always the quality of the environments and reward signals it was trained against. Responsibilities As a Staff Software Engineer, RL Environments, you'll own the technical foundation for how Scale builds, runs, verifies, and delivers RL environments at scale. An RL environment is a real piece of software: a containerized world with real dependencies, real state, real tools, and a grader that has to be correct even when the agent is creative about breaking it. Building one is a full-stack engineering problem. Building thousands of them reproducibly, cheaply, with trustworthy reward signals and throughput measured in millions of rollouts is a systems problem that very few people have solved. You'll work on both. You'll design the platform: sandboxed execution, environment packaging and versioning, rollout orchestration, trajectory capture, verifier frameworks, and the authoring surfaces that let engineers and domain experts produce environments without reinventing infrastructure each time. And you'll go deep on the environments themselves by instrumenting real applications, designing task suites that expose specific capability gaps, and building graders that
From $180K/yr
The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Software Engineer, you will own the development of a vertical feature or a horizontal capability to include defining requirements with stakeholders and implementation until it is accepted by the stakeholders. You will: Design and implement scalable backend systems for Federal customers using cloud-native AI infrastructure. Build features for agentic systems including multi-layered guardrails and data retrieval optimization. Develop data pipelines and machine learning infrastructure to make data sources accessible by agents. Collaborate with cross-functional teams to execute backend solutions for secure environments. Participate in customer engagements to understand requirements and deliver technical solutions. Define requirements with stakeholders and implement features until they are accepted. Contribute to the platform roadmap and product strategy for the Federal business. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in developing and deploying applications in a cloud-native environment. Understanding of containerization (e.g., Docker) and container orchestration (e.g., Kubernetes
Other cities to consider
More places hiring for this role
Get new ai platform engineer jobs in San Francisco, Canada by email
Daily job updates · Unsubscribe anytime