Jobs in Canada

Ai Systems Research And Development Engineer in San Francisco

330 active opportunities · Updated October 2026

Explore current ai systems research and development engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $189.6K/yr

Quick readStrong listing-quality and freshness signals

Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering MLEs, researchers, data scientists and operators for fast and automatic training and evaluation of LLM's, as well as evaluation of data quality. Scale is uniquely positioned at the heart of the field of AI as an indispensable provider of training and evaluation data and end-to-end solutions for the ML lifecycle. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development. You will be building and optimizing the platform to enable our next generation of LLM training, inference and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you! You will: Build, profile and optimize our training and inference framework Collaborate with ML teams to accelerate their research and development and enable them to develop the next generation of models and data curation Research and integrate state-of-the-art technologies to optimize our ML system Ideally you’d have: Strong excitement about system optimization Experience with multi-node LLM training and inference Experience with developing large-scale distributed ML systems Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc. Strong written and verbal communication skills and the ability to operate in a cross functional team environment Nice to haves: Demonstrated expertise in post-training methods &/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the positi

AWSRestAIGo
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $290.4K/yr

Quick readStrong listing-quality and freshness signals

Scale's LLM post-training platform team builds our internal distributed framework for large language model training. The platform powers MLEs, researchers, data scientists, and operators for fast and automatic training and evaluation of LLMs. It also serves as the underlying training framework for the data quality evaluation pipeline. Scale is uniquely positioned at the heart of the field of AI as an indispensable provider of training and evaluation data and end-to-end solutions for the ML lifecycle. You will work closely with Scale’s ML teams and researchers to build the foundation platform which supports all our ML research and development works. You will be building and optimizing the platform to enable our next generation LLM training, inference and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you! You will: Build, profile and optimize our training and inference framework. Collaborate with ML and research teams to accelerate their research and development, and enable them to develop the next generation of models and data curation. Research and integrate state-of-the-art technologies to optimize our ML system. Ideally you’d have: Passionate about system optimization Experience with multi-node LLM training and inference Experience with developing large-scale distributed ML systems Experience with post-training methods like RLHF/RLVR and related algorithms like PPO/GRPO etc. Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc. Strong written and verbal communication skills to operate in a cross functional team environment. Nice to haves: Demonstrated expertise in post-training methods and/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc. Compensation packages at Scale for eligible roles include base salary, equity,

AWSRestAIGo
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including frontier model training, enterprise adoption, defense applications, and autonomous vehicles. Our mission is to develop reliable AI systems for the world’s most important decisions. We’re looking for an AI Product Manager to own the Finance vertical within our Agents Data & Reinforcement Learning Environments team. In this role, you’ll own both the development of RL environments (the realistic, high-fidelity simulations of financial software and workflows that labs use to train and evaluate agents) and the “data as a product” strategy that powers them. You’ll understand where AI is being used in the Finance industry, decide what financial tasks are worth modeling, how to source and structure the underlying data, and how to turn deep domain knowledge into a defensible product. The ideal candidate has lived inside the Finance industry, and is able to pair that domain understanding with a sense for AI research and current agent capabilities in Finance workflows. You’ll translate that expertise into environments and datasets that teach AI agents to perform real financial work, and you’ll be the domain expert Scale’s most important customers and their leading researchers turn to. A strong entrepreneurial & go-to-market mindset will be necessary. What You’ll Do Own the Finance AI roadmap & data strategy: Set product direction for the Finance agents training stack and the data strategy behind it. Establish a vision for where AI is continuing to transform the Finance industry (including investment banking, private equity, public markets, corporate finance, FP&A, etc), driving execution across engineering, operations, and go-to-market teams. Build partnerships with research teams at frontier labs: Work directly with researchers at leading AI labs to understand where their Finance agentic capabilities fall short and shape new product lines and competit

AWSRestAIGo
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $264.8K/yr

Quick readStrong listing-quality and freshness signals

Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. About the General Agents Team The General Agents team, part of Scale’s Enterprise organization, builds robust general agents for customer use cases and applications. The team sits at the intersection of frontier agent development and real-world deployment, translating state-of-the-art reasoning and agentic capabilities into reliable, production-grade systems that drive real economic value. Our agents are scalable systems built around recurring enterprise problem domains, with a strong emphasis on generalization, extensibility, and deployment across many customers. About the Role As a Senior/Staff Machine Learning Engineer (MLE) on the General Agents team, you’ll play a critical role in designing, building, and deploying production-ready AI agents that solve high-impact enterprise problems. You will work across the full agent lifecycle—from model and system design to evaluation, deployment, and iteration—bridging cutting-edge agentic techniques with the constraints and requirements of real customer environments. You will: Design and implement end-to-end agent systems that combine LLM reasoning, tool use, memory, and control logic to solve recurring enterprise use cases. Build scalable, reliable agent architectures that can be deployed across many customers with varying data, tools, and constraints. Develop evaluation frameworks, datasets, environments, and metrics to measure agent performance, reliability, and business impact in production settings. Collaborate closely with product managers, customers, data annotators, and other engineering teams to translate enterprise requirements into robust agent designs. Productionize frontier agent techniques (e.g.,

PythonAWSRestMachine Learning
A
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $170K/yr

Quick readStrong listing-quality and freshness signals

Airtable is the no-code app platform that empowers people closest to the work to accelerate their most critical business processes. More than 500,000 organizations, including 80% of the Fortune 100, rely on Airtable to transform how work gets done. Airtable’s mission is to bring the power of computing and software development to everyone. We are developing a powerful and extensible toolkit that our customers can leverage to solve a variety of different problems and workflows. We’ve seen our most sophisticated customers use the product to run global processes across thousands of employees, coordinate precision manufacturing pipelines, and consolidate previously siloed mission-critical data into a single source of truth. The complexity of these use cases requires us to be extremely thoughtful about how we design and implement new functionality in the product and make sure it’s both easy to use and comprehend for our customers and maintainable for us. As a Full-Stack, Backend engineer at Airtable, you will have the opportunity to work with customers to deeply understand their needs and workflows. You will collaborate with cross-functional partners across product management, design, research and data science to create innovative new features that enable our customers to do their best work. You will be responsible for owning and executing the end-to-end implementation of these new features that will contribute to making our toolkit even more powerful and successful. We currently have openings on: The Admin & Governance Team (Full-Stack/BE) ensures Airtable is secure, compliant, and enterprise-ready. It owns key admin capabilities like the Admin Panel, SSO, and audit systems, as well as foundational features like User Groups. This team's mission is to accelerate organizational value for the largest customers with enterprise-first governance and controls. The Omni Capability & Quality Team (Full-Stack/BE) brings the power of AI directly to Airtable end users—

JavaScriptJavaReactNode.js
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $182K/yr

Quick readStrong listing-quality and freshness signals

At Scale, we develop reliable AI systems for the world’s most important decisions. Scale’s product marketing team is responsible for developing and executing strategies that drive awareness and engagement for Scale’s offerings amongst our core audiences. We take a data-driven approach to understand our customers’ needs and challenges, ensuring that their voices are reflected in product development and messaging. We partner closely with product, engineering, research, sales, comms, and the broader marketing team to create a cohesive customer experience across all our channels. We aim to provide valuable insights and resources that help our customers execute on their AI transformation journey. This role will focus on developing and optimizing vertical-specific messaging and content for Scale’s applications business to ensure our messaging resonates with core buyers across verticals. The ideal candidate combines strategic thinking with hands-on execution. You will: Develop clear, compelling messaging for enterprise offerings tailored to key personas within target verticals Create marketing assets, including slide decks, videos, blog posts, and one-pagers, that effectively communicate our value propositions to vertical leaders and support our sales team’s pursuits Lead positioning and sales enablement efforts to drive awareness and engagement for Scale’s offerings. Drive cross-functional marketing programs around target verticals, in conjunction with product, engineering, sales, and growth marketing to create a cohesive customer experience and contribute to pipeline targets Collaborate with field marketing and events teams to develop vertical-specific event strategies, content, and experiences that drive engagement within customer and target accounts Drive customer marketing efforts, creating both strategy and tactics to maximize the value Scale and customers get from our shared success, including case studies, testimonials, visual assets and event particip

AWSRestAIGo
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including frontier model training, enterprise adoption, defense applications, and autonomous vehicles. Our mission is to develop reliable AI systems for the world's most important decisions. The next frontier for AI is the physical world. We're looking for an AI Product Manager to own the Robotics vertical within our Physical AI team. In this role, you'll own both the development of the data and training environments (the teleoperated demonstrations, real-world collections, simulated tasks, and annotation products that labs use to train and evaluate robot policies) and the "data as a product" strategy that powers them. You'll understand where physical AI is headed, decide what robot tasks and embodiments are worth collecting, how to source and structure the underlying data, and how to turn deep domain knowledge into a defensible product. The ideal candidate has lived inside robotics or physical AI research, and is able to pair that domain understanding with a sense for where current robot policies succeed and fail in real-world workflows. You'll translate that expertise into datasets, environments, and evaluation frameworks that teach robots to do real physical work, and you'll be the domain expert Scale's most important customers and their leading researchers turn to. A strong entrepreneurial & go-to-market mindset will be necessary. What You'll Do Own the Robotics AI roadmap & data strategy: Set product direction for the robotics training stack and the data strategy behind it — what data we collect, on which hardware and embodiments, and what we source internally vs. through our marketplace. Establish a vision for where physical AI is heading, driving execution across engineering, operations, and go-to-market teams. Build partnerships with research teams at frontier labs: Work directly with researchers at leading physical AI labs to understand where their robot p

AWSRestAIGo
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $252K/yr

Quick readStrong listing-quality and freshness signals

About Snorkel At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data. We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler! About Snorkel Snorkel AI is the frontier AI data lab, helping teams build the data and environments behind high-performing frontier and agentic AI. We combine technology with research-driven AI data development to create datasets, benchmarks, evals, and custom solutions for real-world AI systems. Founded out of the Stanford AI Lab in 2019, Snorkel works with leading AI labs and enterprises to move from better data to better outcomes. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler! About The Role Snorkel is hiring a Head of Security to build and lead our security function end-to-end — infrastructure security, application security, and governance, risk & compliance (GRC). You'll own the security function end-to-end — strategy, team, and execution — and operate as the primary security voice with customers, auditors, and the exec team. You'll report to the CTO. This is a builder's role: you'll take security from its current state to a mature, right-sized function as Snorkel scales, hiring and developing the team as needs grow. Key Responsibilities Security Leadership & Team Building Define Snorkel's overall security strategy,

AWSCI/CDAIGo
HI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$160K – $200K/yr

Quick readStrong listing-quality and freshness signals

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About the Role HP IQ is looking for a highly organized Senior Product Design Producer to support and scale Product Design work across key verticals. This role is a liaison between Product Design, Engineering, and Partnerships teams, promoting cross-functional communication and collaboration. This position also manages product pilots and qualitative research schedules. Our ideal candidate is able to effortlessly manage multiple product work streams, competing priorities, and able to adapt to changing circumstances in a fast-paced environment while keeping all design deliverables on track. What You Might Do Partner with a product lead to define scope, develop roadmaps, and prioritize product design deliverables Collaborate with Engineering Project Managers to develop design project schedules, coordinating and tracking designer activity to ensure timely delivery Support a hybrid creative project management methodology in tandem with an Agile software development approach, moving with ease between organizational systems Manage product qualitative research schedules, product pilots, schedules and i

RedisGitAgileAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Scale Labs, Research Scientist — AI Controls and Monitoring As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist focused on AI Controls and Monitoring, you will design methods, systems, and experiments to ensure that advanced AI models and agents remain aligned with intended goals, even in high-stakes or adversarial environments. For example, you might: Develop monitoring techniques and observability methods that track AI behavior in real time to identify and flag deviations, emergent capabilities, or anomalous outputs; Research mechanisms for layered control, including fail-safes, oversight protocols, and intervention methods that can halt or redirect AI systems when risks are detected; Design red-team simulations to probe weaknesses in oversight and control mechanisms, and build mitigations to close identified gaps; Collaborate with policymakers, engineers, and other researchers to establish standards and benchmarks for AI monitoring and escalation. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Practical experience conducting technical research collaboratively. You should be

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Scale Labs, Research Scientist — Safety Post Training As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist working on Safety Post-Training you will develop and apply post-training methods and interpretability techniques to make frontier AI systems safer, and better understood by researchers and policymakers.. For example, you might: Design and run post-training pipelines to study how training choices affect model safety, robustness, and alignment properties; Develop interpretability-informed evaluations that reveal how and why models produce unsafe, deceptive, or otherwise undesirable behaviors, and use those insights to guide targeted mitigations; Collaborate with policymakers, engineers, and other researchers to translate post-training and interpretability findings into actionable safety standards, evaluation benchmarks, and best practices. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar approaches. A track record of published research in machine learning, particularly in generati

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Scale Labs, Research Scientist — Frontier Risk Evaluations As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist focused on Frontier Risk Evaluations, you will design and create evaluation measures, harnesses and datasets for measuring the risks posed by frontier AI systems. For example, you might do any or all of the following: Design and build harnesses to test AI models and systems (including agents) for dangerous capabilities such as security vulnerability exploitation, CBRN uplift, and other high-risk activities; Work with government agencies or other labs to collectively scope and design evaluations to measure and mitigate risks posed by advanced AI systems; Publish evaluation methodologies and write technical reports for policymakers. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Practical experience conducting technical research collaboratively. You should be comfortable building and instrumenting ML pipelines, writing evaluation harnesses, and quickly turning new ideas from the research literature into working prototypes. A track record of published research in m

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with the world's leading enterprises and government organizations to accelerate their AI transformation through frontier AI systems that solve real business problems. Every day, we work with organizations across finance, healthcare, manufacturing, media and telecommunications to build production AI agents that automate complex workflows, help humans, reason over enterprise knowledge, and operate safely at scale. The Opportunity Applied AI is moving faster than ever. New foundation models, reasoning techniques, agent architectures, and research papers emerge every week. Yet building AI systems that reliably solve real-world problems remains one of the hardest engineering challenges. As a Frontier Agent Engineer (Applied AI) , you'll bridge the gap between cutting-edge AI research and production deployment. You'll work directly with enterprise customers to design, evaluate, and deploy intelligent systems that combine frontier models with structured knowledge, retrieval, traditional machine learning, and enterprise software. Unlike traditional ML roles that focus on a single model or product, you'll work across a diverse portfolio of AI challenges spanning multiple industries and use cases. You may build a multi-agent research system and then participate in designing a customer intelligence platform, a healthcare copilot, or an autonomous workflow for a Fortune 100 company. If you enjoy reading new AI papers, experimenting with the latest models, and shipping production systems that create measurable business impact, you'll fit right in. What You'll Build Frontier AI Systems Design and deploy production AI agents that leverage the latest advances in large language models, reasoning, retrieval, memory, and tool use. Architect intelligent systems that combine LLMs, traditional machine learning, structured knowledge, enterprise data

PythonAWSAzureGCP
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with the world's leading enterprises and government organizations to accelerate their AI transformation through frontier AI systems that solve real business problems. Every day, we work with organizations across finance, healthcare, manufacturing, media and telecommunications to build production AI agents that automate complex workflows, help humans, reason over enterprise knowledge, and operate safely at scale. The Opportunity Applied AI is moving faster than ever. New foundation models, reasoning techniques, agent architectures, and research papers emerge every week. Yet building AI systems that reliably solve real-world problems remains one of the hardest engineering challenges. As a Senior Frontier Agent Engineer (Applied AI) , you'll bridge the gap between cutting-edge AI research and production deployment. You'll work directly with enterprise customers to design, evaluate, and deploy intelligent systems that combine frontier models with structured knowledge, retrieval, traditional machine learning, and enterprise software. Unlike traditional ML roles that focus on a single model or product, you'll work across a diverse portfolio of AI challenges spanning multiple industries and use cases. You may build a multi-agent research system and then participate in designing a customer intelligence platform, a healthcare copilot, or an autonomous workflow for a Fortune 100 company. If you enjoy reading new AI papers, experimenting with the latest models, and shipping production systems that create measurable business impact, you'll fit right in. What You'll Build Frontier AI Systems Design and deploy production AI agents that leverage the latest advances in large language models, reasoning, retrieval, memory, and tool use. Architect intelligent systems that combine LLMs, traditional machine learning, structured knowledge, enterpri

PythonAWSAzureGCP
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $250K/yr

Quick readStrong listing-quality and freshness signals

About Scale Scale’s mission is to develop reliable AI systems for the world’s most important decisions. As the leading AI data foundry, we provide the high-quality data and full-stack technologies that power the world’s most advanced models — fueling breakthroughs in generative AI, defense, and autonomous vehicles. We partner with leading enterprises and governments to bring AI into production that performs when it matters most, combining rigorous evaluation with full-stack deployment so our customers can build AI they can trust. About the Team Applied Intelligence Systems (AIS) is part of the Scale Generative AI Platform (SGP), focused on pushing the frontier of what agentic applications can do across diverse enterprise and government use cases. We build the infrastructure and tooling that power agentic AI in production, paired with applied ML research, design, and evaluation to ensure these systems perform reliably at the scale our customers demand. AIS spans multiple workstreams — agent evaluation and oversight, orchestration and tool-use infrastructure, model and systems optimization, and applied research on new agent capabilities — and this role is not scoped to any single one of them. We’re growing fast, with increasing traction across both commercial and public sector customers, and we’re just getting started — this team will define what dependable, production-grade agentic AI looks like. About the Role As a Staff Machine Learning Research Engineer, you will operate across the full breadth of AIS’s technical needs — wherever the hardest ML problem in agentic AI happens to be that quarter. This could mean training and fine-tuning models, designing evaluation and observability systems, building improvement loops from production data, prototyping novel agent architectures, or designing internal systems and tooling that boost productivity across teams. You’re not tied to one team’s roadmap; you’re expected to move to where the technical leverage is highest, and t

AWSRestMachine LearningAI
🔔

Get new ai systems research and development engineer jobs in San Francisco, Canada by email

Daily job updates · Unsubscribe anytime