Jobs in Canada

Scientist 2 in San Francisco

35 active opportunities · Updated October 2026

Explore current scientist 2 jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Scale Labs, Research Scientist — Agent Robustness As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist working on Agent Robustness you will work on the fundamental challenges of building AI agents that are safe and aligned with humans. For example, you might: Research the science of AI agent capabilities with a focus on how they relate to safety, risk factors, and methodologies for benchmarking them; Design and build harnesses to test AI agents’ tendency to take harmful actions when pressured to do so by users or tricked into doing so by elements of their environment; Design and build exploits and mitigations for new and unique failure modes that arise as AI agents gain affordances like coding, web browsing, and computer use; Characterize and design mitigations for potential failure modes or broader risks of systems involving multiple interacting AI agents. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Practical experience conducting technical research collaboratively. You should be comfortable building and leveraging agent scaffolding, designing evaluation harnesses, an

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Scale Labs, Research Scientist — AI Controls and Monitoring As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist focused on AI Controls and Monitoring, you will design methods, systems, and experiments to ensure that advanced AI models and agents remain aligned with intended goals, even in high-stakes or adversarial environments. For example, you might: Develop monitoring techniques and observability methods that track AI behavior in real time to identify and flag deviations, emergent capabilities, or anomalous outputs; Research mechanisms for layered control, including fail-safes, oversight protocols, and intervention methods that can halt or redirect AI systems when risks are detected; Design red-team simulations to probe weaknesses in oversight and control mechanisms, and build mitigations to close identified gaps; Collaborate with policymakers, engineers, and other researchers to establish standards and benchmarks for AI monitoring and escalation. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Practical experience conducting technical research collaboratively. You should be

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Scale Labs, Research Scientist — Safety Post Training As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist working on Safety Post-Training you will develop and apply post-training methods and interpretability techniques to make frontier AI systems safer, and better understood by researchers and policymakers.. For example, you might: Design and run post-training pipelines to study how training choices affect model safety, robustness, and alignment properties; Develop interpretability-informed evaluations that reveal how and why models produce unsafe, deceptive, or otherwise undesirable behaviors, and use those insights to guide targeted mitigations; Collaborate with policymakers, engineers, and other researchers to translate post-training and interpretability findings into actionable safety standards, evaluation benchmarks, and best practices. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar approaches. A track record of published research in machine learning, particularly in generati

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Scale Labs, Research Scientist — Frontier Risk Evaluations As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist focused on Frontier Risk Evaluations, you will design and create evaluation measures, harnesses and datasets for measuring the risks posed by frontier AI systems. For example, you might do any or all of the following: Design and build harnesses to test AI models and systems (including agents) for dangerous capabilities such as security vulnerability exploitation, CBRN uplift, and other high-risk activities; Work with government agencies or other labs to collectively scope and design evaluations to measure and mitigate risks posed by advanced AI systems; Publish evaluation methodologies and write technical reports for policymakers. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Practical experience conducting technical research collaboratively. You should be comfortable building and instrumenting ML pipelines, writing evaluation harnesses, and quickly turning new ideas from the research literature into working prototypes. A track record of published research in m

AWSRestMachine LearningAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $1.5M/yr

Quick readStrong listing-quality and freshness signals

About the Team The Analytics team is looking for experienced Data Scientists and Senior Data Scientists to guide measurement, strategy, and tactical decision-making across the company across a variety of teams and levels. Data Scientists at DoorDash work to uncover insights and turn them into relevant recommendations, driving decisions for the entire organization. Analytics is integral to all operational areas at DoorDash. Please apply here for all non-managerial levels within the following analytics teams: Consumer & Growth Business Operations Dasher & Logistics Customer Experience & Integrity Merchant, Ads & Sales New Verticals International Data Science About the Role As a Data Scientist at DoorDash, you'll use your quantitative background to mentor other scientists and dive into large datasets to guide decision-making. We solve a multitude of exciting challenges including customer acquisition, fraud and support, marketing, balancing supply and demand, new city launches, marketplace efficiency, and more. If you enjoy finding patterns amidst chaos, and have experience using analytics to affect revenue, growth, operations or beyond, we're looking for someone like you! You're excited about this opportunity because you will… Use quantitative analysis and the presentation of data to see beyond the numbers and understand what drives our business Build full-cycle analytics experiments, reports, and dashboards using SQL, R, Python, or other scripting and statistical tools Work with and mentor junior analysts on how to use more advanced methods and solve challenges Produce recommendations and use statistical techniques and hypothesis testing to validate your findings Provide insights to help business and product leaders understand marketplace dynamics, user behaviors, and long-term trends Identify and measure levers to help move essential metrics and make recommendations Work backwards from understanding and sizing problems to ideating solutions Report aga

PythonSQLAWSGit
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $165.6K/yr

Quick readStrong listing-quality and freshness signals

Scale works with the industry’s leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling). This role will focus on optimizing data curation and eval to enhance LLM capabilities in both text and multimodal modalities. In this role, you will develop novel methods to improve the alignment and generalization of large-scale generative models. You will collaborate with researchers and engineers to define best practices in data-driven AI development. You will also partner with top foundation model labs to provide both technical and strategic input on the development of the next generation of generative AI models. You will: Research and develop novel post-training techniques, including SFT, RLHF, and reward modeling, to enhance LLM core capabilities in both text and multimodal modalities. Design and experiment new approaches to preference optimization. Analyze model behavior, identify weaknesses, and propose solutions for bias mitigation and model robustness. Publish research findings in top-tier AI conferences. Ideally you’d have: Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field. Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning. Excellent written and verbal communication skills Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals Previous experience in a customer facing role. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined du

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $165.6K/yr

Quick readStrong listing-quality and freshness signals

Scale works with the industry's leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation. This role is on the evaluation pod within the GenAI Research Organization and will focus on building benchmarks and diagnosing model failure modes in both text and multimodal modalities. In this role, you will develop rigorous evaluations and diagnostic methods that reveal where frontier models fail and why. You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development. You will also partner with top foundation model labs to translate failure analysis into technical and strategic input on the next generation of generative AI models. You will: Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents. You’ll identify everything from capability gaps and reasoning errors to robustness and alignment issues, all focusing on RCA. Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities. Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them. Publish research findings in top-tier AI conferences. Ideally you’d have: Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field. Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development. Excellent written and verbal communication skills. Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals. Previous experience in a customer facing r

AWSRestMachine LearningAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team The Analytics team is looking for Data Scientists to guide measurement, strategy, and tactical decision-making using Advanced Analytics approaches, as we expand our platform across the globe. Data Scientists at DoorDash work to uncover insights and turn them into actionable recommendations, helping drive decisions for the entire organisation. Analytics is very integral to all operational areas at DoorDash. About the Role Data Science at DoorDash involves diving deeper into our data to solve crucial business problems, ideate & run experiments to solve for insights gleaned from this deep dive and work with a cross-functional team to drive real-world operational change. This is a rare operational and actionable data-driven experience. We solve many exciting challenges from all three sides of our marketplace including customer acquisition, balancing supply and demand, fraud and support, marketing, marketplace efficiency, and more. If you enjoy finding patterns amidst chaos, are excited to build a market from 0 to 1, and have experience using analytics to affect revenue, growth, operations or beyond, we're looking for someone like you! You're excited about this opportunity because you will… As a senior Individual Contributor, mentor and influence junior Data Scientists in investigating complex issues and uncovering key drivers of our business Influence the Product and Operations roadmap by making actionable recommendations based on data Interface frequently with senior leadership to showcase your team’s work and tackle complex business problems Drive measurement strategy for the area under scope, defining success metrics and implementing best practices around experiment design and statistical analysis Develop a strategic learning roadmap based on data observations, strategic questions, and hypotheses We're excited about you because you have… A degree in Math, Physics, Statistics, Economics, Computer Science, or a similar domain 8+ years of experi

AWSGitRestAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $302.4K/yr

Quick readStrong listing-quality and freshness signals

About Scale At Scale AI, our mission is to accelerate the development of AI applications. For 8 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including: generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI), and building upon our prior model evaluation work with enterprise customers and governments, to deepen our capabilities and offerings for both public and private evaluations. About the ACE team The Agent Capabilities & Environments (ACE) team, part of Scale’s Research organization, brings together customer-facing Researchers and Applied AI Engineers. Our core mission includes research on agent environments and RL reward signals, benchmarking autonomous agent performance across real-world scenarios and environments, creating robust data programs to improve Large Language Models (LLMs) agentic capabilities and building foundational tools and frameworks for evaluating models as agents. ACE focuses on autonomous agents that dynamically interact with diverse external environments, including code repositories, GUI interfaces, browsers, and more. About This Role This role is at the intersection of cutting-edge AI research and practical application, with a focus on studying the data types essential for building state-of-the-art agents, such as browser and SWE agents. The ideal candidate will explore the data landscape needed to advance intelligent, adaptable AI agents, guiding the data strategy at Scale to drive innovation. This position requires not only expertise in LLM agents and planning algorithms but also creativity in addressing novel challenges related to data, interaction, and evaluation. You will contribute to impactful research publications on agents, collaborate with customer researchers, and work alongside the engineering team to translate t

SQLAWSGCPRest
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $189.6K/yr

Quick readStrong listing-quality and freshness signals

Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering MLEs, researchers, data scientists and operators for fast and automatic training and evaluation of LLM's, as well as evaluation of data quality. Scale is uniquely positioned at the heart of the field of AI as an indispensable provider of training and evaluation data and end-to-end solutions for the ML lifecycle. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development. You will be building and optimizing the platform to enable our next generation of LLM training, inference and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you! You will: Build, profile and optimize our training and inference framework Collaborate with ML teams to accelerate their research and development and enable them to develop the next generation of models and data curation Research and integrate state-of-the-art technologies to optimize our ML system Ideally you’d have: Strong excitement about system optimization Experience with multi-node LLM training and inference Experience with developing large-scale distributed ML systems Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc. Strong written and verbal communication skills and the ability to operate in a cross functional team environment Nice to haves: Demonstrated expertise in post-training methods &/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the positi

AWSRestAIGo
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $290.4K/yr

Quick readStrong listing-quality and freshness signals

Scale's LLM post-training platform team builds our internal distributed framework for large language model training. The platform powers MLEs, researchers, data scientists, and operators for fast and automatic training and evaluation of LLMs. It also serves as the underlying training framework for the data quality evaluation pipeline. Scale is uniquely positioned at the heart of the field of AI as an indispensable provider of training and evaluation data and end-to-end solutions for the ML lifecycle. You will work closely with Scale’s ML teams and researchers to build the foundation platform which supports all our ML research and development works. You will be building and optimizing the platform to enable our next generation LLM training, inference and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you! You will: Build, profile and optimize our training and inference framework. Collaborate with ML and research teams to accelerate their research and development, and enable them to develop the next generation of models and data curation. Research and integrate state-of-the-art technologies to optimize our ML system. Ideally you’d have: Passionate about system optimization Experience with multi-node LLM training and inference Experience with developing large-scale distributed ML systems Experience with post-training methods like RLHF/RLVR and related algorithms like PPO/GRPO etc. Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc. Strong written and verbal communication skills to operate in a cross functional team environment. Nice to haves: Demonstrated expertise in post-training methods and/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc. Compensation packages at Scale for eligible roles include base salary, equity,

AWSRestAIGo
DU
📍 San Francisco, Canada
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team The Customer Experience team serves as a foundational operations pillar at DoorDash, dedicated to resolving friction within the last mile. We architect and oversee an expansive global network of support centers — spanning both teammate-assisted and AI-driven support — obsessing over the user journey to ensure every interaction is seamless and reliable. As the analytics team, our mission is to make every support interaction measurably better: we define what a great resolution looks like, quantify where we fall short, and turn that into a roadmap for product, operations, and AI/ML partners. We are looking for a Manager to lead and grow the analytics team behind our core support experience. About the Role As a Manager on the Customer Experience Analytics team, you'll set the analytical vision for how DoorDash measures and improves customer resolutions across our global network of support teammates and their interactions with our customers. You'll lead and grow a team of data scientists working at the intersection of customer experience quality and operational cost — uncovering opportunities to drive perfect interactions and informing improvements to teammate tooling that leverages AI-driven resolutions. You'll establish a clear measurement framework for resolution quality, own insights to drive strategy and roadmap, and align partners across CX, Product, Engineering, Operations, and AI/ML. This is a high-visibility leadership role: success means better outcomes for customers, a more effective support organization, and a team of data scientists who are growing in their craft. You're excited about this opportunity because you will… Lead, grow, and develop a team of data scientists — providing mentorship, feedback, and clear career development pathways. Set the analytical vision for the core support experience, defining what a great customer resolution looks like and building the metrics to measure it. Uncover opportunities to drive perfect interactions, tr

AIExcelLogisticsRecruitment
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team DoorDash Ads will become the most transparent and effective advertising channel for merchants, brands, and ad buyers/agencies of all sizes to market their offerings to engaged local audiences. We build a variety of products that are easy to use and confidently generate incremental value for advertisers, while also helping consumers discover and engage with brands they love and save money. As the analytics team our goal is to advance product development, understanding of the business, and identify opportunities for the team to drive towards our north stars. About the Role As the leader of a large, high-performing team of data scientists, you’ll own the analytics strategy for DoorDash Ads for Restaurants—one of the company’s fastest-growing and most dynamic businesses. You’ll guide a team spanning multiple levels of seniority to drive insights and decisions across a multi-sided marketplace, optimizing across consumer and merchant outcomes . In your first few months, you’ll establish clear priorities, align cross-functional partners in product, engineering, sales, and strategy, and set the analytical vision for sustainable growth. Success in this role means delivering measurable business impact, elevating the quality of analytics across the organization, and building systems that balance advertiser ROI, consumer experience, and platform health. You will report into the Senior Director of Analytics on our Ads team in our Ads & Promos organization. You’re excited about this opportunity because you will… Lead and develop a high-performing analytics team, providing mentorship, feedback, and clear career development pathways. Define and drive the analytical and product roadmap by setting measurable goals and aligning cross-functional teams around success metrics. Partner closely with product, engineering, and go-to-market teams to uncover insights, optimize funnels, and inform high-impact decisions. Apply advanced analytical methods—such as co

SQLAWSGitRest
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team Merchant Analytics helps DoorDash make better product, business, and go-to-market decisions through high-quality analytics, predictive modeling, experimentation, and strategic thought partnership. We work across some of DoorDash’s most important merchant and marketplace priorities, building the measurement, insights, and decision frameworks that improve outcomes for merchants and drive company impact. About the Role We’re hiring two Data Science Managers, each to lead a pod within Merchant Analytics and help shape high-priority product and business decisions. In this role, you will lead a team of data scientists, partner closely with Strategy & Operations, Product, Engineering, and business leaders, and turn ambiguous questions into clear recommendations that influence roadmap and plan outcomes. Success in this role means building a high-performing team, raising the quality and speed of decision-making, and ensuring analytics work is tightly connected to measurable business impact. You will report into Director, Data Science on our Merchant Analytics team in our Analytics organization. You’re excited about this opportunity because you will… Lead and develop a team of data scientists responsible for high-impact analytics, predictive modeling, and decision support tied to DoorDash’s most important product, business, and GTM priorities. Partner closely with Strategy & Operations, Product, Engineering, and business leaders to shape decisions, influence roadmaps, and improve plan-critical metrics. Define success metrics, build measurement frameworks and predictive models, and use experimentation and analysis to connect product and operational levers to business outcomes. Build a high-performing pod that balances analytical rigor, strong prioritization, and clear storytelling in a fast-moving environment. Scale reusable analytics frameworks, tools, models, and best practices that make the broader organization more effective over time. We’re excited

AWSGitRestAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team As DoorDash continues to expand rapidly, our Core Consumer team plays a pivotal role in shaping the consumer experience through enhancing personalization, search relevance, merchandising strategy, app quality, and the overall ordering experience. Our mission is to implement scalable data solutions and provide insights that directly influence our strategic product direction. We are looking for an Analytics Senior Manager to lead our Discovery team. About the Role As a Senior Manager of Data Science/Analytics, you’ll be leading a team of data scientists who are working on improving the consumer experience. You will develop strategic insights and work closely with the product, engineering, strategy and operations teams to actively build and measure the impact of new features. You will oversee our metrics and analytics strategy to inform strategic product direction, offer technical leadership and build processes to support velocity, and partner with data and product engineers to build robust data foundations. You're excited about this opportunity because you will… Lead and develop a team of Data Scientists in investigating complex issues and uncovering key drivers of our business, along with your own contributions as an Individual Contributor Influence the Product and Operations roadmap by making actionable recommendations based on data Interface frequently with senior leadership to showcase your team’s work and tackle complex business problems Drive measurement strategy for the area under scope, defining success metrics and implementing best practices around experiment design and statistical analysis Develop a strategic learning roadmap based on data observations, strategic questions, and hypotheses We're excited about you because you have… A degree in Math, Physics, Statistics, Economics, Computer Science, or a similar domain Experience managing a team of data scientists, and a track record of delivering impactful analyses 8+ years of experience i

AWSGitRestAI
🔔

Get new scientist 2 jobs in San Francisco, Canada by email

Daily job updates · Unsubscribe anytime