Jobs in Canada

Data Collector And Annotator in San Francisco

225 active opportunities · Updated October 2026

Explore current data collector and annotator jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $165.6K/yr

Quick readStrong listing-quality and freshness signals

Scale works with the industry’s leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling). This role will focus on optimizing data curation and eval to enhance LLM capabilities in both text and multimodal modalities. In this role, you will develop novel methods to improve the alignment and generalization of large-scale generative models. You will collaborate with researchers and engineers to define best practices in data-driven AI development. You will also partner with top foundation model labs to provide both technical and strategic input on the development of the next generation of generative AI models. You will: Research and develop novel post-training techniques, including SFT, RLHF, and reward modeling, to enhance LLM core capabilities in both text and multimodal modalities. Design and experiment new approaches to preference optimization. Analyze model behavior, identify weaknesses, and propose solutions for bias mitigation and model robustness. Publish research findings in top-tier AI conferences. Ideally you’d have: Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field. Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning. Excellent written and verbal communication skills Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals Previous experience in a customer facing role. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined du

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $165.6K/yr

Quick readStrong listing-quality and freshness signals

Scale works with the industry's leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation. This role is on the evaluation pod within the GenAI Research Organization and will focus on building benchmarks and diagnosing model failure modes in both text and multimodal modalities. In this role, you will develop rigorous evaluations and diagnostic methods that reveal where frontier models fail and why. You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development. You will also partner with top foundation model labs to translate failure analysis into technical and strategic input on the next generation of generative AI models. You will: Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents. You’ll identify everything from capability gaps and reasoning errors to robustness and alignment issues, all focusing on RCA. Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities. Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them. Publish research findings in top-tier AI conferences. Ideally you’d have: Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field. Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development. Excellent written and verbal communication skills. Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals. Previous experience in a customer facing r

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

About Scale At Scale AI, our mission is to accelerate the development of AI applications. For 8 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including: generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI), and building upon our prior model evaluation work with enterprise customers and governments, to deepen our capabilities and offerings for both public and private evaluations. About Data Engine Our Generative AI Data Engine powers the world’s most advanced LLMs and generative models through world-class RLHF (Reinforcement Learning with Human Feedback), human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. Our Approach As part of the interview process, you’ll be considered for opportunities across several teams within the GenAI Engineering organization, based on your interests, expertise, and business needs. Potential team placements include Allocation, Growth, Frontier Data, Trust & Safety, Pay, Operator, or Tasking Experience. Together, these teams power Scale’s AI data operations - from building high-impact datasets that push the boundaries of LLM capabilities, to optimizing contributor onboarding and incentives, to safeguarding data integrity through advanced trust, safety, and security measures. They work at the intersection of ML, operations, and analytics to ensure we deliver the highest-quality data at scale. Responsibilities: Design, build, and maintain robust, scalable systems across the full stack, including front-end, back-end, and infrastructure layers Implement high-impact features using modern technologies such as TypeScript, React, Node.js, MongoDB, Elasticsearch, and Temporal Collaborate closely with internal operators (your use

TypeScriptPythonReactNode.js
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team DoorDash Ads will become the most transparent and effective advertising channel for merchants, brands, and ad buyers/agencies of all sizes to market their offerings to engaged local audiences. We build a variety of products that are easy to use and confidently generate incremental value for advertisers, while also helping consumers discover and engage with brands they love and save money. As the analytics team our goal is to advance product development, understanding of the business, and identify opportunities for the team to drive towards our north stars. About the Role As the leader of a large, high-performing team of data scientists, you’ll own the analytics strategy for DoorDash Ads for Restaurants—one of the company’s fastest-growing and most dynamic businesses. You’ll guide a team spanning multiple levels of seniority to drive insights and decisions across a multi-sided marketplace, optimizing across consumer and merchant outcomes . In your first few months, you’ll establish clear priorities, align cross-functional partners in product, engineering, sales, and strategy, and set the analytical vision for sustainable growth. Success in this role means delivering measurable business impact, elevating the quality of analytics across the organization, and building systems that balance advertiser ROI, consumer experience, and platform health. You will report into the Senior Director of Analytics on our Ads team in our Ads & Promos organization. You’re excited about this opportunity because you will… Lead and develop a high-performing analytics team, providing mentorship, feedback, and clear career development pathways. Define and drive the analytical and product roadmap by setting measurable goals and aligning cross-functional teams around success metrics. Partner closely with product, engineering, and go-to-market teams to uncover insights, optimize funnels, and inform high-impact decisions. Apply advanced analytical methods—such as co

SQLAWSGitRest
SC
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$240K – $270K/yr

Quick readStrong listing-quality and freshness signals

About the Role Sigma is transforming how businesses run by delivering a high performance platform on modern data architecture. Hence, we are growing the engineering team and looking for engineers who are excited to solve challenging problems, deliver impactful capabilities throughout our stack to build world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible for all users. What You Will Be Doing You will be responsible for developing elegant and responsive user experience using the latest front-end technologies. You'll own substantial pieces of the product, from design to launch Working with our product, UX design, and backend development teams, you will develop new features and technologies that make our product experience awesome and radically simplify the user experience for non-technical users You will leverage your technical expertise in front-end application development in the creation of novel visualizations for structured and unstructured data and develop new techniques for improving the performance and interactivity of the application Use modern frontend frameworks like React, GraphQL, TypeScript and Node.js Qualifications We Need 10+ years industry experience building and maintaining high-quality software An eye for great design and a passion for building products that provide a great user experience The ability to make the right trade-offs between functionality and delivery speed that supports delivering value to customers, all the while iterating based on feedback and roadmap priorities Desire to be a great teammate and have fun at work without compromising ownership towards your work Strong sense of craftsmanship, and a healthy academic curiosity to solve challenges at sigma Strong computer science fundamentals Qualifications We Want (also, skills you’ll learn!) Experience building software capabilities for analyzing large scale data web applications Prior ex

TypeScriptPythonReactNode.js
SC
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$170K – $240K/yr

Quick readStrong listing-quality and freshness signals

About the Role Sigma is transforming how businesses run by delivering a high performance platform on modern data architecture. Hence, we are growing the engineering team and looking for engineers who are excited to solve challenging problems, deliver impactful capabilities throughout our stack to build world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible for all users. What You Will Be Doing You will be responsible for developing elegant and responsive user experience using the latest front-end technologies. You'll own substantial pieces of the product, from design to launch Working with our product, UX design, and backend development teams, you will develop new features and technologies that make our product experience awesome and radically simplify the user experience for non-technical users You will leverage your technical expertise in front-end application development in the creation of novel visualizations for structured and unstructured data and develop new techniques for improving the performance and interactivity of the application Use modern frontend frameworks like React, GraphQL, TypeScript and Node.js Qualifications We Need 5+ years industry experience building and maintaining high-quality software An eye for great design and a passion for building products that provide a great user experience The ability to make the right trade-offs between functionality and delivery speed that supports delivering value to customers, all the while iterating based on feedback and roadmap priorities Desire to be a great teammate and have fun at work without compromising ownership towards your work Strong sense of craftsmanship, and a healthy academic curiosity to solve challenges at sigma Strong computer science fundamentals Qualifications We Want (also, skills you’ll learn!) Experience building software capabilities for analyzing large scale data web applications Prior exp

TypeScriptPythonReactNode.js
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Scale Labs, Research Scientist — Agent Robustness As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist working on Agent Robustness you will work on the fundamental challenges of building AI agents that are safe and aligned with humans. For example, you might: Research the science of AI agent capabilities with a focus on how they relate to safety, risk factors, and methodologies for benchmarking them; Design and build harnesses to test AI agents’ tendency to take harmful actions when pressured to do so by users or tricked into doing so by elements of their environment; Design and build exploits and mitigations for new and unique failure modes that arise as AI agents gain affordances like coding, web browsing, and computer use; Characterize and design mitigations for potential failure modes or broader risks of systems involving multiple interacting AI agents. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Practical experience conducting technical research collaboratively. You should be comfortable building and leveraging agent scaffolding, designing evaluation harnesses, an

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Scale Labs, Research Scientist — AI Controls and Monitoring As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist focused on AI Controls and Monitoring, you will design methods, systems, and experiments to ensure that advanced AI models and agents remain aligned with intended goals, even in high-stakes or adversarial environments. For example, you might: Develop monitoring techniques and observability methods that track AI behavior in real time to identify and flag deviations, emergent capabilities, or anomalous outputs; Research mechanisms for layered control, including fail-safes, oversight protocols, and intervention methods that can halt or redirect AI systems when risks are detected; Design red-team simulations to probe weaknesses in oversight and control mechanisms, and build mitigations to close identified gaps; Collaborate with policymakers, engineers, and other researchers to establish standards and benchmarks for AI monitoring and escalation. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Practical experience conducting technical research collaboratively. You should be

AWSRestMachine LearningAI
P
📍 San Francisco, Canada· Full-time· Remote
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Prophecy Prophecy is building the next generation AI-powered data prep and analysis platform. Our platform enables business analysts and data teams to transform raw data into reliable, production-ready datasets and insights faster, using modern data infrastructure and AI-driven capabilities. We work with leading enterprises to simplify how organizations prepare, analyze, and operationalize data, while maintaining strong governance, security, and operational control. Our mission is to make it dramatically easier for organizations to turn complex data into trusted insights that drive decisions. About the Roles We are looking for a Director of Strategic Partnerships who can do both: drive revenue through Prophecy's partner ecosystem, and build an effective partner program. This is an early-stage, high-ownership motion, the playbook is still being written, and you'll have real influence over how we engage partners, what good looks like for partner-sourced pipeline, and how we build durable co-sell relationships with Snowflake, Databricks, GCP field and partner teams. You will sit at the intersection of sales, partnerships, and strategy, owning partner performance while building the programs and processes that scale it. You'll work directly with our AEs and SEs to bring partners into deals at the right moments, and you'll serve as the primary point of contact for our strategic cloud and ecosystem partners. What You’ll Own Partner Revenue & Pipeline Own and exceed partner-sourced and partner-influenced revenue targets on a quarterly basis Proactively generate pipeline through Snowflake, Databricks, and GCP field AEs and PDMs — building the relationships that produce qualified, sourced opportunities Drive joint account mapping and target account activation against Prophecy's ICP: enterprises running Alteryx on Snowflake or Databricks Activate co-sell motions through marketplace programs (GCP Marketplace, Snowflake Partner Network, Databricks

AWSAzureGCPAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $189.6K/yr

Quick readStrong listing-quality and freshness signals

Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering MLEs, researchers, data scientists and operators for fast and automatic training and evaluation of LLM's, as well as evaluation of data quality. Scale is uniquely positioned at the heart of the field of AI as an indispensable provider of training and evaluation data and end-to-end solutions for the ML lifecycle. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development. You will be building and optimizing the platform to enable our next generation of LLM training, inference and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you! You will: Build, profile and optimize our training and inference framework Collaborate with ML teams to accelerate their research and development and enable them to develop the next generation of models and data curation Research and integrate state-of-the-art technologies to optimize our ML system Ideally you’d have: Strong excitement about system optimization Experience with multi-node LLM training and inference Experience with developing large-scale distributed ML systems Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc. Strong written and verbal communication skills and the ability to operate in a cross functional team environment Nice to haves: Demonstrated expertise in post-training methods &/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the positi

AWSRestAIGo
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $302.4K/yr

Quick readStrong listing-quality and freshness signals

About Scale At Scale AI, our mission is to accelerate the development of AI applications. For 8 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including: generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI), and building upon our prior model evaluation work with enterprise customers and governments, to deepen our capabilities and offerings for both public and private evaluations. About the ACE team The Agent Capabilities & Environments (ACE) team, part of Scale’s Research organization, brings together customer-facing Researchers and Applied AI Engineers. Our core mission includes research on agent environments and RL reward signals, benchmarking autonomous agent performance across real-world scenarios and environments, creating robust data programs to improve Large Language Models (LLMs) agentic capabilities and building foundational tools and frameworks for evaluating models as agents. ACE focuses on autonomous agents that dynamically interact with diverse external environments, including code repositories, GUI interfaces, browsers, and more. About This Role This role is at the intersection of cutting-edge AI research and practical application, with a focus on studying the data types essential for building state-of-the-art agents, such as browser and SWE agents. The ideal candidate will explore the data landscape needed to advance intelligent, adaptable AI agents, guiding the data strategy at Scale to drive innovation. This position requires not only expertise in LLM agents and planning algorithms but also creativity in addressing novel challenges related to data, interaction, and evaluation. You will contribute to impactful research publications on agents, collaborate with customer researchers, and work alongside the engineering team to translate t

SQLAWSGCPRest
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Scale Labs, Research Scientist — Safety Post Training As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist working on Safety Post-Training you will develop and apply post-training methods and interpretability techniques to make frontier AI systems safer, and better understood by researchers and policymakers.. For example, you might: Design and run post-training pipelines to study how training choices affect model safety, robustness, and alignment properties; Develop interpretability-informed evaluations that reveal how and why models produce unsafe, deceptive, or otherwise undesirable behaviors, and use those insights to guide targeted mitigations; Collaborate with policymakers, engineers, and other researchers to translate post-training and interpretability findings into actionable safety standards, evaluation benchmarks, and best practices. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar approaches. A track record of published research in machine learning, particularly in generati

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Scale Labs, Research Scientist — Frontier Risk Evaluations As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist focused on Frontier Risk Evaluations, you will design and create evaluation measures, harnesses and datasets for measuring the risks posed by frontier AI systems. For example, you might do any or all of the following: Design and build harnesses to test AI models and systems (including agents) for dangerous capabilities such as security vulnerability exploitation, CBRN uplift, and other high-risk activities; Work with government agencies or other labs to collectively scope and design evaluations to measure and mitigate risks posed by advanced AI systems; Publish evaluation methodologies and write technical reports for policymakers. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Practical experience conducting technical research collaboratively. You should be comfortable building and instrumenting ML pipelines, writing evaluation harnesses, and quickly turning new ideas from the research literature into working prototypes. A track record of published research in m

AWSRestMachine LearningAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team At DoorDash, design means making experiences for the people who order, the people who prepare, and the people who deliver. As a Design Manager at DoorDash, you want to build things that matter to real people. You're at your best when you can move from idea to shipped product quickly, bringing experiences to life that reach and influence users at massive scale. You'll care about whether the product you make solved a real problem for real people, or changed how someone experiences their day. You'll work shoulder-to-shoulder with Engineering and Product Management, dig into the data, and use LLM-powered tools alongside traditional design tools. If the LLM powered tools don’t exist yet, you build them. About the Role Trust is the foundation of every interaction on DoorDash. We're looking for a design leader who can define and drive a vision for designing for integrity at scale. In this role, you will join the Integrity organization. You’ll lead a team of designers working across fraud prevention, trust, safety, and compliance — protecting millions of consumers, Dashers, and merchants while keeping the platform seamless and trustworthy. You will report into the Head of Design for our Customer Experience & Integrity organization. This role is hybrid- 1–2 days per week in one of our Design Hubs. You’re excited about this opportunity because you will… Set the technical direction for Design across your product area — decide what gets built, in what order, and why; connect multiple teams around shared platforms so the work compounds instead of duplicating Work on ambiguous problems and turn them into architecture decisions and working systems that teams actually use; earn trust with leadership not through decks, but through prototypes and shipped code that make your point for you Stay close to the code and the craft across multiple projects at once — you're not just reviewing, you're building; the work you ship will move real metrics across the product area

TypeScriptReactAWSGit
SA
📍 San Francisco, Canada· Hybrid
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Snorkel At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data. We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler! In September 2026 we raised a $350 million Series E at a $3.5 billion valuation , and we are scaling our engineering and research teams to meet demand. The role Frontier AI data is expensive to make and hard to measure. Every task we deliver is tested against the strongest models, often through many long-running agent rollouts. Your job is to make that process faster, cheaper, and more rigorous with ML and AI You will be one of the early members of ML & Research Engineering at Snorkel. You will study how frontier-grade data is generated and evaluated, form hypotheses, validate them against real production data, and ship the winners at scale. You will shape the discipline's direction, its standards, and the team that grows around it. What you'll work on Efficient agentic evals. Cut the cost of long-horizon agent evaluation with adaptive sampling, statistically grounded early stopping, model cascades, caching, and cheap-first gating. AI model routing. Route every eval and judge call to the cheapest model that clears the quality bar, with fallback, monitoring, and cost attribution. Fine-tuned small models. Fine-tune and serve open-weight models (LoRA and other

PythonMachine LearningAI
🔔

Get new data collector and annotator jobs in San Francisco, Canada by email

Daily job updates · Unsubscribe anytime