Jobs in Canada

Human Evaluator in San Francisco

89 active opportunities · Updated October 2026

Explore current human evaluator jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $252K/yr

Quick readStrong listing-quality and freshness signals

Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human evaluation and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Platform Engineering team. In this role, you will lead the design and development of core data storage, streaming, caching, and indexing platforms and underlying systems. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the architecture, design, implementation, and reliability of our foundational data platforms and systems, working closely with stakeholders and internal customers to understand and refine requirements. Collaborate with cross-functional teams to define, design, and deliver new features. Proactively identify opportunities for, and driving improvements to, current programming practices, including process enhancements and tool upgrades. Present technical information to teams and stakeholders, providing

MongoDBRedisAWSKubernetes
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Platform Engineering team. In this role, you will support the design and development of shared platforms used across Scale. This includes designing our foundational data platforms and lifecycle, architecting Scale’s core cloud infrastructure and orchestration stack, and redefining how engineers develop, build, test, and deploy software at Scale. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the design, and implementation of our foundational platforms and systems, working closely with stakeholders and internal customers to understand and refine requirements. Collaborating with cross-functional teams to define, design, and deliver new features. Proactively identifying opportunities for, and driving improvements to, current p

SQLMongoDBAWSDocker
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Identity Engineering team. In this role, you will help support the design and development of core software systems specifically focused on identity, access management, authorization, and authentication. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the design, and implementation of our identity infrastructure to ensure secure authentication and authorization across enterprise systems. Build software for authentication mechanisms such as Single Sign-On (SSO), Multi-Factor Authentication (MFA), and federated identity solutions (SAML, OAuth, OpenID Connect). Build software for authorization mechanisms such as Relation-based access control (ReBAC), Attribute-based access control (ABAC), Role-based access cont

PythonJavaNode.jsAWS
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Identity Engineering team. In this role, you will help support the design and development of core software systems specifically focused on identity, access management, authorization, and authentication. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the design, and implementation of our identity infrastructure to ensure secure authentication and authorization across enterprise systems. Build software for authentication mechanisms such as Single Sign-On (SSO), Multi-Factor Authentication (MFA), and federated identity solutions (SAML, OAuth, OpenID Connect). Build software for authorization mechanisms such as Relation-based access control (ReBAC), Attribute-based access control (ABAC), Role-based access cont

PythonJavaNode.jsAWS
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Platform Engineering team. In this role, you will support the design and development of shared platforms used across Scale. This includes designing our foundational data platforms and lifecycle, architecting Scale’s core cloud infrastructure and orchestration stack, and redefining how engineers develop, build, test, and deploy software at Scale. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the design, and implementation of our foundational platforms and systems, working closely with stakeholders and internal customers to understand and refine requirements. Collaborating with cross-functional teams to define, design, and deliver new features. Proactively identifying opportunities for, and driving improvements to, current p

SQLMongoDBAWSDocker
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $252K/yr

Quick readStrong listing-quality and freshness signals

About Scale AI At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. Scale Frontier Data is the organization behind the training and evaluation data that frontier labs depend on. We build the systems, tooling, and expert workflows that turn hard human expertise into signals that models can learn from, across reasoning, coding, agentic tool use, and domain expertise. Reinforcement learning environments are now the center of gravity for that work: the difference between a model that demos well and a model that reliably completes long-horizon work is almost always the quality of the environments and reward signals it was trained against. Responsibilities As a Staff Software Engineer, RL Environments, you'll own the technical foundation for how Scale builds, runs, verifies, and delivers RL environments at scale. An RL environment is a real piece of software: a containerized world with real dependencies, real state, real tools, and a grader that has to be correct even when the agent is creative about breaking it. Building one is a full-stack engineering problem. Building thousands of them reproducibly, cheaply, with trustworthy reward signals and throughput measured in millions of rollouts is a systems problem that very few people have solved. You'll work on both. You'll design the platform: sandboxed execution, environment packaging and versioning, rollout orchestration, trajectory capture, verifier frameworks, and the authoring surfaces that let engineers and domain experts produce environments without reinventing infrastructure each time. And you'll go deep on the environments themselves by instrumenting real applications, designing task suites that expose specific capability gaps, and building graders that

TypeScriptPythonReactAWS
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $179.4K/yr

Quick readStrong listing-quality and freshness signals

About Scale AI At Scale AI, our mission is to accelerate the development of AI applications. For 8 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI) and building upon our prior model evaluation work with enterprise customers and governments to deepen our capabilities and offerings for public and private evaluations. About Data Engine Our Generative AI Data Engine powers the world’s most advanced LLMs and generative models through world-class RLHF (Reinforcement Learning with Human Feedback), human data generation, model evaluation, safety, and alignment. The data we produce is some of the most critical work for how humanity will interact with AI. About Our FDE Team Generating high-quality data is the core problem our business solves. We aim to make producing and delivering high-quality data seamless and efficient for operators and customers. Our Team is building customer and operator-specific infrastructure to provide high-quality data with low turnaround time. You'll be exposed to the cutting edge of the Generative AI industry while directly interfacing with the leading model-building organizations in the space, including the top AI research labs and government agencies. Join us in shaping the future of Artificial General Intelligence. As a Forward Deployed Engineer, you'll be at the forefront of providing the critical data infrastructure that powers the most advanced AI models, directly influencing how humanity interacts with AI. You will work with the world’s leading AI companies and government agencies to solve their most complex AI data-related problems. Responsibilities: Drive Impact: Directly contribute to the advancement of AI by delivering critical data solutions for leading AI innovators and

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. Scale Frontier Data is the organization behind the training and evaluation data that frontier labs depend on. We build the systems, tooling, and expert workflows that turn hard human expertise into signals that models can learn from, across reasoning, coding, agentic tool use, and domain expertise. About our Customer Platform team: Our Customer Platform Team plays a pivotal role in integrating our platform with external systems and ensuring seamless, reliable connectivity for both internal users and customers. As the leader of this team, you’ll drive the strategy, architecture, and development of our connectivity solutions, focusing on API integration, distributed systems, and a robust data platform. Your role will be crucial in maintaining and enhancing our platform’s ability to meet the needs of both our internal and external stakeholders. Responsibilities: Own large areas within our product Comfortable working cross functionally, whether that be internal or external customers Build features end-to-end: front-end, back-end, system design, debugging and testing Deliver experiments at a high velocity and level of quality to engage our customers Work across the entire product lifecycle from conceptualization through production Influence the culture, values, and processes of a growing engineering team Inspire and mentor less experienced engineers Collaborating with cross-functional teams to define, design, and ship new product features and experiences. Requirements: At least 7-10 years of relevant experience is preferred Track record of shipping high-quality products and features at scale Desire to work in a very fast-paced environment Abil

AWSRestAIGo
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

About Scale At Scale AI, our mission is to accelerate the development of AI applications. For 8 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including: generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI), and building upon our prior model evaluation work with enterprise customers and governments, to deepen our capabilities and offerings for both public and private evaluations. About Data Engine Our Generative AI Data Engine powers the world’s most advanced LLMs and generative models through world-class RLHF (Reinforcement Learning with Human Feedback), human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. Our Approach As part of the interview process, you’ll be considered for opportunities across several teams within the GenAI Engineering organization, based on your interests, expertise, and business needs. Potential team placements include Allocation, Growth, Frontier Data, Trust & Safety, Pay, Operator, or Tasking Experience. Together, these teams power Scale’s AI data operations - from building high-impact datasets that push the boundaries of LLM capabilities, to optimizing contributor onboarding and incentives, to safeguarding data integrity through advanced trust, safety, and security measures. They work at the intersection of ML, operations, and analytics to ensure we deliver the highest-quality data at scale. Responsibilities: Design, build, and maintain robust, scalable systems across the full stack, including front-end, back-end, and infrastructure layers Implement high-impact features using modern technologies such as TypeScript, React, Node.js, MongoDB, Elasticsearch, and Temporal Collaborate closely with internal operators (your use

TypeScriptPythonReactNode.js
SF
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $225K/yr

Quick readStrong listing-quality and freshness signals

About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role The Client Experience Product Algorithms team is responsible for the machine learning, AI, experimentation, and product analytics capabilities that power personalized experiences for Stitch Fix clients and stylists. Partnering across Product, Engineering, Design, Styling, Marketing, Merchandising, Finance, Enterprise Analytics, Data Platform, and DSN, the team translates data and algorithms into measurable business impact. As Director, Product Algorithms, you will lead the strategy, execution, and people behind our Growth, Styling, and Fix & Freestyle Algorithms portfolios. You'll define how AI, machine learning, experimentation, and analytics shape the future of personalized shopping while building the operating discipline, technical excellence, and cross-functional alignment needed to deliver scalable business results. Responsibilities Lead the Product Algorithms portfolio across Growth, Styling, and Fix & Freestyle, setting strategy and driving measurable outcomes across acquisition, engagement, retention, styling quality, Fix, Freestyle, outfitting, and related commerce experiences. Define the vision and roadmap for applying data science, machine learning, AI, experimentation, and product analytics to improve client experiences, stylist effectiveness, and business performance. Drive innovation by identifying, evaluating, and scaling modern AI, machine learning, personalization, and experimentation techniques that create meaningful impact while balancing technical feasibility, exec

GitRestMachine LearningAI
HI
17 days ago
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$114K – $200K/yr

Quick readStrong listing-quality and freshness signals

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About the Role We are seeking an experienced IT Engineer to own and enhance the technology experience within HP IQ. This role combines traditional IT operations, SaaS administration, endpoint management, and infrastructure support with a strong focus on automation, scalability, and employee experience. The ideal candidate is a hands-on problem solver who thrives in fast-paced startup environments, takes ownership of systems and processes, and proactively identifies opportunities to improve tooling, security, and operational efficiency. You will serve as a technical escalation point for the IT Support team while helping shape the future of our internal technology stack. What you might do Manage and administer core business platforms including Google Workspace, Microsoft 365, Okta, Slack, and other SaaS applications. Ensure employees have a seamless and productive technology experience across all systems and devices. Evaluate, implement, and optimize tools and workflows that improve employee productivity and reduce operational friction. Develop and maintain IT documentation, standards, processes, and

PythonRedisKubernetesAI
HI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$140K – $225K/yr

Quick readStrong listing-quality and freshness signals

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As the Senior Software Engineer, Tooling and Development Infrastructure, you will play a critical role in shaping the developer productivity tools and automated testing strategy. You’ll collaborate closely with design, development, and quality teams to plan, design, and implement robust automated tools and services that ensure the quality and reliability of our AI software stack. You will be highly hands-on in your work and collaborate closely with stakeholders. This position offers a unique opportunity to influence the development of cutting-edge automation frameworks, foster a culture of quality, and contribute to the long-term success of the organization. What You Might Do Develop and implement automation frameworks and testing strategies that cover the entire software stack, from backend systems to user-facing features. Identify, evaluate, and integrate new tools that streamline development. This includes everything from code quality tools and to Infrastructure-as-Code (IaC) solutions. Lead continuous improvement efforts for our build, release, and test systems, ensuring a robust

PythonRedisCI/CDGit
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $252K/yr

Quick readStrong listing-quality and freshness signals

About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. Role Overview As a Forward Deployed AI Engineering Manager on our Enterprise team, you'll be the technical bridge between Scale AI's cutting-edge AI capabilities and our most strategic customers. You'll work with enterprise clients to understand their unique challenges, lead a team that architects specific AI solutions, and ensure successful deployment and adoption of AI systems in production environments. This is a Management role that combines deep engineering and AI expertise, leading a team, and working on customer-facing problems. You'll work directly with customer engineering teams to integrate AI into their critical workflows. Key Responsibilities Customer Integration & Deployment Partner directly with enterprise customers to understand their technical infrastructure, data pipelines, and business requirements Design and implement custom integrations between Scale AI's platform and customer data environments (cloud platforms, data warehouses, internal APIs) Build robust data connectors and ETL pipelines to ingest, process, and prepare customer data for AI workflows Deploy and configure AI models and agents within customer security and compliance boundaries AI Agent Development Develop production-grade AI agents tailored to customer use cases across domains like customer support, data analysis, content generation, and workflow automation Architect multi-agent systems that orchestrate between different models, tools, and data sources Implement evaluation frameworks to measure agent performance and iterate toward business objectives Design human-in-the-loop workflows and feedback mechanisms for continuous agent improvement Prompt Engineeri

PythonAWSAzureGCP
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $288K/yr

Quick readStrong listing-quality and freshness signals

About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. Role Overview As a Senior Staff Frontier Agents Engineer on our Enterprise team, you'll be the technical bridge between Scale AI's cutting-edge AI capabilities and our most strategic customers. You'll work with enterprise clients to understand their unique challenges, architect custom AI solutions, and ensure successful deployment and adoption of AI systems in production environments. This is a hands-on technical role that combines deep engineering expertise with customer-facing problem solving. You'll work directly with customer engineering teams to integrate AI into their critical workflows. Key Responsibilities Customer Integration & Deployment Partner directly with enterprise customers to understand their technical infrastructure, data pipelines, and business requirements Design and implement custom integrations between Scale AI's platform and customer data environments (cloud platforms, data warehouses, internal APIs) Build robust data connectors and ETL pipelines to ingest, process, and prepare customer data for AI workflows Deploy and configure AI models and agents within customer security and compliance boundaries AI Agent Development Develop production-grade AI agents tailored to customer use cases across domains like customer support, data analysis, content generation, and workflow automation Architect multi-agent systems that orchestrate between different models, tools, and data sources Implement evaluation frameworks to measure agent performance and iterate toward business objectives Design human-in-the-loop workflows and feedback mechanisms for continuous agent improvement Prompt Engineering & Optimization Create sophisticate

PythonAWSAzureGCP
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with the world's leading enterprises and government organizations to accelerate their AI transformation through frontier AI systems that solve real business problems. Every day, we work with organizations across finance, healthcare, manufacturing, media and telecommunications to build production AI agents that automate complex workflows, help humans, reason over enterprise knowledge, and operate safely at scale. The Opportunity Applied AI is moving faster than ever. New foundation models, reasoning techniques, agent architectures, and research papers emerge every week. Yet building AI systems that reliably solve real-world problems remains one of the hardest engineering challenges. As a Frontier Agent Engineer (Applied AI) , you'll bridge the gap between cutting-edge AI research and production deployment. You'll work directly with enterprise customers to design, evaluate, and deploy intelligent systems that combine frontier models with structured knowledge, retrieval, traditional machine learning, and enterprise software. Unlike traditional ML roles that focus on a single model or product, you'll work across a diverse portfolio of AI challenges spanning multiple industries and use cases. You may build a multi-agent research system and then participate in designing a customer intelligence platform, a healthcare copilot, or an autonomous workflow for a Fortune 100 company. If you enjoy reading new AI papers, experimenting with the latest models, and shipping production systems that create measurable business impact, you'll fit right in. What You'll Build Frontier AI Systems Design and deploy production AI agents that leverage the latest advances in large language models, reasoning, retrieval, memory, and tool use. Architect intelligent systems that combine LLMs, traditional machine learning, structured knowledge, enterprise data

PythonAWSAzureGCP
🔔

Get new human evaluator jobs in San Francisco, Canada by email

Daily job updates · Unsubscribe anytime