Jobs in United States

Human Evaluator in United States

1,786 active opportunities · Updated October 2026

Explore current human evaluator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Personal AGI team seeks to empower all of humanity to benefit from frontier intelligence in whatever way they choose. We are responsible for training models to deploy to millions of users globally via ChatGPT, the API, and future products. We aim to evolve ChatGPT from a chatbot to an infinitely capable and personalized superassistant supporting human flourishing. We work on defining, measuring, and improving capabilities across the training stack. Our focus areas include but are not limited to model behavior, personalization, safety, factuality, instruction following, personality, interactivity, multilingual fluency, world interaction, and bringing agents to everyone. We chart the course for what to strive towards. We partner closely with research and product teams across the company ensuring that our models are safe, efficient, and reliable. About the Role You’ll work as a Research Engineer / Scientist on the North Stars team within the broader Personal AGI research org. You will work on bringing the next generation of AI-enabled experiences to all of humanity by closing the capability overhang between power users and the average consumer, including areas like tool-use, feature discovery, connectors, and instruction following. You will think deeply about the current bottlenecks in model behavior, translate these insights into robust evals, training data, reward signals, and model and harness improvements. We're looking for individuals with strong ML engineering skills and research experience passionate about creative, product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve model capability and performance. Collaborate closely with the other research and product teams, allowing customers to optimize their own models. Build robust evaluations for tracking modelin

AWSRestMachine LearningAI
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Personalization-Memory team, within OpenAI's broader Personal AGI organization, is focused on developing agents that can learn from prior interactions in order to become more helpful and efficient over time. We build general-purpose memory and personalization capabilities that transfer across ChatGPT and other agentic products, and we collaborate with applied engineering on the product surfaces that allow users to interact with memory. About the Role As a Research Engineer / Research Scientist on the Personalization-Memory team, you will research and develop improvements to memory usage and personalization in OpenAI's frontier models. Our team works on reinforcement learning, dataset creation, evaluations, and other post-training methods. We partner closely with research and product teams across the company to realize the vision of a truly personalized ChatGPT. We're looking for individuals who have a background in frontier model post-training, are able to iterate quickly, and who are passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda for improving memory use and personalization in frontier models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. Collaborate closely with the research and product teams to influence the shape of technical solutions in the product. You might thrive in this role if you: Are passionate about personalization and building personalized assistants. Have experience working with user signals and human data to turn feedback into reliable signals for training and evaluation. Have a deep understanding of frontier model post-training and machine learning applications. Value principled approaches and research craftsmanship. Are comfortable diving into a lar

AWSRestMachine LearningAI
N
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.6%

$196K – $230K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: We’re seeking an experienced User Researcher to shape growth-related product decisions by delivering insights that fuel AI experiences. This role blends traditional growth research, such as adoption, monetization, and pricing and packaging, with fast-evolving AI features like Notion Agent and Chat that permeate experiences like onboarding and Workspace creation. It requires a unique mix of user research expertise, strategic foresight, and technical fluency in AI and AI product development. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: Design and execute studies that evaluate how people experience Notion with AI in the loop, including how users discover, trust, and get value from AI, and identify barriers to adoption and willingness to pay. Run mixed-methods research (qual + quant), including interviews, concept testing, prototype/usability testing, diary

RestAIGoRust
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -88.6%

From $135K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: Notion is looking for an Early Career Recruiter who wants to reimagine what early career hiring can look like. This is a role for a builder: someone energized by the question of how to find and cultivate exceptional emerging talent, who gets excited about crafting programs and experiences that don't yet exist, and who sees employer brand, candidate experience, and assessment design in an AI era as creative problems worth obsessing over. You'll own Notion's early career strategy end-to-end alongside a team of talented teammates — from how we reach communities to how we evaluate potential, develop talent, and convert interns and new grads into the next generation of Notinos. You'll work closely with our Engineering, Product, Design, Sales, and People teams. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: A differentiated early career engine — design and run

RestAIGoRust
N
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.6%

$98K – $140K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role You'll own the quality bar for Notion AI products. You’ll work with product and engineering teams to build systems to define what “good” looks like, measure our progress, and drive changes to deliver reliable and high-quality AI experiences. Your work directly shapes how Notion's AI products behave for millions of users. This isn't a traditional software engineering role. It’s an art & science role . You won't spend your days writing code. Instead, you'll focus on understanding and shaping how our AI products behave through context engineering, designing evaluation systems, and analyzing data. This team sits in our AI engineering team, working directly with engineering, product, design, and data. This role is a unique blend of ops, strategy, and product thinking. Day to day, you'll live in production data, ship prompt fixes, run evals and, in effect, shape our quality strategy. As part of that you'll shape Notion's model strategy and work directly with frontier AI labs (OpenAI, Anthropic, Google) to evaluate and launch new models. We're looking for problem-seeking generalists interested in 0 → 1 : curious people wi

SQLRestAIGo
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -88.6%

From $152K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role We’re looking for an AI Applications Engineer to help drive Notion’s business transformation efforts. In this role, you’ll be a strategic partner to our internal stakeholders (primarily GTM, Finance and People teams) and deliver and scale creative AI-driven solutions to multi-faceted problems with measurable business impact. You’ll also build reusable components, evaluation patterns, and operational guardrails that make AI delivery repeatable across teams. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You’ll Achieve Work with stakeholders to discover opportunities from ambiguous problem statements, translate them into scoped solutions, and drive iterative releases from idea to adoption Build and ship end‑to‑end AI solutions—from problem framing through data readiness, modeling, evaluation, and production rollout Establish evaluation and production-readiness patterns (met

N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -88.6%

$140K – $195K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: Notion is looking for a Recruiting Program Manager to help design, build, and scale the systems and programs that power how we hire. This is a role for someone who moves fast, takes ownership, and doesn't wait to be told what to do — someone who operates with both strategic altitude and operational precision, and knows how to bring clarity and momentum to complex cross-functional work. We're at an inflection point in how we think about recruiting at Notion, and this role sits at the center of it. You'll partner closely with Recruiters, Recruiting Operations, People Operations, and senior leadership to drive some of our highest-priority initiatives — building the programs and processes that define how Notion finds, evaluates, and brings in exceptional talent. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: The scope of this role will evolve with the team's

N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -12.7%

We're building the platform that lets long-running autonomous agents operate safely inside NVIDIA's enterprise. These are not assistants on a developer's laptop. They are fleets of agents deployed in the cloud, running continuously at scale on shared accelerated compute. They take on real work across enterprise systems, so people get far more done than they could before. This role defines the constructs that agents are built from: the blueprints they start from, the tools, skills, and plugins that power them against enterprise data, the runtime safety harness that keeps them in bounds, and the connections into credential management, sandbox, memory, and observability. The team designs and ships these building blocks so that agent developers across the company can stand up a new agent, wire it in, and run it for days or weeks. Security and safe execution come out of the box, not something each team has to get right on its own. Today an agent runs inside a single harness. Claude, Codex, and open-source agent harnesses each work differently underneath, with their own execution model, tool interface, and telemetry shape. The platform smooths over those differences, so a single skill, safety policy, or trace works the same no matter which harness is running. We want to enable agents that act on a person's behalf, governed and secure, continuously evaluated and self-improving. These agents coordinate and hand work off to each other, with identity and policy following every hop. They route and tune themselves across harnesses from live eval signals, and get better from their own production telemetry instead of waiting on a human to retrain them. Have you run agents on a harness like Claude or Codex and hit the walls that show up when they run for real, for days, against live systems — and wanted them to learn from it on their own? We're building the platform that solves those problems once, for every team. What you'll b

PythonAIFinanceHR
MT
📍 Boise, ID - Main Site, United States
✓ Quality checkedCompany trend +1216.7%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As an AI Reimagination Engineer in Micron's Generative AI Center of Excellence (GenAI COE), you have an outstanding chance to transform how businesses function. You will partner with different business areas to break down large, manual, multi-step processes and redesign them into effective, autonomous systems. This position combines process reimagination with AI systems engineering, making you a key part of the transformation journey. You will collaborate closely with the GenAI COE, IT architecture, security, and project teams. You will lead projects from the initial redesign to the final build, delivering solutions that business teams can adopt and scale. Responsibilities: Decompose end-to-end processes: Map current-state flows, quantify effort and risk, and lead eliminate/simplify/agentify analysis before automation. Architect the agentic solution: Build future-state flows and AI architecture, including task and agent decomposition, orchestration patterns, tool and data access, memory and context strategy, and human-in-the-loop controls. Translate inventions into buildable solutions by developing agent workflows, composing prompt and context strategies, MCP/connector and integration requirements, and evaluation criteria. Follow Micron's “Secure by Design”

PythonSQLAIC#
MT
📍 Boise, ID - Main Site, United States
✓ Quality checkedCompany trend +1216.7%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As an AI Reimagination Engineer in Micron's Generative AI Center of Excellence (GenAI COE), you have an outstanding chance to transform how businesses function. You will partner with different business areas to break down large, manual, multi-step processes and redesign them into effective, autonomous systems. This position combines process reimagination with AI systems engineering, making you a key part of the transformation journey. You will collaborate closely with the GenAI COE, IT architecture, security, and project teams. You will lead projects from the initial redesign to the final build, delivering solutions that business teams can adopt and scale. Responsibilities: Decompose end-to-end processes: Map current-state flows, quantify effort and risk, and lead eliminate/simplify/agentify analysis before automation. Architect the agentic solution: Build future-state flows and AI architecture, including task and agent decomposition, orchestration patterns, tool and data access, memory and context strategy, and human-in-the-loop controls. Translate inventions into buildable solutions by developing agent workflows, composing prompt and context strategies, MCP/connector and integration requirements, and evaluation criteria. Follow Micron's “Secure by Design”

PythonSQLAIC#
MT
📍 Boise, ID - Main Site, United States
✓ Quality checkedCompany trend +1216.7%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As an AI Reimagination Engineer in Micron's Generative AI Center of Excellence (GenAI COE), you have an outstanding chance to transform how businesses function. You will partner with different business areas to break down large, manual, multi-step processes and redesign them into effective, autonomous systems. This position combines process reimagination with AI systems engineering, making you a key part of the transformation journey. You will collaborate closely with the GenAI COE, IT architecture, security, and project teams. You will lead projects from the initial redesign to the final build, delivering solutions that business teams can adopt and scale. Responsibilities: Decompose end-to-end processes: Map current-state flows, quantify effort and risk, and lead eliminate/simplify/agentify analysis before automation. Architect the agentic solution: Build future-state flows and AI architecture, including task and agent decomposition, orchestration patterns, tool and data access, memory and context strategy, and human-in-the-loop controls. Translate inventions into buildable solutions by developing agent workflows, composing prompt and context strategies, MCP/connector and integration requirements, and evaluation criteria. Follow Micron's “Secure by Design”

PythonSQLAIC#
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -83.9%

$230K – $325K/yr

Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Safety teams work to ensure our products are safe, trusted, and resilient as frontier AI systems scale globally. We tackle some of the company’s most important challenges across understanding and preventing misuse and misalignment, intercepting fraud and abuse, and protecting vulnerable users. We are hiring Data Scientists to help build the analytical foundations that allow OpenAI to deploy increasingly capable AI responsibly. We are hiring Data Scientists across several teams that contribute to safety in different ways, including: Safety Systems Integrity Product Policy This is a high-impact role operating at the intersection of product, safety, policy, and research. About the Role As a Data Scientist, Safety, you will help solve complex and ambiguous problems where rigorous analysis directly informs critical decisions. Depending on your background and team alignment, you may work on areas such as: Measure harmful or abusive behavior across OpenAI’s products Detect fraud, manipulation, coordinated misuse Evaluate and improve safety classifiers, rules systems, mitigation systems, and human review workflows Design experiments and causal analyses to understand product, policy, and mitigation impacts Build prevalence estimators, dashboards, monitoring systems, and executive decision frameworks Diagnose gaps in safety and integrity systems using behavioral and product data, and help quantify and navigate false positive / false negative tradeoffs Translate ambiguous safety risks into measurable problems and evidence-based recommendations Partner with Product, Engineering, Policy, Research, and Operations teams to improve safety outcomes Build zero-to-one analytical systems in rapidly evolving domains Ideal Candidate We’re looking for strong Data Scientists who thrive in ambiguous, high-leverage environments. You may be a fit if you have: Strong statistical reasoning and analytical judgment Experience with experimentation, causal inference, or obse

PythonSQLAWSRest
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Personal AGI team is responsible for training and improving pre-trained models to be deployed into ChatGPT, the API, and potential future products. The team partners closely with research and product teams across the company, and conducts research as a final step to prepare for real world deployment to millions of users, ensuring that our models are safe, efficient, and reliable. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models. Our team works in research areas combining reinforcement learning and products. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve model capability and performance. Collaborate closely with the other research and product teams, allowing customers to optimize their own models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. You might thrive in this role if you: Have a deep understanding of machine learning and machine learning applications. Have a working knowledge of relevant models, and building evaluations for model capability improvement. Are comfortable diving into a large ML codebase to debug. Thrive in a dynamic and technically complex environment. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and

AWSRestMachine LearningAI
H
📍 Indiana, United States· Remote
✓ Quality checkedCompany trend +310%

Become a part of our caring community The Transition Coordinator (Care Coach 2) evaluates member's needs and requirements. This evaluation aims to achieve and/or maintain an optimal wellness state. The Coordinator does this by guiding members/families toward resources and facilitating interaction with them. These resources are appropriate for the care and wellbeing of members. The Care Coach 2 work assignments are varied and frequently require interpretation and independent determination of the appropriate courses of action. Position Responsibilities: Support the ongoing member transitions in and out of the Indiana Medicaid programs, the Contractor's enrollment, and among care settings. Complete transitions and assists with the planning and preparation for them, and the follow-up care after. Works with the Member Advocate Coordinator and other member-focused departments of the plan. This collaboration ensures continuity and coordination of care and member and provider communication through the initial transition, ongoing benefit plan, and MCE transfers. Ensure the transfer and receipt of all outstanding prior authorization decisions, utilization management data, and clinical information such as prevention and wellness programs(s), care management and complex case management notes. Help with transitions from the custodial setting to the home and community-based setting. We ask that you have telephonic and in-person meetings within an assigned region. The purpose of these meetings is to work with various stakeholders, including long-term care members, hospital/rehab staff discharge planners, family members/POA's, PCP's, and other healthcare professionals. The ultimate goal is to prevent custodial placements whenever possible. Assess and evaluate member's needs to establish a member specific car

Recruitment
H
📍 Tril Ft Myers, United States
✓ Quality checkedCompany trend +310%

Become a part of our caring community The Care Coach 1 assesses and evaluates member's needs and requirements to achieve and/or maintain optimal wellness state by guiding members/families toward and facilitate interaction with resources appropriate for the care and wellbeing of members. The Care Coach 1 work assignments are often straightforward and of moderate complexity. Reports to the Regional Care Coach Manager. Looking for motivated Care Coach in COLLIER county FLORIDA!! We are looking for dynamic case managers that enjoy making a difference in the lives of others! You must live in Collier county in Florida. This rewarding role allows you to spend time connecting with our members to ensure they receive the services they need. The Care Coach 1 employs a variety of strategies, approaches and techniques to support a member's optimal wellness state by coordinating services & resources. Identifies and resolves barriers that hinder effective care. Ensures patient is progressing towards desired outcomes by continuously monitoring patient care through use of assessment, data, conversations with member, and active care planning. Understands own work area professional concepts/standards, regulations, strategies and operating standards. Work is managed and often guided by precedent and/or documented procedures/regulations/professional standards with some interpretation. The Care Coach 1 Visit Medicaid members in their homes, Assisted Living Facilities, and/or Long Term Care Facilities and other care settings – 75-90% local travel Assesses and evaluates member's needs and requirements in order to establish a member specific care plan Ensures members are receiving services in the least restrictive setting in order to achieve and/or maintain optimal well-being Planning and implementing interven

Recruitment
🔔

Get new human evaluator jobs in United States by email

Daily job updates · Unsubscribe anytime