Jobiba hiring network

Human Evaluator Jobs

3,920 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current human evaluator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

R
Roblox
📍 United States• Full-time
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Role Overview We are searching for a self-motivated individual who is passionate about Roblox and thrives in detail. You will be responsible for classifying content based on a specific set of guidelines to help us improve Roblox systems. As a Human Evaluator you will have the opportunity to provide us direct feedback on the performance of various products across the company, but primarily for features within the studio and creator store. You will: Classify content based on a set of instructions; Content classification includes the review and classification of text, image, video, scripts, and audio Track and document insights and trends related to annotation projects Test out new features within studio and provide detailed feedback Become an expert in a variety of topics to enable more accurate evaluations Dedicate between 25-29 hours per week with a work schedule from Monday to Friday You have: At least 1+ years of experience developing on Roblox or other creation platforms. Solid knowledge of the technical aspects of Roblox Studio. Proficiency with Lua Strong verbal and written communication skills Demonstrated patience for repetitive tasks and attention to detail Effective time mana

awsgitai
View job →
F
Figma
📍 Ca New York• Full-time• From $258K/yr
1mo ago

Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! The Figma Research team is hiring an Director, Research - AI Evals to own how we measure the quality of Figma's AI-powered experiences. As Figma ships more AI capabilities across our products, the question "is this really good?" has never mattered more — and answering it rigorously is what this role exists to do. You'll define what "good" means for our AI features, build the frameworks and quality bars to measure it, and turn that into trusted signal that product teams rely on to decide what to ship. The ideal candidate brings deep, hands-on experience evaluating AI/LLM-powered products — blending human evaluation with automated, model-based approaches — along with the product instinct and communication skills to make evaluation genuinely useful. Partnering with Product, Design, Engineering, and Data Science, you'll sit upstream of nearly every AI shipping decision at Figma and directly shape the quality of features used by millions of people. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: Own AI evaluation methods and operations for Figma's AI-powered experiences — define quality dimensions, design how we measure them, and turn results into decision-ready signal Build and maintain evaluation frameworks, rubrics, golden datasets, and quality bars, combining human evaluation with automated/model-based approaches (e.g., LLM-as-judge) where appropriate Partner with engineering to stand up repeatable, reproducible evaluati

SA
Scale AI
📍 San Francisco• Full-time• From $252K/yr
17 days ago

Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human evaluation and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Platform Engineering team. In this role, you will lead the design and development of core data storage, streaming, caching, and indexing platforms and underlying systems. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the architecture, design, implementation, and reliability of our foundational data platforms and systems, working closely with stakeholders and internal customers to understand and refine requirements. Collaborate with cross-functional teams to define, design, and deliver new features. Proactively identify opportunities for, and driving improvements to, current programming practices, including process enhancements and tool upgrades. Present technical information to teams and stakeholders, providing

mongodbredisaws
View job →
H
Hp
📍 Texas• $105.1K – $161.8K/yr
1mo ago

Human Factors Engineer Description - This role is responsible for shaping the usability and overall user experience of physical products or systems. The role requires a strong blend of human factors expertise, research rigor, and user empathy. The role collaborates closely with cross-functional teams, including designers, product managers, engineers, and stakeholders, to create user-centered design solutions that meet both user needs and business objectives. The role conducts thorough user research to uncover insights, behaviors, and pain points, translating findings into actionable design solutions. Responsibilities Plans, designs, and conducts usability studies and human factors evaluations across multiple stages of product development. Executes formative and summative usability testing, including protocol development, participant recruitment, data collection, and analysis. Leads human factors validation activities in alignment with relevant standards. Applies quantitative and qualitative research methods to generate actionable insights. Analyzes research data and translates findings into clear, practical design recommendations. Prepares and delivers research reports and presentations for cross-functional stakeholders. Collaborates with industrial designers, engineers, and product managers to integrate user insights into product development. Benchmarks product usability and experience against key competitors. Identifies and surfaces usability risks early in development to reduce downstream cost and rework. Supports continuous improvement of research processes, tools, and best practices. Education & Experience Recommended Four-year or Graduate Degree in Design, Human Fac

recruitment
View job →
P
16 days ago

Human Data Quality Engineer (Founding Team) Prolific Prolific isn’t just enabling AI innovation – we’re redefining it. While foundational AI technologies are becoming commoditized, Prolific’s human data infrastructure provides the high-quality, diverse data required to train the next generation of AI models. Through our platform, we empower researchers and companies to access a global, ethically curated participant base, ensuring cutting-edge AI research and training grounded in inclusivity and precision. The Role As one of the founding members of Prolific's newly formed AI Data Services team, you'll help build the quality systems behind some of the world's most advanced AI models. Data quality is a strategic priority for Prolific, so this is a high-visibility role with direct exposure to senior stakeholders. This isn't a traditional QA role. We are not looking for someone to review data against a predefined checklist. We are looking for an innovative thinker that can leverage their expertise to define what good means where no definition exists yet. Acting as a strategic thought partner, you’ll work at the intersection of human data, machine learning, evaluation across frontier use-cases that define what high-quality human data looks like for the next generation of advanced AI. This means that much of the work involves novel problems with no established answer, so you’ll be comfortable working through ambiguity.. . Your primary focus is working directly with clients and alongside frontier AI labs, translating what their models need into robust human data and evaluation strategies. Rather than checking quality at the end of a project, you'll engineer quality into every stage of the lifecycle, from study design and participant strategy through to evaluation, launch readiness and client delivery.You will also work alongside our product engineering, and supply teams to define and build the quality infrastructure that will enable us to deliver h

pythonsqlmachine learning
View job →
N
Notion
📍 San Francisco• Full-time• $196K – $230K/yr
1mo ago

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered experiences—focusing on what “good” looks like not only for model output quality, but for the end-to-end product experience where people discover, set goals, delegate work, review results, and build trust over time with AI. This role sits at the intersection of research craft and evaluation operations: you’ll run studies that uncover user mental models, expectations, and failure/recovery behaviors, then translate those insights into reusable rubrics, workflows, and measurement approaches that product, design, engineering, and data science can apply consistently. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: Define what “good” looks like (frameworks & rubrics): Establish clear, reusable evaluation criteria that reflect real user expectations—helpfulness,

pythonsqlgit
View job →

Test & Evaluation Lab Tech (Electrical/Electronic Equipment) Company: The Boeing Company The Boeing Company has an exciting opportunity for Test & Evaluation Lab Tech to join Boeing Defense, Space & Security (BDS) in Berkeley, MO. As a Boeing employee you’ll be part of a winning team that does great things every day BDS is a global leader in the development, production, maintenance and enhancement of fixed-wing and rotary wing aircraft, commercial and government satellites, human spaceflight programs and weapons. Key markets include aeronautics, space and weapons. Core capabilities are in development, production and mission enabling upgrades of integrated solutions. BDS delivers the most digitally advanced, simply and efficiently produced and intelligently supported solutions to its customers. Position Responsibilities: • Installing and maintaining Flight Simulation equipment consisting of computers, monitors, audio systems, visual systems, crew stations and rack equipment • Troubleshoot and resolve electronic hardware issues with minimum down time • Plans, conducts and documents tests on products, systems, components, materials, and manufacturing processes and technologies • Designs, fabricates, builds and analyzes test systems, materials, components, software and methodologies • Coordinates test requirements, determines test methods and capabilities • Conducts system test preparations by coordinating system integration, fabrication, modification, setup, checkout and tear-down • Troubleshoots and resolves laboratory and test-related problems • Coordinates test laboratory activities and maintenance and repairs • Coaches and trains others • Works under general direction </

recruitment
View job →

Test & Evaluation Lab Tech (Electrical/Electronic Equipment) Company: The Boeing Company Job Description The Boeing Company has an exciting opportunity for Test & Evaluation Lab Tech to join Boeing Defense, Space & Security (BDS) in Berkeley, MO. As a Boeing employee you’ll be part of a winning team that does great things every day BDS is a global leader in the development, production, maintenance and enhancement of fixed-wing and rotary wing aircraft, commercial and government satellites, human spaceflight programs and weapons. Key markets include aeronautics, space and weapons. Core capabilities are in development, production and mission enabling upgrades of integrated solutions. BDS delivers the most digitally advanced, simply and efficiently produced and intelligently supported solutions to its customers. Position Responsibilities: • Installing and maintaining Flight Simulation equipment consisting of computers, monitors, audio systems, visual systems, crew stations and rack equipment • Troubleshoot and resolve electronic hardware issues with minimum down time • Plans, conducts and documents tests on products, systems, components, materials, and manufacturing processes and technologies • Designs, fabricates, builds and analyzes test systems, materials, components, software and methodologies <span style

recruitment
View job →
P
16 days ago

Human Data Quality Analyst, AI Business Prolific Prolific isn’t just enabling AI innovation – we’re redefining it. While foundational AI technologies are becoming commoditized, Prolific’s human data infrastructure provides the high-quality, diverse data required to train the next generation of AI models. Through our platform, we empower researchers and companies to access a global, ethically curated participant base, ensuring cutting-edge AI research and training grounded in inclusivity and precision. The Role Prolific provides the human data that powers the next generation of AI models, working with frontier labs to capture the complex human judgments researchers need to train, evaluate and improve them. As a Human Data Quality Analyst, you'll be on the front line of making sure that data captures the right signal and is genuinely good enough to do its job. This isn’t traditional, back-office QA. You’ll be doing real analytical work: digging into datasets, identifying patterns and failure modes, investigating why quality has shifted, and turning complex findings into clear insights that help us improve how data is collected, reviewed and delivered. You'll spend real time reading annotations closely, but that is how you gather evidence, not what you produce. What you produce is analysis, practical recommendations and better quality controls. You'll get hands-on exposure to human data, annotation, machine learning pipelines and AI evaluation, working alongside Quality, Engineering, Operations and Delivery on new and evolving problems. There won't always be an established playbook. You'll be guided by our quality engineers, but you'll also need to run your own analysis, test your assumptions and recognise when you need input. It is a role with a steep learning curve from day one and a strong opportunity for someone early in their career to build deep, practical experience in a fast-moving area of AI. What You’ll Be Doing Run day-to-day qualit

pythonsqlmachine learning
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Human Data team turns human feedback into reliable signals for training and evaluation. We design and run end-to-end programs that capture the depth of human intent behind everyday and high-stakes uses of our models. Our remit spans bespoke data campaigns, scalable synthetic data generation, and product-embedded signals. We partner closely across all research teams to translate these signals into training datasets, novel evaluations, and feedback loops that push the frontier of our models and advance their applications. About the Role As a Program Manager (PGM) in the Human Data team you will partner with our research teams, operations and engineering to execute complex programs for collecting high-quality data. You will be a key interface between our external vendors and AI trainers, ensuring human data campaigns are successfully completed. Your work will play a key role in enabling OpenAI to train safe models that will land in the real world This role is based in our San Francisco HQ. In this role, you will: Work in a high velocity environment, where the outcome of your work will have a direct impact on the models that OpenAI deploy in the real world Work closely with external vendors, trainers and internal researchers to collect, review, and deliver high-quality data Gather requirements, write instructions, define success criteria, and calibrate the AI trainers Use internal tooling to assess labeled data and provide feedback to AI trainers Think critically and share recommendations on tooling and process improvements, optimizing for quality, throughput, and AI trainer experience You’ll thrive in this role if: You thrive in dynamic environments. You are comfortable navigating ambiguity, managing shifting priorities, and adapting to fast-paced changes without missing a beat. You’re curious about AI, LLMs, Agents. While not required, an interest or background in these areas will help you connect the dots in our broader mission. You have a can-do a

awsrestai
View job →
R
Roblox
📍 San Mateo• Full-time• From $226.4K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As Senior Technical Program Manager for the Human-in-the-Loop (HITL) Operations , you will own the end-to-end lifecycle of Roblox's AI data labeling and model evaluation ecosystem. This is a high-leverage, high-ownership role at the intersection of Research, Engineering, and Operations. You will scale our data operations that is handling million of items evaluated annually today with projection of 4x the current volume by 2030, managing a multi-million annual budget and a distributed workforce of 100+ remote contractors — all while driving the platform evolution from manual workflows to AI-assisted, LLM-accelerated operations. This is a rare opportunity to build the data infrastructure that directly determines the quality of Roblox's AI models across creator tools, content discovery, in-experience AI, and 3D generative content — at the exact moment when human judgment is the critical ingredient for getting these models right. You Will Lead data programs end-to-end. Own the full lifecycle of labeling and model evaluation workflows across Roblox's AI teams — from translating ambiguous ML requirements into structured annotation tasks, to overseeing contractor execution, quality review, a

pythonsqlaws
View job →
B
12 days ago

Human Engineer (Associate or Experienced) Company: The Boeing Company Are you ready to join a team of innovators, strategic disruptors, and dreamers who dare to redefine the future of the defense industry? Are you driven to create the unimaginable? Are you passionate about addressing human performance within the design and development of cutting-edge aerospace systems? Boeing Defense, Space & Security (BDS ) has an exciting opportunity for an associate or experienced Human Engineer to join our Human Engineering team in Colorado Spings, CO working on the front end of the next generation of Space Mission Systems technology. As a Human Factors Engineer, you will play a pivotal role in the design, development, and integration of complex aerospace systems, ensuring that all components function cohesively to meet user needs and operational requirements. You will collaborate with cross-functional teams, including systems engineers, software developers, and hardware specialists, to create systems that are not only technically robust but also intuitive and user-friendly, enhancing overall efficiency and user satisfaction. Position Responsibilities : Apply Human Engineering knowledge and principles to the analysis, design, and evaluation of complex systems Define system performance requirements to ensure safe and successful human operation of physical, functional, and program interfaces for all system operators Apply human performance principles, methodologies, and technologies to the design of complex systems Develop and implement research methodologies and analysis plans to test and evaluate developmental prototypes Apply a knowledge of military standards (such as MIL-STD-1472, MIL-STD-1474,

recruitment
View job →
O
1mo ago

About the Team OpenAI's Human Data Team creates custom data solutions driving groundbreaking research. Our work enhances and evaluates our flagship models and products like ChatGPT, GPT-5, and Sora, and contributes to safety initiatives through collaboration with our Preparedness and Safety Systems teams. About the Role As a Research Program Manager (RPM) in the Human Data team you will partner with research and engineering to design and implement pragmatic solutions for collecting high-quality data. You will be a key interface between our research roadmap, external vendors, AI trainers, and the Human Data engineering team. This role is based in our San Francisco HQ. In this role, you will: Collaborate with Research: Partner with researchers to scope data collection needs, define success metrics, and establish quality measurement frameworks. Design & Execute Data Collection Campaigns: Translate research needs into actionable plans and accelerate execution by leveraging existing tooling and iterating to reach the desired outcome. In many cases, you will need to implement scrappy new solutions while partnering with engineering to design robust/scalable solutions. Unblock Yourself: You must be deeply uncomfortable with the idea of sitting around waiting for external dependencies, and have the technical acumen and drive to figure out how to achieve at least partial success in the interim. Optimize Systems & Processes: Build and optimize dashboards to track campaign performance, leveraging SQL and Python for data analysis and actionable insights. Drive Technical Roadmaps: Collaborate with engineers to enhance data platforms, resolve blockers, and ensure security best practices such as access management. Scale Your Impact : Advise and empower program managers and vendors to drive day-to-day execution so that you can focus on addressing high priority opportunities. You might thrive in this role if you: Are proficient in SQL and Python for data analysis, including q

pythonsqlaws
View job →

The AI platform is responsible for all AI infrastructure across Datadog. Our mission is to provide tools and platforms that enable data scientists and engineers to conduct large-scale training and inference with ease. We support products such as Bits AI , LLMObs and all our AI research . As an engineering manager for the Evaluation & Annotation team, you’ll join a new and fast growing team and organization. You will support building and scaling the team, define our technical vision and help shape the roadmap. Your team will lead the charge on multiple critical technical challenges: AI model evaluation both offline and online, designing tooling and processes around human annotation, and establishing the standard around synthetics and AI generated datasets. You’ll work closely with sister teams in the AI platform organization ensuring a seamless AI development cycle. You’ll also partner with the Applied AI org and with Datadog infrastructure & tooling teams to build out systems from the ground up. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Manage and grow the Evaluation & Annotation team, directly managing 4-6 engineers Define our technical roadmap in alignment with AI platform goals and the Applied AI team roadmap. Work with our core platform teams to tailor Datadog's storage and data pipelines to our needs Create a strong team culture aligned with our engineering standards and our customer focus Participate in hands-on work: Code reviews, design reviews and some coding Who You Are: A Software Engineer at heart with a previous experience leading software engineering teams, as a tech lead or people manager Excellent leader with strong interpersonal skills, and the

restaigo
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s mission is to ensure that artificial general intelligence (AGI) benefits all of humanity. A key part of achieving that mission is training models that deeply understand and reflect human preferences — the Human Data team is at the heart of that effort. The Human Data engineering team creates the systems that enable scalable, high-quality human feedback. These systems are essential to how OpenAI trains and improves its most advanced models. Engineers on this team collaborate closely with world-class researchers to bring alignment techniques to life — from experimental ideas to production-ready feedback loops. About the Role We’re looking for software engineers to join the Human Data team and build the platforms, prototypes, tools, and infrastructure that power how our AI models are trained, aligned, and evaluated. You’ll partner with researchers and cross-functional teams to bring alignment ideas to life, influence future model training, and shape how models interact with the real world. We’re looking for people who are excited by technical ownership, enjoy working across the stack, and are eager to solve ambiguous problems in a high-impact, fast-paced environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build and maintain robust full-stack systems for feedback collection, data labeling, and evaluation pipelines, while maintaining high levels of security. Translate experimental alignment research into scalable production infrastructure, including inference and model training stacks. Design and iterate on user-facing tools and backend services to support high-quality data workflows Partner with researchers, engineers, and program leads to shape feedback loops and model interaction paradigms Drive infrastructure improvements that enable faster iteration and scaling across OpenAI’s frontier models, from internal r

awsrestai
View job →
🔔

Get new human evaluator jobs by email

Daily job updates · Unsubscribe anytime