Jobiba hiring network

Human Evaluator Jobs

3,920 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current human evaluator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

C
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? This role will focus on evaluating data science and coding tasks, requiring you to review and debug code, analyze model trajectories, and assess data visualization script development and logical flow implementation. Your work will contribute to our model development efforts and the logic our models apply when completing task requests. Please note : This is a part-time independent contractor position available within Canada. We seek candidates who are able to commit to 16 hours per week minimum at a 40 CAD/hour contract rate. This role is BYOD 💻 - Bring Your Own Device (laptop). Remote work within Canada. 12 month contract. Performance incentives included! As a Data Annotation Specialist, you will: Evaluate the model's ability to respond to coding requests, workflows, and code base-related questions using available tools. Assess agent trajectories and model capabilities for code generation, tabular and graphic manipulation, and debugging requests. Prompt models to complete complex data science tasks and review the accuracy of generated responses. Label, proofread, and improve machine-written and human-written soft

pythonsqlai
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? This role will focus on evaluating coding tasks, requiring you to review and debug code, navigate repository architecture, and analyze model trajectories. Your work will contribute to our model development efforts and the logic our models apply when completing task requests. Please note : This is a part-time independent contractor position available within Canada. We seek candidates who are able to commit to 16 hours per week minimum at a 40 CAD/hour contract rate. This role is BYOD 💻 - Bring Your Own Device (laptop). Remote work within Canada. 12 month contract. Performance incentives included! As a Data Annotation Specialist, you will: Evaluate the model's ability to respond to coding requests, workflows, and code base-related questions using available tools. Assess agent trajectories and model capabilities for code generation and debugging requests. Prompt models to complete complex coding tasks and review the accuracy of generated responses. Label, proofread, and improve machine-written and human-written software engineering-related outputs. Report quality and performance trends related to model/agent behavio

javascriptpythonjava
View job →
BA
Bolna AI
📍 Bengaluru• Full-time
1mo ago

About Bolna Bolna is Voice AI infrastructure built for India - and now for the world. We help businesses deploy intelligent voice agents that can call, converse, and convert in any language, at scale. From collections to customer support to sales, our agents handle millions of conversations so humans don’t have to. We’re a YC F25 company, backed by General Catalyst, with 1,050+ paying customers and growing fast. Our team of ~25 is based in Bengaluru. The Role Every voice AI agent Bolna deploys makes real-time judgment calls - when to speak, when to go silent, when a customer is done talking, when to hand off. We’re building automated systems to grade these calls at scale, using LLMs as judges of call quality. But before you trust a model’s judgment, you verify it against a human’s. That’s this role. You’ll listen to real calls, annotate what actually happened, and check whether our automated systems - LLM-as-judge evals and quantitative signal detection - got it right. It’s precise, high-attention work, and it sits right at the center of how we know our voice agents are actually working. This is an internship role for someone early in their career who wants hands-on exposure to how a voice AI company builds trust in its own AI. What You’ll Do Annotation Listen to and annotate real customer calls - transcription review, issue tagging, labeling - using tools like Label Studio Follow (and help sharpen) annotation guidelines for a multilingual environment (Hindi, English, Hinglish, ) Verifying LLM-as-Judge Evaluations For calls flagged by our automated eval pipeline, verify whether the model’s call was actually correct - for example, confirming whether a detected barge-in (agent/customer talking over each other) genuinely happened by listening to the audio Mark agreements and disagreements clearly, with reasoning, so we can measure and improve model accuracy over time All tools needed for this will be provided Verifying Quantitative Measures Check system-flagged quantit

N
Nuro
📍 Mountain View• Full-time• From $160.4K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role As a Senior Software Engineer, Collision Avoidance Testing, you will work closely with the onboard autonomy, evaluation infrastructure, data science, operations, and simulation teams to evaluate the performance of the autonomous vehicle in collision avoidance scenarios. You will be responsible for identifying gaps in onroad and simulation tests, designing appropriate tests to fill those gaps, reviewing test outcomes, and iterating on test design. You will analyse the results, recommend product changes, and report on residual risk both to leadership and in the safety case. About the Work Using human benchmarks from naturalistic data, develop statistical models that drive strategic decisions as a function of test volume, system performance, and confidence level. Identify and design realistic scenarios to drive test coverage of collision avoidance scenarios and collision detection scenarios using replay simulations of logs an

pythonaic++
View job →
M
Mongodb
📍 Palo Alto• Full-time• From $1.2M/yr
1mo ago

Role Overview & Executive Impact Serving as the West Coast anchor for MongoDB's Global VIP Support, this position acts as the “single-threaded DRI” for the technology needs of M3 — the CEO's direct reports — in the US Pacific time zone. You hold end-to-end accountability for their digital journey, from two weeks before Day 1 through daily operations, high-stakes events, global travel, and eventual offboarding, ensuring a seamless experience without handoff friction. The population you support is intentionally scoped and governed by HR and the Office of the CEO, not by IT — this is Phase 1 of a model designed to extend, tier by tier, to SVP+, then VP+, and ultimately the full Senior Leadership Team. The standards, runbooks, and judgment you establish now become the template each later phase reuses. This role prioritizes proactive engagement over traditional ticketing. By monitoring system health and resolving issues before they impact the executive, you act as a protective barrier for critical workflows. You combine high-level IT proficiency with AV expertise to lead high-stakes events and Board meetings, while also empowering the Executive Assistant community through specialized technical training. The way of working behind this role is built on a simple framework: routine diagnostics, health checks, and reporting are handled through automation and intelligent tooling wherever possible, freeing your time for the judgment calls, relationship-building, and in-person moments that actually require a human. Because this framework depends on staying ahead of what's possible, you are expected to stay current on emerging technology and AI trends and proactively evaluate how they can improve the executive support experience. Innovation through AI and automation is fundamental to this model. You will implement intelligent solutions to handle routine diagnostics and triage, eliminating manual toil so that human intervention is reserved for high-impact moments. These automa

pythonmongodbaws
View job →
F
Figma
📍 Ca New York• Full-time• From $153K/yr
1mo ago

Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! Figma is seeking a versatile and experienced Machine Learning / AI Engineer to join our growing AI team, working at the intersection of applied machine learning, infrastructure, and product innovation. Whether you’re building intelligent search systems, crafting scalable data pipelines, or enhancing AI-powered creativity tools, your work will drive user productivity, shape new product experiences, and advance the state of AI at Figma. You’ll collaborate closely with engineers, researchers, designers, and product managers across multiple teams to deliver high-quality ML-driven features and infrastructure. This is a high-impact, cross-functional role where you’ll shape both foundational systems and user-facing capabilities. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you’ll do at Figma: Design, build, and productionize ML models for Search, Discovery, Ranking, Retrieval-Augmented Generation (RAG), and generative AI features. Build and maintain scalable data pipelines to collect high-quality training and evaluation datasets, including annotation systems and human-in-the-loop workflows. Collaborate with AI researchers to iterate on datasets, evaluation metrics, and model architectures to improve quality and relevance. Work with product engineers to define and deliver impactful AI features across Figma’s platform. Partner with infrastructure engineers to develop and optimize systems for training, inference, monitoring, and deployment. Explo

pythonawsci/cd
View job →
C
Coinbase
📍 - India• Full-time• Remote• From ₹94.2L/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Coinbase’s Consumer Engagement and Experience (CEE) Surfaces team owns the front door to support for millions of customers: help.coinbase.com, social support, and Athena, the knowledge platform used by CX agents. We are rebuilding these experiences to be AI-native, with retrieval, generation, and agentic resolution at the core. We’re hiring an Engineering Manager to lead and scale the team responsible for these surfaces. You’ll set technical strategy and 12-month direction for AI-assisted self-service support, improving automated resolution quality and customer satisfaction while meeting Coinbase’s standards for safety, privacy, reliability, and cost. What you’ll do Lead a high-ownership engineering team building Help Center, social support, and agent knowledge experiences. Own end-to-end delivery of self-service discovery, conversational and agentic resolution, intelligent routing, escalation, and human handoff. Define architecture for secure, reliable, scalable systems spanning search, knowledge retrieval, chat orchestration, CRM, and ticketing integrations. Establish evaluation harnesses, quality metrics, feedback loops, and safeguards for accuracy, groundedness, privacy, and safety. Partner with Product, Design, CX Operations, Data, Compliance, Legal, Privacy, and Security on roadmaps, metrics, and execution. Evaluate build-versus-buy options and operationalize t

REMOTEsqlawskubernetes
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $218K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Staff Machine Learning Engineer, Identity Verification As a Staff Machine Learning Engineer on the Identity Verification team within the Platform group, you'll own the ML systems that determine whether a person, document, and capture session are legitimate. Every signup, account recovery, and high-risk action at Coinbase depends on these models. You'll lead the technical strategy for IDV ML end-to-end, from architecture through production enforcement, protecting the integrity of millions of accounts. What you'll do: Own the full IDV ML stack, including document authenticity models, 1:1 and 1:N face-match, liveness detection, presentation-attack detection, and deepfake/injection detection from feature pipeline through threshold tuning and production enforcement. Build identity-graph systems using GNNs that cluster accounts sharing biometric, device, and document signals to detect synthetic-identity rings and coordinated fraud at onboarding. Develop behavioral and device-intelligence models for capture-session anomaly detection, bot-vs-human classification, and device-fingerprint-based risk scoring at real-time latency. Drive vendor ML strategy by benchmarking external models against a Coinbase-owned evaluation set, designing dynamic routing logic across providers and geographies, and building the in-house evaluation layer that catches regressions before they reac

REMOTEpythonawsgit
View job →

The Cyber Deployment Manager partners with customers throughout the full lifecycle—from technical discovery and solution design through implementation, deployment, and adoption. You’ll work closely with Sales and Solutions Engineering during the pre-sales process to understand customer security priorities, assess technical requirements, develop solution architectures, and support demonstrations, workshops, and proofs of concept. After a customer commits, you’ll remain engaged as the technical deployment lead, translating the proposed solution into a production-ready implementation. You’ll guide integrations, establish success criteria, manage technical risks, and help customers operationalize AI across security workflows such as secure code review, vulnerability management, threat detection, incident response, SOC operations, and GRC automation. In this role, you will: Lead technical discovery with security executives, practitioners, architects, and engineering teams. Partner with Sales and Solutions Engineering on solution design, demonstrations, workshops, technical validation, and proofs of concept. Translate customer requirements into clear architectures, deployment plans, success criteria, and implementation milestones. Own the transition from pre-sales solution design into post-sales deployment and adoption. Serve as the primary technical partner during implementation, coordinating customer stakeholders and internal Product, Engineering, Security, and GTM teams. Build and troubleshoot integrations involving APIs, agents, security tools, cloud platforms, data sources, and enterprise workflows. Identify deployment risks, technical blockers, and product gaps, and drive them toward resolution. Help customers establish evaluation frameworks, governance controls, guardrails, monitoring, and human-review processes. Measure adoption and business impact, ensuring deployed solutions deliver meaningful security outcomes. Turn successful customer deployments into reusable

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time• $230K – $325K/yr
1mo ago

About the Team OpenAI’s Safety teams work to ensure our products are safe, trusted, and resilient as frontier AI systems scale globally. We tackle some of the company’s most important challenges across understanding and preventing misuse and misalignment, intercepting fraud and abuse, and protecting vulnerable users. We are hiring Data Scientists to help build the analytical foundations that allow OpenAI to deploy increasingly capable AI responsibly. We are hiring Data Scientists across several teams that contribute to safety in different ways, including: Safety Systems Integrity Product Policy This is a high-impact role operating at the intersection of product, safety, policy, and research. About the Role As a Data Scientist, Safety, you will help solve complex and ambiguous problems where rigorous analysis directly informs critical decisions. Depending on your background and team alignment, you may work on areas such as: Measure harmful or abusive behavior across OpenAI’s products Detect fraud, manipulation, coordinated misuse Evaluate and improve safety classifiers, rules systems, mitigation systems, and human review workflows Design experiments and causal analyses to understand product, policy, and mitigation impacts Build prevalence estimators, dashboards, monitoring systems, and executive decision frameworks Diagnose gaps in safety and integrity systems using behavioral and product data, and help quantify and navigate false positive / false negative tradeoffs Translate ambiguous safety risks into measurable problems and evidence-based recommendations Partner with Product, Engineering, Policy, Research, and Operations teams to improve safety outcomes Build zero-to-one analytical systems in rapidly evolving domains Ideal Candidate We’re looking for strong Data Scientists who thrive in ambiguous, high-leverage environments. You may be a fit if you have: Strong statistical reasoning and analytical judgment Experience with experimentation, causal inference, or obse

pythonsqlaws
View job →

About the Team The Personal AGI team is responsible for training and improving pre-trained models to be deployed into ChatGPT, the API, and potential future products. The team partners closely with research and product teams across the company, and conducts research as a final step to prepare for real world deployment to millions of users, ensuring that our models are safe, efficient, and reliable. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models. Our team works in research areas combining reinforcement learning and products. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve model capability and performance. Collaborate closely with the other research and product teams, allowing customers to optimize their own models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. You might thrive in this role if you: Have a deep understanding of machine learning and machine learning applications. Have a working knowledge of relevant models, and building evaluations for model capability improvement. Are comfortable diving into a large ML codebase to debug. Thrive in a dynamic and technically complex environment. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and

awsrestmachine learning
View job →
H
Humana
📍 Indiana• Remote
10 days ago

Become a part of our caring community The Transition Coordinator (Care Coach 2) evaluates member's needs and requirements. This evaluation aims to achieve and/or maintain an optimal wellness state. The Coordinator does this by guiding members/families toward resources and facilitating interaction with them. These resources are appropriate for the care and wellbeing of members. The Care Coach 2 work assignments are varied and frequently require interpretation and independent determination of the appropriate courses of action. Position Responsibilities: Support the ongoing member transitions in and out of the Indiana Medicaid programs, the Contractor's enrollment, and among care settings. Complete transitions and assists with the planning and preparation for them, and the follow-up care after. Works with the Member Advocate Coordinator and other member-focused departments of the plan. This collaboration ensures continuity and coordination of care and member and provider communication through the initial transition, ongoing benefit plan, and MCE transfers. Ensure the transfer and receipt of all outstanding prior authorization decisions, utilization management data, and clinical information such as prevention and wellness programs(s), care management and complex case management notes. Help with transitions from the custodial setting to the home and community-based setting. We ask that you have telephonic and in-person meetings within an assigned region. The purpose of these meetings is to work with various stakeholders, including long-term care members, hospital/rehab staff discharge planners, family members/POA's, PCP's, and other healthcare professionals. The ultimate goal is to prevent custodial placements whenever possible. Assess and evaluate member's needs to establish a member specific car

REMOTErecruitment
View job →

Become a part of our caring community The Care Coach 1 assesses and evaluates member's needs and requirements to achieve and/or maintain optimal wellness state by guiding members/families toward and facilitate interaction with resources appropriate for the care and wellbeing of members. The Care Coach 1 work assignments are often straightforward and of moderate complexity. Reports to the Regional Care Coach Manager. Looking for motivated Care Coach in COLLIER county FLORIDA!! We are looking for dynamic case managers that enjoy making a difference in the lives of others! You must live in Collier county in Florida. This rewarding role allows you to spend time connecting with our members to ensure they receive the services they need. The Care Coach 1 employs a variety of strategies, approaches and techniques to support a member's optimal wellness state by coordinating services & resources. Identifies and resolves barriers that hinder effective care. Ensures patient is progressing towards desired outcomes by continuously monitoring patient care through use of assessment, data, conversations with member, and active care planning. Understands own work area professional concepts/standards, regulations, strategies and operating standards. Work is managed and often guided by precedent and/or documented procedures/regulations/professional standards with some interpretation. The Care Coach 1 Visit Medicaid members in their homes, Assisted Living Facilities, and/or Long Term Care Facilities and other care settings – 75-90% local travel Assesses and evaluates member's needs and requirements in order to establish a member specific care plan Ensures members are receiving services in the least restrictive setting in order to achieve and/or maintain optimal well-being Planning and implementing interven

recruitment
View job →
H
Humana
📍 South Carolina• Remote
10 days ago

Become a part of our caring community The Field Care Manager Nurse 2 assesses and evaluates member's needs and requirements to achieve and/or maintain optimal wellness state by guiding members/families toward and facilitate interaction with resources appropriate for the care and wellbeing of members. You will report to the Manager, Care Management of Behavioral Health. The Field Care Manager Nurse 2 employs a variety of strategies, approaches, and techniques to manage a member's physical, environmental, and psycho-social health issues. Identifies and resolves barriers that hinder effective care. Ensures patient is progressing towards desired outcomes by continuously monitoring patient care through assessments and/or evaluations. May create member care plans. Understands department, segment, and organizational strategy and operating objectives, including their linkages to related areas. In this role, you will travel up to 50% of the time to support collaboration, conduct face-to-face meetings, and engage directly with staff, providers, members, and their families. NOTE: You should reside close to the Midlands OR Upstate area where your region will be. Use your skills to make an impact Required Qualifications Bachelor's in nursing (BSN) and have an active license in the state of South Carolina without disciplinary action. Must reside in the State of South Carolina 2 or more years of experience of case/care management 2 or more years working with the behavioral health population Knowledge of community health and social service agencies and additional community resources Use a variety of electronic information applications/software programs including electronic medical recor

REMOTErecruitment
View job →
H
Humana
📍 Indiana• Remote
13 days ago

Become a part of our caring community The Field Service Coordinator (Care Coach 1) assesses and evaluates member's needs and requirements. This is done to achieve and/or maintain optimal wellness state by guiding members/families toward resources appropriate for their care and wellbeing. The Service Coordinator work assignments are often straightforward and of moderate complexity. Your role will involve meeting members in their location, spending quality time assessing their needs and barriers and then connecting our members with quality services to promote their ultimate well-being and guide health outcomes. Responsibilities include: Administer ongoing long-term services and support (LTSS) related assessments through person-centered thinking approaches. Contacts members both telephonically and/or in-person to establish goals and priorities. This involves evaluating resources, developing a plan of care, and identifying LTSS providers and community partnerships. The goal is to provide a combination of services and supports that best meet the needs and goals of the member and caregiver through person-centered thinking approaches. Development and modification of Service Plan and involve applicable members of the care team in care planning (Informal caregiver coach, PCP) Support members through navigation of their LTSS and related environmental and social needs Use available information about member to prevent the need for administration of duplicative assessments. Focus on supporting members or caregivers in accessing long-term services and support, social, housing, educational and other services, regardless of funding sources to meet their needs. Assist members in maintaining Medicaid eligibility Collaborate with Medical Director/Geriatrician/Care Coordinator as deemed necessary to ensure cohesive, holist

REMOTEExcelrecruitment
View job →
🔔

Get new human evaluator jobs by email

Daily job updates · Unsubscribe anytime