About DataCamp DataCamp's mission is to empower everyone with the data and AI skills essential for 21st-century success. By providing practical, engaging learning experiences, DataCamp equips learners and organizations of all sizes to harness the power of data and AI. As a trusted partner to over 17 million learners and 6,000+ companies , including 80% of the Fortune 1000, DataCamp is leading the charge in addressing the critical data and AI skills shortage. About the role Research has consistently shown that one-on-one tutoring dramatically outperforms traditional classroom instruction—a finding known as Bloom's 2 Sigma Problem. Despite decades of awareness, no solution has successfully delivered personalized tutoring at scale. Until now. We’re building an AI Creator that turns learning into action—empowering users to instantly build, experiment, and solve real-world problems, transforming knowledge into tangible impact from day one.. You will work on the core AI system that delivers real-time, adaptive learning experiences to learners. If you're excited about building the future of human learning, this role is for you. About you At DataCamp, we seek individuals who embody our core values of data-driven decision-making, action, transparency, ownership, and customer focus. You thrive in a fast-paced, high-performing environment and are driven by a passion for making a meaningful impact. You're adaptable, embracing change and ambiguity with enthusiasm. Your initiative and entrepreneurial spirit push you beyond just meeting targets—you aim to understand the "why" behind our goals and take ownership to drive the business forward. You’re a collaborative team player who values transparency and always seeks to improve and innovate. If this sounds like you, we encourage you to apply! Responsibilities Build and evolve the core AI tutoring system , including prompt architectures, agentic workflows, and real-time adaptive behaviors. Run rigorous experimentation and evaluation
Jobiba hiring network
Human Evaluator Jobs
3,920 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current human evaluator jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About DataCamp DataCamp's mission is to empower everyone with the data and AI skills essential for 21st-century success. By providing practical, engaging learning experiences, DataCamp equips learners and organizations of all sizes to harness the power of data and AI. As a trusted partner to over 17 million learners and 6,000+ companies , including 80% of the Fortune 1000, DataCamp is leading the charge in addressing the critical data and AI skills shortage. About the role Research has consistently shown that one-on-one tutoring dramatically outperforms traditional classroom instruction—a finding known as Bloom's 2 Sigma Problem. Despite decades of awareness, no solution has successfully delivered personalized tutoring at scale. Until now. We’re building an AI Creator that turns learning into action—empowering users to instantly build, experiment, and solve real-world problems, transforming knowledge into tangible impact from day one.. You will work on the core AI system that delivers real-time, adaptive learning experiences to learners. If you're excited about building the future of human learning, this role is for you. About you At DataCamp, we seek individuals who embody our core values of data-driven decision-making, action, transparency, ownership, and customer focus. You thrive in a fast-paced, high-performing environment and are driven by a passion for making a meaningful impact. You're adaptable, embracing change and ambiguity with enthusiasm. Your initiative and entrepreneurial spirit push you beyond just meeting targets—you aim to understand the "why" behind our goals and take ownership to drive the business forward. You’re a collaborative team player who values transparency and always seeks to improve and innovate. If this sounds like you, we encourage you to apply! Responsibilities Build and evolve the core AI tutoring system , including prompt architectures, agentic workflows, and real-time adaptive behaviors. Run rigorous experimentation and evaluation
About DataCamp DataCamp's mission is to empower everyone with the data and AI skills essential for 21st-century success. By providing practical, engaging learning experiences, DataCamp equips learners and organizations of all sizes to harness the power of data and AI. As a trusted partner to over 17 million learners and 6,000+ companies , including 80% of the Fortune 1000, DataCamp is leading the charge in addressing the critical data and AI skills shortage. About the role Research has consistently shown that one-on-one tutoring dramatically outperforms traditional classroom instruction—a finding known as Bloom's 2 Sigma Problem. Despite decades of awareness, no solution has successfully delivered personalized tutoring at scale. Until now. We’re building an AI Creator that turns learning into action—empowering users to instantly build, experiment, and solve real-world problems, transforming knowledge into tangible impact from day one.. You will work on the core AI system that delivers real-time, adaptive learning experiences to learners. If you're excited about building the future of human learning, this role is for you. About you At DataCamp, we seek individuals who embody our core values of data-driven decision-making, action, transparency, ownership, and customer focus. You thrive in a fast-paced, high-performing environment and are driven by a passion for making a meaningful impact. You're adaptable, embracing change and ambiguity with enthusiasm. Your initiative and entrepreneurial spirit push you beyond just meeting targets—you aim to understand the "why" behind our goals and take ownership to drive the business forward. You’re a collaborative team player who values transparency and always seeks to improve and innovate. If this sounds like you, we encourage you to apply! Responsibilities Build and evolve the core AI tutoring system , including prompt architectures, agentic workflows, and real-time adaptive behaviors. Run rigorous experimentation and evaluation
About DataCamp DataCamp's mission is to empower everyone with the data and AI skills essential for 21st-century success. By providing practical, engaging learning experiences, DataCamp equips learners and organizations of all sizes to harness the power of data and AI. As a trusted partner to over 17 million learners and 6,000+ companies , including 80% of the Fortune 1000, DataCamp is leading the charge in addressing the critical data and AI skills shortage. About the role Research has consistently shown that one-on-one tutoring dramatically outperforms traditional classroom instruction—a finding known as Bloom's 2 Sigma Problem. Despite decades of awareness, no solution has successfully delivered personalized tutoring at scale. Until now. We’re building an AI Creator that turns learning into action—empowering users to instantly build, experiment, and solve real-world problems, transforming knowledge into tangible impact from day one.. You will work on the core AI system that delivers real-time, adaptive learning experiences to learners. If you're excited about building the future of human learning, this role is for you. About you At DataCamp, we seek individuals who embody our core values of data-driven decision-making, action, transparency, ownership, and customer focus. You thrive in a fast-paced, high-performing environment and are driven by a passion for making a meaningful impact. You're adaptable, embracing change and ambiguity with enthusiasm. Your initiative and entrepreneurial spirit push you beyond just meeting targets—you aim to understand the "why" behind our goals and take ownership to drive the business forward. You’re a collaborative team player who values transparency and always seeks to improve and innovate. If this sounds like you, we encourage you to apply! Responsibilities Build and evolve the core AI tutoring system , including prompt architectures, agentic workflows, and real-time adaptive behaviors. Run rigorous experimentation and evaluation
About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. Role Overview As a Forward Deployed AI Engineering Manager on our Enterprise team, you'll be the technical bridge between Scale AI's cutting-edge AI capabilities and our most strategic customers. You'll work with enterprise clients to understand their unique challenges, lead a team that architects specific AI solutions, and ensure successful deployment and adoption of AI systems in production environments. This is a Management role that combines deep engineering and AI expertise, leading a team, and working on customer-facing problems. You'll work directly with customer engineering teams to integrate AI into their critical workflows. Key Responsibilities Customer Integration & Deployment Partner directly with enterprise customers to understand their technical infrastructure, data pipelines, and business requirements Design and implement custom integrations between Scale AI's platform and customer data environments (cloud platforms, data warehouses, internal APIs) Build robust data connectors and ETL pipelines to ingest, process, and prepare customer data for AI workflows Deploy and configure AI models and agents within customer security and compliance boundaries AI Agent Development Develop production-grade AI agents tailored to customer use cases across domains like customer support, data analysis, content generation, and workflow automation Architect multi-agent systems that orchestrate between different models, tools, and data sources Implement evaluation frameworks to measure agent performance and iterate toward business objectives Design human-in-the-loop workflows and feedback mechanisms for continuous agent improvement Prompt Engineeri
About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. Role Overview As a Senior Staff Frontier Agents Engineer on our Enterprise team, you'll be the technical bridge between Scale AI's cutting-edge AI capabilities and our most strategic customers. You'll work with enterprise clients to understand their unique challenges, architect custom AI solutions, and ensure successful deployment and adoption of AI systems in production environments. This is a hands-on technical role that combines deep engineering expertise with customer-facing problem solving. You'll work directly with customer engineering teams to integrate AI into their critical workflows. Key Responsibilities Customer Integration & Deployment Partner directly with enterprise customers to understand their technical infrastructure, data pipelines, and business requirements Design and implement custom integrations between Scale AI's platform and customer data environments (cloud platforms, data warehouses, internal APIs) Build robust data connectors and ETL pipelines to ingest, process, and prepare customer data for AI workflows Deploy and configure AI models and agents within customer security and compliance boundaries AI Agent Development Develop production-grade AI agents tailored to customer use cases across domains like customer support, data analysis, content generation, and workflow automation Architect multi-agent systems that orchestrate between different models, tools, and data sources Implement evaluation frameworks to measure agent performance and iterate toward business objectives Design human-in-the-loop workflows and feedback mechanisms for continuous agent improvement Prompt Engineering & Optimization Create sophisticate
AI Software Engineer, Agent Harness Location: Bengaluru, Karnataka (or throughout India remote-friendly with travel) About EnCharge AI EnCharge AI is building the next generation AI platform. Our novel in-memory-computing architecture delivers a 10x step-function improvement in compute energy efficiency and performance for AI inference workloads. As the demands of artificial intelligence move beyond today's models, we believe fundamental underlying infrastructure must evolve. We are an experienced team of AI researchers, silicon & systems engineers, and architects backed by leading investors, poised to become the essential platform for the next wave of AI innovation. The Opportunity We serve open-weight models and our own bespoke checkpoints on EnCharge hardware. The models change often, and the harness around them needs to keep up. You own this layer that runs agents against files, tools, documents with permissions, memory, unattended execution, and real outputs. It will be assembled from a combination of open-source and bespoke code. Key Responsibilities Own the harness architecture end to end — agent loop, safe execution, context management, knowledge base, memory, permissions, orchestration, outputs, interfaces, observability — one component per layer, with clear interfaces so layers can be swapped. Build the pieces with no open-source equivalent e.g. session semantics, enforced permissions, memory in a human-editable file, orchestrator, and outputs. Keep pace with the models: adapters, prompt formats, tool-call schemas, stop conditions, benchmarking and evaluation. Make tool use reliable across models of uneven tool-calling quality — validation, repair, retries, fallbacks. Develop agents, tools, and MCP servers for internal and customer use cases, and review them for security before they ship. Build the evaluation harness: task suites, regression runs on every model or harness change, cost and latency per task alongside quality. Define the interfaces:
At Dscout, we’re building the most flexible and powerful UX research platform on the market—trusted by the world’s top brands in finance (JP Morgan Chase, Intuit, Charles Schwab, PayPal), healthcare (Aya, Headspace), consumer goods (Keen, Verizon, Target, Northface), and tech (Google, Amazon, Facebook, Meta, Spotify, AirBnB). Our tools help teams deeply understand the humans behind their products, so they can build better ones. We are expanding our smart and driven team and would love for you to join us. AI native product is fundamentally different engineering problem than building deterministic software: the same input won't always produce the same output, and "working" means the agent behaves well across the full distribution of real world scenarios, not that it passes a fixed test suite. We're looking for an Applied AI Engineer with 2-5 years of experience building and shipping AI systems used by professionals at enterprise. You're comfortable working with modern LLM-based systems and agentic workflows, and you know how to turn powerful models into reliable product features. You have strong product judgment and think deeply about tradeoffs between LLM approaches and traditional ML when designing solutions. You care about evaluation, iteration speed, and making sure AI systems actually drive measurable business impact reliably . What you'll do Own the production improvement loop across agent behavior, customer and operator feedback, evaluation, experimentation, and verified business outcomes Instrument agent workflows so model interactions, tool use, decisions, failures, human edits, and downstream outcomes can be understood in context Define meaningful quality standards, representative evaluation datasets, regression coverage, and production monitoring. Investigate why agents underperform across context, knowledge, instructions, tools, routing, guardrails, or workflow design Design and ship targeted behavior improvements, including changes to prompting, cont
At Dscout, we’re building the most flexible and powerful UX research platform on the market—trusted by the world’s top brands in finance (JP Morgan Chase, Intuit, Charles Schwab, PayPal), healthcare (Aya, Headspace), consumer goods (Keen, Verizon, Target, Northface), and tech (Google, Amazon, Facebook, Meta, Spotify, AirBnB). Our tools help teams deeply understand the humans behind their products, so they can build better ones. We are expanding our smart and driven team and would love for you to join us. We're looking for an AI Product Manager who brings the full standard PM toolkit — user research, market and competitive analysis, roadmap and strategy, cross-functional delivery, and strong collaboration instincts - and applies it to products where the model underneath doesn't behave the same way twice. You have a real, hands-on feel for what different LLMs are actually good and bad at, and you use that to prototype ideas yourself, sometimes shipping small, production-quality AI features directly. You default to ownership: of the roadmap, of outcomes, of the quality bar that decides whether something's actually ready to ship, and of what an agent should be trusted to do on its own versus when a human needs to stay in the loop. That's because building AI-native products means quality is a distribution, not a pass/fail — a feature can work correctly most of the time and still need a real answer for the failure tail, since the underlying system is non-deterministic, not just complex. What you'll do Lead the roadmap and strategy for your product area, from problem discovery through delivery and post-launch iteration Lead the evaluation bar for the model-powered surfaces you're responsible for: define eval sets, identify failure modes, and know the rollback plan before a change ships Prototype product ideas directly using your own understanding of model strengths and weaknesses, and ship small, production-quality AI features yourself when that's the fastest path to
About Bazaarvoice At Bazaarvoice, we create smart shopping experiences. Through our expansive global network, product-passionate community & enterprise technology, we connect thousands of brands and retailers with billions of consumers. Our solutions enable brands to connect with consumers and collect valuable user-generated content, at an unprecedented scale. This content achieves global reach by leveraging our extensive and ever-expanding retail, social & search syndication network. And we make it easy for brands & retailers to gain valuable business insights from real-time consumer feedback with intuitive tools and dashboards. The result is smarter shopping: loyal customers, increased sales, and improved products. The problem we are trying to solve : Brands and retailers struggle to make real connections with consumers. It's a challenge to deliver trustworthy and inspiring content in the moments that matter most during the discovery and purchase cycle. The result? Time and money spent on content that doesn't attract new consumers, convert them, or earn their long-term loyalty. Our brand promise : closing the gap between brands and consumers. Founded in 2005, Bazaarvoice is headquartered in Austin, Texas with offices in North America, Europe, Asia and Australia. It’s official: Bazaarvoice is a Great Place to Work in the US , Australia, India, Lithuania, France, Germany and the UK! The Human Resources function is responsible for providing human resources (HR) support to a major business unit or functional area. This includes any issues relating to talent acquisition, compensation and benefits, employee relations, training and development, and HR operations. The Analytics, Modeling, & Reporting area is responsible for analyzing and reporting on data sets for a particular area to evaluate, recommend and support the implementation of business strategies and operational processes. The Human Resources Information System (HRIS) and Reporting focus specializes
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We are hiring a AI Platform Engineer to build the execution layer for Supabase's internal AI systems. Supabase is building an AI-native internal operating system: a common way of working across the company where AI carries a meaningful share of the operational load rather than sitting alongside it as an assistant. We are standing up a new central team to build those systems, enable the teams, and embed AI operations throughout the organization. You are the engineer on that team. The execution layer is yours. You will build the platform that actually runs agents: an event-triggered queue, a headless model-agnostic runtime, durable state so work survives a restart, a human review gate, atomic rollback, and full logging of every prompt, tool call and decision so any run can be reconstructed. You will build the evaluation layer that makes any of it trustworthy, because an agent that cannot be measured cannot be trusted with anything beyond reading. And you will build the agents themselves, across everything from executive reporting down to a layer of agents that watch the platform and improve it. This is a governance-heavy environment by design, and that is the interesting part of the problem. Agents are risk-tiered from read-only internal data through to external-facing output, with review depth, evaluation requirements and human approval scaling by tier. Some capabilities are permanently off limits: an agent may read and report, it may carry a human-authored update into a system of record once a human consents, and it may initiate contact within a strict budget, but it may never autonomously write a commitment (an owner, a due date, a status) into a shared work system. Your job is
Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. About the role As an Applied Research Engineer in our Video team, you will help build the next generation of production-grade foundation models for human-centric video generation. You will join a highly focused team working at the intersection of large-scale generative modeling, distributed systems, and production engineering. Our mission is to develop and optimize video base models that power realistic, controllable, and emotionally expressive synthetic humans at scale. This is not pure research. This is applied research with direct product impact. You will work on advancing training recipes, scaling distributed systems, improving evaluation frameworks, and optimizing inference to ensure our models are high quality, stable, and efficient enough for real-world deployment. Your work will directly influence models used by tens of thousands of businesses worldwide. What you’ll do You will own and execute end-to-end research and engineering projects, from hypothesis to production impact. This includes: Developing and scaling latent video diffusion models tailored for human-centric video generation Designing conditioning mechanisms to improve control (pose, emotion, script, camera) without sacrificing fidelity Advanc
Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. About the role As a Research Engineer in our Video team, you will help build the next generation of production-grade foundation models for human-centric video generation. You will join a highly focused team working at the intersection of large-scale generative modeling, distributed systems, and production engineering. Our mission is to develop and optimize video base models that power realistic, controllable, and emotionally expressive synthetic humans at scale. This is not pure research. This is applied research with direct product impact. You will work on advancing training recipes, scaling distributed systems, improving evaluation frameworks, and optimizing inference to ensure our models are high quality, stable, and efficient enough for real-world deployment. Your work will directly influence models used by tens of thousands of businesses worldwide. What you’ll do You will own and execute end-to-end research and engineering projects, from hypothesis to production impact. This includes: Developing and scaling latent video diffusion models tailored for human-centric video generation Designing conditioning mechanisms to improve control (pose, emotion, script, camera) without sacrificing fidelity Advancing distr
Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. About the role The Data team manages the complete lifecycle of data for researchers - from sourcing and large-scale processing to delivering datasets that power our models. Data sits at the heart of our Research efforts and enables all other teams. As part of the Data team, you’ll work with over a million hours of video and audio data. This role exists at the intersection of applied research, data engineering, and ML infrastructure rather than being a traditional research position . You’ll build the world’s best human-centric data lake by collaborating closely with our model training teams. By understanding their requirements, you’ll extract new features and annotations that elevate our datasets. You should be passionate about enhancing model performance through high-quality, accurate datasets. Our infrastructure and pipelines are in great shape, and this role provides room to not only enhance them but also influence the team’s longer-term strategy. What we're looking for: A strong background in data-centric, applied Machine Learning, with hands-on experience improving model performance through data quality, curation, labeling, and evaluation rather than model architecture alone Experience working on the data la
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? This role will focus on breaking down engineering tasks such as reporting, component search and handling of material billing/information into individual steps that can be tackled through tool use. Your breakdown will teach the Cohere model the logic needed to complete each task. Please Note: This is a part-time independent contractor position available within Canada . We seek candidates who are able to commit to 16 hours per week minimum at a 30 CAD/hour contract rate. This role is BYOD 💻 - Bring Your Own Device (laptop). Remote work within Canada. 12 month contract. Performance incentives included! As a Data Annotation Specialist, you will: Evaluate the model's ability to respond to engineering procedures and workflows Task models to complete engineering tasks along with verified information to evaluate the accuracy of the model's responses Label, proofread, and improve machine-written and human-written engineering-related outputs. Follow our style guide, and make recommendations on unique situations that fall outside of its scope. You may be a good fit if you have: 3+ years of industry experience working in eng
Get new human evaluator jobs by email
Daily job updates · Unsubscribe anytime