Jobiba hiring network

Human Evaluator Jobs

3,920 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current human evaluator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Safety Systems manages the complete lifecycle of safety efforts for OpenAI’s frontier models, ensuring our models are deployed responsibly and have a positive impact on society. Our work spans diverse research and engineering initiatives—from system-level safeguards and model training to evaluation and red-teaming—all aimed at mitigating misuse and maintaining our high bar for safety. We lead OpenAI's commitment to developing and deploying safe Artificial General Intelligence (AGI), fostering a culture of trust, responsibility, and transparency. Our goal is to continuously learn from deployments, distribute AI’s benefits widely, and ensure that powerful tools remain aligned with human values and safety considerations. About the Role The Safety Measurement Product Manager owns OpenAI's approach to measuring harm and safeguard efficacy in production, including driving the strategy for our suite of safety measurement platforms and products used across the company. You will partner closely with our safety research and engineering teams to determine what we measure, where we measure it, and how we measure it, feeding those insights directly into critical leadership decisions and back into our safety work. You will also represent the company's topline safety metric as well as prioritize incoming requests from partner teams to expand our safety measurement platform to more use cases. This position is based in San Francisco, CA, with relocation assistance available. In this role, you will: Partner closely with data science, research, engineering, policy teams, and other stakeholders to craft a vision for understanding safety outcomes and prevalence on our platforms. Define strategic priorities and product roadmaps focused on improving safety measurement approaches will scaling our measurement platform to more use cases, products, and cross-functional team needs. Establish repeatable processes to integrate cutting-edge AI safety research into OpenAI’s safety m

awsrestai
View job →
S
10 days ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are hiring a Staff Software Engineer for our Frontier Security AI team. Snowflake's Frontier Security AI teams develop production-grade LLM applications, intelligent agents, AI infrastructure, and evaluation systems for enterprise customers — products that must meet a high bar for quality, security, reliability, and efficiency while operating over sensitive data at large scale. In this role, you will lead the design and development of our Agentic Harness and agent evaluation platform, working across product, infrastructure, applied AI, security, and modeling teams to take new capabilities from prototype to dependable customer value. AS A STAFF SOFTWARE ENGINEER AT SNOWFLAKE, YOU WILL: Architect and build the Agentic Harness that executes complex, multi-step AI workflows across models, tools, data, and services. Design stable interfaces for tool execution, context construction, state management, memory, permissions, retries, fallbacks, and human review. Own agent quality end to end by building evaluation harnesses, representative datasets, automated graders, experiment pipelines, and release gates. Convert ambiguous reports such as "the agent feels worse" into measurable failure modes, reproducible tests, and durable fixes. Analyze production agent trajectories to identif

REMOTEtypescriptpythonjava
View job →
W
Writer
📍 London• Full-time• Remote
13 days ago

🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About this role We’re looking for a collaborative and builder-oriented enterprise sales rep experienced at helping prospective customers at large companies navigate the evaluation, business case development, and procurement of transformative technology. Your objective will be to help convert enterprise prospects (3K-8K FTE, although tilted towards Fortune 1000) who are active in our trials or who request a sales demo from our website or within our product. While most of your pipeline will come inbound, you'll also be responsible for generating pipeline from ideal customer profile (ICP) accounts within your account set. Your positivity, sense of curiosity, and ability to create champions from early adopters in the AI space will help shape our entire culture. This is a hybrid role based out of our London hub. You'll be reporting to our RVP. 🦸🏻‍♀️ Your responsibilities Develop a deep understanding of our users and why they are exploring WRITER Become a trusted product speciali

REMOTEaiprocurement
View job →

About the Team: Tubi's Internal Tools team is at the forefront of AI integration, developing everything from developer resources to production-grade AI for business operations. We are the group responsible for turning AI from an experiment into an operating capability: training, infrastructure, developer agents, and AI-powered business systems. Engineers operate with high ownership and autonomy, collaborating on shared architectural decisions and AI infrastructure. What You'll Do: Own systems end to end — design them, build them, and support them in production. Lead the projects you own: sequence the work, decide what lands first, and set technical direction for the engineers working with you. Sit with the people who use what you build, and turn what you learn there into a system. Design the service boundaries, contracts and schema evolution that let our platforms grow without breaking the teams depending on them. Make our AI systems dependable in production: evaluation harnesses, human approval steps before an agent acts, retries that handle a model returning something unexpected, and cost tracking that tells you what a task costs before you run it. Build what other engineers build on — agent skills, tool and MCP integrations, shared libraries — and raise the bar through code review, design discussion and mentoring. Spot the platform work nobody has asked for yet, make the case for it, and build it. Your Background: 5+ years of professional experience building and operating production systems, from design through production ownership. A system you designed and can walk us through end to end — where its boundaries sit, what constrained it, and what you chose against. Strong programming proficiency in a statically typed language such as Rust, Go, C++, Java, Kotlin, C#, or TypeScript. Production Rust is a plus rather than a requirement. You have owned a service in production: you wrote the runbooks, you knew what it cost, and you were the one paged when it broke. Expe

typescriptjavaai
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $252K/yr
17 days ago

About Scale AI At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. Scale Frontier Data is the organization behind the training and evaluation data that frontier labs depend on. We build the systems, tooling, and expert workflows that turn hard human expertise into signals that models can learn from, across reasoning, coding, agentic tool use, and domain expertise. Reinforcement learning environments are now the center of gravity for that work: the difference between a model that demos well and a model that reliably completes long-horizon work is almost always the quality of the environments and reward signals it was trained against. Responsibilities As a Staff Software Engineer, RL Environments, you'll own the technical foundation for how Scale builds, runs, verifies, and delivers RL environments at scale. An RL environment is a real piece of software: a containerized world with real dependencies, real state, real tools, and a grader that has to be correct even when the agent is creative about breaking it. Building one is a full-stack engineering problem. Building thousands of them reproducibly, cheaply, with trustworthy reward signals and throughput measured in millions of rollouts is a systems problem that very few people have solved. You'll work on both. You'll design the platform: sandboxed execution, environment packaging and versioning, rollout orchestration, trajectory capture, verifier frameworks, and the authoring surfaces that let engineers and domain experts produce environments without reinventing infrastructure each time. And you'll go deep on the environments themselves by instrumenting real applications, designing task suites that expose specific capability gaps, and building graders that

typescriptpythonreact
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $179.4K/yr
17 days ago

About Scale AI At Scale AI, our mission is to accelerate the development of AI applications. For 8 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI) and building upon our prior model evaluation work with enterprise customers and governments to deepen our capabilities and offerings for public and private evaluations. About Data Engine Our Generative AI Data Engine powers the world’s most advanced LLMs and generative models through world-class RLHF (Reinforcement Learning with Human Feedback), human data generation, model evaluation, safety, and alignment. The data we produce is some of the most critical work for how humanity will interact with AI. About Our FDE Team Generating high-quality data is the core problem our business solves. We aim to make producing and delivering high-quality data seamless and efficient for operators and customers. Our Team is building customer and operator-specific infrastructure to provide high-quality data with low turnaround time. You'll be exposed to the cutting edge of the Generative AI industry while directly interfacing with the leading model-building organizations in the space, including the top AI research labs and government agencies. Join us in shaping the future of Artificial General Intelligence. As a Forward Deployed Engineer, you'll be at the forefront of providing the critical data infrastructure that powers the most advanced AI models, directly influencing how humanity interacts with AI. You will work with the world’s leading AI companies and government agencies to solve their most complex AI data-related problems. Responsibilities: Drive Impact: Directly contribute to the advancement of AI by delivering critical data solutions for leading AI innovators and

awsrestmachine learning
View job →
SA
17 days ago

About Scale AI At Scale AI, our mission is to accelerate the development of AI applications. For 10 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI) and building upon our prior model evaluation work with enterprise customers and governments to deepen our capabilities and offerings for public and private evaluations. About Data Engine Our Generative AI Data Engine powers the world’s most advanced LLMs and generative models through world-class RLHF (Reinforcement Learning with Human Feedback), human data generation, model evaluation, safety, and alignment. The data we produce is some of the most critical work for how humanity will interact with AI. About Our FDE Team Generating high-quality data is the core problem our business solves. We aim to make producing and delivering high-quality data seamless and efficient for operators and customers. Our Team is building customer and operator-specific infrastructure to provide high-quality data with low turnaround time. You'll be exposed to the cutting edge of the Generative AI industry while directly interfacing with the leading model-building organizations in the space, including the top AI research labs and government agencies. Join us in shaping the future of Artificial General Intelligence. As a Forward Deployed Engineer, you'll be at the forefront of providing the critical data infrastructure that powers the most advanced AI models, directly influencing how humanity interacts with AI. You will work with the world’s leading AI companies and government agencies to solve their most complex AI data-related problems. Responsibilities: Drive Impact: Directly contribute to the advancement of AI by delivering critical data solutions for leading AI innovators an

awsrestmachine learning
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $216K/yr
17 days ago

At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. Scale Frontier Data is the organization behind the training and evaluation data that frontier labs depend on. We build the systems, tooling, and expert workflows that turn hard human expertise into signals that models can learn from, across reasoning, coding, agentic tool use, and domain expertise. About our Customer Platform team: Our Customer Platform Team plays a pivotal role in integrating our platform with external systems and ensuring seamless, reliable connectivity for both internal users and customers. As the leader of this team, you’ll drive the strategy, architecture, and development of our connectivity solutions, focusing on API integration, distributed systems, and a robust data platform. Your role will be crucial in maintaining and enhancing our platform’s ability to meet the needs of both our internal and external stakeholders. Responsibilities: Own large areas within our product Comfortable working cross functionally, whether that be internal or external customers Build features end-to-end: front-end, back-end, system design, debugging and testing Deliver experiments at a high velocity and level of quality to engage our customers Work across the entire product lifecycle from conceptualization through production Influence the culture, values, and processes of a growing engineering team Inspire and mentor less experienced engineers Collaborating with cross-functional teams to define, design, and ship new product features and experiences. Requirements: At least 7-10 years of relevant experience is preferred Track record of shipping high-quality products and features at scale Desire to work in a very fast-paced environment Abil

awsrestai
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $180K/yr
17 days ago

About Scale At Scale AI, our mission is to accelerate the development of AI applications. For 8 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including: generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI), and building upon our prior model evaluation work with enterprise customers and governments, to deepen our capabilities and offerings for both public and private evaluations. About Data Engine Our Generative AI Data Engine powers the world’s most advanced LLMs and generative models through world-class RLHF (Reinforcement Learning with Human Feedback), human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. Our Approach As part of the interview process, you’ll be considered for opportunities across several teams within the GenAI Engineering organization, based on your interests, expertise, and business needs. Potential team placements include Allocation, Growth, Frontier Data, Trust & Safety, Pay, Operator, or Tasking Experience. Together, these teams power Scale’s AI data operations - from building high-impact datasets that push the boundaries of LLM capabilities, to optimizing contributor onboarding and incentives, to safeguarding data integrity through advanced trust, safety, and security measures. They work at the intersection of ML, operations, and analytics to ensure we deliver the highest-quality data at scale. Responsibilities: Design, build, and maintain robust, scalable systems across the full stack, including front-end, back-end, and infrastructure layers Implement high-impact features using modern technologies such as TypeScript, React, Node.js, MongoDB, Elasticsearch, and Temporal Collaborate closely with internal operators (your use

typescriptpythonreact
View job →
F
Flexport
📍 Manila• Full-time
19 days ago

About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. The Opportunity: Flexport is looking for an experienced, full-cycle Recruiter to lead hiring for a portfolio of roles across APAC. You'll partner directly with hiring managers to understand team needs, build strong candidate pipelines, and deliver a great candidate experience from first outreach through offer. This role suits someone with roughly 5 years of recruiting experience who is comfortable owning requisitions end-to-end and operating with a good degree of independence. You will: Manage full-cycle recruiting for assigned roles — intake, sourcing, screening, interview coordination, offer, and close. Partner with hiring managers to define role requirements, build hiring plans, and set realistic timelines. Source active and passive candidates through job boards, LinkedIn Recruiter, employee referrals, and direct outreach. Screen resumes and conduct initial phone/video interviews to assess candidate qualifications and fit. Guide hiring managers and interview panels on structured interviewing and consistent, bias-aware evaluation. Manage candidate communication and experience throughout the process, keeping candidates informed at every stage. Negotiate offers and

aigoexcel
View job →
S
Smartsheet
📍 Bengaluru• Full-time
20 days ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. You Will: Data Architecture and Design: Designing and overseeing the architecture of scalable and reliable data platforms, including data pipelines, storage solutions, and processing systems Data Modelling and Management:Developing and implementing data models, ensuring data quality, and establishing data governance policies Data Pipeline Development: Building and optimising data pipelines for ingesting, processing, and transforming large datasets from various sources Performance Optimisation: Identifying and resolving performance bottlenecks in data pipelines and systems, ensuring efficient data retrieval and processing Technology Evaluation and Innovation: Staying abreast of emerging data technologies and exploring opportunities for innovation to improve the organisation’s data infrastructure Troubleshooting and Problem Solving: Diagnosing and resolving complex data-related issues, ensuring the stability and reliability of the data platform Data Security and Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data Perform other duties as assigned You Have: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field. 10+ years of experience in data engineering or a similar role. Enterprise SaaS software solutions with high availability and scalability Solution handling large scale structured and unstructured data from varied data sources Experience in building and maintaining data platform systems such as distributed compute,

pythonjavasql
View job →
S
21 days ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet is looking for a highly technical Sr. Manager to lead our India-based Service Desk and IT Automation operations. This is a working-manager role: you'll personally execute technical work — scripting, automation builds, escalated ticket resolution — while also leading a team. If you're looking for a purely strategic or delegation-only leadership role, this isn't it. You'll manage a team spanning Service Desk, Tier 1 SOC support, and IT Automation Engineering, driving the technical maturity of our global support infrastructure through automation, scripting, deep systems integration, and increasingly, AI-augmented workflows. This role reports to the Director of End User Systems, with close, day-to-day partnership across US-based Desktop Support managers. For the right candidate, this role has a clear path to Director as we continue scaling our India GCC — with growth into that level anticipated within 12–18 months. This is also a strong fit for a Director-level leader currently at a SaaS startup who's looking to move into a larger, enterprise SaaS organization — and who is open to a step down in title now in exchange for a step up in scale, platform, and long-term trajectory. You Will IT Automation & Architecture Drive the IT automation function, identifying repetitive Tier 1 tasks and owning their automation end-to-end — from scoping through build, testing, and measurement Evaluate and build AI-augmented automation where it meaningfully reduces toil (e.g. intelligent ticket triage, SOC alert summarisation, AI

pythonawsai
View job →
S
Smartsheet
📍 USA-• Full-time• Remote• From $1.5M/yr
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. We are looking to hire a Product Manager II to join our Product Organization. You'll own a set of experiences and help define what "great" looks like as AI reshapes how customers work across Smartsheet. You’ll help identify the needs of our customers, and empower customers to drive meaningful change in their organizations. You’ll report to our Director, Product Management. You may work from our Bellevue, WA office, or remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Own a set of features end-to-end — defining and shipping them against clear success criteria Define a set of features/experiences based on customer needs, market trends, and business outcomes Make the feature-level release calls (ship, block, or defer) and own quality and readiness for what you ship, including AI behavior where it's involved Use internal and external data to guide experimentation, decisions, opportunities and evaluation of success Partner with research, product design, data science and analytics teams to validate hypotheses Communicate strategy, product insights, winning and losing experimentation, etc. to a wide audience You Have: A technical degree, or equivalent experience working closely with engineers to make product decisions on behalf of customers Two or more years of product management experience, ideally building tools that people use every day A data-informed mindset: you define and track feature metrics, dig into the "why" behind them, and separate real customer value from hype Experience partne

REMOTEvueawsagile
View job →
SF
Stitch Fix
📍 Remote• Full-time• Remote• From $225K/yr
1mo ago

About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role The Client Experience Product Algorithms team is responsible for the machine learning, AI, experimentation, and product analytics capabilities that power personalized experiences for Stitch Fix clients and stylists. Partnering across Product, Engineering, Design, Styling, Marketing, Merchandising, Finance, Enterprise Analytics, Data Platform, and DSN, the team translates data and algorithms into measurable business impact. As Director, Product Algorithms, you will lead the strategy, execution, and people behind our Growth, Styling, and Fix & Freestyle Algorithms portfolios. You'll define how AI, machine learning, experimentation, and analytics shape the future of personalized shopping while building the operating discipline, technical excellence, and cross-functional alignment needed to deliver scalable business results. Responsibilities Lead the Product Algorithms portfolio across Growth, Styling, and Fix & Freestyle, setting strategy and driving measurable outcomes across acquisition, engagement, retention, styling quality, Fix, Freestyle, outfitting, and related commerce experiences. Define the vision and roadmap for applying data science, machine learning, AI, experimentation, and product analytics to improve client experiences, stylist effectiveness, and business performance. Drive innovation by identifying, evaluating, and scaling modern AI, machine learning, personalization, and experimentation techniques that create meaningful impact while balancing technical feasibility, exec

REMOTEgitrestmachine learning
View job →
S
Smartsheet
📍 USA-• Full-time• Remote• From $1.5M/yr
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. As the Marketing Platform Operations Manager, you will own data quality and data flow across Smartsheet's marketing technology stack with a particular focus on how lead and account data moves between systems, where it breaks down, and how it can move better in an AI-first environment. You'll dig into the mechanics of our lead flow: where data originates, how it's enriched and scored, how it syncs across platforms, and where gaps or failures cost the business pipeline. This is a hands-on, systems-fluent role. You'll partner closely with Marketing Platform Operations, Campaign Operations, Analytics Engineering, and Sales teams to diagnose data issues, improve data architecture, and support the evaluation and rollout of new tools that improve enrichment, orchestration, and data integrity across the funnel. You'll also contribute to broader Marketing Operations initiatives as they arise. This role reports to the Director, Marketing Platform Operations and can be based in our Bellevue, WA office or remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Own data quality standards across CRM and Marketing Automation Platform (MAP), including field completeness, accuracy, and consistency. Lead deduplication efforts and ongoing hygiene initiatives across contact, lead, and account records. Diagnose root causes of recurring data quality issues and design scalable fixes rather than one-off cleanups. Establish and monitor data quality metrics and reporting to track health over time. Map and docum

REMOTEvuesqlaws
View job →
🔔

Get new human evaluator jobs by email

Daily job updates · Unsubscribe anytime