Jobiba hiring network

Human Evaluator Jobs

3,920 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current human evaluator jobs. Use filters to narrow by work mode, employment type, experience and date posted.

SF
Stitch Fix
📍 San Francisco• Full-time• From $225K/yr
1mo ago

About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role The Client Experience Product Algorithms team is responsible for the machine learning, AI, experimentation, and product analytics capabilities that power personalized experiences for Stitch Fix clients and stylists. Partnering across Product, Engineering, Design, Styling, Marketing, Merchandising, Finance, Enterprise Analytics, Data Platform, and DSN, the team translates data and algorithms into measurable business impact. As Director, Product Algorithms, you will lead the strategy, execution, and people behind our Growth, Styling, and Fix & Freestyle Algorithms portfolios. You'll define how AI, machine learning, experimentation, and analytics shape the future of personalized shopping while building the operating discipline, technical excellence, and cross-functional alignment needed to deliver scalable business results. Responsibilities Lead the Product Algorithms portfolio across Growth, Styling, and Fix & Freestyle, setting strategy and driving measurable outcomes across acquisition, engagement, retention, styling quality, Fix, Freestyle, outfitting, and related commerce experiences. Define the vision and roadmap for applying data science, machine learning, AI, experimentation, and product analytics to improve client experiences, stylist effectiveness, and business performance. Drive innovation by identifying, evaluating, and scaling modern AI, machine learning, personalization, and experimentation techniques that create meaningful impact while balancing technical feasibility, exec

gitrestmachine learning
View job →
SF
Stitch Fix
📍 United States• Full-time• Remote• From $125K/yr
1mo ago

About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As a Lead Engineer on the Product Catalog Manager Team, you will help set the technical direction for the systems that power Stitch Fix’s product data ecosystem. You will work on the tools, workflows, and data models that support the full product lifecycle, from new style creation and catalog enrichment to product readiness, validation, and downstream product experiences. You will own complex problem spaces from discovery through delivery, translate business and merchandising needs into scalable technical solutions, and lead execution across ambiguous, cross-functional initiatives. This role requires strong technical judgment, deep ownership, clear communication, and the ability to influence partners across Engineering, Product, Merchandising, Data Science, and Operations. Your work will directly impact product data quality, catalog accuracy, merchandising efficiency, product readiness, and the client experience. Responsibilities: Own and evolve critical catalog systems, including product onboarding, attribute management, data enrichment, validation workflows, and product readiness tooling. Design and operate scalable services and data models that ensure product information is accurate, complete, consistent, and available to downstream systems. Drive discovery with Product, Merchandising, Data Science, and Operations partners to identify high-impact problems, evaluate tradeoffs, and define clear technical roadmaps. Independently lead initiatives from concept through production rollout, including tec

REMOTEpythonjavasql
View job →
R
1mo ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Red Team team’s mission is to proactively identify and simulate real-world threats against Robinhood’s platforms, properties, and people. Through red teaming and adversarial simulations, the team evaluates security controls, uncovers vulnerabilities, and helps continuously strengthen Robinhood’s overall security posture in close partnership with Detection & Response, Physical Security, and Engineering. As a Staff Offensive Security Engineer, you will take a hands-on role in designing and executing stealthy adversarial simulations to validate assumptions and uncover gaps in detection and response. You’ll leverage threat modeling, penetration testing, and research-driven techniques to emulate sophisticated attackers, while collaborating cross-functionally to improve defenses and shape more secure systems. This role is based in our Toronto, Canada office(s), with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. What you’ll

javascriptpythonjava
View job →
R
1mo ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the Team + Role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Proactive Capabilities team builds novel tooling and leverages state-of-the-art automation to eliminate business impact from threat actors! Our mission is to close critical attack paths before they can be exploited and to mitigate active threats with speed and precision. As a Software Engineer on the Proactive Capabilities team , you will design, build, and scale engineering tools that empower Security Engineers, eliminate operational toil, and expand security coverage across Robinhood. This is a creative, high-impact role where you will identify high-priority security challenges, evaluate buy-vs-build opportunities, and turn one-off solutions into robust, company-wide security capabilities! This role is based in our Bellevue, WA and Menlo Park, CA offices, with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. What You'll Do Build and scale flexible, reliable security tooling and services that improve visibility and expand scan c

pythonvueaws
View job →
R
1mo ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. Robinhood's Technical Accounting team sits at the intersection of innovation and rigor — translating novel products and complex transactions into sound, defensible accounting positions. We're looking for a collaborative, analytically sharp problem-solver to join us as a Technical Accounting Manager supporting new products and strategic transactions. In this role, you'll evaluate emerging and often ambiguous fact patterns, form well-reasoned accounting conclusions, and help shape how new products and deals are structured from an accounting perspective. You'll work cross-functionally with Product, Legal, Controllership, Tax, Finance, Financial Reporting, Internal Controls, and our external auditors to ensure Robinhood's accounting keeps pace with its fast-evolving business. This role is based in our Menlo Park, CA; or New York, NY offices, with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. What you’ll do Own technical accounting res

vueawsgit
View job →
R
Robinhood
📍 Menlo Park• Full-time
1mo ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Finance & Strategy team partners closely with Finance leadership and teams across Product, Engineering, Recruiting, Procurement, and Accounting to guide key business decisions. The team focuses on financial planning, investment analysis, and portfolio performance tracking to support strategic initiatives, including venture fund activities. The work centers on developing clear financial insights that help leadership understand performance trends and evaluate opportunities. The team values thoughtful analysis, clear communication, and strong collaboration across functions! As a Finance & Strategy Senior Analyst, this role supports venture fund initiatives through financial modeling, portfolio analysis, and planning and reporting processes. The position works closely with senior leaders to deliver data-driven insights that inform capital allocation and strategic priorities. Responsibilities span both structured financial reporting and flexible analysis across business areas. The role also contributes to improving financial processes through standardization and automation, along with advancing AI-driven workflow improvements! This role is based in our Menlo Park, CA

vueawsai
View job →
R
1mo ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Financial Crimes Team at Robinhood works across multiple verticals to protect our customers and ensure strict compliance with all relevant Anti-Money Laundering (AML) and Sanctions regulations. As our team of passionate professionals continues to evolve the program, we are intensely focused on leveraging innovative thinking, advanced analytics, and cutting-edge technology to build a best-in-class Financial Crimes Program. As the Senior Specialist, Model Validation & Analytics, you will play a key role in this mission by executing the validation, testing, tuning, and optimization of various transaction monitoring and filtering programs across Financial Crimes. You will partner closely with the AML Surveillance Team, engineering partners, and cross-functional verticals to ensure our financial crimes models and systems are accurately implemented, rigorously evaluated, and effectively calibrated to meet the diverse regulatory requirements of our domestic and international lines of business. This role is based in our Denver, CO, New York, NY, or Westlake, TX office(s), with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power

pythonvuesql
View job →
R
Robinhood
📍 Menlo Park• Full-time
1mo ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. Robinhood’s mission has always been to democratize finance for all. With the launch of Robinhood Ventures, we’re taking a bold step forward - bringing access to private market investing to everyone, not just accredited investors. This new product line is designed to unlock opportunities traditionally limited to a select few, helping more people participate in the growth of private companies. As a Deal Lead for Robinhood Ventures , you will be at the forefront of sourcing and executing high-impact investments for our flagship publicly traded fund. You will be responsible for identifying new opportunities, conducting thorough financial and operational diligence, and supporting closing transactions in a competitive market. You will manage the full investment lifecycle—from sourcing and evaluation through investment committee materials, negotiation support, and portfolio engagement. This role reports to the Head of Robinhood Ventures and offers a rare opportunity to help shape a new product from its early stages while contributing to long-term platform growth. This role is based in our Menlo Park, CA office, with in-person attendance expected at least 3 days per week

vueawsrest
View job →
BA
Bolna AI
📍 India• Full-time
1mo ago

At Bolna, we’re building tools that change the way teams leverage Voice AI. We’re looking for a Founding Machine Learning Engineer to own the end-to-end lifecycle of building, evaluating, deploying, and improving models that power millions of production conversations. This is a high-impact, high ownership role where you won’t just work on Bolna’s ML stack—you’ll help build the foundation it scales on. Our team includes IIT alumni with experience at Bain, Atlassian, Uber, Zomato, and LinkedIn, and is backed by leading investors. Responsibilities: Build the data engine - Design pipelines to source and clean conversational voice data across Indian languages, accents, and telephony conditions. Fine-tune models that ship - Fine tune and train models to improve accuracy, speed, and reliability across different use-cases. Define what "good" means - Build evaluation datasets and benchmarks for transcription accuracy, voice naturalness, interruption handling, latency, and end-to-end conversation quality. Set up human-in-the-loop pipelines to capture subjective quality at scale. Ship to production - Work with the engineering team to deploy models into a latency-sensitive, high-volume system. Monitor performance in the wild, debug regressions, and iterate fast. Required Skills: 3+ years of hands-on ML experience with deep practical real-world experience in training models. Strong Python and PyTorch fundamentals with exposure in distributed training, and modern fine-tuning techniques (LoRA, QLoRA, DPO, RLHF, etc.). Training data as a first-class problem. Experience designing data pipelines from collection, cleaning, labeling, deduplication, augmentation and treating data quality as a core engineering discipline. Rigorous about evaluation. You know that "looks good in a demo" is not a benchmark. You build the evals before you trust the model. Speech model experience is a plus with real-time / streaming inference experience where you would have contributed to latency optimization

pythonmachine learningai
View job →
O
1mo ago

About the Team GTM Growth Engineering builds AI-native products that help OpenAI's go-to-market and B2B marketing organizations scale with greater speed, intelligence, and operational effectiveness. We apply OpenAI models to real business workflows and build the systems that make those applications useful and dependable: customer context, agent behavior, feedback, evaluation, experimentation, and appropriate human oversight. Our work brings together software engineering, applied AI, product, data, and GTM operations. We measure success through the quality of customer engagement, pipeline, conversion, and the effectiveness of our sales and marketing teams. About the Role We're looking for an Applied AI Engineer to build production systems that help AI-powered go-to-market workflows improve over time. You will connect agent behavior, customer and operator feedback, evaluation, experimentation, and business outcomes to make these systems more effective, reliable, and responsive to evolving customer needs. This is a deeply technical, cross-functional role with end-to-end ownership of the agent improvement loop: understand production behavior, identify failure modes, improve how the system decides or acts, and validate the resulting impact. You will partner with Engineering, Product, Data Science, Sales, and B2B Marketing to turn real-world signals into safer, more effective agent behavior and measurable improvements in customer engagement, conversion, qualified pipeline, and team productivity. In this role, you will: Own the production improvement loop across agent behavior, customer and operator feedback, evaluation, experimentation, and verified business outcomes. Instrument agent workflows so model interactions, tool use, decisions, failures, human edits, and downstream outcomes can be understood in context. Define meaningful quality standards, representative evaluation datasets, regression coverage, and production monitoring for real GTM workflows. Investigate why a

pythonawsrest
View job →

About the Team The Personal AGI team is responsible for training and improving pre-trained models to be deployed into ChatGPT, the API, and potential future products. In the Model Experience team, we shape the default character and behavior of ChatGPT: how the model communicates, responds to users, uses its capabilities, and behaves across different contexts and languages. Our goal is to make every interaction with ChatGPT thoughtful, helpful, and trustworthy. We take an opinionated view of what good human–AI interaction should look like, then turn that vision into real model behavior through human data, evaluations, reward models, and post-training. Our work sits at the intersection of research, product, and model design. We partner closely with teams across OpenAI to conduct research and ensure our models are thoughtful, safe, reliable to serve millions of users. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models. Our team works in research areas combining reinforcement learning and products. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research and the quality of human-AI interaction. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve model capability and performance. Collaborate closely with the other research and product teams, allowing customers to optimize their own models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. You might thrive in this role if you: Have a deep understanding of machine learning and machine learning applications. Have good judgment about model behavior and can communicate this judgment effec

awsrestmachine learning
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Alignment team at OpenAI is dedicated to ensuring that our AI systems are safe, trustworthy, and consistently aligned with human values, even as they scale in complexity and capability. Our work is at the cutting edge of AI research, focusing on developing methodologies that enable AI to robustly follow human intent across a wide range of scenarios, including those that are adversarial or high-stakes. We concentrate on the most pressing challenges, ensuring our work addresses areas where AI could have the most significant consequences. By focusing on risks that we can quantify and where our efforts can make a tangible difference, we aim to ensure that our models are ready for the complex, real-world environments in which they will be deployed. The two pillars of our approach are: (1) harnessing improved capabilities into alignment, making sure that our alignment techniques improve, rather than break, as capabilities grow, and (2) centering humans by developing mechanisms and interfaces that enable humans to both express their intent and to effectively supervise and control AIs, even in highly complex situations. About the Role As a Research Engineer / Research Scientist on the Alignment team, you will be at the forefront of ensuring that our AI systems consistently follow human intent, even in complex and unpredictable scenarios. Your role will involve designing and implementing scalable solutions that ensure the alignment of AI as their capabilities grow and that integrate human oversight into AI decision-making. This role is especially well suited for someone who can move from an ambiguous model-behavior question to a concrete experimental setup: formulate the hypothesis, build the evaluation or intervention, run the experiment, analyze the result, and decide what the evidence supports. This role may be based in San Francisco or London, subject to team needs and location approval. In this role, you will: We are seeking research engineers and res

typescriptpythonaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. Within Safety Systems, the Model Policy team aligns model behavior with desired human values and norms. We co-design policy with models and for models by driving rapid policy taxonomy iteration based on data and defining evaluation criteria for foundational models’ ability to reason about safety. About the Role Frontier AI systems are rapidly expanding what is possible in cybersecurity and software engineering. These capabilities create major defensive opportunities, but they also raise serious dual-use and misuse risks across areas such as malware development, exploit discovery, vulnerability chaining, credential abuse, cyber intrusion, and autonomous offensive operations. In this role, you will help define how OpenAI’s models should behave in high-risk cybersecurity contexts. You will develop policy frameworks, threat models, taxonomies, evaluations, and behavioral specifications that guide model behavior across training, deployment, and monitoring systems. This role sits at the intersection of cybersecurity, AI safety, threat modeling, evaluation science, and policy implementation. You will work closely with research, engineering, safety training, preparedness, and product teams to build policies that are technically grounded, measurable, enforceable, and responsive to real-world cyber risk. Your Responsibilities: Design and maintain model policies for cybersecurity and frontier-risk domains, especially dual-use and high-risk cyber capabilities. Translate cybersecurity threat models into clear behavioral specifications, evaluation criteria, grading guidance, and system-level mitigations. Define practical boundaries between legitimate security research, defensive workflows, and assistance that could materially enable harmful activity. Build policy artifacts that support i

awsgitrest
View job →
O
OpenAI
📍 India• Full-time
1mo ago

About the Team Our Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. Within Safety Systems, the Model Policy team aligns model behavior with desired human values and norms. We co-design policy with models and for models by driving rapid policy taxonomy iteration based on data and defining evaluation criteria for foundational models’ ability to reason about safety. About the Role If you have a specific expertise or speciality related to this work, please note it in your application via your resume, cover letter or application note. Frontier AI systems are expanding what people can do across domains, creating both enormous opportunities and difficult safety questions: when should a model help, when should it refuse, and how do we make those boundaries clear enough to train, evaluate, and enforce? In this role, you will help define how OpenAI’s models should behave in high-risk or high-ambiguity contexts, such as agentic systems, multimodal systems, user safety, privacy, and other emerging risk domains. This is an ideal role for someone who can move across unfamiliar topics, reason from first principles, and turn ambiguity into practical model behavior. You will work closely with research, engineering, product, preparedness, and operations teams to build policies that are technically grounded, measurable, and responsive to real-world risk. In this role, you will: Design and maintain model policies across safety-relevant domains, including dual-use, agentic, and emerging frontier-risk areas. Translate risk and harm models into clear behavioral specifications, evaluation criteria, grading guidance, and system-level safeguards. Define practical boundaries between beneficial uses of AI and assistance that could materially enable harm, exploitation, misuse, or unsafe outcomes. Build policy artifacts that support model training, evaluation, and deplo

awsgitrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Safety Systems team is responsible for various safety work to ensure our best models can be safely deployed to the real world to benefit the society, and is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Safety Oversight Research team aims to fundamentally advance our capabilities to maintain oversight over frontier AI models, and leverage these advances to ensure OpenAI’s deployed models are safe and beneficial. This requires a breadth of new ML research in the areas of human-AI collaboration, reasoning, robustness, and scalable oversight to keep pace with model capabilities. We invest heavily in developing novel model and system-level methods of identifying and mitigating AI misuse and misalignment. Our goal is to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. About the Role OpenAI is seeking a senior researcher with a passion for AI safety and experience in safety research. Your role will set directions for research to maintain effective oversight of safe AGI and work on research projects to identify and mitigate misuse and misalignment in our AI systems. You will play a critical role in defining how a safe AI system should look in the future at OpenAI, making a significant impact on our mission to build and deploy safe AGI. In this role, you will: Develop and refine AI monitor models to detect and mitigate known and emerging patterns of misuse and misalignment. Set research directions and strategies to make our AI systems safer, more aligned, and more robust. Evaluate and design effective red-teaming pipelines to examine the end-to-end robustness of our safety systems, and identify areas for future improvement. Conduct research to improve models’ ability to reason about questions of human values, and apply these improved models to practical safety challen

pythonawsrest
View job →
🔔

Get new human evaluator jobs by email

Daily job updates · Unsubscribe anytime