Jobs in United States

Human Evaluator in United States

1,785 active opportunities · Updated October 2026

Explore current human evaluator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -83.9%

About the Team The Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Model Policy team aligns model behavior with desired human values and norms. We co-design policy with models and for models by driving rapid policy taxonomy iteration based on data and defining evaluation criteria for foundational models’ ability to reason about safety. Key focus areas include: catastrophic risk, mental health, teen safety and multimodal safety. About the Role Providing access to frontier AI systems raises complex questions around dual-use science and catastrophic risk. How should models respond to requests involving chemical synthesis, biological experimentation, or pathogen research? Where is the boundary between legitimate scientific inquiry and information that could enable misuse? How do we design policies that meaningfully reduce risk without unnecessarily restricting beneficial research? This is a senior role in which you’ll help shape policy creation and development at OpenAI for addressing biological and chemical risks. You will develop structured policy frameworks and taxonomies to guide safe model behavior. This role sits at the intersection of biosecurity expertise, AI safety research, and policy design. You will help ensure that frontier AI systems can support beneficial life sciences research, such as drug discovery, public health, and biosafety, while reducing the risk that these capabilities could be misused. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you’ll: Design and maintain model policies governing chemical and biological risk, defining how models should safely handle dual-use scenarios. Develop structured taxonomies of chemical and biological risk that inform model training data, evaluation benchmarks, and safet

AWSGitRestAI
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Our Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Model Policy team aligns model behavior with desired human values and norms. We co-design policy with models and for models by driving rapid policy taxonomy iteration based on data and defining evaluation criteria for foundational models’ ability to reason about safety. Key focus areas include: catastrophic risk, mental health, teen safety and multimodal safety. About the Role Providing access to frontier AI systems raises complex questions around dual-use science and catastrophic risk. How should models respond to requests involving chemical synthesis, biological experimentation, or pathogen research? Where is the boundary between legitimate scientific inquiry and information that could enable misuse? How do we design policies that meaningfully reduce risk without unnecessarily restricting beneficial research? This is a senior role in which you’ll help shape policy creation and development at OpenAI for addressing biological and chemical risks. You will develop structured policy frameworks and taxonomies to guide safe model behavior. This role sits at the intersection of biosecurity expertise, AI safety research, and policy design. You will help ensure that frontier AI systems can support beneficial life sciences research, such as drug discovery, public health, and biosafety, while reducing the risk that these capabilities could be misused. Our relevant publications: Preparedness framework Preparing for future AI capabilities in biology Safety evaluations hub OpenAI GPT5 System Card Evaluating Fairness in ChatGPT Improving Model Safety Behavior with Rule-Based Rewards OpenAI Model Spec Your Responsibilities: Design and maintain model policies governing chemical and biological risk, defining how models should safely handle dual-use scenarios. Develop structured taxonomi

AWSGitRestMachine Learning
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Safety Systems manages the complete lifecycle of safety efforts for OpenAI’s frontier models, ensuring our models are deployed responsibly and have a positive impact on society. Our work spans diverse research and engineering initiatives—from system-level safeguards and model training to evaluation and red-teaming—all aimed at mitigating misuse and maintaining our high bar for safety. We lead OpenAI's commitment to developing and deploying safe Artificial General Intelligence (AGI), fostering a culture of trust, responsibility, and transparency. Our goal is to continuously learn from deployments, distribute AI’s benefits widely, and ensure that powerful tools remain aligned with human values and safety considerations. About the Role The Safety Measurement Product Manager owns OpenAI's approach to measuring harm and safeguard efficacy in production, including driving the strategy for our suite of safety measurement platforms and products used across the company. You will partner closely with our safety research and engineering teams to determine what we measure, where we measure it, and how we measure it, feeding those insights directly into critical leadership decisions and back into our safety work. You will also represent the company's topline safety metric as well as prioritize incoming requests from partner teams to expand our safety measurement platform to more use cases. This position is based in San Francisco, CA, with relocation assistance available. In this role, you will: Partner closely with data science, research, engineering, policy teams, and other stakeholders to craft a vision for understanding safety outcomes and prevalence on our platforms. Define strategic priorities and product roadmaps focused on improving safety measurement approaches will scaling our measurement platform to more use cases, products, and cross-functional team needs. Establish repeatable processes to integrate cutting-edge AI safety research into OpenAI’s safety m

AWSRestAIGo
M
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

ABOUT THE TEAM The AI Foundations Team at Mural is pioneering how generative AI transforms visual collaboration and decision-making. We’re a remote-first group of engineers, designers, and product thinkers focused on helping teams work together more effectively. Our goal isn’t to replace human creativity. It’s to amplify it, building AI that enhances how people align, communicate, and make decisions visually. YOUR MISSION You will design and build the core AI systems and platforms that enable Mural’s next wave of agentic, AI-driven collaboration experiences. Rather than building isolated AI features, you’ll work on the core backend systems that power Mural’s agent platform, including agent orchestration, durable execution, contextual memory, tool integration, observability, and evaluation. Your work will enable intelligent agents to reason over product context, act on behalf of users, and operate reliably and safely at scale. Our stack at Mural includes Azure OpenAI, React, Node, MongoDB. WHAT YOU'LL DO Build the core backend systems that power Mural’s agent platform, including orchestration, durable execution, tool execution, memory, observability, and evaluation infrastructure Design scalable services and APIs that allow AI agents to retrieve context, coordinate multi-step workflows, interact with Mural data, and act reliably on behalf of users Develop the agent memory layer, including systems for conversation context, product context, retrieval, summarization, compaction, and long-term context management Create infrastructure to monitor, debug, and improve agent behavior through traces, metrics, feedback loops, and offline evaluation Translate complex, open-ended product needs into clear backend architectures, service boundaries, data models, and implementation plans that align technical capabilities with user value Help define the technical direction for agentic AI at Mural, contributing to long-term architecture and strategy Champion engineering excellence, men

ReactMongoDBAzureAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $525.5K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Senior Director, Generative AI About the Role Roblox Build is our generative creation product, the platform where creators design, build, and publish 3D experiences. We are looking for a Senior Director of Generative AI to lead the Applied AI organization inside Build, responsible for turning state-of-the-art foundation models into high-quality, reliable creation systems at Roblox scale. This leader will own the full applied AI stack: model strategy and routing, model adaptation and fine-tuning, code generation (CodeGen), 3D layout generation (LayoutGen), and the evaluation science and infrastructure that tells us what actually works. You Will Own model strategy and routing for Build. Design and build an intelligent model layer that selects the right model for each creation task based on quality, capability, latency, cost, and safety, leveraging both frontier models and Roblox-adapted open-source models. Lead model adaptation across the Applied AI org, including fine-tuning, distillation, synthetic data generation, human feedback pipelines, and preference optimization for Roblox-specific creation tasks such as Luau code generation and 3D scene understanding. Drive CodeGen capabilities

AWSGitRestMachine Learning
S
📍 Bellevue, WA, United States· Full-time
✓ High-confidence listingCompany trend -92%

From $1.6M/yr

Quick readStrong listing-quality and freshness signals

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Role Overview: As the Sr. CE Program Manager, you will serve as a strategic partner to the Sr. Director, Customer Adoption Programs & Strategy in leading customer excellence programs. This role combines day-to-day program management with strategic planning, cross-functional collaboration, and performance analytics. You will drive strategic customer experience initiatives, evaluate new tools and processes, own performance reporting, and build frameworks that inform how we optimize for adoption, retention, and customer value. This role requires both tactical execution excellence and the ability to think strategically about how business changes impact our customer experience programs, while partnering closely with BI, Customer Success, Scale, Product Marketing, and other teams to deliver unified customer experiences. You Will: Strategic Leadership & Representation Represent the customer experience in field meetings and cross-functional forums, ensuring customer experience strategies align with broader Customer Success objectives Serve as a strategic thought partner to the Customer Experience and Customer Success teams on opportunities and decision-making that further enhance the customer journey Support the evaluation of tools, processes, and workflows Partner with business partner leadership on strategic initiatives that bridge digital and scale motions Present on behalf of leadership in internal and external forums Program Strategy & Execution Design, build, and iterate multi-channel customer engagement

S
📍 Bellevue, WA, United States· Full-time· Remote
✓ High-confidence listingCompany trend -92%

From $1.9M/yr

Quick readStrong listing-quality and freshness signals

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet is seeking an experienced sales leader to oversee a team of sales managers as a Regional Vice President (RVP) of Commercial Sales . The ideal candidate will scale, operate and continuously improve our commercial sales business through high performing teams, strategies and execution. This leadership role is based in the US and reports directly to a VP of Commercial Sales. You Will: Recruit, hire, coach, and develop sales talent across 35+ individual contributors and 4+ managers Ensure the team exceeds sales and operational targets on a monthly and quarterly basis Engage with our SMB customer base to directly work on strategic sales opportunities Provide day-to-day leadership of sales team and ensure proper application of sales methodology and departmental standards Develop, execute, evaluate and refine aspects of sales strategy, process, and workflow to deliver maximum customer engagement, bookings and revenue growth Monitor performance of sales representatives and teams for quota attainment, adherence to policies, and development of action plans, as needed. Develop and deliver business operational reviews summarizing key performance drivers and critical insight to increase future sales opportunities and required strategies and tactics. Create a high-performing, energized, and supportive culture, emphasizing reward and recognition You Have: Minimum of 7+ years of sales management experience, preferably in a multi-contact, high volume inbound and outbound environment 2+ years of experienc

SF
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -100%

From $125K/yr

Quick readStrong listing-quality and freshness signals

About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As a Lead Engineer on the Product Catalog Manager Team, you will help set the technical direction for the systems that power Stitch Fix’s product data ecosystem. You will work on the tools, workflows, and data models that support the full product lifecycle, from new style creation and catalog enrichment to product readiness, validation, and downstream product experiences. You will own complex problem spaces from discovery through delivery, translate business and merchandising needs into scalable technical solutions, and lead execution across ambiguous, cross-functional initiatives. This role requires strong technical judgment, deep ownership, clear communication, and the ability to influence partners across Engineering, Product, Merchandising, Data Science, and Operations. Your work will directly impact product data quality, catalog accuracy, merchandising efficiency, product readiness, and the client experience. Responsibilities: Own and evolve critical catalog systems, including product onboarding, attribute management, data enrichment, validation workflows, and product readiness tooling. Design and operate scalable services and data models that ensure product information is accurate, complete, consistent, and available to downstream systems. Drive discovery with Product, Merchandising, Data Science, and Operations partners to identify high-impact problems, evaluate tradeoffs, and define clear technical roadmaps. Independently lead initiatives from concept through production rollout, including tec

PythonJavaSQLMySQL
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team GTM Growth Engineering builds AI-native products that help OpenAI's go-to-market and B2B marketing organizations scale with greater speed, intelligence, and operational effectiveness. We apply OpenAI models to real business workflows and build the systems that make those applications useful and dependable: customer context, agent behavior, feedback, evaluation, experimentation, and appropriate human oversight. Our work brings together software engineering, applied AI, product, data, and GTM operations. We measure success through the quality of customer engagement, pipeline, conversion, and the effectiveness of our sales and marketing teams. About the Role We're looking for an Applied AI Engineer to build production systems that help AI-powered go-to-market workflows improve over time. You will connect agent behavior, customer and operator feedback, evaluation, experimentation, and business outcomes to make these systems more effective, reliable, and responsive to evolving customer needs. This is a deeply technical, cross-functional role with end-to-end ownership of the agent improvement loop: understand production behavior, identify failure modes, improve how the system decides or acts, and validate the resulting impact. You will partner with Engineering, Product, Data Science, Sales, and B2B Marketing to turn real-world signals into safer, more effective agent behavior and measurable improvements in customer engagement, conversion, qualified pipeline, and team productivity. In this role, you will: Own the production improvement loop across agent behavior, customer and operator feedback, evaluation, experimentation, and verified business outcomes. Instrument agent workflows so model interactions, tool use, decisions, failures, human edits, and downstream outcomes can be understood in context. Define meaningful quality standards, representative evaluation datasets, regression coverage, and production monitoring for real GTM workflows. Investigate why a

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Personal AGI team is responsible for training and improving pre-trained models to be deployed into ChatGPT, the API, and potential future products. In the Model Experience team, we shape the default character and behavior of ChatGPT: how the model communicates, responds to users, uses its capabilities, and behaves across different contexts and languages. Our goal is to make every interaction with ChatGPT thoughtful, helpful, and trustworthy. We take an opinionated view of what good human–AI interaction should look like, then turn that vision into real model behavior through human data, evaluations, reward models, and post-training. Our work sits at the intersection of research, product, and model design. We partner closely with teams across OpenAI to conduct research and ensure our models are thoughtful, safe, reliable to serve millions of users. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models. Our team works in research areas combining reinforcement learning and products. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research and the quality of human-AI interaction. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve model capability and performance. Collaborate closely with the other research and product teams, allowing customers to optimize their own models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. You might thrive in this role if you: Have a deep understanding of machine learning and machine learning applications. Have good judgment about model behavior and can communicate this judgment effec

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Alignment team at OpenAI is dedicated to ensuring that our AI systems are safe, trustworthy, and consistently aligned with human values, even as they scale in complexity and capability. Our work is at the cutting edge of AI research, focusing on developing methodologies that enable AI to robustly follow human intent across a wide range of scenarios, including those that are adversarial or high-stakes. We concentrate on the most pressing challenges, ensuring our work addresses areas where AI could have the most significant consequences. By focusing on risks that we can quantify and where our efforts can make a tangible difference, we aim to ensure that our models are ready for the complex, real-world environments in which they will be deployed. The two pillars of our approach are: (1) harnessing improved capabilities into alignment, making sure that our alignment techniques improve, rather than break, as capabilities grow, and (2) centering humans by developing mechanisms and interfaces that enable humans to both express their intent and to effectively supervise and control AIs, even in highly complex situations. About the Role As a Research Engineer / Research Scientist on the Alignment team, you will be at the forefront of ensuring that our AI systems consistently follow human intent, even in complex and unpredictable scenarios. Your role will involve designing and implementing scalable solutions that ensure the alignment of AI as their capabilities grow and that integrate human oversight into AI decision-making. This role is especially well suited for someone who can move from an ambiguous model-behavior question to a concrete experimental setup: formulate the hypothesis, build the evaluation or intervention, run the experiment, analyze the result, and decide what the evidence supports. This role may be based in San Francisco or London, subject to team needs and location approval. In this role, you will: We are seeking research engineers and res

TypeScriptPythonAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Our Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. Within Safety Systems, the Model Policy team aligns model behavior with desired human values and norms. We co-design policy with models and for models by driving rapid policy taxonomy iteration based on data and defining evaluation criteria for foundational models’ ability to reason about safety. About the Role Frontier AI systems are rapidly expanding what is possible in cybersecurity and software engineering. These capabilities create major defensive opportunities, but they also raise serious dual-use and misuse risks across areas such as malware development, exploit discovery, vulnerability chaining, credential abuse, cyber intrusion, and autonomous offensive operations. In this role, you will help define how OpenAI’s models should behave in high-risk cybersecurity contexts. You will develop policy frameworks, threat models, taxonomies, evaluations, and behavioral specifications that guide model behavior across training, deployment, and monitoring systems. This role sits at the intersection of cybersecurity, AI safety, threat modeling, evaluation science, and policy implementation. You will work closely with research, engineering, safety training, preparedness, and product teams to build policies that are technically grounded, measurable, enforceable, and responsive to real-world cyber risk. Your Responsibilities: Design and maintain model policies for cybersecurity and frontier-risk domains, especially dual-use and high-risk cyber capabilities. Translate cybersecurity threat models into clear behavioral specifications, evaluation criteria, grading guidance, and system-level mitigations. Define practical boundaries between legitimate security research, defensive workflows, and assistance that could materially enable harmful activity. Build policy artifacts that support i

AWSGitRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Safety Systems team is responsible for various safety work to ensure our best models can be safely deployed to the real world to benefit the society, and is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Safety Oversight Research team aims to fundamentally advance our capabilities to maintain oversight over frontier AI models, and leverage these advances to ensure OpenAI’s deployed models are safe and beneficial. This requires a breadth of new ML research in the areas of human-AI collaboration, reasoning, robustness, and scalable oversight to keep pace with model capabilities. We invest heavily in developing novel model and system-level methods of identifying and mitigating AI misuse and misalignment. Our goal is to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. About the Role OpenAI is seeking a senior researcher with a passion for AI safety and experience in safety research. Your role will set directions for research to maintain effective oversight of safe AGI and work on research projects to identify and mitigate misuse and misalignment in our AI systems. You will play a critical role in defining how a safe AI system should look in the future at OpenAI, making a significant impact on our mission to build and deploy safe AGI. In this role, you will: Develop and refine AI monitor models to detect and mitigate known and emerging patterns of misuse and misalignment. Set research directions and strategies to make our AI systems safer, more aligned, and more robust. Evaluate and design effective red-teaming pipelines to examine the end-to-end robustness of our safety systems, and identify areas for future improvement. Conduct research to improve models’ ability to reason about questions of human values, and apply these improved models to practical safety challen

PythonAWSRestMachine Learning
I
📍 California, United States· Hybrid
✓ High-confidence listingCompany trend -23.5%

$82.5K – $123.7K/yr

Quick readStrong listing-quality and freshness signals

What if the work you did every day could impact the lives of people you know? Or all of humanity? At Illumina, we are expanding access to genomic technology to realize health equity for billions of people around the world. Our efforts enable life-changing discoveries that are transforming human health through the early detection and diagnosis of diseases and new treatment options for patients. Working at Illumina means being part of something bigger than yourself. Every person, in every role, has the opportunity to make a difference. Surrounded by extraordinary people, inspiring leaders, and world changing projects, you will do more and become more than you ever thought possible. Position Title: Senior Financial Analyst, R&D Finance Location: San Diego, CA (Hybrid) Reports To: Associate Director, Finance – R&D FLSA Status: Exempt Position Summary: The R&D Finance Business Partner provides finance support for a complex portfolio of R&D functions, programs, strategic initiatives, and R&D leadership team members. The role works with R&D and Finance leadership to translate technical and operational activity into clear financial insight, improve forecast accuracy, evaluate investment trade-offs, and support resource decisions across operating expense, headcount, and capital investment. This position operates as an individual contributor in a matrixed environment with limited day-to-day direction. Success requires strong financial rigor, sound business judgment, clear communication, and the ability to build trusted partnerships while maintaining appropriate financial challenge and discipline. *This is a full-time role, Monday through Friday, with an expectation of 2–3 in-office days per week and additional on-site presence as

Artificial IntelligenceAISapExcel
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -83.9%
Quick readStrong listing-quality and freshness signals

About the role As a Pricing Strategist focused on GTM, you will help shape pricing strategy for B2B enterprise customers across our product portfolio. Working within Finance and partnering with GTM, Sales, and Product, you will focus on pricing analytics, price performance, and the effectiveness of credit programs while contributing to broader commercial strategy. In this role, you will Build regular pricing-performance reviews, analyze discounting and concessions, and turn findings into better pricing guidance. Develop commercial strategies and pricing programs for customer segments such as education, startups, and government. Define objectives and success measures for credit programs and promotions, and evaluate their return on investment. Translate segment goals and product pricing strategy into scalable pricing frameworks and commercial structures. Partner on strategic enterprise deals to align commercial proposals with sound deal economics and pricing strategy. You might thrive in this role if you Approach problems from first principles, identify root causes, and develop clear options for stakeholders. Work effectively through ambiguity and differing perspectives. Understand how sales organizations operate and collaborate well across GTM strategy, Finance, and Product. Bring relevant experience from consulting, strategy and operations, or pricing, particularly work involving analytical teams. Experience with pricing analytics, pricing frameworks, discounting guardrails, or new-product pricing is helpful, but a strictly pricing-specific background is not required. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we mu

Artificial IntelligenceAIFinance
🔔

Get new human evaluator jobs in United States by email

Daily job updates · Unsubscribe anytime