Human Data Quality Engineer (Founding Team) Prolific Prolific isn’t just enabling AI innovation – we’re redefining it. While foundational AI technologies are becoming commoditized, Prolific’s human data infrastructure provides the high-quality, diverse data required to train the next generation of AI models. Through our platform, we empower researchers and companies to access a global, ethically curated participant base, ensuring cutting-edge AI research and training grounded in inclusivity and precision. The Role As one of the founding members of Prolific's newly formed AI Data Services team, you'll help build the quality systems behind some of the world's most advanced AI models. Data quality is a strategic priority for Prolific, so this is a high-visibility role with direct exposure to senior stakeholders. This isn't a traditional QA role. We are not looking for someone to review data against a predefined checklist. We are looking for an innovative thinker that can leverage their expertise to define what good means where no definition exists yet. Acting as a strategic thought partner, you’ll work at the intersection of human data, machine learning, evaluation across frontier use-cases that define what high-quality human data looks like for the next generation of advanced AI. This means that much of the work involves novel problems with no established answer, so you’ll be comfortable working through ambiguity.. . Your primary focus is working directly with clients and alongside frontier AI labs, translating what their models need into robust human data and evaluation strategies. Rather than checking quality at the end of a project, you'll engineer quality into every stage of the lifecycle, from study design and participant strategy through to evaluation, launch readiness and client delivery.You will also work alongside our product engineering, and supply teams to define and build the quality infrastructure that will enable us to deliver h
Jobiba hiring network
Human Data Quality Engineer Jobs
15 active opportunities · Updated for September 2026
Fresh results
15 shown
Explore current human data quality engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Human Data Quality Analyst, AI Business Prolific Prolific isn’t just enabling AI innovation – we’re redefining it. While foundational AI technologies are becoming commoditized, Prolific’s human data infrastructure provides the high-quality, diverse data required to train the next generation of AI models. Through our platform, we empower researchers and companies to access a global, ethically curated participant base, ensuring cutting-edge AI research and training grounded in inclusivity and precision. The Role Prolific provides the human data that powers the next generation of AI models, working with frontier labs to capture the complex human judgments researchers need to train, evaluate and improve them. As a Human Data Quality Analyst, you'll be on the front line of making sure that data captures the right signal and is genuinely good enough to do its job. This isn’t traditional, back-office QA. You’ll be doing real analytical work: digging into datasets, identifying patterns and failure modes, investigating why quality has shifted, and turning complex findings into clear insights that help us improve how data is collected, reviewed and delivered. You'll spend real time reading annotations closely, but that is how you gather evidence, not what you produce. What you produce is analysis, practical recommendations and better quality controls. You'll get hands-on exposure to human data, annotation, machine learning pipelines and AI evaluation, working alongside Quality, Engineering, Operations and Delivery on new and evolving problems. There won't always be an established playbook. You'll be guided by our quality engineers, but you'll also need to run your own analysis, test your assumptions and recognise when you need input. It is a role with a steep learning curve from day one and a strong opportunity for someone early in their career to build deep, practical experience in a fast-moving area of AI. What You’ll Be Doing Run day-to-day qualit
About the Team OpenAI’s mission is to ensure that artificial general intelligence (AGI) benefits all of humanity. A key part of achieving that mission is training models that deeply understand and reflect human preferences — the Human Data team is at the heart of that effort. The Human Data engineering team creates the systems that enable scalable, high-quality human feedback. These systems are essential to how OpenAI trains and improves its most advanced models. Engineers on this team collaborate closely with world-class researchers to bring alignment techniques to life — from experimental ideas to production-ready feedback loops. About the Role We’re looking for software engineers to join the Human Data team and build the platforms, prototypes, tools, and infrastructure that power how our AI models are trained, aligned, and evaluated. You’ll partner with researchers and cross-functional teams to bring alignment ideas to life, influence future model training, and shape how models interact with the real world. We’re looking for people who are excited by technical ownership, enjoy working across the stack, and are eager to solve ambiguous problems in a high-impact, fast-paced environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build and maintain robust full-stack systems for feedback collection, data labeling, and evaluation pipelines, while maintaining high levels of security. Translate experimental alignment research into scalable production infrastructure, including inference and model training stacks. Design and iterate on user-facing tools and backend services to support high-quality data workflows Partner with researchers, engineers, and program leads to shape feedback loops and model interaction paradigms Drive infrastructure improvements that enable faster iteration and scaling across OpenAI’s frontier models, from internal r
About Dot Collective We are a new generation consultancy based across UK and EU and founded on the premises of the engineering excellence and empowering people to make an impact. We work with all modern tech stacks and typically run agile scrum on all our projects. About you Are you passionate about data and its transformational powers? Do you like being able to make a huge difference in a limited period of time? We might be just the right place for you. Your key skills and capabilities: Designing and implementing robust python testing automation frameworks using BDD (Behave) ETL Testing: Proven experience in testing ETL pipelines, data validation, and ensuring data quality Scripting automated tests and working collaboratively with other engineers in a continuous build environment Familiarity with CI/CD tools such as AWS pipelines, Github actions Understanding of cloud architecture principles preferably AWS and/or GCP Data Testing experience - Ability to understand data requirements and perform comprehensive data testing Understanding of data engineering Excellent PyTest and SQL skills Nice to have: Behave BDD experience We expect you to have some knowledge about best practices in designing and building scalable and performant cloud-native data platforms and be comfortable with testing them. Our promise to you We will always see you as a human being and will do our very best to support your needs and wellbeing – well-designed co-working and collaboration spaces, remote working patterns that work for you, parenting leave, sabbaticals and ability to work on personal projects. We believe that a geled team is worth its weight in gold – we will do everything we can to avoid breaking well-performing teams. Whilst continuity across every project is not always possible, we thoughtfully assemble high-performing, blended teams with the appropriate levels of exper
About the Team The Human Data team turns human feedback into reliable signals for training and evaluation. We design and run end-to-end programs that capture the depth of human intent behind everyday and high-stakes uses of our models. Our remit spans bespoke data campaigns, scalable synthetic data generation, and product-embedded signals. We partner closely across all research teams to translate these signals into training datasets, novel evaluations, and feedback loops that push the frontier of our models and advance their applications. About the Role As a Program Manager (PGM) in the Human Data team you will partner with our research teams, operations and engineering to execute complex programs for collecting high-quality data. You will be a key interface between our external vendors and AI trainers, ensuring human data campaigns are successfully completed. Your work will play a key role in enabling OpenAI to train safe models that will land in the real world This role is based in our San Francisco HQ. In this role, you will: Work in a high velocity environment, where the outcome of your work will have a direct impact on the models that OpenAI deploy in the real world Work closely with external vendors, trainers and internal researchers to collect, review, and deliver high-quality data Gather requirements, write instructions, define success criteria, and calibrate the AI trainers Use internal tooling to assess labeled data and provide feedback to AI trainers Think critically and share recommendations on tooling and process improvements, optimizing for quality, throughput, and AI trainer experience You’ll thrive in this role if: You thrive in dynamic environments. You are comfortable navigating ambiguity, managing shifting priorities, and adapting to fast-paced changes without missing a beat. You’re curious about AI, LLMs, Agents. While not required, an interest or background in these areas will help you connect the dots in our broader mission. You have a can-do a
Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. About the role The Data team manages the complete lifecycle of data for researchers - from sourcing and large-scale processing to delivering datasets that power our models. Data sits at the heart of our Research efforts and enables all other teams. As part of the Data team, you’ll work with over a million hours of video and audio data. This role exists at the intersection of applied research, data engineering, and ML infrastructure rather than being a traditional research position . You’ll build the world’s best human-centric data lake by collaborating closely with our model training teams. By understanding their requirements, you’ll extract new features and annotations that elevate our datasets. You should be passionate about enhancing model performance through high-quality, accurate datasets. Our infrastructure and pipelines are in great shape, and this role provides room to not only enhance them but also influence the team’s longer-term strategy. What we're looking for: A strong background in data-centric, applied Machine Learning, with hands-on experience improving model performance through data quality, curation, labeling, and evaluation rather than model architecture alone Experience working on the data la
About the Team OpenAI's Human Data Team creates custom data solutions driving groundbreaking research. Our work enhances and evaluates our flagship models and products like ChatGPT, GPT-5, and Sora, and contributes to safety initiatives through collaboration with our Preparedness and Safety Systems teams. About the Role As a Research Program Manager (RPM) in the Human Data team you will partner with research and engineering to design and implement pragmatic solutions for collecting high-quality data. You will be a key interface between our research roadmap, external vendors, AI trainers, and the Human Data engineering team. This role is based in our San Francisco HQ. In this role, you will: Collaborate with Research: Partner with researchers to scope data collection needs, define success metrics, and establish quality measurement frameworks. Design & Execute Data Collection Campaigns: Translate research needs into actionable plans and accelerate execution by leveraging existing tooling and iterating to reach the desired outcome. In many cases, you will need to implement scrappy new solutions while partnering with engineering to design robust/scalable solutions. Unblock Yourself: You must be deeply uncomfortable with the idea of sitting around waiting for external dependencies, and have the technical acumen and drive to figure out how to achieve at least partial success in the interim. Optimize Systems & Processes: Build and optimize dashboards to track campaign performance, leveraging SQL and Python for data analysis and actionable insights. Drive Technical Roadmaps: Collaborate with engineers to enhance data platforms, resolve blockers, and ensure security best practices such as access management. Scale Your Impact : Advise and empower program managers and vendors to drive day-to-day execution so that you can focus on addressing high priority opportunities. You might thrive in this role if you: Are proficient in SQL and Python for data analysis, including q
About Scale AI At Scale AI, our mission is to accelerate the development of AI applications. For 10 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI) and building upon our prior model evaluation work with enterprise customers and governments to deepen our capabilities and offerings for public and private evaluations. About Data Engine Our Generative AI Data Engine powers the world’s most advanced LLMs and generative models through world-class RLHF (Reinforcement Learning with Human Feedback), human data generation, model evaluation, safety, and alignment. The data we produce is some of the most critical work for how humanity will interact with AI. About Our FDE Team Generating high-quality data is the core problem our business solves. We aim to make producing and delivering high-quality data seamless and efficient for operators and customers. Our Team is building customer and operator-specific infrastructure to provide high-quality data with low turnaround time. You'll be exposed to the cutting edge of the Generative AI industry while directly interfacing with the leading model-building organizations in the space, including the top AI research labs and government agencies. Join us in shaping the future of Artificial General Intelligence. As a Forward Deployed Engineer, you'll be at the forefront of providing the critical data infrastructure that powers the most advanced AI models, directly influencing how humanity interacts with AI. You will work with the world’s leading AI companies and government agencies to solve their most complex AI data-related problems. Responsibilities: Drive Impact: Directly contribute to the advancement of AI by delivering critical data solutions for leading AI innovators an
About Scale AI At Scale AI, our mission is to accelerate the development of AI applications. For 8 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI) and building upon our prior model evaluation work with enterprise customers and governments to deepen our capabilities and offerings for public and private evaluations. About Data Engine Our Generative AI Data Engine powers the world’s most advanced LLMs and generative models through world-class RLHF (Reinforcement Learning with Human Feedback), human data generation, model evaluation, safety, and alignment. The data we produce is some of the most critical work for how humanity will interact with AI. About Our FDE Team Generating high-quality data is the core problem our business solves. We aim to make producing and delivering high-quality data seamless and efficient for operators and customers. Our Team is building customer and operator-specific infrastructure to provide high-quality data with low turnaround time. You'll be exposed to the cutting edge of the Generative AI industry while directly interfacing with the leading model-building organizations in the space, including the top AI research labs and government agencies. Join us in shaping the future of Artificial General Intelligence. As a Forward Deployed Engineer, you'll be at the forefront of providing the critical data infrastructure that powers the most advanced AI models, directly influencing how humanity interacts with AI. You will work with the world’s leading AI companies and government agencies to solve their most complex AI data-related problems. Responsibilities: Drive Impact: Directly contribute to the advancement of AI by delivering critical data solutions for leading AI innovators and
About Scale At Scale AI, our mission is to accelerate the development of AI applications. For 8 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including: generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI), and building upon our prior model evaluation work with enterprise customers and governments, to deepen our capabilities and offerings for both public and private evaluations. About Data Engine Our Generative AI Data Engine powers the world’s most advanced LLMs and generative models through world-class RLHF (Reinforcement Learning with Human Feedback), human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. Our Approach As part of the interview process, you’ll be considered for opportunities across several teams within the GenAI Engineering organization, based on your interests, expertise, and business needs. Potential team placements include Allocation, Growth, Frontier Data, Trust & Safety, Pay, Operator, or Tasking Experience. Together, these teams power Scale’s AI data operations - from building high-impact datasets that push the boundaries of LLM capabilities, to optimizing contributor onboarding and incentives, to safeguarding data integrity through advanced trust, safety, and security measures. They work at the intersection of ML, operations, and analytics to ensure we deliver the highest-quality data at scale. Responsibilities: Design, build, and maintain robust, scalable systems across the full stack, including front-end, back-end, and infrastructure layers Implement high-impact features using modern technologies such as TypeScript, React, Node.js, MongoDB, Elasticsearch, and Temporal Collaborate closely with internal operators (your use
Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. Role Summary: We are looking for a talented and detail-oriented Sr Data Engineer to tackle data challenges. You will design, build, and maintain critical data pipelines and datasets, supporting areas like recruiting, compensation, talent management, and learning and development. Your work will enhance data accessibility and empower the People Team and business leaders to make informed decisions with high-quality, reliable data. Key Responsibilities: Develop and maintain robust data pipelines and datasets. Build foundational data products for key business areas. Enhance self-service data capabilities for the People Team. Ensure high standards in ETL/ELT operations, data quality, and pipeline reliability. Join us to drive impactful change and support SoFi's mission of fostering a thriving workplace through data excellence. What you’ll do: Design and build production dbt models in Snowflake that integrate Workday and other People systems into well-modeled, documented datasets, including slowly changing dimensions for People history. Build and operate Airflow DAGs that ingest People systems data and orchestrate dbt runs, keeping loads reliable and re-runnable. Own data quality and observability: dbt tests, freshness checks, row-count validation, and monitoring so issues are caught before stakeholders see them.
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? This role will focus on breaking down engineering tasks such as reporting, component search and handling of material billing/information into individual steps that can be tackled through tool use. Your breakdown will teach the Cohere model the logic needed to complete each task. Please Note: This is a part-time independent contractor position available within Canada . We seek candidates who are able to commit to 16 hours per week minimum at a 30 CAD/hour contract rate. This role is BYOD 💻 - Bring Your Own Device (laptop). Remote work within Canada. 12 month contract. Performance incentives included! As a Data Annotation Specialist, you will: Evaluate the model's ability to respond to engineering procedures and workflows Task models to complete engineering tasks along with verified information to evaluate the accuracy of the model's responses Label, proofread, and improve machine-written and human-written engineering-related outputs. Follow our style guide, and make recommendations on unique situations that fall outside of its scope. You may be a good fit if you have: 3+ years of industry experience working in eng
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? This role will focus on evaluating coding tasks, requiring you to review and debug code, navigate repository architecture, and analyze model trajectories. Your work will contribute to our model development efforts and the logic our models apply when completing task requests. Please note : This is a part-time independent contractor position available within Canada. We seek candidates who are able to commit to 16 hours per week minimum at a 40 CAD/hour contract rate. This role is BYOD 💻 - Bring Your Own Device (laptop). Remote work within Canada. 12 month contract. Performance incentives included! As a Data Annotation Specialist, you will: Evaluate the model's ability to respond to coding requests, workflows, and code base-related questions using available tools. Assess agent trajectories and model capabilities for code generation and debugging requests. Prompt models to complete complex coding tasks and review the accuracy of generated responses. Label, proofread, and improve machine-written and human-written software engineering-related outputs. Report quality and performance trends related to model/agent behavio
About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. About the role: We are looking for a Senior Data Engineer to help build and evolve the data platform that powers critical business decisions and customer-facing experiences. You will own significant parts of our data architecture that is main powerhouse of DevRev Computer’s memory for accurate and efficient Answers. As a part of data team, you will design and operate scalable data systems, and work closely with Software Engineering, AI Agent teams, Data Science, and Product teams to turn complex data requirements into reliable, high-quality data products. This role is ideal for an experienced engineer who enjoys solving challenging problems involving large-scale data, distributed systems, database architecture, and performance optimization . You will have significant technical ownership and the opportunity to influence the direction of our agentic data platform while helping raise the engineering bar across the team. What you'll do: Own data architecture for large-scale, high-impact projects, making thoughtful tradeoffs across scalability, reliability, performance, maintainability, and operational cost. Design, build, and operate scalable data pipelines and data systems that reliably ingest, transform, store, and se
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet is looking for an experienced Data Engineer to support and accelerate the productivity of our Go-To Market organization. In this role, you will facilitate Smartsheet's growth strategy through insights produced with data analysis. You will work directly with senior leadership to inform strategic decision-making as part of a collaborative, motivated team. Your work will be instrumental in helping our Sales partners optimize their pipeline, increase retention, and close deals. You Will: Help to create, evolve, and maintain the reporting infrastructure for the Sales organization to promote more intelligent discussion at all stages of the sales funnel. Partner with engineering, data science, and other analytics teams to build scalable analytics solutions with a focus on reporting capabilities. Mine large datasets to ensure data accuracy and completeness for reporting purposes. Analyze and monitor data for anomalies, with a focus on data quality and consistency within reports. Guide commission and territory planning processes with analytical support, providing the necessary reports and data. Develop and maintain key performance indicators (KPIs) to track sales performance, pipeline health, and other critical metrics. Automate report generation and distribution to improve efficiency. Provide technical expertise and guidance to junior analysts on data extraction, transformatio
Get new human data quality engineer jobs by email
Daily job updates · Unsubscribe anytime