Jobs in Canada

Quality Manager in Canada

333 active opportunities · Updated October 2026

Explore current quality manager jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.

G
📍 Mountain View, Canada· Full-time
✓ High-confidence listing

$140K – $265K/yr

Quick readStrong listing-quality and freshness signals

About Glean: Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With over 100 enterprise SaaS connectors, flexible LLM choice, and robust APIs, Glean gives organizations the infrastructure to govern, scale, and customize AI across their entire business - without vendor lock-in or costly implementation cycles. At its core, Glean is redefining how enterprises find, use, and act on knowledge. Its Enterprise Graph and Personal Knowledge Graph map the relationships between people, content, and activity, delivering deeply personalized, context-aware responses for every employee. This foundation powers Glean’s agentic capabilities - AI agents that automate real work across teams by accessing the industry’s broadest range of data: enterprise and world, structured and unstructured, historical and real-time. The result: measurable business impact through faster onboarding, hours of productivity gained each week, and smarter, safer decisions at every level. Recognized by Fast Company as one of the World’s Most Innovative Companies (Top 10, 2025), by CNBC’s Disruptor 50, Bloomberg’s AI Startups to Watch (2026), Forbes AI 50, and Gartner’s Tech Innovators in Agentic AI, Glean continues to accelerate its global impact. With customers across 50+ industries and 1,000+ employees in more than 25 countries, we’re helping the world’s largest organizations make every employee AI-fluent, and turning the superintelligent enterprise from concept into reality. If you’re excited to shape how the world works, you’ll help build systems used daily across Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub, and many more - deeply embedded where people get things done. You’ll ship agentic capabilities on an open, extensible stack, with the craf

PythonJavaAWSGit
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -73.4%

From C$136K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our core philosophy is to empower developers to self-serve rather than be bottlenecked by a central quality team. We believe that by creating smart, automated tooling, we can eliminate common roadblocks, making developers happier and more productive. The Rider Quality team is focused on elevating quality, testing, and accessibility across our mobile platforms. We are currently shifting our strategy to heavily leverage automation and AI to revolutionize how we approach testing. We're not doing traditional QA - we're building the future of quality engineering. This is a chance to step into a Senior role where you won't just write code; you'll design systems that define how hundreds of engineers deliver product faster, happier, and with fewer bugs. We are fundamentally shifting away from manual processes and existing automation frameworks that struggle to keep up, betting heavily on Artificial Intelligence and agentic frameworks to drive a massive "shift left" in our organization. Success in this role is measured by tangible impact, including increased developer satisfaction, engineering hours saved, and bugs/incidents avoided in production. We are seeking a highly skilled and innovative Senior Software Engineer to join our team. You will be instrumental in building the next generation of our quality assurance platform. You will be responsible for designing and developing advanced AI-powered tooling for test case generation, review, and execution. Your work will directly impact our ability to provide fast, actionable feedback to developers, "shifting left" to ensure quality from the earliest stages of development. You will play a key role in integrating these tools into our existing developer workflows, ensuring seamless adoption and maximum impact. We are looking for someone with a strong background in

PythonCI/CDRestMachine Learning
P
📍 Toronto, ON, CA· Full-time
✓ Quality checkedCompany trend -100%

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is the world's leading visual search and discovery platform, serving over 500 million monthly active users globally on their journey from inspiration to action. As we scale experiences in a complicated ecosystem, ensuring they are safe, fair, and trustworthy is paramount. We are looking for a Senior Data Scientist to help lead Pinterest's Trust and Safety mandate by designing the foundations for measuring the prevalence of unsafe content across the platform. In this role, you will design and build sampling frameworks, complex data aggregations, and measurement methodologies to track Trust & Safety policy violations across complex, multi-component user interactions. You will work in a highly collaborative and cross-functional environment, partnering with ML Engineers, Trust & Safety Ops, subject matter expert teams, and Product

PythonSQLAWSRest
D
12 days ago
📍 Vancouver, Canada
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Dialpad Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage. Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, Dialpad was built to resolve. Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved. Market-leading brands, including Randstad, Motorola Solutions, Netflix, the San Diego Padres, the Colorado Rockies Baseball Club, and Cal Athletics, trust Dialpad. Dialpad is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile. Being a Dialer At Dialpad, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more. We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves. We look for people who are intensely curious and hold themselves to a high bar. Our ambition is significant, and achieving it requires a team that operates at the highest level. We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic . About the team Dialpad's Contact Center QA team owns the quality of the AI-driven services behind our customer engagement platform. We validate that complex features like AI Routing, Digital Agents, and real-time Analytics are genuinely enterprise-grade in scale, reliability, and performance. And we go further, working cross-fun

JavaScriptPythonJavaVue
D
📍 Vancouver, Canada· Full-time
✓ High-confidence listing

From C$150.5K/yr

Quick readStrong listing-quality and freshness signals

About Dialpad Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage. Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, Dialpad was built to resolve. Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved. Market-leading brands, including Randstad, Motorola Solutions, Netflix, the San Diego Padres, the Colorado Rockies Baseball Club, and Cal Athletics, trust Dialpad. Dialpad is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile. Being a Dialer At Dialpad, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more. We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves. We look for people who are intensely curious and hold themselves to a high bar. Our ambition is significant, and achieving it requires a team that operates at the highest level. We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic . Your role As a Sr. SDET in Agentic QA, you will own the test automation and quality frameworks that support Dialpad’s AI Voice Agent services. You will develop automated tests for end-to-end product experiences, from frontend UI to backend services to APIs to audio/text interactions. You will test orchestration flows, agent conf

JavaScriptPythonJavaReact
C
📍 Canada· Contract
✓ Quality checkedCompany trend -91.5%

Who are we? Our mission is to scale intelligence to serve humanity. We’re training and deploying frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic search, RAG, and agents. We believe that our work is instrumental to the widespread adoption of AI. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. We like to work hard and move fast to do what’s best for our customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate about their craft. Each person is one of the best in the world at what they do. We believe that a diverse range of perspectives is a requirement for building great products. Join us on our mission and shape the future! Why this role? We are on a mission to build machines that understand the world and make them safely accessible to all. Data quality is foundational to this process. Machines (or Large Language Models to be exact) learn in similar ways to humans - by way of feedback. By labelling, ranking, auditing, and correcting text output, you will improve Large Language Model’s performance for iterations to come, thus having a lasting impact on Cohere’s tech. Cohere seeks dynamic and dedicated Data Annotators with exceptional Modern Standard German Writing, Translation, and Data Manipulation skill. Please Note: This is a part-time independent contractor position available within Canada . We seek candidates who are able to commit to 16 hours per week minimum at a 30 CAD/hour contract rate. This role is BYOD 💻 - Bring Your Own Device (laptop). Remote work within Canada. 12 months contract. Performance incentives included! As a Data Annotation Specialist, you will: Label and Rank: Accurately label and rank machine learning data with advanced proficiency in German, ensuring data integrity and quality. Audit and Correct: Sc

Machine LearningAIGo
C
📍 Canada· Contract
✓ High-confidence listingCompany trend -91.5%

From C$30/hr

Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We are on a mission to build machines that understand the world and make them safely accessible to all. Data quality is foundational to this process. Machines (or Large Language Models, to be exact) learn in similar ways to humans, by way of feedback. By labelling, ranking, auditing, and correcting model output, you will improve Large Language Models' performance for iterations to come, thus having a lasting impact on Cohere's technology. We are hiring Generalist professionals with broad backgrounds that span multiple consumer-facing or personal domains. This is a judgment-driven role, not passive data entry. You will review, assess, and provide structured feedback across a broad and evolving range of tasks, evaluating, stress-testing, and improving our models on English-language data spanning multiple modalities (text, image, and structured formats such as JSON, CSV/TSV, and Markdown). This is a great opportunity for professionals with strong analytical skills to contribute to high-impact annotation projects. Please Note: This is a part-time independent contractor position available within Canada only. We seek ca

C
📍 Canada· Contract
✓ Quality checkedCompany trend -91.5%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? This role will focus on breaking down engineering tasks such as reporting, component search and handling of material billing/information into individual steps that can be tackled through tool use. Your breakdown will teach the Cohere model the logic needed to complete each task. Please Note: This is a part-time independent contractor position available within Canada . We seek candidates who are able to commit to 16 hours per week minimum at a 30 CAD/hour contract rate. This role is BYOD 💻 - Bring Your Own Device (laptop). Remote work within Canada. 12 month contract. Performance incentives included! As a Data Annotation Specialist, you will: Evaluate the model's ability to respond to engineering procedures and workflows Task models to complete engineering tasks along with verified information to evaluate the accuracy of the model's responses Label, proofread, and improve machine-written and human-written engineering-related outputs. Follow our style guide, and make recommendations on unique situations that fall outside of its scope. You may be a good fit if you have: 3+ years of industry experience working in eng

C
📍 Canada· Contract
✓ Quality checkedCompany trend -91.5%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? This role will focus on evaluating data science and coding tasks, requiring you to review and debug code, analyze model trajectories, and assess data visualization script development and logical flow implementation. Your work will contribute to our model development efforts and the logic our models apply when completing task requests. Please note : This is a part-time independent contractor position available within Canada. We seek candidates who are able to commit to 16 hours per week minimum at a 40 CAD/hour contract rate. This role is BYOD 💻 - Bring Your Own Device (laptop). Remote work within Canada. 12 month contract. Performance incentives included! As a Data Annotation Specialist, you will: Evaluate the model's ability to respond to coding requests, workflows, and code base-related questions using available tools. Assess agent trajectories and model capabilities for code generation, tabular and graphic manipulation, and debugging requests. Prompt models to complete complex data science tasks and review the accuracy of generated responses. Label, proofread, and improve machine-written and human-written soft

PythonSQLAIGo
C
📍 Canada· Contract
✓ Quality checkedCompany trend -91.5%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? This role will focus on evaluating coding tasks, requiring you to review and debug code, navigate repository architecture, and analyze model trajectories. Your work will contribute to our model development efforts and the logic our models apply when completing task requests. Please note : This is a part-time independent contractor position available within Canada. We seek candidates who are able to commit to 16 hours per week minimum at a 40 CAD/hour contract rate. This role is BYOD 💻 - Bring Your Own Device (laptop). Remote work within Canada. 12 month contract. Performance incentives included! As a Data Annotation Specialist, you will: Evaluate the model's ability to respond to coding requests, workflows, and code base-related questions using available tools. Assess agent trajectories and model capabilities for code generation and debugging requests. Prompt models to complete complex coding tasks and review the accuracy of generated responses. Label, proofread, and improve machine-written and human-written software engineering-related outputs. Report quality and performance trends related to model/agent behavio

JavaScriptPythonJavaSQL
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $165.6K/yr

Quick readStrong listing-quality and freshness signals

Scale works with the industry’s leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling). This role will focus on optimizing data curation and eval to enhance LLM capabilities in both text and multimodal modalities. In this role, you will develop novel methods to improve the alignment and generalization of large-scale generative models. You will collaborate with researchers and engineers to define best practices in data-driven AI development. You will also partner with top foundation model labs to provide both technical and strategic input on the development of the next generation of generative AI models. You will: Research and develop novel post-training techniques, including SFT, RLHF, and reward modeling, to enhance LLM core capabilities in both text and multimodal modalities. Design and experiment new approaches to preference optimization. Analyze model behavior, identify weaknesses, and propose solutions for bias mitigation and model robustness. Publish research findings in top-tier AI conferences. Ideally you’d have: Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field. Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning. Excellent written and verbal communication skills Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals Previous experience in a customer facing role. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined du

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $165.6K/yr

Quick readStrong listing-quality and freshness signals

Scale works with the industry's leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation. This role is on the evaluation pod within the GenAI Research Organization and will focus on building benchmarks and diagnosing model failure modes in both text and multimodal modalities. In this role, you will develop rigorous evaluations and diagnostic methods that reveal where frontier models fail and why. You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development. You will also partner with top foundation model labs to translate failure analysis into technical and strategic input on the next generation of generative AI models. You will: Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents. You’ll identify everything from capability gaps and reasoning errors to robustness and alignment issues, all focusing on RCA. Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities. Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them. Publish research findings in top-tier AI conferences. Ideally you’d have: Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field. Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development. Excellent written and verbal communication skills. Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals. Previous experience in a customer facing r

AWSRestMachine LearningAI
T-
📍 Toronto, Canada· Full-time
✓ High-confidence listing

From C$908.4K/yr

Quick readStrong listing-quality and freshness signals

About the Role: Tubi is seeking a highly skilled and experienced QA Automation Engineer to lead quality assurance initiatives for our cutting-edge streaming and AI-driven product features. This pivotal role involves ensuring exceptional end-to-end user experiences, robust streaming playback, and the accuracy and integrity of our AI/ML features across web, mobile, and OTT platforms. We're looking for a candidate with a strong background in streaming QA and deep technical knowledge of media workflows. You'll be instrumental in collaborating with engineering, product, and data science teams to define comprehensive QA strategies that guarantee both functional excellence and data-level quality This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: Design and lead test strategies for streaming workflows, playback systems, and AI-powered features. Test across platforms (web, mobile, and connected TV) to ensure functional parity and playback stability. Validate streaming performance—including ABR logic, encoding pipelines, and DRM integrations—under diverse real-world conditions. Debug with precision using tools like Charles Proxy, Chrome DevTools, ADB, and Xcode. Collaborate with data and ML teams to validate AI model updates, recommendations, and personalization accuracy. Leverage AI-assisted QA tools to enhance regression coverage, UI validation, and anomaly detection. Contribute to automation and CI/CD frameworks, driving faster, more reliable releases. Oversee QA deliverables for multiple concurrent releases and ensure seamless sign-off for production launches. Monitor live environments for playback or recommendation anomalies post-release and escalate issues promptly. Continuously improve QA processes, metrics, and reporting for streaming and AI validation. Your Background: Bachelor’s degree in Computer Science, Software Engineering, or related field, or equivalent hands-on experi

JavaScriptTypeScriptPythonJava
R
📍 Toronto, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com . Reddit’s Ads Data Science team is looking for a Senior Staff Data Scientist to lead the scientific strategy behind Reddit’s ads measurement and signal systems. This role sits at the center of how Reddit proves advertiser value, improves signal quality, builds privacy-aware measurement systems, and connects ads exposure to real advertiser outcomes. As a Senior Staff Data Scientist for Ads Measurement, you will be the principal architect of our technical vision and the driving force behind the next generation of our measurement ecosystem. In a landscape rapidly shifting due to privacy regulations, browser changes, and evolving platform dynamics, you will spearhead innovation across experimental design (incrementality/lift), identity, signals, and privacy-safe 1P/3P measurement. This is a high-visibility, high-impact role requiring a rare blend of deep experimentation & causal inference expertise, strategic foresight, and the ability to influence cross-functional executives and industry standards. You will not just adapt to the changing ad-tech environment, you will redefine how we measure value. Responsibilities Technical Vision & Strategy: Define the long-term data science vision & strategy across ads measurement, signal quality, identity, attribution, and privacy. Establish how Reddit should evaluate advertiser value, measurement quality, and signal utility across first-party and third-party products. Set the Cross-Pillar Measurement Science Strategy: Define the long-term data science strategy across a

PythonSQLRestMachine Learning
SA
📍 San Francisco, Canada· Hybrid
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Snorkel At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data. We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler! In September 2026 we raised a $350 million Series E at a $3.5 billion valuation , and we are scaling our engineering and research teams to meet demand. The role Frontier AI data is expensive to make and hard to measure. Every task we deliver is tested against the strongest models, often through many long-running agent rollouts. Your job is to make that process faster, cheaper, and more rigorous with ML and AI You will be one of the early members of ML & Research Engineering at Snorkel. You will study how frontier-grade data is generated and evaluated, form hypotheses, validate them against real production data, and ship the winners at scale. You will shape the discipline's direction, its standards, and the team that grows around it. What you'll work on Efficient agentic evals. Cut the cost of long-horizon agent evaluation with adaptive sampling, statistically grounded early stopping, model cascades, caching, and cheap-first gating. AI model routing. Route every eval and judge call to the cheapest model that clears the quality bar, with fallback, monitoring, and cost attribution. Fine-tuned small models. Fine-tune and serve open-weight models (LoRA and other

PythonMachine LearningAI
🔔

Get new quality manager jobs in Canada by email

Daily job updates · Unsubscribe anytime