Jobs in Canada

Quality Assurance Manager in San Francisco

94 active opportunities · Updated October 2026

Explore current quality assurance manager jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team The Support Quality team’s vision is to “Provide teammates, team leads, and AI Agents with actionable feedback on 100% of cases so they can get 1% better every day. Deliver actionable business insight to wider S&O to help improve business policy and processes / Product teams to increase our customer experience.” About the Role You will report to the Director, Teammate Enablement, in our Customer Experience organization. The Teammate Enablement Team is responsible for ensuring that our Teammates have the resources (tooling, quality measurements, and knowledge base) and training to effectively and empathetically support DoorDash’s customers. Where opportunities are identified, we drive feedback to our business partners in order to empower Teammates to be 1% better every day for our customers. The team’s primary role will be to: i) build automated QA metrics that have high coverage, precision, and low false-positive rates; ii) support all manual QA efforts, including incubation and scale-up of new measurements; iii) deliver actionable insights to our partner teams to help improve the processes and policies on our support experience, as well as insights for our vendors to drive better agent performance management; and iv) manage our QA tech providers to ensure they deliver against our ambitious roadmap. This is a leadership role, with the expectation of managing the priorities of a small team while remaining accountable to key initiatives oneself. You’re excited about this opportunity because you will… Lead a small team centered around identifying and fixing top opportunities in our quality assurance programs that support customer service at DoorDash. Design, launch, and evolve our conversational intelligence program, utilizing AI technology to monitor, measure, and drive continuous improvements in our customer service interactions. Drive the strategy of our quality assurance program, identifying the right strategies to balance manual quality assuranc

AWSGitRestAI
SL
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $230K/yr

Quick readStrong listing-quality and freshness signals

Location - Hybrid This is a hybrid role based out of our San Francisco, Corporate Headquarter office, 3 days in the office, 2 days remote OR one of our Hub locations, Boston, MA or Raleigh, NC, which means you may be expected to work from a designated co-working space from time to time, and will otherwise work remotely from home, until such time as a dedicated office is established. About Us Sauce Labs is the world’s largest full-lifecycle, test automation platform, and the company behind Selenium. Trusted by 80% of the world’s top ten largest financial institutions and over 300,000 enterprise users, Sauce Labs provides the only AI platform capable of turning business intent into autonomous testing and quality assurance. With a proprietary dataset of 8.7 billion test runs, Sauce Labs empowers the Fortune 2000 to bridge the gap between AI-driven code generation and enterprise-grade software quality. Learn more at http://saucelabs.com . Release Assurance at the Speed of AI | Meet the new Sauce Labs The Role The Senior Director, Growth is an AI-first revenue leader responsible for setting the macro strategy, revenue architecture, and cross-functional alignment of an intelligent, end-to-end pipeline engine. Reporting to executive leadership, this role bridges marketing and sales executive strategy—focusing on expanding our outbound/inbound SDR function, driving predictive pipeline models, and enabling functional marketing leads to scale their teams. Rather than managing day-to-day tactical execution, you will focus on high-level channel architecture, budget optimization, sales leadership alignment, and scaling our revenue operations to maximize global pipeline output. Responsibilities Pipeline Strategy & Executive Revenue Alignment Align on global pipeline target setting, forecasting models, and macro growth strategy across inbound and outbound channels. Champion an agentic AI revenue stack, equipping functional leaders and teams with predictive tools, d

SQLAIGoRust
SL
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $110K/yr

Quick readStrong listing-quality and freshness signals

Location - Hybrid This is a hybrid role based out of our San Francisco, Corporate Headquarter office, 3 days in the office, 2 days work from home. About Us Sauce Labs is the world’s largest full-lifecycle, test automation platform, and the company behind Selenium. Trusted by 80% of the world’s top ten largest financial institutions and over 300,000 enterprise users, Sauce Labs provides the only AI platform capable of turning business intent into autonomous testing and quality assurance. With a proprietary dataset of 8.7 billion test runs, Sauce Labs empowers the Fortune 2000 to bridge the gap between AI-driven code generation and enterprise-grade software quality. Learn more at http://saucelabs.com . Release Assurance at the Speed of AI | Meet the new Sauce Labs The Role The Senior GTM Strategy & Planning Analyst will play a key role in supporting go-to-market (GTM) strategy through sales forecasting, churn forecasting, revenue analytics, and capacity planning across the GTM teams. This role will partner closely with Sales, Marketing, Customer Success, and Finance to align GTM initiatives with company growth objectives and contribute to executive-level insights for board presentations. Responsibilities GTM Board Materials: Support preparation of sales forecasting, churn forecasting, and revenue insight materials for executive and board review, ensuring data accuracy and alignment with company strategy. Executive Reporting: Help develop executive-level reporting materials, including presentations that translate complex data into clear, actionable insights. Forecasting: Build and maintain sales forecasting models in partnership with Sales Leadership, modeling pipeline generation, deal velocity, close rates, and other key metrics. Build predictive models to forecast customer churn, leveraging historical data and customer behavior insights to inform retention and renewal strategies. GTM Capacity Planning: Build and maintain capacity models to help ensure GTM t

AIGoRustExcel
SL
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $230K/yr

Quick readStrong listing-quality and freshness signals

Location - Hybrid This is a hybrid role based out of our San Francisco, Corporate Headquarter office, 3 days in the office, 2 days remote OR one of our Hub locations, Boston, MA or Raleigh, NC, which means you may be expected to work from a designated co-working space from time to time, and will otherwise work remotely from home, until such time as a dedicated office is established. About Us Sauce Labs is the world’s largest full-lifecycle, test automation platform, and the company behind Selenium. Trusted by 80% of the world’s top ten largest financial institutions and over 300,000 enterprise users, Sauce Labs provides the only AI platform capable of turning business intent into autonomous testing and quality assurance. With a proprietary dataset of 8.7 billion test runs, Sauce Labs empowers the Fortune 2000 to bridge the gap between AI-driven code generation and enterprise-grade software quality. Learn more at http://saucelabs.com . Release Assurance at the Speed of AI | Meet the new Sauce Labs The Role: As the Head of Product Marketing (Director Level), you will be a strategic leader responsible for defining and executing the comprehensive product vision for Sauce Labs. You will own the overall product marketing strategy, guiding a high-performing team to drive leadership, accelerate product adoption, and significantly impact revenue growth. This role requires a blend of strategic thinking, hands-on leadership, and a deep understanding of the continuous testing and software development lifecycle market. You will serve as a key bridge between product development, sales, and broader marketing functions, ensuring our market narrative is compelling, differentiated, and aligned with business objectives. Responsibilities: Strategic Leadership & Vision: Define and articulate the overarching product marketing strategy that aligns with company goals and market opportunities. Drive the narrative, positioning, and messagi

AIGoRustDevOps
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team The Code Quality team sits within the Developer Platform organization and owns the systems that keep DoorDash's codebase healthy and secure as it scales: static analysis, quality gates, test frameworks, regression infrastructure, and tooling. Our job is to make sure the signals engineers rely on before shipping — test results, coverage, performance feedback etc — are fast and trustworthy. The decisions we make about tooling and standards directly shape how confidently and quickly engineering teams at DoorDash can ship to production. About the Role We're looking for Software Engineers to help build and maintain the systems that validate code quality across DoorDash's engineering org, treating our tooling as a critical product for the engineers who rely on it every day: static analysis and quality gates, test frameworks and regression infrastructure. You’ll design the tooling and automation that will help derive trustworthy quality signals, integrate them into the development lifecycle, and make it easy for engineers to execute reliable, repeatable workflows. You will collaborate across the engineering org, partnering directly with the teams who use what you build to understand the accuracy, reliability and performance of their functionality. You will report into the Engineering Manager on our Code Quality team in our Developer Platform organization. You must be located in either San Francisco, CA, Sunnyvale, CA, Los Angeles, CA, Seattle, WA, or New York, NY. You're excited about this opportunity because you will… Build and maintain quality tooling — static analysis, quality gates, coverage reporting, test frameworks, regression infrastructure — and integrate it directly into our developer workflows and CI/CD pipelines Define and derive quality signals - flakiness, pass rate, coverage, performance, scale readiness etc - Build tooling that improves everyday engineering workflows, including local development, CI/CD, debugging, and rollou

AWSCI/CDGitRest
G
📍 San Francisco, Canada
✓ High-confidence listing

$140K – $265K/yr

Quick readStrong listing-quality and freshness signals

About Glean: Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With over 100 enterprise SaaS connectors, flexible LLM choice, and robust APIs, Glean gives organizations the infrastructure to govern, scale, and customize AI across their entire business - without vendor lock-in or costly implementation cycles. At its core, Glean is redefining how enterprises find, use, and act on knowledge. Its Enterprise Graph and Personal Knowledge Graph map the relationships between people, content, and activity, delivering deeply personalized, context-aware responses for every employee. This foundation powers Glean’s agentic capabilities - AI agents that automate real work across teams by accessing the industry’s broadest range of data: enterprise and world, structured and unstructured, historical and real-time. The result: measurable business impact through faster onboarding, hours of productivity gained each week, and smarter, safer decisions at every level. Recognized by Fast Company as one of the World’s Most Innovative Companies (Top 10, 2025), by CNBC’s Disruptor 50, Bloomberg’s AI Startups to Watch (2026), Forbes AI 50, and Gartner’s Tech Innovators in Agentic AI, Glean continues to accelerate its global impact. With customers across 50+ industries and 1,000+ employees in more than 25 countries, we’re helping the world’s largest organizations make every employee AI-fluent, and turning the superintelligent enterprise from concept into reality. If you’re excited to shape how the world works, you’ll help build systems used daily across Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub, and many more - deeply embedded where people get things done. You’ll ship agentic capabilities on an open, extensible stack, with the craf

PythonJavaMachine LearningAI
G
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$180K – $205K/yr

Quick readStrong listing-quality and freshness signals

About Glean: Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With over 100 enterprise SaaS connectors, flexible LLM choice, and robust APIs, Glean gives organizations the infrastructure to govern, scale, and customize AI across their entire business - without vendor lock-in or costly implementation cycles. At its core, Glean is redefining how enterprises find, use, and act on knowledge. Its Enterprise Graph and Personal Knowledge Graph map the relationships between people, content, and activity, delivering deeply personalized, context-aware responses for every employee. This foundation powers Glean’s agentic capabilities - AI agents that automate real work across teams by accessing the industry’s broadest range of data: enterprise and world, structured and unstructured, historical and real-time. The result: measurable business impact through faster onboarding, hours of productivity gained each week, and smarter, safer decisions at every level. Recognized by Fast Company as one of the World’s Most Innovative Companies (Top 10, 2025), by CNBC’s Disruptor 50, Bloomberg’s AI Startups to Watch (2026), Forbes AI 50, and Gartner’s Tech Innovators in Agentic AI, Glean continues to accelerate its global impact. With customers across 50+ industries and 1,000+ employees in more than 25 countries, we’re helping the world’s largest organizations make every employee AI-fluent, and turning the superintelligent enterprise from concept into reality. If you’re excited to shape how the world works, you’ll help build systems used daily across Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub, and many more - deeply embedded where people get things done. You’ll ship agentic capabilities on an open, extensible stack, with the craf

PythonJavaAWSGit
G
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Glean: Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With over 100 enterprise SaaS connectors, flexible LLM choice, and robust APIs, Glean gives organizations the infrastructure to govern, scale, and customize AI across their entire business - without vendor lock-in or costly implementation cycles. At its core, Glean is redefining how enterprises find, use, and act on knowledge. Its Enterprise Graph and Personal Knowledge Graph map the relationships between people, content, and activity, delivering deeply personalized, context-aware responses for every employee. This foundation powers Glean’s agentic capabilities - AI agents that automate real work across teams by accessing the industry’s broadest range of data: enterprise and world, structured and unstructured, historical and real-time. The result: measurable business impact through faster onboarding, hours of productivity gained each week, and smarter, safer decisions at every level. Recognized by Fast Company as one of the World’s Most Innovative Companies (Top 10, 2025), by CNBC’s Disruptor 50, Bloomberg’s AI Startups to Watch (2026), Forbes AI 50, and Gartner’s Tech Innovators in Agentic AI, Glean continues to accelerate its global impact. With customers across 50+ industries and 1,000+ employees in more than 25 countries, we’re helping the world’s largest organizations make every employee AI-fluent, and turning the superintelligent enterprise from concept into reality. If you’re excited to shape how the world works, you’ll help build systems used daily across Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub, and many more - deeply embedded where people get things done. You’ll ship agentic capabilities on an open, extensible stack, with the craf

AWSGitAIGo
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

C$40 – C$55/hr

Quick readStrong listing-quality and freshness signals

You will: Source and acquire quality candidates for high volume positions to meet hiring and activation goals Review applications and screen candidates Support candidates throughout the recruiting process Partner with other growth recruiters, ops leads and cross-functional leaders to understand Scale’s business needs, then devise and execute on the recruiting strategy needed to deliver on them. Follow and manage processes that will evolve over time as the company grows. Represent and champion the brand, promoting its value proposition to candidates and driving engagement. Represent Scale at occasional onsite growth recruiting events. Ideally you have: 2+ years of recruiting/sourcing experience in a fast-paced, high-growth environment. Excellent written and verbal communication skills, with the ability to tailor messaging to diverse audiences Experience sourcing candidates and filling high volume roles Experience acting as the primary candidate liaison, delivering white-glove service through timely updates and personalized support throughout the interview process Nice to have: Previous experience in a start-up. Either experience managing a full-cycle recruiting experience, or a passion and ability to learn to do so. Experience with various tools such as Linkedin Recruiter, Clay, Modernloop, Gem Extreme attention to detail AI builder, AI-forward operator, or strong interest in agentic coding and experimenting with AI tooling The hourly salary range for this position is $40 to $55/hr. This is a remote position with occasional travels required for onsite growth recruiting events. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position, determined by work location and additional factors, including job-related skills, experience, interview performance, and relevant education or training. PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allow

AWSRestAIGo
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $165.6K/yr

Quick readStrong listing-quality and freshness signals

Scale works with the industry’s leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling). This role will focus on optimizing data curation and eval to enhance LLM capabilities in both text and multimodal modalities. In this role, you will develop novel methods to improve the alignment and generalization of large-scale generative models. You will collaborate with researchers and engineers to define best practices in data-driven AI development. You will also partner with top foundation model labs to provide both technical and strategic input on the development of the next generation of generative AI models. You will: Research and develop novel post-training techniques, including SFT, RLHF, and reward modeling, to enhance LLM core capabilities in both text and multimodal modalities. Design and experiment new approaches to preference optimization. Analyze model behavior, identify weaknesses, and propose solutions for bias mitigation and model robustness. Publish research findings in top-tier AI conferences. Ideally you’d have: Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field. Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning. Excellent written and verbal communication skills Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals Previous experience in a customer facing role. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined du

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $165.6K/yr

Quick readStrong listing-quality and freshness signals

Scale works with the industry's leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation. This role is on the evaluation pod within the GenAI Research Organization and will focus on building benchmarks and diagnosing model failure modes in both text and multimodal modalities. In this role, you will develop rigorous evaluations and diagnostic methods that reveal where frontier models fail and why. You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development. You will also partner with top foundation model labs to translate failure analysis into technical and strategic input on the next generation of generative AI models. You will: Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents. You’ll identify everything from capability gaps and reasoning errors to robustness and alignment issues, all focusing on RCA. Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities. Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them. Publish research findings in top-tier AI conferences. Ideally you’d have: Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field. Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development. Excellent written and verbal communication skills. Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals. Previous experience in a customer facing r

AWSRestMachine LearningAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. Scale Frontier Data is the organization behind the training and evaluation data that frontier labs depend on. We build the systems, tooling, and expert workflows that turn hard human expertise into signals that models can learn from, across reasoning, coding, agentic tool use, and domain expertise. About our Customer Platform team: Our Customer Platform Team plays a pivotal role in integrating our platform with external systems and ensuring seamless, reliable connectivity for both internal users and customers. As the leader of this team, you’ll drive the strategy, architecture, and development of our connectivity solutions, focusing on API integration, distributed systems, and a robust data platform. Your role will be crucial in maintaining and enhancing our platform’s ability to meet the needs of both our internal and external stakeholders. Responsibilities: Own large areas within our product Comfortable working cross functionally, whether that be internal or external customers Build features end-to-end: front-end, back-end, system design, debugging and testing Deliver experiments at a high velocity and level of quality to engage our customers Work across the entire product lifecycle from conceptualization through production Influence the culture, values, and processes of a growing engineering team Inspire and mentor less experienced engineers Collaborating with cross-functional teams to define, design, and ship new product features and experiences. Requirements: At least 7-10 years of relevant experience is preferred Track record of shipping high-quality products and features at scale Desire to work in a very fast-paced environment Abil

AWSRestAIGo
SA
📍 San Francisco, Canada· Hybrid
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Snorkel At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data. We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler! In September 2026 we raised a $350 million Series E at a $3.5 billion valuation , and we are scaling our engineering and research teams to meet demand. The role Frontier AI data is expensive to make and hard to measure. Every task we deliver is tested against the strongest models, often through many long-running agent rollouts. Your job is to make that process faster, cheaper, and more rigorous with ML and AI You will be one of the early members of ML & Research Engineering at Snorkel. You will study how frontier-grade data is generated and evaluated, form hypotheses, validate them against real production data, and ship the winners at scale. You will shape the discipline's direction, its standards, and the team that grows around it. What you'll work on Efficient agentic evals. Cut the cost of long-horizon agent evaluation with adaptive sampling, statistically grounded early stopping, model cascades, caching, and cheap-first gating. AI model routing. Route every eval and judge call to the cheapest model that clears the quality bar, with fallback, monitoring, and cost attribution. Fine-tuned small models. Fine-tune and serve open-weight models (LoRA and other

PythonMachine LearningAI
SA
📍 San Francisco, Canada
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

At Scale, our mission is to develop reliable AI systems for the world's most important decisions. For 10 years, Scale has provided the high-quality data and full-stack technologies that power the world's leading models, and has helped enterprises and governments build, deploy, and oversee AI applications that deliver real impact. We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force. Public Sector engineers build the core product including the systems required to ingest and process federal datasets that support real-time decision-making in contested environments. As a New Grad Software Engineer on this team, you will own meaningful, mission-facing work from day one: shipping features, sitting with the government stakeholders who use them, and iterating fast. Example Projects Build multi-layered guardrails that keep agents safe and predictable in high-stakes federal environments Optimize data retrieval for agents, including RAG pipelines over large, heterogeneous federal datasets Build orchestration for fleets of asynchronous agents running long-horizon tasks Develop systems that automatically alert users to deviations and anomalies in incoming data Create interfaces that illustrate how an agent reached a decision, so operators can audit and trust its output Develop data pipelines and ML infrastructure that make previously siloed government data sources accessible to agents Build evaluation infrastructure that measures model reliability against mission requirements Ship full-stack tooling that lets analysts query, visualize, and explore mission data Deploy and harden applications into secure, air-gapped, and cloud-native government environments Requirements A graduation date in Fall 2026 or Spring 2027 with a Bachelor's degree (or equivalent) in a relevant field (Computer Science, EECS, Computer Engineering, Statistics) Product engineering expe

TypeScriptPythonReactMongoDB
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $250K/yr

Quick readStrong listing-quality and freshness signals

About Scale Scale’s mission is to develop reliable AI systems for the world’s most important decisions. As the leading AI data foundry, we provide the high-quality data and full-stack technologies that power the world’s most advanced models — fueling breakthroughs in generative AI, defense, and autonomous vehicles. We partner with leading enterprises and governments to bring AI into production that performs when it matters most, combining rigorous evaluation with full-stack deployment so our customers can build AI they can trust. About the Team Applied Intelligence Systems (AIS) is part of the Scale Generative AI Platform (SGP), focused on pushing the frontier of what agentic applications can do across diverse enterprise and government use cases. We build the infrastructure and tooling that power agentic AI in production, paired with applied ML research, design, and evaluation to ensure these systems perform reliably at the scale our customers demand. AIS spans multiple workstreams — agent evaluation and oversight, orchestration and tool-use infrastructure, model and systems optimization, and applied research on new agent capabilities — and this role is not scoped to any single one of them. We’re growing fast, with increasing traction across both commercial and public sector customers, and we’re just getting started — this team will define what dependable, production-grade agentic AI looks like. About the Role As a Staff Machine Learning Research Engineer, you will operate across the full breadth of AIS’s technical needs — wherever the hardest ML problem in agentic AI happens to be that quarter. This could mean training and fine-tuning models, designing evaluation and observability systems, building improvement loops from production data, prototyping novel agent architectures, or designing internal systems and tooling that boost productivity across teams. You’re not tied to one team’s roadmap; you’re expected to move to where the technical leverage is highest, and t

AWSRestMachine LearningAI
🔔

Get new quality assurance manager jobs in San Francisco, Canada by email

Daily job updates · Unsubscribe anytime