Jobs in Canada

Back End Td Reliability Lab Manager in San Francisco

37 active opportunities · Updated October 2026

Explore current back end td reliability lab manager jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

Scale GP is Scale's enterprise Generative AI platform—APIs and infrastructure for knowledge retrieval, inference, evaluation, and intelligent automation. We power mission-critical workflows for leading enterprises, helping teams turn complex data and models into reliable, production-ready AI systems. We're building a new AI Enablement team to create the next generation of agent-powered tools that ground AI in real operational workflows. Our goal: help internal teams demystify their own workflows, then deploy agentic systems that reason over data, take action, and deliver measurable outcomes. We don't build in a vacuum. You'll use our own platform to solve real business problems internally—then selectively commercialize that same stack for customers. What we run on is what we sell. This is a 0→1 team. We're looking for a sharp, product-minded engineer who thrives in ambiguity, moves fast, and loves building systems from scratch alongside customers and cross-functional partners. You'll work closely with product, forward-deployed engineers, data scientists, and applied AI teams to turn real-world problems into scalable production solutions. If you like shipping fast, owning outcomes, and working across the stack—from polished frontends to distributed backends to LLM integrations—this role is for you. What You’ll Do Own full-stack features and projects end-to-end — from design through production deployment — within a larger product area Sample surfaces - Accounting Agents, Finance Copilots, GTM Agents, Agentic Experimentation Platforms Develop reliable backend services in Typescript/Python, work with distributed systems, data pipelines, and AI/ML infrastructure Integrate LLMs, vector databases, and agentic frameworks to power intelligent workflows Ship quickly through tight experimentation loops while maintaining high quality and reliability Adapt across the stack and learn new tools as needed to solve real problems end-to-end Ideal Experience 3+ years of full-tim

TypeScriptPythonAWSRest
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $1.5M/yr

Quick readStrong listing-quality and freshness signals

About the Team At DoorDash, we’re reimagining how people connect with the things they need — whether it’s a meal, a grocery run, and anything in between. Our audiences — Consumers, Dashers, and Merchants — are at the heart of everything the Design org builds. Our content design team plays a big part in this, with each content designer shaping the experience and bringing our vision to life through clear, thoughtful language that makes our three-sided marketplace easier, faster, and more human. About the Role We’re looking for a full stack UX Design Engineer to be the technical backbone of our global Content Design org and leader who builds the platforms, tooling, and ML systems that make world-class, localized product content effortless across DoorDash, Wolt, and Deliveroo. You’ll work at the intersection of design and engineering to build content tooling and embed LLM-driven workflows into product experiences. You’ll accelerate content design AI tooling so that CD’s can focus on the highest leverage strategic work, while your tools enable product designers and cross functional partners to ship high-quality content faster across DoorDash, Wolt, and Deliveroo. You're excited about this opportunity because... Design and build internal content tools that: Help PMs and designers generate, manage, localize, and deploy product content at scale. Plug custom GPTs and other LLMs into everyday product workflows for content iteration and improvements. Own full-stack development of these tools: Build and maintain frontend experiences using React and modern JavaScript/TypeScript. Design and implement backend services and APIs (e.g., Kotlin/Java) to support content tooling, experimentation, and automation. Integrate tooling with experimentation and infra: Embed content tools into A/B testing platforms so teams can test and ship variants with minimal engineering dependency. Build pipelines from variant generation → experiment → automated deployment of winners. Operationalize

JavaScriptTypeScriptJavaReact
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, leading the design and architecture of our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly. You’re excited about this opportunity because you will… Lead the design of infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Own and evolve our open-weights serving stack — real-time GPU endpoints, high-thr

PythonAWSGCPKubernetes
HI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$140K – $225K/yr

Quick readStrong listing-quality and freshness signals

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As the Senior Software Engineer, Tooling and Development Infrastructure, you will play a critical role in shaping the developer productivity tools and automated testing strategy. You’ll collaborate closely with design, development, and quality teams to plan, design, and implement robust automated tools and services that ensure the quality and reliability of our AI software stack. You will be highly hands-on in your work and collaborate closely with stakeholders. This position offers a unique opportunity to influence the development of cutting-edge automation frameworks, foster a culture of quality, and contribute to the long-term success of the organization. What You Might Do Develop and implement automation frameworks and testing strategies that cover the entire software stack, from backend systems to user-facing features. Identify, evaluate, and integrate new tools that streamline development. This includes everything from code quality tools and to Infrastructure-as-Code (IaC) solutions. Lead continuous improvement efforts for our build, release, and test systems, ensuring a robust

PythonRedisCI/CDGit
SC
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$135K – $180K/yr

Quick readStrong listing-quality and freshness signals

Solution Architect Sigma Computing The SA role has evolved. Here’s the version we’re hiring for. The SA job in 2026 is not the SA job in 2023. Three things now sit at the center of how we evaluate this role. This hire has to do all three at a senior level, with the architectural depth to back it up. 1. Use AI every day to do the job better. If you are not using Claude, ChatGPT, Cursor, or equivalents to accelerate your account prep, architecture diagramming, prototype builds, RFP responses, and discovery synthesis, you are getting outworked by SAs who are. We expect this hire to treat AI tooling as default infrastructure, not novelty. Come with a point of view on what you run, why, and how you use it to compress weeks of work into days. 2. Sell AI into the account. Buyers want to talk about agents, MCP, A2A, context engineering, and which model is powering what. You have to be fluent. You know Sigma’s AI surface cold: Sigma Assistant in build, analyze, and plan modes, AI functions, input tables with LLM enrichment, MCP integration, and warehouse-native agent patterns. You can architect Sigma agents and warehouse agents into a customer’s stack and explain the tradeoffs to a head of data and a CISO in the same call. You also speak credibly about Claude, OpenAI, Gemini, and the broader stack the customer already runs. 3. Sell against AI. Every enterprise deal has AI competition in it. Sometimes it is Databricks Genie. Sometimes it is Snowflake Cortex Analyst. Sometimes it is a systems integrator pitching a bespoke agent built over the weekend. You know where each of these breaks at scale, where Sigma’s warehouse-native architecture wins on governance, freshness, and cost, and how to draw the line for a skeptical CDO without hand-waving. You can defend that position in an architecture review, on a security questionnaire, and across three follow-up calls. About Sigma Sigma is the AI runtime environment for the modern enterprise. Teams build apps, agents, an

PythonSQLAIGo
N
📍 San Francisco, Canada
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

POSITION SUMMARY: A Phlebotomist serves patients by identifying the best method for retrieving blood; preparing specimens for laboratory testing; and performing screening procedures. The Phlebotomist will also act as an operations manager for the designated patient service center (PSC) and oversee Natera’s phlebotomy program at the specified location. Depending upon growth opportunities, this role may also require oversight of other phlebotomists as needed to support patient volume growth. *** IF YOUR STATE REQUIRES A PHLEBOTOMY LICENSE, IT MUST BE SENT IN WITH YOUR RESUME WITH YOUR APPLICATION *** PRIMARY RESPONSIBILITIES: Verifies test requisitions by comparing information with orders and requisition documentation; brings discrepancies to the attention of Natera product management leadership. Verifies patient by reading patient identification. Obtains blood specimens by performing venipunctures and finger sticks. Maintains specimen integrity by using aseptic technique, following Natera[KC1] procedures; observes isolation procedures. Tracks collected specimens by initialing, dating, and noting times of collection; maintaining daily tallies of collections performed[MB2] in Natera provided system. Maintains quality results by following Natera procedures and testing schedule; recording results in the quality-control log; identifying and reporting needed changes reporting KPIs to product management leadership on biweekly basis. Maintains safe, secure, and healthy work environment by following standards and procedures; complies with legal regulations. Resolves unusual test orders by contacting the physician, pathologist, nursing station, or reference laboratory; referring unresolved orders back to the originator for further clarification; notifying internal Natera team [MB3] of unresolved orders. Updates job knowledge by participating in educational opportunities; reading professional publications; main

A
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Amplitude is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Winter 2026 Report, Amplitude is the best-in-class solution for product, data, and marketing teams. Learn more at amplitude.com . As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do. Amplitude’s Commitment to Diversity Equity & Inclusion (DEI): Amplitude believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive. About The Role & Team Amplitude is only as useful as the data inside it. The Data Warehouse and Integrations team owns how that data gets in and out — importing behavioral and customer data from cloud data warehouses like Snowflake, Databricks, and BigQuery, from cloud object storage like S3 and Azure Blob Storage, and pushing enriched event data back out to warehouses, object storage, streaming destinations, and downstream advertising and marketing platforms. That means batch and streaming pipelines moving billions of events a day, connections that have to keep working across dozens of customer-controlled systems, credentials and configuration that have to stay correct and secure, and latency and reliability targets that customers build their own pipelines on top of. Recent work includes launching new warehouse export destinations, migrating our import

JavaAWSAzureGCP
🔔

Get new back end td reliability lab manager jobs in San Francisco, Canada by email

Daily job updates · Unsubscribe anytime