Jobs in Canada

Back End Td Reliability Lab Manager in San Francisco

37 active opportunities · Updated October 2026

Explore current back end td reliability lab manager jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $252K/yr

Quick readStrong listing-quality and freshness signals

The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Staff Software Engineer, you will orchestrate the implementation of vertical features and horizontal capabilities to include mentoring other engineers on defining requirements with stakeholders and communication tradeoffs of technical implementations on feature and capabilities until they are accepted by the stakeholders. You will: Orchestrate feature implementation across the Federal engineering team to ensure architectural consistency. Define technical strategy for agentic guardrails, explainability, and fleet orchestration. Ensure system reliability and performance across multiple security classifications and network types. Mentor engineers in the process of defining requirements with stakeholders and gathering acceptance. Communicate high-level technical trade-offs and implementation strategies to senior government stakeholders and Scale C-Suite members. Influence the long-term product strategy and technical roadmap for the Federal business unit. Consult on the architecture of AI-powered solutions for large-scale federal contracts. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in

AWSAzureGCPDocker
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. Scale Frontier Data is the organization behind the training and evaluation data that frontier labs depend on. We build the systems, tooling, and expert workflows that turn hard human expertise into signals that models can learn from, across reasoning, coding, agentic tool use, and domain expertise. About our Customer Platform team: Our Customer Platform Team plays a pivotal role in integrating our platform with external systems and ensuring seamless, reliable connectivity for both internal users and customers. As the leader of this team, you’ll drive the strategy, architecture, and development of our connectivity solutions, focusing on API integration, distributed systems, and a robust data platform. Your role will be crucial in maintaining and enhancing our platform’s ability to meet the needs of both our internal and external stakeholders. Responsibilities: Own large areas within our product Comfortable working cross functionally, whether that be internal or external customers Build features end-to-end: front-end, back-end, system design, debugging and testing Deliver experiments at a high velocity and level of quality to engage our customers Work across the entire product lifecycle from conceptualization through production Influence the culture, values, and processes of a growing engineering team Inspire and mentor less experienced engineers Collaborating with cross-functional teams to define, design, and ship new product features and experiences. Requirements: At least 7-10 years of relevant experience is preferred Track record of shipping high-quality products and features at scale Desire to work in a very fast-paced environment Abil

AWSRestAIGo
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Software Engineer, you will own the development of a vertical feature or a horizontal capability to include defining requirements with stakeholders and implementation until it is accepted by the stakeholders. You will: Design and implement scalable backend systems for Federal customers using cloud-native AI infrastructure. Build features for agentic systems including multi-layered guardrails and data retrieval optimization. Develop data pipelines and machine learning infrastructure to make data sources accessible by agents. Collaborate with cross-functional teams to execute backend solutions for secure environments. Participate in customer engagements to understand requirements and deliver technical solutions. Define requirements with stakeholders and implement features until they are accepted. Contribute to the platform roadmap and product strategy for the Federal business. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in developing and deploying applications in a cloud-native environment. Understanding of containerization (e.g., Docker) and container orchestration (e.g., Kubernetes

AWSAzureGCPDocker
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $216K/yr

Quick readStrong listing-quality and freshness signals

The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Senior Software Engineer, you will lead the development of a vertical feature or a horizontal capability to include defining requirements with stakeholders and implementation until it is accepted by the stakeholders. You will: Lead the design and implementation of scalable backend systems and distributed architectures for Federal customers. Manage the full lifecycle of feature development from requirement definition to deployment on classified networks. Direct the orchestration of asynchronous agent fleets to meet mission requirements. Lead customer engagements to translate mission needs into technical requirements. Own the communication with stakeholders to ensure implementation meets defined acceptance criteria. Conduct technical reviews and identify risks within machine learning infrastructure and model serving. Drive the platform roadmap by providing technical specifications for Federal product offerings. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in developing and deploying applications in a cloud-native environment. Understanding of containerization (e.g., Docker) and contai

AWSAzureGCPDocker
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

About Scale At Scale AI, our mission is to accelerate the development of AI applications. For 8 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including: generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI), and building upon our prior model evaluation work with enterprise customers and governments, to deepen our capabilities and offerings for both public and private evaluations. About Data Engine Our Generative AI Data Engine powers the world’s most advanced LLMs and generative models through world-class RLHF (Reinforcement Learning with Human Feedback), human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. Our Approach As part of the interview process, you’ll be considered for opportunities across several teams within the GenAI Engineering organization, based on your interests, expertise, and business needs. Potential team placements include Allocation, Growth, Frontier Data, Trust & Safety, Pay, Operator, or Tasking Experience. Together, these teams power Scale’s AI data operations - from building high-impact datasets that push the boundaries of LLM capabilities, to optimizing contributor onboarding and incentives, to safeguarding data integrity through advanced trust, safety, and security measures. They work at the intersection of ML, operations, and analytics to ensure we deliver the highest-quality data at scale. Responsibilities: Design, build, and maintain robust, scalable systems across the full stack, including front-end, back-end, and infrastructure layers Implement high-impact features using modern technologies such as TypeScript, React, Node.js, MongoDB, Elasticsearch, and Temporal Collaborate closely with internal operators (your use

TypeScriptPythonReactNode.js
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $184K/yr

Quick readStrong listing-quality and freshness signals

Scale AI is seeking a highly skilled and motivated Software Engineer, Frontier AI Infrastructure to join our dynamic Public Sector Engineering team. As a part of this team, you will own the model inference layer - enabling state of the art models, debugging the latest AI tools, managing networking, debugging latency, and tracking pricing/usage metrics for AI models. You will lead technical discussions on the frontlines with cloud vendors and customers to deliver on critical contracts and to debug platform issues. You will also work upstream with Product to understand features before they break, moving us from "infra-only debugging" to proactive integration testing. You will: Design and implement secure scalable backend systems for Public Sector customers, leveraging Scale's modern and cloud-native AI infrastructure. Own services or systems and define their long-term health goals, while also improving the health of surrounding components Re-architect the stack to run in compliant or restrictive environments. This requires designing swappable components (auth, storage, logging) to meet government/security mandates without breaking the product. You will work with Product to build integration tests that catch issues early, shifting the focus from "infra-only debugging" to preventing failures upstream. Participate actively in customer engagements, working closely with stakeholders to understand requirements and deliver innovative solutions. Contribute to the platform roadmap and product strategy for Scale AI's Public Sector business, playing a key role in shaping the future direction of our offerings. Must have: At least an active secret clearance and the ability & willingness to up level to TS/SCI with CI Poly. This is a requirement and candidates will not be considered who do not hold at least a secret clearance Ideally you'd have: Full Stack Development: Proficiency in both front-end and back-end development, including experience with modern web develo

AWSAzureGCPDocker
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team The Storage organization builds and operates the online stateful systems and abstractions that DoorDash Engineering depends on: reliable, efficient, secure, and easy to use. Within Storage, the Distributed Caching team owns every caching offering at DoorDash end to end, including ElastiCache (Redis/Valkey), Boulder (our KVRocks-based key-value store for high-QPS feature serving), Entity Cache (read Bill Shen’s engineering blog post, “ High-Performance Proxy Cache for DoorDash Services ”), and the Distributed Lock Service, plus the smart clients (asgard-redis, valkey-go) that sit in front of them. These systems back critical product surfaces across DoorDash, Wolt, and Deliveroo: the team runs roughly 400 ElastiCache clusters serving hundreds of millions of GET requests per second in aggregate, and Boulder, our offline-to-online feature store, serves billions of feature lookups per second at peak. About the Role The team owns provisioning of clusters and the smart clients that sit in front of them, baking in sensible defaults so that other engineering teams get a turnkey caching solution instead of having to run their own. You'll help drive Boulder's evolution to scale further, improve cost efficiency, enhance performance, and support real-time updates; re-platform the Distributed Lock Service onto a strongly consistent backend; and build the self-serve tooling and recommendation engine that let customers describe a workload (QPS, TTL, payload size, latency profile) and get the right backend without talking to a human. You'll go deep on cache invalidation, replication, sharding, compaction, and failover, while shipping the guardrails, automation, and observability that keep this scale operable by a small team. You must be located in San Francisco, Seattle, or the New York Metro Area for this hybrid position. You will report to the Engineering Manager on the Distributed Caching team within the Storage organization. You’re excited about this opportunity b

JavaRedisAWSKubernetes
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $1.3M/yr

Quick readStrong listing-quality and freshness signals

About the Team DoorDash Labs is an independent team within DoorDash. We explore robotics and automation to transform last-mile logistics in the long term. If you want to work on commercializing autonomy and robotics in a service used by millions of people — and on bringing the merchant partners who power that service along with us — then we want to talk to you! About the Role Come help us redefine last-mile logistics through robotics, automation, and other advanced technologies. Autonomy only works when merchants — restaurants, retailers, and other partners — can reliably interact with our robots: handing off orders, troubleshooting edge cases, and trusting the experience enough to keep using it. This role owns that side of the equation. We're looking for a Merchant Success & Growth lead to build the strategy and operational mechanisms that get merchants onboard, keep them performing, and turn their day-to-day reality into a tight feedback loop for product and engineering. You're excited about this opportunity because you will… Own merchant adoption and performance KPIs for autonomy end-to-end — defining what "successful merchant interaction with a robot" means, instrumenting it, and driving improvement against it. Build the playbooks and operational mechanisms to onboard merchants to autonomy — from first conversation through training, go-live, and steady-state ops — and scale them across markets. Partner with sales, account management, and field ops to recruit and ramp the right merchant cohorts for each stage of the product, and design experiments that test new merchant-facing features and handoff models. Define merchant performance benchmarks (handoff success rate, dwell time, dasher/robot interaction quality, merchant CSAT) and run the cadence that holds partners and internal teams accountable to them. Stand up the feedback loop from the field back to product and engineering — turning merchant complaints, edge cases, and frontline observations into prioriti

AWSGitRestAI
L
📍 San Francisco, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%
Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Airports team is part of a mission-critical endeavor that keeps travelers moving smoothly. As an engineer on our team, your role will be essential in making sure drivers and riders enjoy a dependable experience at airports. You'll work hand in hand with various teams across Lyft, fostering collaboration and driving innovation to tackle the unique challenges of the travel and airports industry. Your responsibilities will also involve managing real-time communication with airports, ensuring our technology integrates seamlessly into their operations. Your skills will be the driving force behind enhancing the airport journey for millions. Responsibilities: Drive high-impact projects and innovate new solutions to provide the best user experience. Lead large features from idea to positive execution and launch Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and distribute knowledge Participate in our teams on call rotation. Identify, triage, debug and resolve issues/bugs across our various applications and platforms Have the ability to explain the various trade offs made in decisions Manage project priorities, deadlines, and deliverables. Experience: BS/MS or equivalent in Computer Engineering, Computer Science, or related field or equivalent practical experience 2-5+ years of software engineering/production infrastructure industry experience Experience with Python, Go Proficiency in object-oriented programming Experience working with data structures or algorithms Ability to work with a low-ego, highly collaborative, and cross-functional team Bonus points: experience pursuing side projects or open-source project Benefits: Great medical, dental, and vision insurance options with additional programs available when enrolled M

PythonAIGoHR
TI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$136.3K – $273.9K/yr

Quick readStrong listing-quality and freshness signals

We believe that communication belongs to everyone. TextNow is redefining how people connect by combining simplicity, intelligence, and accessibility. We’re a team of curious builders using technology to make communication more affordable and powerful for millions of users every day. As a Software Developer - Backend , you won’t just build services—you’ll shape the systems, architecture, and tooling that make them possible. At TextNow, Members of Technical Staff combine leadership and hands-on coding to enable the highest leverage opportunities possible. Being able to operate strategically, as well as diving into the lowest-level details, is a must. You’ll take technical ownership of key backend domains and work across mobile, web, and data to create faster, smarter, more reliable systems. AI and automation are at the core of how we build. You’ll use them to accelerate development, improve performance, detect and resolve issues faster, and continuously raise the bar for backend development excellence. We’re hiring Members of Technical Staff across multiple levels (intermediate/senior/staff+). Whether you’re an experienced developer looking to lead complex systems, or an early-career developer eager to grow, we’ll align your title and scope based on experience and impact. This role is about impact at scale. You’ll shape how TextNow builds and operates its backend systems—using AI and automation to make development faster, decisions smarter, and experiences more seamless for millions of users worldwide. What You'll Do Design, develop, and sustain high-performance, scalable backend services using Go microservices and modern cloud-native tooling. Lead architectural modernization and modularization to improve scalability, observability, and developer velocity. Define and own the entire lifecycle of your systems: API design, data modeling, deployment (CI/CD), live-traffic monito

AWSKubernetesCI/CDMicroservices
SC
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$240K – $270K/yr

Quick readStrong listing-quality and freshness signals

About the Role Sigma is transforming how businesses run by delivering a high performance platform on the modern data architecture. Hence, we are growing the engineering team and looking for engineers who are excited to solve challenging problems, deliver impactful capabilities throughout our stack to build world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible for all users. What You Will Be Doing Solve challenging problems that arise in providing an interactive experience on data warehouses for data exploration and analysis Build with modern tools and languages like Rust, Go, GraphQL, Node, and Kubernetes Build backend distributed services, new algorithms and modern API to support a cloud application Triage product or system issues and debug/track/resolve by analyzing the sources of issues Design and implement new software features to support our fast growing user base Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies Qualifications We Need 5+ years industry experience building and maintaining high-quality software Experience building and deploying robust and secure web applications in a continuous deployment environment Desire to be a great teammate and have fun at work without compromising ownership towards your work Strong sense of craftsmanship, and a healthy academic curiosity to solve challenges at sigma Strong Computer Science fundamentals Qualifications We Want (also, skills you’ll learn!) Data driven aptitude and its application to solve distributed system problems Data model design, and API development experience SQL query optimization and database internals Administered cloud service infrastructure (GCP, AWS, Azure) Prior experience working at high growth company solving technical problems to enable continued success Additional Job details The base salary range for this position is $240k - $270

PythonSQLAWSAzure
SC
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$170K – $240K/yr

Quick readStrong listing-quality and freshness signals

About the Role Sigma is transforming how businesses run by delivering a high performance platform on the modern data architecture. Hence, we are growing the engineering team and looking for engineers who are excited to solve challenging problems, deliver impactful capabilities throughout our stack to build world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible for all users. What You Will Be Doing Solve challenging problems that arise in providing an interactive experience on data warehouses for data exploration and analysis Build with modern tools and languages like Rust, Go, GraphQL, Node, and Kubernetes Build backend distributed services, new algorithms and modern API to support a cloud application Triage product or system issues and debug/track/resolve by analyzing the sources of issues Design and implement new software features to support our fast growing user base Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies Qualifications We Need 5+ years industry experience building and maintaining high-quality software Experience building and deploying robust and secure web applications in a continuous deployment environment Desire to be a great teammate and have fun at work without compromising ownership towards your work Strong sense of craftsmanship, and a healthy academic curiosity to solve challenges at sigma Strong Computer Science fundamentals Qualifications We Want (also, skills you’ll learn!) Data driven aptitude and its application to solve distributed system problems Data model design, and API development experience SQL query optimization and database internals Administered cloud service infrastructure (GCP, AWS, Azure) Prior experience working at high growth company solving technical problems to enable continued success Additional Job details The base salary range for this position is $170k - $240

PythonSQLAWSAzure
DU
📍 San Francisco, Canada
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team DoorDash Labs is an independent team within DoorDash. We're hiring a backend software engineer to work at the intersection of software engineering and robotics to solve key business problems with elegant technical solutions. If you have a passion for applying robotics solutions to a service loved by millions of people, then we want to talk to you! About the Role We’re looking for Backend Engineers to work on both Product and Product Platform based teams in DoorDash Labs. Product focused Engineers work at the intersection of product and infrastructure to solve key business problems with elegant technical solutions. You'll operate our backend services and architecture that support all product functionality and will be challenged to consider the big picture -- collaborating cross-functionally, as well as evaluating and executing on trade-offs to maximize business impact for the company. You're excited about this opportunity because you will... Design and implement backend services for IoT that integrates with core DoorDash data, focused on reliability, and future extensibility Create a well documented APIs for other departments to integrate with Improve performance, reliability, scalability and security for our backend systems Introduce tools and best practices to accelerate our development process Design and implement backend services for autonomous delivery system that integrate with core DoorDash data. We're excited about you because you have... B.S., M.S., or PhD. in Computer Science or equivalent 6+ years of industry experience as a software engineer Experience with backend for frontend architecture Ability to improve efficiency, scalability, and stability of multiple system resources Experience with service oriented architecture, writing REST API’s, unit testing, and architectural design Understanding of modern web stacks and architecture (HTTP, REST) Experience with SQL Experience with either Java or Kotlin Nice to Have Experience with

JavaSQLPostgreSQLRedis
A
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $170K/yr

Quick readStrong listing-quality and freshness signals

Airtable is the no-code app platform that empowers people closest to the work to accelerate their most critical business processes. More than 500,000 organizations, including 80% of the Fortune 100, rely on Airtable to transform how work gets done. Airtable’s mission is to bring the power of computing and software development to everyone. We are developing a powerful and extensible toolkit that our customers can leverage to solve a variety of different problems and workflows. We’ve seen our most sophisticated customers use the product to run global processes across thousands of employees, coordinate precision manufacturing pipelines, and consolidate previously siloed mission-critical data into a single source of truth. The complexity of these use cases requires us to be extremely thoughtful about how we design and implement new functionality in the product and make sure it’s both easy to use and comprehend for our customers and maintainable for us. As a Full-Stack, Backend engineer at Airtable, you will have the opportunity to work with customers to deeply understand their needs and workflows. You will collaborate with cross-functional partners across product management, design, research and data science to create innovative new features that enable our customers to do their best work. You will be responsible for owning and executing the end-to-end implementation of these new features that will contribute to making our toolkit even more powerful and successful. We currently have openings on: The Admin & Governance Team (Full-Stack/BE) ensures Airtable is secure, compliant, and enterprise-ready. It owns key admin capabilities like the Admin Panel, SSO, and audit systems, as well as foundational features like User Groups. This team's mission is to accelerate organizational value for the largest customers with enterprise-first governance and controls. The Omni Capability & Quality Team (Full-Stack/BE) brings the power of AI directly to Airtable end users—

JavaScriptJavaReactNode.js
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform providing APIs for knowledge retrieval, inference, evaluation, and more. We are seeking a strong Senior Full-Stack Engineer to help us build, scale, and refine our rapidly growing product. The ideal candidate is deeply grounded in software engineering best practices and experienced in developing and scaling modern web applications end-to-end. You will work across the stack—from React/TypeScript frontends to Python-based backends—while integrating with LLMs and machine learning systems. You will solve complex challenges in scalability, reliability, and product experience while owning significant product areas in a fast-paced environment. What You’ll Do Own major full-stack product areas , driving features from design through production deployment. Build modern frontend experiences using React and TypeScript, ensuring performance, usability, and responsiveness. Develop reliable backend services in Python, working with distributed systems, data pipelines, and ML/LLM components. Integrate with LLMs, vector databases, and AI infrastructure to power intelligent product experiences. Deliver experiments and new features quickly , maintaining high quality and tight feedback loops with customers. Collaborate across product, ML, and infrastructure teams to shape the direction of Scale GP. Adapt quickly —learning new technologies, frameworks, and tools as needed across the stack. Ideal Experience 5+ years of full-time engineering experience , post-graduation. Strong experience developing full-stack applications using React, TypeScript, and Python . Experience scaling or shipping products at high-growth startups . Familiarity with LLMs, vector databases, embeddings, or other modern AI tooling (tinkering or production experience welcome). Proficiency with SQL and modern API development. Experience with Kubernetes , containerization, and microservice architectures. Experience working with at leas

TypeScriptPythonReactSQL
🔔

Get new back end td reliability lab manager jobs in San Francisco, Canada by email

Daily job updates · Unsubscribe anytime