Jobs in Canada

Reliability Engineer in San Francisco

40 active opportunities · Updated October 2026

Explore current reliability engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

TI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

C$113.4K – C$162K/yr

Quick readStrong listing-quality and freshness signals

We believe communication belongs to everyone. We exist to democratize phone service. TextNow is evolving the way the world connects and that's because we're made up of people with curious minds who bring an optimistic, yet critical lens into the work we do. We're the largest provider of free phone service in the nation. And we're just getting started. Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between! This role is about impact at scale. You’ll shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded into everyday engineering practice. Using AI is not optional, it’s expected. From design and architecture to implementation, testing, debugging, documentation, and operational analysis, you will actively leverage AI tools to increase velocity, improve code quality, and make better technical decisions. We provide a robust suite of AI-powered development tools and workflows to support you, and we expect you to continuously evolve how you use them to raise the bar for efficiency, clarity, and product excellence across the organization. What You'll Do Ensure System Reliability: Design, build, and maintain scalable, resilient, and highly available systems to support TextNow’s infrastructure and services. Automation & Infrastructure as Code: Develop and maintain automation using Terraform, Ansible, and other tools to enable efficient deployment, scaling, and operations of cloud-based systems (AWS preferred). Incident Response & On-Call Support: Participate in an on-call rotation, troubleshoot issues, and drive incident resolution to minimize downtime and improve syste

AWSCI/CDGitAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team DoorDash Labs, established in 2018, serves as the innovation hub for DoorDash, focusing on developing automation and robotics solutions to enhance last-mile logistics. The team's mission is to create technologies that support and augment human networks, aiming to improve efficiency for Dashers, merchants, and consumers alike. We’re ruthlessly focused on business impact. We are a highly senior team composed of former pioneers from a variety of different robotics industries. As of 2025, DoorDash has completed 10B lifetime deliveries. We’re focused on how to do the next 10B even better. About the Role We are seeking a highly motivated Senior Reliability & Test Engineer to join our team. This individual will play a key role in the development and validation of our unmanned platforms at the system and component levels. You will partner closely with EE, ME, and Autonomy teams to translate mission needs into robust, reliable hardware. The ideal candidate thrives in a fast-moving, cross-functional environment where reliability and test rigor determine program success. You will be hands-on in developing test methods and equipment to uncover failures before they happen in the field. You will partner closely with EE, ME, and Autonomy teams to translate mission needs into robust, reliable hardware. The ideal candidate thrives in a fast-moving, cross-functional environment where reliability and test rigor determine program success. You’re excited about this opportunity because you will… Architect and implement rigorous validation strategies, utilizing Python scripts for automation while leveraging CAD and shop tools to engineer bespoke test fixtures and hardware rigs. Oversee experimental execution across internal facilities and external laboratories, maintaining technical mastery over vibration tables, environmental chambers, DAQ systems, and ingress protection testing. Translate high-level vehicle reliability requirements into granula

PythonAWSGitRest
TI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$136.3K – $273.9K/yr

Quick readStrong listing-quality and freshness signals

TextNow is on a mission to make communications affordable and accessible for everyone. As a full MVNO operating our own mobile core network over LTE and 5G NSA, we have the unique advantage of controlling our network infrastructure end-to-end. We operate the HSS, PGW, and other critical network functions, giving us the flexibility to innovate and deliver exceptional service to millions of users. About the Role Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. T extNow is looking for a new SecOps team member to secure, monitor , and enable automated response within our infrastructure. What You’ll Do Ensure Secure & Reliable Systems: Design, implement, and maintain security-focused infrastructure to protect TextNow’s services while ensuring reliability and scalability. Security Automation & Infrastructure as Code: Develop and enforce best practices using Terraform, Ansible, Crowdstrike , and AWS security tools , ensuring secure configurations, automated compliance checks, and infrastructure as code. Threat Detection & Incident Response: Participate in an on-call rotation to respond to security incidents, investigate vulnerabilities, and implement proactive measures to prevent future threats. Work closely with engineering teams to remediate security risks. Monitoring & Logging for Security: Improve observability by implementing security monitoring solutions, logging best practices, and alerting mechanisms to detect anomalies and suspicious activity. Access Control & Identity Management: Manage IAM roles, permissions, and policies to ensure least privilege access and enforce security controls across cloud and internal systems. Collaboration & Security Advocacy: Wo

AWSCI/CDGitAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team DoorDash Labs, established in 2018, serves as the innovation hub for DoorDash, focusing on developing automation and robotics solutions to enhance last-mile logistics. The team's mission is to create technologies that support and augment human networks, aiming to improve efficiency for Dashers, merchants, and consumers alike. We’re ruthlessly focused on business impact. We are a highly senior team composed of former pioneers from a variety of different robotics industries. As of 2025, DoorDash has completed 10B lifetime deliveries. We’re focused on how to do the next 10B even better. About the Role We are seeking a highly motivated Senior/Staff Test Engineer to join our team. This individual will play a key role in the development and validation of our unmanned platforms at the system and component levels. The ideal candidate has a strong background in test development, test execution, and root cause analysis with a proven track record of collaboratively managing risk throughout a fast paced development process. You’re excited about this opportunity because you will… Run and monitor tests within our facility as well as at outside test labs. Collaborate with a tight knit team to identify and understand test failures. Be hands-on in developing test methods and equipment to uncover failures before they happen in the field. Find clarity through root cause analysis of lab and field failures and suggest design changes to prevent them. Use your creativity to create novel and scaled tests for autonomous systems. We’re excited about you because you have… A bachelors or advanced degree in a relevant engineering discipline. Mastery of test equipment such as environmental chambers, vibration tables, water testers, DAQs, etc.. Ability to bring order to complex test and development programs via clear technical communication and documentation. Experience designing and building testers and equipment. Ability to write Python scripts to automate

PythonAWSGitRest
SC
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$170K – $240K/yr

Quick readStrong listing-quality and freshness signals

Senior Software Engineer - Observability and Reliability About the Role We are growing the engineering team and looking for engineers who have the chops to build and deliver world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible. What You Will Be Doing Build observability tools and platforms, including: metrics, logging, distributed tracing, dashboarding, alerting, application performance management Build with modern tools and languages like Go, Open Telemetry and Kubernetes Participate in on-call rotation and ensure uptime of services Create runtime tools/processes that optimize cloud triaging and limit downtime Define best practices around making our systems and services measurable Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies. We expect successful candidates to be coding a majority of their time Qualifications We Need Strong Computer Science fundamentals 5+ years industry experience building and maintaining high-quality software, especially software other engineers use You apply a product mindset to infrastructure systems and feel accomplished enabling others Desire to be a great teammate and have fun at work Strong sense of craftsmanship, and a healthy academic curiosity Qualifications We Want (also, skills you’ll learn!) Experience building systems for data analytics Distributed systems monitoring and profiling skills Knowledge of cloud application security models Administered cloud service infrastructure (GCP, AWS, Azure) Startup experience Additional Job details Additional Job details The base salary range for this position is $170k - $240k annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience. Base pay is one part of the Total Package that is provided to compensate and recognize e

PythonSQLAWSAzure
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $184K/yr

Quick readStrong listing-quality and freshness signals

Scale AI is seeking a highly skilled and motivated Software Engineer, ARC (Architecture, Reliability, & Compute) to join our dynamic Public Sector Engineering team. As a part of this team, you will define how the company ships software, establishing the patterns for deploying into complex government and high-security environments, rather than just running Terraform scripts. You will build and maintain internal CLIs/tools that standardize testing, deployment, environment management and are tools that engineering relies on to prevent downstream breakages. You will execute on automated deployment efforts to pay down tech debt, creating fully functional staging/testing environments, and defining the company's standard for safe deployments. You will: Design and implement secure scalable backend systems for Public Sector customers, leveraging Scale's modern and cloud-native AI infrastructure. Own services or systems and define their long-term health goals, while also improving the health of surrounding components. Re-architect the stack to run in compliant or restrictive environments. This requires designing swappable components (auth, storage, logging) to meet government/security mandates without breaking the product. Collaborate with cross-functional teams to define and execute the vision for backend solutions, ensuring they meet the unique needs of government agencies operating in secure environments. Participate actively in customer engagements, working closely with stakeholders to understand requirements and deliver innovative solutions. Contribute to the platform roadmap and product strategy for Scale AI's Public Sector business, playing a key role in shaping the future direction of our offerings. Must have: At least an active secret clearance and the ability & willingness to up level to TS/SCI with CI Poly. This is a requirement and candidates will not be considered who do not hold at least a secret clearance Ideally you'd have: Full Stack Development: Prof

AWSAzureGCPDocker
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus

PythonJavaSQLAWS
DU
📍 San Francisco, Canada
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team DoorDash Labs is an independent team within DoorDash. We're hiring a backend software engineer to work at the intersection of software engineering and robotics to solve key business problems with elegant technical solutions. If you have a passion for applying robotics solutions to a service loved by millions of people, then we want to talk to you! About the Role We’re looking for Backend Engineers to work on both Product and Product Platform based teams in DoorDash Labs. Product focused Engineers work at the intersection of product and infrastructure to solve key business problems with elegant technical solutions. You'll operate our backend services and architecture that support all product functionality and will be challenged to consider the big picture -- collaborating cross-functionally, as well as evaluating and executing on trade-offs to maximize business impact for the company. You're excited about this opportunity because you will... Design and implement backend services for IoT that integrates with core DoorDash data, focused on reliability, and future extensibility Create a well documented APIs for other departments to integrate with Improve performance, reliability, scalability and security for our backend systems Introduce tools and best practices to accelerate our development process Design and implement backend services for autonomous delivery system that integrate with core DoorDash data. We're excited about you because you have... B.S., M.S., or PhD. in Computer Science or equivalent 6+ years of industry experience as a software engineer Experience with backend for frontend architecture Ability to improve efficiency, scalability, and stability of multiple system resources Experience with service oriented architecture, writing REST API’s, unit testing, and architectural design Understanding of modern web stacks and architecture (HTTP, REST) Experience with SQL Experience with either Java or Kotlin Nice to Have Experience with

JavaSQLPostgreSQLRedis
SC
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

$170K – $235K/yr

Quick readStrong listing-quality and freshness signals

About the Role Sigma Computing is redefining business intelligence by making complex data analysis accessible through a high-performance platform built for the modern data stack. The Compiler Team plays a foundational role in this mission by transforming user-driven spreadsheet interactions into highly optimized SQL queries, enabling seamless exploratory analytics on cloud data warehouses. As a member of the Compiler Team, you will join a group of engineers dedicated to building the core systems and abstractions that power Sigma’s intuitive spreadsheet interface, ensuring speed, reliability, and scalability for all users. What You Will Be Doing Tackle core challenges at the intersection of data modeling, query compilation, and large-scale interactive analytics—making it possible for end-users to query data warehouses efficiently without deep technical knowledge Design, build, and maintain sophisticated compiler infrastructure and intermediate representations that translate spreadsheet operations into optimized query plans Apply advanced optimization strategies to improve performance and accuracy across a wide range of query workloads and data architectures Contribute to both backend (Rust) and key frontend foundations (TypeScript), evolving critical abstractions that enable end-to-end workflow optimizations and new features Debug, analyze, and resolve complex issues, ensuring robustness and maintainability in a rapidly evolving product Collaborate with engineers and product stakeholders to review designs and code, driving technical best practices and architectural decisions throughout the team and company Qualifications We Need 5+ years experience engineering high-quality software systems Demonstrated success building and maintaining complex infrastructure or core platform services Deep understanding of Computer Science fundamentals, particularly in compilers, algorithms, SQL Optimization Passion for teamwork, technical ownership, and continually

TypeScriptPythonSQLAWS
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team Data is at the foundation of DoorDash success. The Data Engineering team builds database solutions for various use cases including reporting, product analytics, marketing optimization and financial reporting. By implementing pipelines, data structures, and data warehouse architectures; this team serves as the foundation for decision-making at DoorDash. About the Role DoorDash is looking for a Softare Engineer II to be a technical powerhouse to help us scale our data infrastructure, automation and tools to meet growing business needs. You're excited about this opportunity because you will… Work with business partners and stakeholders to understand data requirements Work with engineering, product teams and 3rd parties to collect required data Design, develop and implement large scale, high volume, high performance data models and pipelines for Data Lake and Data Warehouse Develop and implement data quality checks, conduct QA and implement monitoring routines Improve the reliability and scalability of our ETL processes Manage a portfolio of data products that deliver high-quality, trustworthy data Help onboard and support other engineers as they join the team We're excited about you because… 3+ years of professional experience working in data engineering, business intelligence, or a similar role You have proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software Proficiency in programming languages such as Python/Java 3+ years of experience in ETL orchestration and workflow management tools like Airflow, Flink, Oozie and Azkaban using AWS/GCP Expert in Database fundamentals, SQL and distributed computing 3+ years of experience with the Distributed data/similar ecosystem (Spark, Hive, Druid, Presto) and streaming technologies such as Kafka/Flink. Experience working with Snowflake, Redshift, PostgreSQL and/or other DBMS

PythonJavaSQLPostgreSQL
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team The Storage teams build and operate online stateful systems and abstractions that are reliable, efficient, secure and easy to use for DoorDash Engineering. The teams are responsible for understanding Product Engineering’s evolving needs and developing platform and infrastructure capabilities to serve them. The team currently supports CockroachDB, Cassandra, Kafka and Redis as well as data abstraction services to reduce the complexity of interacting with storage systems for Product Engineers. About the Role The Storage team is building and operating a high-performance, scalable, and reliable data abstraction layer that optimizes both efficiency and reliability. Our goal is to create a platform that manages itself and fades into the background—empowering engineers to focus on delivering product experiences our customers love. This role is available across two teams within Storage, each solving unique and high-impact challenges: One team is building the orchestration layer for DoorDash’s storage platform—unifying lifecycle management, operations, and self-serve APIs for databases and streaming systems, turning complex, stateful infrastructure into reliable, developer-friendly services used across the company. One team builds and operates the distributed data platform powering DoorDash's largest stateful workloads -- including Cassandra, which backs critical product surfaces across DoorDash, Wolt, and Roo. You'll design high-throughput data abstractions, smart clients, and platform services that make distributed data reliable and easy to work with at multi-petabyte, multi-million-QPS scale, with opportunities to go deep on distributed systems internals and contribute to the open-source Cassandra ecosystem. If you're passionate about distributed systems, developer experience, and building foundational infrastructure at scale, we'd love to hear from you. You must be located in San Francisco, Sunnyvale, Seattle, or the New York Metro Area for this hybrid pos

JavaSQLRedisAWS
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the technical direction for our in-house Spark deployment and shape the architecture that will run DoorDash's data, analytics, and ML compute for the next five years and beyond. You will own the deep, cross-cutting problems that span the runtime, the shuffle service, the scheduler, and the overall service reliability — making the architectural calls that compound across the platform's lifetime. You will partner with the Engineering Manager on technical roadmap, hiring, and team shape, and act as the senior technical voice in cross-team partnerships with Data Engineering, ML Platform, and product engineering teams that depend on the platform. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Set the multi-year technical direction for an in-house Spark-on-Kubernetes platform — runtime, shuffle, scheduler, reliability — and make the architectural calls that compound for years. Own the deepest distributed-systems problems on the team: shuffle architecture, multi-tenant scheduling, runtime performance, and the failure modes that only show up at scale. Partner with the Engineering Manager on technical roadmap, hiring, inte

PythonJavaSQLAWS
SA
📍 San Francisco, Canada
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

At Scale, our mission is to develop reliable AI systems for the world's most important decisions. For 10 years, Scale has provided the high-quality data and full-stack technologies that power the world's leading models, and has helped enterprises and governments build, deploy, and oversee AI applications that deliver real impact. We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force. Public Sector engineers build the core product including the systems required to ingest and process federal datasets that support real-time decision-making in contested environments. As a New Grad Software Engineer on this team, you will own meaningful, mission-facing work from day one: shipping features, sitting with the government stakeholders who use them, and iterating fast. Example Projects Build multi-layered guardrails that keep agents safe and predictable in high-stakes federal environments Optimize data retrieval for agents, including RAG pipelines over large, heterogeneous federal datasets Build orchestration for fleets of asynchronous agents running long-horizon tasks Develop systems that automatically alert users to deviations and anomalies in incoming data Create interfaces that illustrate how an agent reached a decision, so operators can audit and trust its output Develop data pipelines and ML infrastructure that make previously siloed government data sources accessible to agents Build evaluation infrastructure that measures model reliability against mission requirements Ship full-stack tooling that lets analysts query, visualize, and explore mission data Deploy and harden applications into secure, air-gapped, and cloud-native government environments Requirements A graduation date in Fall 2026 or Spring 2027 with a Bachelor's degree (or equivalent) in a relevant field (Computer Science, EECS, Computer Engineering, Statistics) Product engineering expe

TypeScriptPythonReactMongoDB
HI
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

C$45 – C$51/hr

Quick readStrong listing-quality and freshness signals

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Software Engineering Intern on the Systems team at HP IQ, you’ll work on low-level software that sits close to the hardware and helps power intelligent experiences across our products. This role is ideal for students who enjoy understanding how complex systems work under the hood. You’ll have the opportunity to work across multiple layers of the software stack, investigate performance bottlenecks, optimize system behavior, and build software that interacts closely with hardware and system resources. We’re looking for engineers who are curious about more than whether something works — you want to understand how it works, why it performs the way it does, and how to make it better. What You Might Do Build and optimize low-level systems software using languages such as C and C++. Investigate performance bottlenecks and improve the speed, efficiency, and reliability of existing systems. Work on data processing and sensor pipelines that connect software with underlying hardware. Analyze and improve memory usage, resource management, and system performance. Work across multiple lay

RedisLinuxRestAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $252K/yr

Quick readStrong listing-quality and freshness signals

The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Staff Software Engineer, you will orchestrate the implementation of vertical features and horizontal capabilities to include mentoring other engineers on defining requirements with stakeholders and communication tradeoffs of technical implementations on feature and capabilities until they are accepted by the stakeholders. You will: Orchestrate feature implementation across the Federal engineering team to ensure architectural consistency. Define technical strategy for agentic guardrails, explainability, and fleet orchestration. Ensure system reliability and performance across multiple security classifications and network types. Mentor engineers in the process of defining requirements with stakeholders and gathering acceptance. Communicate high-level technical trade-offs and implementation strategies to senior government stakeholders and Scale C-Suite members. Influence the long-term product strategy and technical roadmap for the Federal business unit. Consult on the architecture of AI-powered solutions for large-scale federal contracts. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in

AWSAzureGCPDocker
🔔

Get new reliability engineer jobs in San Francisco, Canada by email

Daily job updates · Unsubscribe anytime