Jobs in Canada

Reliability Engineer Iii in Canada

138 active opportunities · Updated October 2026

Explore current reliability engineer iii jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.

T-
📍 Toronto, Canada· Full-time
✓ High-confidence listing

From C$1.2M/yr

Quick readStrong listing-quality and freshness signals

About the Role: We are looking for a talented Automation Engineer to join our Automation Engineering team in Toronto. In this role, you will be responsible for designing and implementing automated tests for Mobile development. You will collaborate closely with QA engineers and developers to build scalable test frameworks, improve automation coverage, and contribute to the efficiency of our multi-platform release process. You will also design data-driven end-to-end checks around playback and ad insertion , integrate them into CI/CD pipelines as quality gates, and operate a reliable device lab to prevent regressions from shipping. Your work will directly accelerate testing and release velocity while improving revenue-critical reliability across Tubi’s Android and IOS apps. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: Design, implement, and maintain automated tests for mobile development (Android & iOS) Contribute to the development and optimization of cross-platform automation frameworks. Write and maintain test scripts in JavaScript/TypeScript , using frameworks such as Puppeteer, Appium, WebDriverIO, Selenium. Ensure test cases are integrated into CI/CD pipelines and provide reliable feedback on product quality. Help identify flaky tests, investigate root causes, and improve test stability. Collaborate with developers and QA engineers to clarify requirements and improve test strategies. Participate in code reviews and follow best practices for test automation . Your Background: Bachelor’s degree or above in a technical field (e.g., Computer Science, Engineering, Mathematics), or equivalent industry experience. 3+ years of hands-on experience in automation testing for mobile devices Strong programming skills in JavaScript/TypeScript (preferred), or Python/Java. Experience with automation frameworks (e.g. Puppeteer, Appium,, WebDriverIO, Selenium, Playwright, T

JavaScriptTypeScriptPythonJava
L
📍 Ontario, Canada· Full-time
✓ High-confidence listing

From $158K/yr

Quick readStrong listing-quality and freshness signals

Lithic is the modern card issuing and processing platform empowering ambitious financial companies to build the future of payments. Our infrastructure powers card programs for 100+ innovative clients, from fintechs reimagining credit and digital banking to platforms transforming disbursements and spend management. Companies like Mercury, Flex, and Novo rely on Lithic's developer-friendly APIs, direct network connections, and flawless reconciliation to launch and scale card programs in weeks, not years. We're building a future where access to better financial products materially improves people's lives, free from the constraints of 30-year-old mainframes and legacy processors. We're proud to be backed by world-class investors who share that vision, including Bessemer Venture Partners, Index Ventures, Spark Capital, Stripes, and Mastercard, along with many others. We're a team of 170+ across 26 states and 7 countries, headquartered in New York City. We are hiring for our Treasury team Software Engineers at various levels (II and Senior) who are curious and willing to dive deep and understand our technology and domain in order to solve interesting and hard problems.The Treasury team maintains and builds the backend services that manage the flow of funds between Lithic and third parties. This includes our ledger, ACH and wire infrastructure, and associated reconciliation. The systems we maintain have high standards of reliability and correctness. You will become an expert in the card payments space. The Treasury team primarily uses Python for their tech stack. What You'll Do: Ensure high reliability and correctness for Lithic’s ledger and orchestrated funds flows Develop new features to better serve Lithic customers Ensure that the team is delivering reliable, secure, and scalable code with minimal tech debt Own initiatives from planning to launch, keeping stakeholders informed and aligned along the way Lead efforts to improve systems and processes wit

PythonGitRestAI
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $252K/yr

Quick readStrong listing-quality and freshness signals

The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Staff Software Engineer, you will orchestrate the implementation of vertical features and horizontal capabilities to include mentoring other engineers on defining requirements with stakeholders and communication tradeoffs of technical implementations on feature and capabilities until they are accepted by the stakeholders. You will: Orchestrate feature implementation across the Federal engineering team to ensure architectural consistency. Define technical strategy for agentic guardrails, explainability, and fleet orchestration. Ensure system reliability and performance across multiple security classifications and network types. Mentor engineers in the process of defining requirements with stakeholders and gathering acceptance. Communicate high-level technical trade-offs and implementation strategies to senior government stakeholders and Scale C-Suite members. Influence the long-term product strategy and technical roadmap for the Federal business unit. Consult on the architecture of AI-powered solutions for large-scale federal contracts. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in

AWSAzureGCPDocker
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $252K/yr

Quick readStrong listing-quality and freshness signals

Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale our product in a fast-paced environment. The ideal candidate will have a strong understanding of software engineering principles and practices, as well as experience with large-scale distributed systems. You will be responsible for owning large new areas within our product, working across backend, frontend, and interacting with LLMs and ML models. You will solve hard engineering problems in scalability and reliability. You will: Own large new areas within our product Work across backend, frontend, and interacting with LLMs and ML models Deliver experiments at a high velocity and level of quality to engage our customers Work across the entire product lifecycle from conceptualization through production Be able, and willing, to multi-task and learn new technologies quickly Ideally you'd have: 7+ years of full-time engineering experience, post-graduation Experience scaling products at hyper growth startups Experience tinkering with or productizing LLMs, vector databases, and the other latest AI technologies Proficient in Python or Javascript/Typescript, and SQL Experience with Kubernetes Experience with major cloud providers (AWS, Azure, GCP) Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval

JavaScriptTypeScriptPythonJava
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform providing APIs for knowledge retrieval, inference, evaluation, and more. We are seeking a strong Senior Full-Stack Engineer to help us build, scale, and refine our rapidly growing product. The ideal candidate is deeply grounded in software engineering best practices and experienced in developing and scaling modern web applications end-to-end. You will work across the stack—from React/TypeScript frontends to Python-based backends—while integrating with LLMs and machine learning systems. You will solve complex challenges in scalability, reliability, and product experience while owning significant product areas in a fast-paced environment. What You’ll Do Own major full-stack product areas , driving features from design through production deployment. Build modern frontend experiences using React and TypeScript, ensuring performance, usability, and responsiveness. Develop reliable backend services in Python, working with distributed systems, data pipelines, and ML/LLM components. Integrate with LLMs, vector databases, and AI infrastructure to power intelligent product experiences. Deliver experiments and new features quickly , maintaining high quality and tight feedback loops with customers. Collaborate across product, ML, and infrastructure teams to shape the direction of Scale GP. Adapt quickly —learning new technologies, frameworks, and tools as needed across the stack. Ideal Experience 5+ years of full-time engineering experience , post-graduation. Strong experience developing full-stack applications using React, TypeScript, and Python . Experience scaling or shipping products at high-growth startups . Familiarity with LLMs, vector databases, embeddings, or other modern AI tooling (tinkering or production experience welcome). Proficiency with SQL and modern API development. Experience with Kubernetes , containerization, and microservice architectures. Experience working with at leas

TypeScriptPythonReactSQL
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

Scale GP is Scale's enterprise Generative AI platform—APIs and infrastructure for knowledge retrieval, inference, evaluation, and intelligent automation. We power mission-critical workflows for leading enterprises, helping teams turn complex data and models into reliable, production-ready AI systems. We're building a new AI Enablement team to create the next generation of agent-powered tools that ground AI in real operational workflows. Our goal: help internal teams demystify their own workflows, then deploy agentic systems that reason over data, take action, and deliver measurable outcomes. We don't build in a vacuum. You'll use our own platform to solve real business problems internally—then selectively commercialize that same stack for customers. What we run on is what we sell. This is a 0→1 team. We're looking for a sharp, product-minded engineer who thrives in ambiguity, moves fast, and loves building systems from scratch alongside customers and cross-functional partners. You'll work closely with product, forward-deployed engineers, data scientists, and applied AI teams to turn real-world problems into scalable production solutions. If you like shipping fast, owning outcomes, and working across the stack—from polished frontends to distributed backends to LLM integrations—this role is for you. What You’ll Do Own full-stack features and projects end-to-end — from design through production deployment — within a larger product area Sample surfaces - Accounting Agents, Finance Copilots, GTM Agents, Agentic Experimentation Platforms Develop reliable backend services in Typescript/Python, work with distributed systems, data pipelines, and AI/ML infrastructure Integrate LLMs, vector databases, and agentic frameworks to power intelligent workflows Ship quickly through tight experimentation loops while maintaining high quality and reliability Adapt across the stack and learn new tools as needed to solve real problems end-to-end Ideal Experience 3+ years of full-tim

TypeScriptPythonAWSRest
T
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a skilled Software Engineer with a passion for building high-performance, low-level systems software. In this role, you’ll contribute to the development and optimization of the infrastructure that powers our cutting-edge processors, with a primary focus on C/C++ development and low-level programming. You'll work closely with large inference and training model development to further drive Scale Out software and hardware performance. This role is hybrid, based out of Toronto, ON. Who You Are Strong C or C++ systems engineer with a deep understanding of memory, threading, I/O, and low-level execution models. Experienced building low-level software, drivers, embedded systems, or performance-critical infrastructure. Comfortable working close to hardware and curious about how systems behave under the hood. Proficient with Linux systems programming and debugging tools such as gdb, strace, and perf. Structured problem solver who thrives in fast-paced, highly technical environments. What We Need Design, develop, and maintain core infrastructure software that interfaces directly with Tenstorrent hardware. Build low-level libraries and APIs for communication and synchronization across compute nodes. Optimize system-level software for performance, scalability, and reliability in distributed environments. Support hardware

AWSLinuxAIC++
T
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is building next-generation CPU and AI silicon. You’ll work at the forefront of hardware innovation, diagnosing complex issues across chips, systems, firmware, and software while collaborating with some of the brightest engineers in the industry. This role offers the opportunity to solve challenging technical problems, build impactful debug solutions, and directly influence the reliability and performance of cutting-edge AI compute platforms. This role is hybrid, based out of Toronto, Canada. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced in hardware debug and post-silicon bring-up for CPU, SoC, or ASIC systems. Strong understanding of processor architecture and microarchitecture (RISC-V, x86, or ARM) with familiarity in debug and trace methodologies (e.g., iJTAG). Hands-on engineer who excels at diagnosing complex hardware, firmware, and software issues through root-cause analysis. Comfortable working in the lab with a passion for building debug tools, automation, and scalable methodologies. Collaborative team player with experience partnering across ASIC, firmware, software, and validation teams. What We Need

PythonAWSAIExcel
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team The Code Quality team sits within the Developer Platform organization and owns the systems that keep DoorDash's codebase healthy and secure as it scales: static analysis, quality gates, test frameworks, regression infrastructure, and tooling. Our job is to make sure the signals engineers rely on before shipping — test results, coverage, performance feedback etc — are fast and trustworthy. The decisions we make about tooling and standards directly shape how confidently and quickly engineering teams at DoorDash can ship to production. About the Role We're looking for Software Engineers to help build and maintain the systems that validate code quality across DoorDash's engineering org, treating our tooling as a critical product for the engineers who rely on it every day: static analysis and quality gates, test frameworks and regression infrastructure. You’ll design the tooling and automation that will help derive trustworthy quality signals, integrate them into the development lifecycle, and make it easy for engineers to execute reliable, repeatable workflows. You will collaborate across the engineering org, partnering directly with the teams who use what you build to understand the accuracy, reliability and performance of their functionality. You will report into the Engineering Manager on our Code Quality team in our Developer Platform organization. You must be located in either San Francisco, CA, Sunnyvale, CA, Los Angeles, CA, Seattle, WA, or New York, NY. You're excited about this opportunity because you will… Build and maintain quality tooling — static analysis, quality gates, coverage reporting, test frameworks, regression infrastructure — and integrate it directly into our developer workflows and CI/CD pipelines Define and derive quality signals - flakiness, pass rate, coverage, performance, scale readiness etc - Build tooling that improves everyday engineering workflows, including local development, CI/CD, debugging, and rollou

AWSCI/CDGitRest
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team DoorDash is a data driven organization and relies on timely, accurate and reliable data to drive many business and product decisions. The Core Data Platform organization owns all the infrastructure necessary to run an operationally efficient analytical data stack. About the Roles The Data Platform team spans data mobility frameworks, ingestion, infrastructure, tools, and governance. Together, they design and operate scalable compute and ingestion frameworks using technologies such as Spark, Flink, Kafka, Airflow, and modern lakehouse solutions, while also building abstractions and tools that simplify data workflows for engineers, analysts, and ML practitioners. In parallel, these teams establish strong data quality, cataloging, privacy, and compliance standards to ensure trust in analytics and regulatory adherence. As relatively high-impact teams, they offer engineers the opportunity to shape the roadmap, influence core platform decisions, and directly enable DoorDash’s business-critical insights and real-time personalization capabilities. You must be located in San Francisco, CA, Sunnyvale, CA, Seattle, WA, or New York, NY. You're excited about this opportunity because you will… Drive vision & strategy for building the frameworks charter and position it to handle the challenges of a rapidly growing business. Scale the analytical platform for the increasing amounts of data and use cases. You will bring your expertise in building and operating high scale systems with a focus on reliability, scalability and cost efficiency. Collaborate with stakeholders building solutions on top of the platform Foster a positive and supportive work culture, upleveling others. We're excited about you because you have… B.S., M.S., or PhD. in Computer Science or equivalent. 2+ years of industry experience at our I4 level, 5+ years of industry experience at our I5 level Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in th

AWSGitRestAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team DoorDash Labs is an independent team within DoorDash. We explore robotics and automation to transform last mile logistics in the long term. If you have a passion for applying robotics solutions in a service used by millions of people, then we want to talk to you! About the Role We're hiring a Robotics Infrastructure Engineer in our Autonomy Software team. In this role, you'll own, build, and manage the infrastructure that makes aerial autonomy development possible. You'll work on the onboard systems that keep a drone alive (process management, health monitoring, parameterization) and the development environment that makes the team fast (build systems, CI/CD, logging, debugging, regression testing). This is not cloud infrastructure. This is real-time, fault-tolerant, onboard software for vehicles that cannot gracefully restart at 50 meters altitude. You're excited about this opportunity because you will… Play an integral role on a small and focused team Develop and own critical onboard components: process management, health monitoring, configuration management, and message passing Own the build system (C++, Python) and middleware layer (ROS2), including cross-compilation for Jetson targets Design and maintain CI/CD pipelines and regression testing infrastructure Build and manage the parameterization system, including schema definition, validation, migration, and deployment Build robotics logging, plotting, and debugging tools that make the entire team more productive Work closely with the simulation team to support SIL/HIL development workflows Define reliability standards for onboard software: watchdogs, failover, and graceful degradation We're excited about you because… You have prior experience at a robotics company in a similar infrastructure role You have experience with robotics middleware (ROS2, LCM, eCal, Apex.AI) You have experience with build systems and package managers (CMake, Bazel, Nix, Conan) You have experience with NVidia Jetson and Je

PythonAWSCI/CDGit
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team Data is at the foundation of DoorDash success. The Data Engineering team builds database solutions for various use cases including reporting, product analytics, marketing optimization and fi nancial reporting. Team serves as the foundation for decision-making at DoorDash. About the Role DoorDash is looking for a Sta ff Software Engineer,Data to be a technical lead and help architect and scale our data reliability, data infrastructure, automation and tools to meet growing business needs. You’re excited about this opportunity because you will... Own critical data systems that support multiple products/teams Develop, implement and enforce best practices for data infrastructure and automation Design, develop and implement large scale, high volume, high performance data models and pipelines for Data Lake and Data Warehouse Improve the reliability and scalability of our Ingestion, data processing, ETLs, Reporting tools and data ecosystem services Manage a portfolio of data products that deliver high-quality, trustworthy data Help onboard and support other engineers as they join the team We’re excited about you because... 8+ years of professional experience as a hands-on engineer and technical leader leading multiple projects 6+ years experience working in data platform and data engineering or a similar role You have proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software Pro fi ciency in programming languages such as Python/Kotlin/Scala 4+ years of experience in ETL orchestration and work fl ow management tools like Air fl ow Expert in database fundamentals, SQL, data reliability practices and distributed computing 4+ years of experience with the Distributed data/similar ecosystem (Spark, Presto) and streaming technologies such as Kaa/Flink/Spark Streaming Excellent communication skills and experience working

PythonSQLAWSGit
G
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Overview: Qsight is a high-growth division of Guidepoint focused on building data intelligence solutions for the healthcare sector. Qsight leverages proprietary datasets and rigorous analysis of alternative data sources to generate actionable insights for top-tier institutional investors, medical device manufacturers, and pharmaceutical companies. The Qsight team develops market intelligence products designed to be highly relevant, accurate, and scalable – delivering superior insights to a diverse, global client base. We are seeking an experienced, motivated Tehnical Operations Engineer to join our growing team. This is a multiple-hats role focused on SaaS/platform operations and tier-2 support for client-facing systems. You will own the administration and reliability of key tools, troubleshoot and resolve escalations with clear documentation, and build lightweight automation and reporting to reduce manual work as we scale. You will partner closely with Customer Success, Product, and Engineering to proactively monitor, support, and improve critical systems. Through practical, creative problem-solving, you will strengthen reliability, accelerate time to resolution, and increase operational visibility. Day to day, you will triage and resolve client technical questions, manage vendor license administration and renewals, and produce reporting that informs operational decisions. This role is a launchpad toward an SRE/Platform Engineering track as you grow into deeper automation, reliability engineering, and systems design work. This is a hybrid position based out of our Toronto office. What You’ll Do: Platform Support Own routine ops and configuration changes for critical SaaS platforms – Including Tableau, Freshdesk, Datadog, and our own client facing and internal portals Configure and maintain Freshdesk portals, routing, SLAs, permissions, integrations, etc. based on business requirements. Automate manual operations with Python, PowerAutomate, and shell scri

PythonSQLRestAI
EA
📍 Remote - US, Canada· Full-time· Remote
✓ High-confidence listing

$200K – $250K/yr

Quick readStrong listing-quality and freshness signals

EnCharge AI is a leader in advanced AI hardware and software systems for edge-to-cloud computing. EnCharge’s robust and scalable next-generation in-memory computing technology provides orders-of-magnitude higher compute efficiency and density compared to today’s best-in-class solutions. The high-performance architecture is coupled with seamless software integration and will enable the immense potential of AI to be accessible in power, energy, and space constrained applications. EnCharge AI launched in 2022 and is led by veteran technologists with backgrounds in semiconductor design and AI systems. Lead DFT Engineer Job Description: Developing silicon for edge-to-cloud computing isn't just about speed; it’s about balancing high-performance data processing with extreme power efficiency and reliability in remote environments. As the Design for Test (DFT) Lead, you will be the architect of our testing strategy, ensuring our data center chips are flawlessly manufacturable and resilient enough for edge deployment. Key Responsibilities: Architectural Leadership: Define and implement the end-to-end DFT architecture for complex SoCs, including Hierarchical DFT, Scan compression, Boundary Scan and MBIST. Edge-Specific Reliability: Develop strategies for In-System Test (IST) and power-on self-test (POST) to ensure chip health in remote edge data centers. Implementation & Flow: Oversee scan insertion, ATPG (Stuck-at, Transition, Path Delay), and Memory /Logic BIST. Cross-Functional Synergy: Collaborate with Design, Physical Design, and Yield teams to ensure high test coverage while minimizing area overhead and power impact as well as timing analysis . Post-Silicon Validation: Lead the bring-up and debug phase on ATE (Automated Test Equipment) to root-cause silicon failures and optimize test time. Technical Requirements: Experience: 12+ years in DFT, with at least 2 years in a leadership or principal role. Bachelor’s degree in a related field. Tools:

C
📍 Canada· Full-time
✓ Quality checkedCompany trend -91.5%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The internal infrastructure team is responsible for building world-class infrastructure and tools used to train, evaluate and serve Cohere's foundational models. By joining our team, you will work in close collaboration with AI researchers to support their AI workload needs on the cutting edge, with a strong focus on stability, scalability, and observability. You will be responsible for building and operating superclusters across multiple clouds. Your work will directly accelerate the development of industry-leading AI models that power Cohere's platform North. Please Note: All of our infrastructure roles require participating in a 24x7 on-call rotation, where you are compensated for your on-call schedule. As a Staff Software Engineer, you will: Build and scale ML-optimized HPC infrastructure : Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads. Optimize for AI/ML training : Collaborate with cloud providers to fine-tune infrastructure for cost efficiency, reliability, and performance , leveraging technologies like R

PythonKubernetesGitLinux
🔔

Get new reliability engineer iii jobs in Canada by email

Daily job updates · Unsubscribe anytime