Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. As a Site Reliability Engineer you will: Build self-service systems that automate managing, deploying and operating services. This includes our custom Kubernetes operators that support language model deployments. Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems. Take steps required to ensure we hit defined SLOs, including pa
Jobs in Canada
Reliability Engineer Iii in Toronto
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current reliability engineer iii jobs in Toronto. Filter by work mode, employment type, experience, department, date posted and distance.
From C$1.4M/yr
About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission is to engineer resilience from the ground up, enabling our product teams to innovate rapidly while ensuring our users have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through a culture of data-driven decision-making, blameless learning, and relentless automation. As a Senior Site Reliability Engineer, you are a hands-on engineer who blends deep software development expertise with a passion for operational excellence. You will be responsible for designing, building, and running the resilient, scalable, and increasingly self-healing systems that power our products. You will apply sound engineering principles to solve our most complex reliability challenges, with a mandate to automate everything, eliminate toil, and write robust, maintainable code. You will be a force multiplier, mentoring other engineers and elevating the site reliability bar for the entire organization. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: System Architecture & Design: Design, build, and maintain scalable, highly available, and fault-tolerant distributed systems. Partner with development teams as a reliability consultant, reviewing designs and influencing architectural decisions to ensure new services are built with reliability, observability, and performance as core principles, not afterthoughts. Automation & Software Development: Write robust, performant, and maintainable code to automate operational tasks, and CI/CD pipelines. Build the internal tools, libraries, and frameworks that enable engineering teams to self-service their
From C$160K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Staff Software Reliability Engineer - Data Platform About the Team The Data Platform team is responsible for the foundational data services, systems, and data products for Okta that benefit our users. Today, the Data Platform team solves challenges and enables: Streaming analytics Interactive end-user reporting Data and ML platform for Okta to scale Telemetry of our products and data Our elite team is fast, creative and flexible. We encourage ownership. We expect great things from our engineers and reward them with stimulating new projects, new technologies and the chance to have significant equity in a company. Okta is about to change the cloud computing landscape forever. About the Position This is an opportunity for experienced Software Reliability Engineers to join our fast growing Data Platform organization that is passionate about scaling high volume, low-latency, distributed data-platform services & data products. In this role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Okta to scale for years to come. As a member of the Data Platform team, you will be responsible for designing, building, and deploying the systems that power our data analytics and ML. Our analytics infrastructure stack sits on top of many modern technologies, including Kinesis, Flink, ElasticSearch, and Snowflake. We are looking for experienced Software Engineers who can help desi
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation of AI computing systems. In this highly visible technical leadership role, you'll drive reliability from architecture through production, partnering across hardware, software, and manufacturing teams to build high-performance AI platforms that set the standard for uptime, durability, and quality. If you're passionate about solving complex engineering challenges and influencing products at scale, you'll have the opportunity to shape technology powering the future of AI. This role is hybrid, based out of Toronto, Canada. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You've spent 8+ years in reliability engineering, ideally in high-performance computing, AI hardware, or data center systems. You're comfortable with the statistical side of the job, HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA are all familiar territory. You can work through a technical problem in a thermal lab and then explain the risks and trade-offs clearly to leadership. You're good at bringing people together, mechanical, electrical, thermal, softw
From C$40/hr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Interns work side-by-side with top engineers in the industry while having autonomy from the get-go. They contribute to user-facing products and are able to see their work go live quickly. Lyft fosters a collaborative environment in the office, so there's always a sharp mind eager to hear about your next idea. So what's yours? Responsibilities: Own your project, while checking in with other team members throughout the day with questions and updates You leave the code in a better state than when you found it (progressive refactor) You value reliability, ensured by testing (unit, integration and load tests) Participate in code reviews to ensure code quality and distribute knowledge Continuous integration and deployment Go home knowing that your work today is meaningfully improving the lives of every Lyft driver and every Lyft passenger! Experience: Currently pursuing a Bachelor's or Master's degree in Computer Science from a university in Canada (required) , with a graduation date between December 2027 and Summer 2028 (required). For any candidates who are master's students who worked between their bachelor's and master's programs: candidates should also have less than 2 years of relevant full-time work experience Available during Summer 2027 for the internship in Toronto Strong knowledge of CS fundamentals Excellent communication skills Passion for community, sustainability, and/or transportation Ability to thrive in a startup environment Experience with real-time technology problems Contributions to open source projects Experience working with databases Experience solving real-time technology problems Experience with mobile development Benefits: Mental health benefits In addition to holidays, interns receive 2 days paid time off and 3 days sick time off Subsidized commuter benefits and Lyft ride credi
From C$40/hr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Interns work side-by-side with top engineers in the industry while having autonomy from the get-go. They contribute to user-facing products and are able to see their work go live quickly. Lyft fosters a collaborative environment in the office, so there's always a sharp mind eager to hear about your next idea. So what's yours? Responsibilities: Own your project, while checking in with other team members throughout the day with questions and updates You leave the code in a better state than when you found it (progressive refactor) You value reliability, ensured by testing Participate in code reviews to ensure code quality and distribute knowledge Continuous integration and deployment Go home knowing that your work today is meaningfully improving the lives of every Lyft driver and every Lyft passenger! Experience: Currently pursuing a Bachelor's or Master's degree in Computer Science, Data Science or related major from a university in Canada (required) , with a graduation date between December 2027 and Summer 2028 (required) . For any candidates who are master's students who worked between their bachelor's and master's programs: candidates should also have less than 2 years of relevant full-time work experience Available during Summer 2027 for an internship in Toronto Strong knowledge of CS fundamentals Knowledge of SQL and data modeling fundamentals Experience working with databases Excellent communication skills Interest in solving large scale data problems in a real world scenario Passion for community, sustainability, and/or transportation Benefits: Mental health benefits In addition to holidays, interns receive 2 days paid time off and 3 days sick time off Subsidized commuter benefits and Lyft ride credits Lyft is committed to creating an inclusive workforce that fosters belonging. Lyft belie
From C$40/hr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Interns work side-by-side with top engineers in the industry while having autonomy from the get-go. They contribute to user-facing products and are able to see their work go live quickly. Lyft fosters a collaborative environment in the office, so there's always a sharp mind eager to hear about your next idea. So what's yours? Responsibilities: Own your project, while checking in with other team members throughout the day with questions and updates You leave the code in a better state than when you found it (progressive refactor) You value reliability, ensured by testing (unit, integration and load tests) Participate in code reviews to ensure code quality and distribute knowledge Continuous integration and deployment Go home knowing that your work today is meaningfully improving the lives of every Lyft driver and every Lyft passenger! Experience: Currently pursuing a Bachelor's or Master's degree in Computer Science from a university in Canada (required) , with a graduation date between December 2027 and Summer 2028 (required). For any candidates who are master's students who worked between their bachelor's and master's programs: candidates should also have less than 2 years of relevant full-time work experience Available during Summer 2027 for an internship in Toronto Strong knowledge of CS fundamentals Knowledge of Python, JavaScript, CSS, and HTML Experience working with leading JavaScript frameworks, like React Experience with modern frontend testing tools, such as Webpack, Babel, Jest, Jasmine, Protractor, and WebDriver Understanding of how browsers and DOM work Experience with Git or other distributed version control systems Experience with browser developer tools Experience with the Unix command line interface Solid understanding of web performance Experience with TypeScript Experience with CSS
From C$1.3M/yr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Safety and Customer Care (SCC) team at Lyft manages over 1.7 million monthly human and AI interactions and serves as Lyft's primary direct touchpoint with riders and drivers. We handle critical infrastructure that powers both human associates and AI agents to make riders and drivers feel safe and comfortable while riding or driving with Lyft, transforming every support interaction into a moment of genuine connection. As a Data Engineer on the SCC team, you will have ownership over the data modeling and pipelines that power SCC’s Associate and AI Agent Platform . Your efforts will be critical to the reliability of our pipelines, execution of third party data integrations, accurate reporting of agents performance, and efficiency improvements that can save millions of dollars / year. You will work cross-functionally to bridge Lyft's business goals with data engineering. Your efforts will allow access to business and user behavior insights, using huge amounts of Lyft data to fuel several teams such as Analytics, Data Science, Engineering, and many others. Responsibilities: Owner of the core data pipeline, responsible for scaling up data processing flow to meet the rapid data growth at Lyft Evolve data model and data schema based on business and engineering needs Implement systems tracking data quality and consistency Develop tools supporting self-service data pipeline management (ETL) SQL and MapReduce job tuning to improve data processing performance Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and distribute knowledge Collaborate cross-functionally with product, engineering, data science, and marketing teams to understand business problems and align on prioritization and solutions Experience: Bachelor's degree in Compute
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Okta Privileged Access Management (PAM) is an identity-centric approach to a common and critical privileged access use case. Our elegant Zero Trust architecture is purpose-built for the modern cloud and helps customers solve challenging security and operations pain points at scale. We are looking for a software engineer to join our fast-growing team with a focus on scalability, reliability, and enhancing the core building blocks of the product. In this role you will: Be deeply involved in evolving the core architecture of PAM. Work in our product development teams to build scalable, composable components of our platform. Be responsible for designing and implementing scalable architecture patterns. Delight our customers by providing world class UX using our React-based design system Design and build APIs that customers rely on for access to production infrastructure. Work on backend components written in Go and frontend components written in React. You might be a good fit if you: Have 3-5 years of software development experience with a background in Golang or similar programming languages. Proficient in React or similar front-end UI stacks. Experienced working with relational databases like PostgreSQL or similar RDBMS technologies. have the ability to complete a feature end to end from designing database models to backend APIs and frontend UI components. Experienced working with any cloud provider such as AWS, GCP or Azure. Thrive in a collaborativ
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Driver team is dedicated to fostering a platform of high-quality service by empowering drivers to perform their best. We are looking for a product-minded engineer who wants to build and improve products that sit at the center of the core driver experience. Products you drive will solve pain points that matter most to drivers by streamlining key interactions, reducing friction, and creating systems that feel intuitive, fair, and supportive. As an engineer at Lyft, you'll collaborate with teams like product, data science, analytics, and operations on code that empower us to iterate quickly, while focusing on delighting our passengers and drivers. Responsibilities: Design and implement backend features end-to-end with clear ownership, delivering well-scoped work from technical design through to production with moderate guidance from senior engineers Write clean, reliable, well-tested code that meets team standards and holds up in code review Participate actively in code reviews, giving specific and constructive feedback while continuing to develop your own review instincts Debug and resolve issues across backend services including performance bottlenecks, reliability problems, and data integrity issues Collaborate with product managers, designers, and partner engineering teams to clarify requirements and surface technical constraints early Contribute to technical discussions and help evaluate implementation approaches for new features Write unit and integration tests for your own code and develop familiarity with the team's broader testing and observability practices Address technical debt and make incremental improvements to existing services as part of regular development work Participate in on-call rotations, respond to production incidents, and support teammates in mitigating custome
From C$1.2M/yr
About the Role: We are looking for a talented Automation Engineer to join our Automation Engineering team in Toronto. In this role, you will be responsible for designing and implementing automated tests for Mobile development. You will collaborate closely with QA engineers and developers to build scalable test frameworks, improve automation coverage, and contribute to the efficiency of our multi-platform release process. You will also design data-driven end-to-end checks around playback and ad insertion , integrate them into CI/CD pipelines as quality gates, and operate a reliable device lab to prevent regressions from shipping. Your work will directly accelerate testing and release velocity while improving revenue-critical reliability across Tubi’s Android and IOS apps. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: Design, implement, and maintain automated tests for mobile development (Android & iOS) Contribute to the development and optimization of cross-platform automation frameworks. Write and maintain test scripts in JavaScript/TypeScript , using frameworks such as Puppeteer, Appium, WebDriverIO, Selenium. Ensure test cases are integrated into CI/CD pipelines and provide reliable feedback on product quality. Help identify flaky tests, investigate root causes, and improve test stability. Collaborate with developers and QA engineers to clarify requirements and improve test strategies. Participate in code reviews and follow best practices for test automation . Your Background: Bachelor’s degree or above in a technical field (e.g., Computer Science, Engineering, Mathematics), or equivalent industry experience. 3+ years of hands-on experience in automation testing for mobile devices Strong programming skills in JavaScript/TypeScript (preferred), or Python/Java. Experience with automation frameworks (e.g. Puppeteer, Appium,, WebDriverIO, Selenium, Playwright, T
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a skilled Software Engineer with a passion for building high-performance, low-level systems software. In this role, you’ll contribute to the development and optimization of the infrastructure that powers our cutting-edge processors, with a primary focus on C/C++ development and low-level programming. You'll work closely with large inference and training model development to further drive Scale Out software and hardware performance. This role is hybrid, based out of Toronto, ON. Who You Are Strong C or C++ systems engineer with a deep understanding of memory, threading, I/O, and low-level execution models. Experienced building low-level software, drivers, embedded systems, or performance-critical infrastructure. Comfortable working close to hardware and curious about how systems behave under the hood. Proficient with Linux systems programming and debugging tools such as gdb, strace, and perf. Structured problem solver who thrives in fast-paced, highly technical environments. What We Need Design, develop, and maintain core infrastructure software that interfaces directly with Tenstorrent hardware. Build low-level libraries and APIs for communication and synchronization across compute nodes. Optimize system-level software for performance, scalability, and reliability in distributed environments. Support hardware
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is building next-generation CPU and AI silicon. You’ll work at the forefront of hardware innovation, diagnosing complex issues across chips, systems, firmware, and software while collaborating with some of the brightest engineers in the industry. This role offers the opportunity to solve challenging technical problems, build impactful debug solutions, and directly influence the reliability and performance of cutting-edge AI compute platforms. This role is hybrid, based out of Toronto, Canada. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced in hardware debug and post-silicon bring-up for CPU, SoC, or ASIC systems. Strong understanding of processor architecture and microarchitecture (RISC-V, x86, or ARM) with familiarity in debug and trace methodologies (e.g., iJTAG). Hands-on engineer who excels at diagnosing complex hardware, firmware, and software issues through root-cause analysis. Comfortable working in the lab with a passion for building debug tools, automation, and scalable methodologies. Collaborative team player with experience partnering across ASIC, firmware, software, and validation teams. What We Need
Overview: Qsight is a high-growth division of Guidepoint focused on building data intelligence solutions for the healthcare sector. Qsight leverages proprietary datasets and rigorous analysis of alternative data sources to generate actionable insights for top-tier institutional investors, medical device manufacturers, and pharmaceutical companies. The Qsight team develops market intelligence products designed to be highly relevant, accurate, and scalable – delivering superior insights to a diverse, global client base. We are seeking an experienced, motivated Tehnical Operations Engineer to join our growing team. This is a multiple-hats role focused on SaaS/platform operations and tier-2 support for client-facing systems. You will own the administration and reliability of key tools, troubleshoot and resolve escalations with clear documentation, and build lightweight automation and reporting to reduce manual work as we scale. You will partner closely with Customer Success, Product, and Engineering to proactively monitor, support, and improve critical systems. Through practical, creative problem-solving, you will strengthen reliability, accelerate time to resolution, and increase operational visibility. Day to day, you will triage and resolve client technical questions, manage vendor license administration and renewals, and produce reporting that informs operational decisions. This role is a launchpad toward an SRE/Platform Engineering track as you grow into deeper automation, reliability engineering, and systems design work. This is a hybrid position based out of our Toronto office. What You’ll Do: Platform Support Own routine ops and configuration changes for critical SaaS platforms – Including Tableau, Freshdesk, Datadog, and our own client facing and internal portals Configure and maintain Freshdesk portals, routing, SLAs, permissions, integrations, etc. based on business requirements. Automate manual operations with Python, PowerAutomate, and shell scri
From C$118.8K/yr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Machine Learning is at the heart of Lyft’s products and decision-making. Machine Learning Engineers at Lyft operate in dynamic environments, moving quickly to build the world’s best transportation solutions. We tackle a wide range of challenges, from pricing and marketplace frameworks that ensure reliability and competitiveness, to agentic AI platforms that automate analytical workflows, to behavioral detection systems that protect the integrity of our network. We operate at the intersection of applied ML and real business impact, shipping models that directly influence revenue, rider experience, and partner trust. Lyft Business builds products that help organizations move the people who matter most—employees, customers, patients, and guests—easily and efficiently. Our offerings include Business Travel, Lyft Pass, and Concierge (for healthcare and non-healthcare rides), enabling companies to manage transportation at scale through APIs, integrations (e.g., Concur, Expensify), and dedicated tools. These platforms power high-impact B2B use cases across corporate travel, healthcare access, customer experience, and community programs. We're looking for a Machine Learning Engineer to design, build, and deploy ML systems across Lyft Business. This is a high-scope role: you won't be siloed into one problem area. Instead, you'll move across pricing algorithms, fraud and behavior detection, agentic AI systems, and emerging ML applications as the business evolves. You'll write production-quality code, own models end-to-end from prototyping through deployment, and collaborate closely with Data Scientists, Product Managers, and Software Engineers to translate complex business problems into scalable ML solutions. This role is ideal for someone who is technically versatile, energized by variety, and wants to see th
Other cities to consider
More places hiring for this role
Get new reliability engineer iii jobs in Toronto, Canada by email
Daily job updates · Unsubscribe anytime