Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe processes over $1T in payments volume per year, which is roughly 1% of the world’s GDP. The tremendous amount of data makes Stripe one of the best places to do machine learning. The ML Infra team builds services and tools that power every step in the ML lifecycle, including data exploration, feature generation, experimentation, training, deploying, serving ML models, and building LLM applications. With the phenomenal developments happening in the field of AI, we are positioned to accelerate the adoption of AI/ML across all parts of the company by building highly scalable and reliable foundational infrastructure. What you’ll do You will work closely with machine learning engineers, data scientists, and product engineering teams to enable seamless end-to-end experience in building solutions across data, analytics, and AI/ML platforms. You will build the next generation of ML Infra services and major new capabilities that substantially improve ML development velocity and MLOps maturity across the company. Responsibilities Designing and building scalable, reliable, and secure services for notebooks, ML model training, experimentation, serving, and LLM applications across multiple regions. Creating services and libraries that enable ML engineers at Stripe to seamlessly transition from experimentation to production across Stripe’s systems. Working directly with product teams and ML engineers to improve their day-to-day pr
Jobs in Canada
Infrastructure And Mlops Engineer in Canada
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current infrastructure and mlops engineer jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.
C$135K – C$210K/yr
Overview: Guidepoint seeks an experienced Data/AI Engineer as an integral member of the Toronto-based AI team. The Toronto Technology Hub serves as the base of our Data/AI/ML team, dedicated to building a modern data infrastructure for advanced analytics and the development of responsible AI. This strategic investment is integral to Guidepoint’s vision for the future, aiming to develop cutting-edge Generative AI and analytical capabilities that will underpin Guidepoint’s Next-Gen research enablement platform and data products. This role demands exceptional leadership and technical prowess to drive the development of next-generation research enablement platforms and AI-driven data products. You will develop and scale Generative AI-powered systems, including large language model (LLM) applications and research agents, while ensuring the integration of responsible AI and best-in-class MLOps. The Senior AI/ML Engineer will be a primary contributor to building scalable AI/ML capabilities using Databricks and other state-of-the-art tools across all of Guidepoint’s products. Guidepoint’s Technology team thrives on problem-solving and creating happier users. As Guidepoint works to achieve its mission of making individuals, businesses, and the world smarter through personalized knowledge-sharing solutions, the engineering team is taking on challenges to improve our internal application architecture and create new AI-enabled products to optimize the seamless delivery of our services. This is a hybrid position based in Toronto. What You'll Do: Architect and Build Production Systems: Design, build, and operate scalable, low-latency backend services and APIs that serve Generative AI features, from retrieval-augmented generation (RAG) pipelines to complex agentic systems. Own the AI Application Lifecycle: Own the end-to-end lifecycle of AI-powered applications, including system design, development, deployment (CI/CD), monitoring, and optimization
C$135K – C$210K/yr
Overview: Guidepoint seeks an experienced AI Engineer as an integral member of the Toronto-based AI team. The Toronto Technology Hub serves as the base of our Data/AI/ML team, dedicated to building a modern data infrastructure for advanced analytics and the development of responsible AI. This strategic investment is integral to Guidepoint’s vision for the future, aiming to develop cutting-edge Generative AI and analytical capabilities that will underpin Guidepoint’s Next-Gen research enablement platform and data products. This role demands exceptional leadership and technical prowess to drive the development of next-generation research enablement platforms and AI-driven data products. You will develop and scale Generative AI-powered systems, including large language model (LLM) applications and research agents, while ensuring the integration of responsible AI and best-in-class MLOps. The AI/ML Engineer will be a primary contributor to building scalable AI/ML capabilities using Databricks and other state-of-the-art tools across all of Guidepoint’s products. Guidepoint’s Technology team thrives on problem-solving and creating happier users. As Guidepoint works to achieve its mission of making individuals, businesses, and the world smarter through personalized knowledge-sharing solutions, the engineering team is taking on challenges to improve our internal application architecture and create new AI-enabled products to optimize the seamless delivery of our services. This is a hybrid position based in Toronto. What You'll Do: Architect and Build Production Systems: Design, build, and operate scalable, low-latency backend services and APIs that serve Generative AI features, from retrieval-augmented generation (RAG) pipelines to complex agentic systems. Own the AI Application Lifecycle: Own the end-to-end lifecycle of AI-powered applications, including system design, development, deployment (CI/CD), monitoring, and optimization in production
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft Infrastructure builds the systems engineers depend on to ship stable, scalable, and efficient services. We're hiring a Senior Technical Program Manager to run cross-functional programs across our infrastructure and data platform teams. This role blends program delivery with product sense: you'll own the roadmap for your area, set priorities, and act as the voice of the customer back into how we build. Responsibilities Run infrastructure programs end to end, from kickoff through delivery Own the roadmap for your platform area: shape the strategy, sequence the work, and make the prioritization calls Drive data platform migration and modernization work, coordinating across engineering, data, and platform teams to keep dependencies and timelines under control Be the voice of the customer: partner with engineering teams across Lyft, surface their pain points, and feed that back into priorities and roadmaps Define success metrics and adoption goals, gather and document customer requirements, and make sure what ships actually solves the problem Build feedback loops with customer teams and turn what you hear into concrete improvements Partner with engineering and infrastructure leads to build plans, call out risks early, and keep stakeholders aligned Own program health: track milestones, surface blockers before they slip, and keep decision-makers in the loop Use your technical background in distributed systems and data infrastructure to ask sharp questions and build plans the team believes in Share in the team's release oncall rotation Experience 5+ years in Technical Program Management or a TPM/PM hybrid role A background in software, data, or systems engineering, enough to go deep with engineers Experience owning a roadmap: setting strategy, prioritizing across competing demands, and defining what suc
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We are a software engineering team with expertise in enabling ML models in production. We deploy AI models to run in variety of environments: air-gapped government networks, forward-deployed defense environments, edge nodes, and enterprises with strict data sovereignty requirements. Our customers rely on us for frontier AI capabilities running on hardware they control, often with constrained GPU resources and limited direct access. Rising to that challenge and meeting those expectations is what Palantir's excels at. We treat models like any other software: continuously tested, continually delivered, packaged for reproducible deployment, and built for long-term maintainability. You will own services end-to-end, and work across the full stack, from inference engines, GPU scheduling to deployment pipelines, observability, and integration with Palantir's platform. The goal is to deliver new models and capabilities quickly and continuously. Join us if you want to solve problems at the intersection of infrastructure and machine learning that directly enable critical customers.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role As a Platform Engineer on Palantir's Identity Platform team, you will design, build, and operate secure-by-design identity infrastructure and tooling. You will make identity governance and access management easier and more secure to implement for Palantirians and customers worldwide. As part of Palantir's best-in-class Information Security organization, you will research, implement, and scale innovative solutions that help Palantir stay ahead of a dynamic threat landscape. You will help the team build the next-generation identity platform that treats agents and workload identities as first-class principals: the access graph that makes entitlements legible, the policy engine that makes authorization enforceable and auditable, and the token-issuance layer that keeps access short-lived, scoped, and least-privilege by default. Your work will shape the arc of these components from design to production. You will join the high-performing Identity Platform team, engineers who are passionate about delivering identity outcomes at scale that reduce risk and friction. The team builds and operates the identity platforms that serve both corporate and production (customer-facing) infrastructure, and the paved-road tooling and secure baselines that teams across Palantir deploy on. Your goal will be to make the secure path the easy path. Your work will directly strengthen the identity substrate beneath Palantir's most critical deployments, from a globally distributed workforce to regulated and air-gapped environments.
From C$107K/yr
Location Details: Canada, Remote At GoDaddy, the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) , and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Contribute to the development of GoDaddy’s eCommerce and SSO infrastructure and Kubernetes systems on AWS. On a day-to-day basis you will be working on the team who designs, writes, tests and deploys the infrastructure and application management software for GoDaddy’s eCommerce applications. Expect to learn every day. What you'll get to do... Work as a polyglot engineer, writing and maintaining Infrastructure as code with frameworks/ ecosystems such as Java, Unix CLI, and NodeJS Build and operate infrastructure workflows and deployment pipelines using Kubernetes, Argo Workflows, Argo CD, and GitOps practices Design, build, and own services and APIs in Java, running on Kubernetes-based platforms across AWS and distributed systems Develop and support application and infrastructure delivery pipelines, enabling reliable releases of eComm, Auth and Infrastructure services Collaborate closely with other GoDaddy departments to help advance security and technical standards, maintain regulatory compliances while operating eComm & Auth platforms Your experience should include... 5+ years of strong backend software engineering experience in Java Hands-on experience with Kubernetes, including Helm, Kustomize, or equivalent tools to deploy and manage backend services Experience building and operating high-volume, mission-critical production systems on AWS with continuous deployment (CD) practices Strong experience with infrastructure as code, supporting backend applications and services Experience with observability a
From C$136K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Auth0 is growing rapidly and looking for exceptional new team members to help take us to the next level. One team, one score. We never compromise on identity. You should never compromise yours either. We want you to bring your whole self to Auth0. If you're passionate, practice radical transparency to build trust and respect, and thrive when you're collaborating, experimenting and learning – this may be your ideal work environment. We are looking for team members that want to help us build upon what we have accomplished so far and make it better every day. N+1 > N. About the Team Here at Auth0 we’re focused on securing the world’s identities so innovators can innovate. We’re currently hiring a senior Full Stack Software Engineer to join our Acquisitions and Activation Team . This team owns the crucial first impressions of the Auth0 ecosystem. We apply rigorous, data-driven experimentation and A/B testing to optimize the entire early customer journey—from the moment a developer signs up, to the precise second they hit their "ah-ha" moment. We bridge the gap between deep technical infrastructure and user psychology, building primarily on a JavaScript ecosystem (Node.js and React) . If you are a full-stack engineer who loves blending robust software architecture with rapid, metrics-driven experimentation, this is the team for you. What You Will Do Optimize the Acquisition Funnel: Own, build, and continuously scale the entire signup experience to reduce fric
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The internal infrastructure team is responsible for building world-class infrastructure and tools used to train, evaluate and serve Cohere's foundational models. By joining our team, you will work in close collaboration with AI researchers to support their AI workload needs on the cutting edge, with a strong focus on stability, scalability, and observability. You will be responsible for building and operating superclusters across multiple clouds. Your work will directly accelerate the development of industry-leading AI models that power Cohere's platform North. Please Note: All of our infrastructure roles require participating in a 24x7 on-call rotation, where you are compensated for your on-call schedule. As a Staff Software Engineer, you will: Build and scale ML-optimized HPC infrastructure : Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads. Optimize for AI/ML training : Collaborate with cloud providers to fine-tune infrastructure for cost efficiency, reliability, and performance , leveraging technologies like R
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Load and Fault team sits within Robinhood's Developer Infrastructure organization, with a mission to give every engineering team the tools they need to test their services under real-world conditions before those conditions test them in production. We build the platforms and frameworks that enable load testing, fault injection, and resilience validation at scale — treating reliability as a developer productivity problem, not just an operations one. Our work directly raises the quality bar for every service Robinhood ships, and we partner closely with engineering teams across the organization to make resilience testing a seamless part of the development workflow. As a Senior Software Engineer on the Load and Fault Environments team, you will design and build the infrastructure that lets Robinhood's engineers simulate load, inject faults, and validate system behavior under stress — at the scale of a fast-growing financial platform. You'll own meaningful components of the load testing and fault injection platform, write production-quality code, and collaborate with engineers across infrastructure and product teams to ensure the tooling you build gets adopted and drives real
About the Team The Code Quality team sits within the Developer Platform organization and owns the systems that keep DoorDash's codebase healthy and secure as it scales: static analysis, quality gates, test frameworks, regression infrastructure, and tooling. Our job is to make sure the signals engineers rely on before shipping — test results, coverage, performance feedback etc — are fast and trustworthy. The decisions we make about tooling and standards directly shape how confidently and quickly engineering teams at DoorDash can ship to production. About the Role We're looking for Software Engineers to help build and maintain the systems that validate code quality across DoorDash's engineering org, treating our tooling as a critical product for the engineers who rely on it every day: static analysis and quality gates, test frameworks and regression infrastructure. You’ll design the tooling and automation that will help derive trustworthy quality signals, integrate them into the development lifecycle, and make it easy for engineers to execute reliable, repeatable workflows. You will collaborate across the engineering org, partnering directly with the teams who use what you build to understand the accuracy, reliability and performance of their functionality. You will report into the Engineering Manager on our Code Quality team in our Developer Platform organization. You must be located in either San Francisco, CA, Sunnyvale, CA, Los Angeles, CA, Seattle, WA, or New York, NY. You're excited about this opportunity because you will… Build and maintain quality tooling — static analysis, quality gates, coverage reporting, test frameworks, regression infrastructure — and integrate it directly into our developer workflows and CI/CD pipelines Define and derive quality signals - flakiness, pass rate, coverage, performance, scale readiness etc - Build tooling that improves everyday engineering workflows, including local development, CI/CD, debugging, and rollou
C$113.4K – C$162K/yr
We believe communication belongs to everyone. We exist to democratize phone service. TextNow is evolving the way the world connects and that's because we're made up of people with curious minds who bring an optimistic, yet critical lens into the work we do. We're the largest provider of free phone service in the nation. And we're just getting started. Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between! This role is about impact at scale. You’ll shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded into everyday engineering practice. Using AI is not optional, it’s expected. From design and architecture to implementation, testing, debugging, documentation, and operational analysis, you will actively leverage AI tools to increase velocity, improve code quality, and make better technical decisions. We provide a robust suite of AI-powered development tools and workflows to support you, and we expect you to continuously evolve how you use them to raise the bar for efficiency, clarity, and product excellence across the organization. What You'll Do Ensure System Reliability: Design, build, and maintain scalable, resilient, and highly available systems to support TextNow’s infrastructure and services. Automation & Infrastructure as Code: Develop and maintain automation using Terraform, Ansible, and other tools to enable efficient deployment, scaling, and operations of cloud-based systems (AWS preferred). Incident Response & On-Call Support: Participate in an on-call rotation, troubleshoot issues, and drive incident resolution to minimize downtime and improve syste
$110K – $150K/yr
Location: San Francisco, CA (hybrid) What is Verse? The race to AI has become the race to power. Every breakthrough in artificial intelligence depends on one thing: access to electricity. But across the country, aging grid infrastructure and years-long interconnection queues are slowing the deployment of the data centers that will power the next generation of innovation. Solving this challenge isn't just about energy—it's about unlocking the future of AI. At Verse, we're building the energy intelligence platform for the AI economy. Our software helps the world's largest energy consumers achieve faster, cheaper, and cleaner power by combining real-time control of energy assets with complete visibility into their energy portfolio. Backed by Bessemer Venture Partners, GV, Coatue, and NVIDIA, and built by pioneers in grid-scale batteries, energy markets, and enterprise software, we're redefining how the world's most ambitious organizations access and manage energy. The Role We're looking for a highly analytical Senior Business Operations Analyst to support the growth and execution of our Dispatch Intelligence product. This role sits at the intersection of business operations, customer, product, engineering, and data science and has two core areas of responsibility. First, you will help drive overall program execution for Dispatch Intelligence: bringing structure to complex cross-functional initiatives, improving processes, tracking progress, and ensuring teams stay aligned on priorities and timelines. Second, you will help build and operate the processes through which flexible energy assets are onboarded onto the Verse platform and continuously improve their operational performance. The ideal candidate combines strong analytical problem-solving with exceptional project management and is comfortable working across both technical and commercial teams. You will play a critical role in helping Verse scale Dispatch Intelligence from individual projects and assets
From C$1.5K/yr
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Software Platform team accelerates developer velocity and increases system reliability by building the foundational platforms and tools that power Robinhood engineering. Within this group, the Kubernetes Compute team focuses on building and operating a highly available, scalable Kubernetes-powered container platform. We ensure that our infrastructure seamlessly supports reliable application deployments, integrates core platform capabilities, and enables multi-region scalability. We are expanding our core container systems to support our next phase of technical growth! As a Senior Software Develope r, you will focus heavily on building, operating, and expanding our container provisioning platforms. You will be responsible for designing resilient container infrastructure and contributing to our technical migration to Amazon EKS to improve platform reliability. In this position, you will collaborate with engineering teams across Robinhood to deliver reliable platform integrations for core capabilities like networking and security. Your work will directly help our infrastructure scale efficiently while maintaining a high standard of safety and system uptime. This role is bas
From C$136K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Platform Network Engineering Team Auth0 by Okta is an easy-to-implement authentication and authorization platform designed by developers for developers. We make access to applications safe, secure, and seamless for over 100 million daily logins worldwide. Our modern approach to identity enables this Tier 0 global service to deliver convenience, privacy, and security so customers can focus on innovation. The Senior Software Engineer Opportunity You will be part of the Platform Network engineering team responsible for all connectivity of Auth0. You will play a key engineering role as we evolve our network architecture to meet the demands of enormous growth and support the hundreds of millions of users who rely on us to provide uninterrupted access. You will get to work with engineers throughout the engineering organization. What you’ll be doing Implement internal and edge networking infrastructure and design solutions that work at global scale and with multi-cloud and multi-region constraints. Carry cross-team initiatives from end to end: code reviews, design reviews, operational robustness, security hygiene, etc. Design and develop new services, tools, and automation to expose network functionality to other Okta engineering and operations teams. Research and implement solutions addressing cross-cutting concerns such as routing, failover, and scaling. Participate in the team’s on-call rotation. What you’ll bring to the role Have 3+ years of
Other cities to consider
More places hiring for this role
Get new infrastructure and mlops engineer jobs in Canada by email
Daily job updates · Unsubscribe anytime