Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe processes over $1T in payments volume per year, which is roughly 1% of the world’s GDP. The tremendous amount of data makes Stripe one of the best places to do machine learning. The ML Infra team builds services and tools that power every step in the ML lifecycle, including data exploration, feature generation, experimentation, training, deploying, serving ML models, and building LLM applications. With the phenomenal developments happening in the field of AI, we are positioned to accelerate the adoption of AI/ML across all parts of the company by building highly scalable and reliable foundational infrastructure. What you’ll do You will work closely with machine learning engineers, data scientists, and product engineering teams to enable seamless end-to-end experience in building solutions across data, analytics, and AI/ML platforms. You will build the next generation of ML Infra services and major new capabilities that substantially improve ML development velocity and MLOps maturity across the company. Responsibilities Designing and building scalable, reliable, and secure services for notebooks, ML model training, experimentation, serving, and LLM applications across multiple regions. Creating services and libraries that enable ML engineers at Stripe to seamlessly transition from experimentation to production across Stripe’s systems. Working directly with product teams and ML engineers to improve their day-to-day pr
Jobs in Canada
Infrastructure And Mlops Engineer in Toronto
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current infrastructure and mlops engineer jobs in Toronto. Filter by work mode, employment type, experience, department, date posted and distance.
C$135K – C$210K/yr
Overview: Guidepoint seeks an experienced Data/AI Engineer as an integral member of the Toronto-based AI team. The Toronto Technology Hub serves as the base of our Data/AI/ML team, dedicated to building a modern data infrastructure for advanced analytics and the development of responsible AI. This strategic investment is integral to Guidepoint’s vision for the future, aiming to develop cutting-edge Generative AI and analytical capabilities that will underpin Guidepoint’s Next-Gen research enablement platform and data products. This role demands exceptional leadership and technical prowess to drive the development of next-generation research enablement platforms and AI-driven data products. You will develop and scale Generative AI-powered systems, including large language model (LLM) applications and research agents, while ensuring the integration of responsible AI and best-in-class MLOps. The Senior AI/ML Engineer will be a primary contributor to building scalable AI/ML capabilities using Databricks and other state-of-the-art tools across all of Guidepoint’s products. Guidepoint’s Technology team thrives on problem-solving and creating happier users. As Guidepoint works to achieve its mission of making individuals, businesses, and the world smarter through personalized knowledge-sharing solutions, the engineering team is taking on challenges to improve our internal application architecture and create new AI-enabled products to optimize the seamless delivery of our services. This is a hybrid position based in Toronto. What You'll Do: Architect and Build Production Systems: Design, build, and operate scalable, low-latency backend services and APIs that serve Generative AI features, from retrieval-augmented generation (RAG) pipelines to complex agentic systems. Own the AI Application Lifecycle: Own the end-to-end lifecycle of AI-powered applications, including system design, development, deployment (CI/CD), monitoring, and optimization
C$135K – C$210K/yr
Overview: Guidepoint seeks an experienced AI Engineer as an integral member of the Toronto-based AI team. The Toronto Technology Hub serves as the base of our Data/AI/ML team, dedicated to building a modern data infrastructure for advanced analytics and the development of responsible AI. This strategic investment is integral to Guidepoint’s vision for the future, aiming to develop cutting-edge Generative AI and analytical capabilities that will underpin Guidepoint’s Next-Gen research enablement platform and data products. This role demands exceptional leadership and technical prowess to drive the development of next-generation research enablement platforms and AI-driven data products. You will develop and scale Generative AI-powered systems, including large language model (LLM) applications and research agents, while ensuring the integration of responsible AI and best-in-class MLOps. The AI/ML Engineer will be a primary contributor to building scalable AI/ML capabilities using Databricks and other state-of-the-art tools across all of Guidepoint’s products. Guidepoint’s Technology team thrives on problem-solving and creating happier users. As Guidepoint works to achieve its mission of making individuals, businesses, and the world smarter through personalized knowledge-sharing solutions, the engineering team is taking on challenges to improve our internal application architecture and create new AI-enabled products to optimize the seamless delivery of our services. This is a hybrid position based in Toronto. What You'll Do: Architect and Build Production Systems: Design, build, and operate scalable, low-latency backend services and APIs that serve Generative AI features, from retrieval-augmented generation (RAG) pipelines to complex agentic systems. Own the AI Application Lifecycle: Own the end-to-end lifecycle of AI-powered applications, including system design, development, deployment (CI/CD), monitoring, and optimization in production
From C$136K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Auth0 is growing rapidly and looking for exceptional new team members to help take us to the next level. One team, one score. We never compromise on identity. You should never compromise yours either. We want you to bring your whole self to Auth0. If you're passionate, practice radical transparency to build trust and respect, and thrive when you're collaborating, experimenting and learning – this may be your ideal work environment. We are looking for team members that want to help us build upon what we have accomplished so far and make it better every day. N+1 > N. About the Team Here at Auth0 we’re focused on securing the world’s identities so innovators can innovate. We’re currently hiring a senior Full Stack Software Engineer to join our Acquisitions and Activation Team . This team owns the crucial first impressions of the Auth0 ecosystem. We apply rigorous, data-driven experimentation and A/B testing to optimize the entire early customer journey—from the moment a developer signs up, to the precise second they hit their "ah-ha" moment. We bridge the gap between deep technical infrastructure and user psychology, building primarily on a JavaScript ecosystem (Node.js and React) . If you are a full-stack engineer who loves blending robust software architecture with rapid, metrics-driven experimentation, this is the team for you. What You Will Do Optimize the Acquisition Funnel: Own, build, and continuously scale the entire signup experience to reduce fric
From C$1.5K/yr
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Software Platform team accelerates developer velocity and increases system reliability by building the foundational platforms and tools that power Robinhood engineering. Within this group, the Kubernetes Compute team focuses on building and operating a highly available, scalable Kubernetes-powered container platform. We ensure that our infrastructure seamlessly supports reliable application deployments, integrates core platform capabilities, and enables multi-region scalability. We are expanding our core container systems to support our next phase of technical growth! As a Senior Software Develope r, you will focus heavily on building, operating, and expanding our container provisioning platforms. You will be responsible for designing resilient container infrastructure and contributing to our technical migration to Amazon EKS to improve platform reliability. In this position, you will collaborate with engineering teams across Robinhood to deliver reliable platform integrations for core capabilities like networking and security. Your work will directly help our infrastructure scale efficiently while maintaining a high standard of safety and system uptime. This role is bas
From C$136K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Platform Network Engineering Team Auth0 by Okta is an easy-to-implement authentication and authorization platform designed by developers for developers. We make access to applications safe, secure, and seamless for over 100 million daily logins worldwide. Our modern approach to identity enables this Tier 0 global service to deliver convenience, privacy, and security so customers can focus on innovation. The Senior Software Engineer Opportunity You will be part of the Platform Network engineering team responsible for all connectivity of Auth0. You will play a key engineering role as we evolve our network architecture to meet the demands of enormous growth and support the hundreds of millions of users who rely on us to provide uninterrupted access. You will get to work with engineers throughout the engineering organization. What you’ll be doing Implement internal and edge networking infrastructure and design solutions that work at global scale and with multi-cloud and multi-region constraints. Carry cross-team initiatives from end to end: code reviews, design reviews, operational robustness, security hygiene, etc. Design and develop new services, tools, and automation to expose network functionality to other Okta engineering and operations teams. Research and implement solutions addressing cross-cutting concerns such as routing, failover, and scaling. Participate in the team’s on-call rotation. What you’ll bring to the role Have 3+ years of
About the Role: As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams to build world-class machine learning inference platforms. These platforms power essential services like personalized recommendations, search, and content understanding across Tubi. A core responsibility of this team is developing and maintaining low-latency ML model serving systems that support Deep Learning, LLM, and Search models. This involves building self-service infrastructure and critical components such as the inference engine, feature store, vector store, and experimentation engine. You will improve the way we deploy and operate our services and even contribute to open-source projects. This role grants the architectural freedom to explore new frameworks, lead critical cross-functional projects, and transform the capabilities of our ML and Product teams. Responsibilities: Design and build scalable, high throughput, and low latency distributed systems using Scala Build reusable components and services that serve various ML applications like Personalization, Search, Ads and Exploration Partner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art. Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessary Mentor other engineers on the team on system design, effective incident management, interviewing, leveraging LLMs for work, etc. Collaborate with ML, Product, and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi. Your Background: Experience designing and building scalable, distributed systems in any modern backend language (e.g., Scala, Java, Python, Go, C++); experience with Scala or JVM b
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta's Core Platform team provides the foundational infrastructure that powers all engineering at Vanta. We're expanding upmarket to support enterprise customers, which requires strategic investment in platform systems that ensure security, reliability, and developer productivity at scale. As we expand upmarket to support enterprise and regulated customers, we’re investing heavily in platform capabilities that scale securely while reducing cognitive load for product teams. As the Engineering Manager, Core Platform at Vanta, you'll own the foundational infrastructure that every engineer builds on, ensuring it scales with company growth while remaining fast, simple, and reliable. This team’s ownership spans shared services infrastructure, observability and monitoring, datastore management, and async work systems. Our Engineering Managers develop and grow high-performing teams that deliver significant value to our customers and enable our business to scale. This role sits at the intersection of technical architecture and team development, with real authority to set direction and grow a world-class platform team. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll do as an Engineering Manager at Vanta: Lead and grow high-performing platform engineering teams that deliver reliable, scalable infrastructure and operational excellence for Vanta’s products and customers Set technical direction and drive multi-quarter platform initiatives spanning infrastructure reliability, security, scalability, and developer experience across shared systems and services Partner closely with product engineerin
About the Role: The Machine Learning team at Tubi drives the innovation behind personalized user experiences. With the largest inventory in the industry and hundreds of millions of viewers, we tackle problems in the space of recommendations, search, content understanding, and ads optimization that shape the future of streaming. We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction and execution for scaling our machine learning capabilities while ensuring our distributed systems and infrastructure can support innovation at massive scale. You will combine technical depth with leadership excellence to guide teams that deliver both foundational ML systems and high-performance distributed services. This is a hybrid role for our Toronto office. What You'll Do: Lead and manage high-performing teams across ML engineering and ML infrastructure, fostering a culture of innovation, collaboration, and growth. Define and execute the strategic roadmap for ML systems, including recommendation, personalization, and ads optimization. Oversee the design, development, and deployment of scalable ML pipelines: data ingestion, feature engineering, model training, evaluation, and serving. Architect distributed systems to support ML workloads at scale, ensuring reliability, observability, and operational excellence. Partner closely with Product, Engineering, and Content teams to align on business goals and deliver impactful ML-driven experiences. Support best practices in experimentation, evaluation, and ML system monitoring. Ensure cost efficiency, scalability, and performance in ML infrastructure investments. Your Background: 10+ years of industry experience spanning machine learning engineering and distributed systems. 3+ years of leadership and management experience, with a proven ability to build and lead strong t
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Core Infrastructure team's mission is to build and evolve the foundational platform that every Robinhood engineering team builds on — owning the systems, primitives, and developer-facing abstractions that power 24/7 trading, crypto, and global expansion. We treat infrastructure as a product: reliable, fast to provision, and invisible to the teams above it. As a Senior Staff Software Developer on Core Infrastructure, you will own the architectural evolution of three deeply interconnected domains: service mesh and connectivity, compute platform, and infrastructure provisioning. Your decisions will directly shape engineering velocity, operational reliability, and Robinhood's ability to expand to new regions and markets. This is not an operations role — it's a once-in-a-platform-lifecycle opportunity to redesign the foundation before complexity becomes permanent! This role is based in our Toronto, ON office, with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-p
From C$1.5K/yr
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Software Platform team accelerates developer velocity and increases system reliability by building the foundational platforms and tools that power Robinhood engineering. Within this group, the Kubernetes Compute team focuses on building and operating a highly available, scalable Kubernetes-powered container platform. We ensure that our infrastructure seamlessly supports reliable application deployments, integrates core platform capabilities, and enables multi-region scalability. We are expanding our core container systems to support our next phase of technical growth! As a Software Developer, you will focus on building, maintaining, and scaling our container provisioning platforms. Working alongside senior engineers, you will write code to improve our infrastructure capabilities and actively participate in our technical transition to Amazon EKS. In this role, you will collaborate with teams across the organization to ensure robust platform integrations for everyday application needs like security and networking. Your efforts will directly improve system visibility, automation, and reliability across the platform. This role is based in our Toronto office(s), with in-perso
About Forma.ai: Forma.ai is a Series B startup that's revolutionizing how sales compensation is designed, managed and optimized. We handle billions in annual managed commissions for market leaders like Edmentum, Stryker, and Autodesk. Our growth has been fuelled by our passion for fundamentally changing and shaping how companies use sales intelligence to drive business strategy. We’re welcoming equally driven individuals who are excited about creating something big! The Opportunity As a Staff Security Engineer, you will be a hands-on technical leader strengthening security across Forma's application, cloud infrastructure, development lifecycle, internal systems, and incident-response practices. Security today is shared across Engineering and DevOps. You'll work closely with both teams and have real room to shape how Forma approaches security as we grow. Depending on your interests and the needs of the business, the role could develop into a deeper individual-contributor position or help build a dedicated security team. You'll work directly with Engineering, DevOps, IT, Product, Legal, and Privacy to identify risks, design practical controls, automate security processes, and help teams ship secure and reliable software. What you'll do Cloud and infrastructure security Design and implement security controls across Forma's AWS environments, with a focus on IAM, least-privilege access, service identities, and account boundaries. Embed security requirements into Terraform and other Infrastructure as Code, and improve secrets, certificate, encryption-key, and credential management. Build automated checks for insecure configurations, excessive permissions, exposed resources, and configuration drift across Kubernetes, containers, serverless workloads, networking, and data services. Application, data, and AI security Run threat modelling and security architecture reviews for new products, services, APIs, data pipelines, and third-party integrations
C$162K – C$420K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. About the role As a Senior Software Engineer, you will architect and scale the core systems that power Sentry's billing infrastructure, ensuring accuracy and reliability at massive scale. You will collaborate on building the next generation of Sentry’s usage tracking pipeline, processing hundreds of billions of events daily with low latency and financial-grade accuracy. You will help design flexible pricing primitives that support everything from per-event usage billing to complex enterprise contracts, enabling product and sales teams to experiment rapidly while maintaining revenue accuracy and reduced time-to-market for new products. You will contribute to technical decisions on data consistency challenges unique to billing—like handling event delays, retroactive pricing changes, and distributed count reconciliation across our infrastructure. You'll love this job if you Want to solve the "easy to explain, hard to build" problems—like ensuring a customer's bill matches their usage perfectly, even when processing hundreds of billions of events daily across distributed
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Evaluation is critical to making progress in scaling intelligence. As models continue to become superhuman in many real-world use cases, we must continue to develop new evaluation techniques that accurately reflect what models are already capable of, as well as set the agenda for what future models should be capable of. In this role, you are responsible for creating these next-generation evaluation methods and infrastructure to measure LLM progress. As a Senior Research Scientist, Model Evaluation, you will: Create ambitious new evaluation benchmarks that push the limits of what our models can accomplish. Work on highly cross-functional teams to translate model feedback into trustworthy, repeatable evaluations. Conduct research to advance the state-of-the-art in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency. Build scalable and reusable tools for digging into model performance. You may be a good fit if: You enjoy rapidly building prototypes that demonstrate the boundaries of what LLMs are capable of, and you have developed res
From C$108K/yr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. To fulfill Lyft’s mission of creating the real-time transportation network of the future, Lyft needs engineers from a scope of disciplines. We need engineers to help us intelligently scale our backend infrastructure, to increase pricing in real time given localized supply and demand data, to predict future localized demand, to create delightful and intuitive UX flows for passengers and drivers, to improve our cloud infrastructure orchestration and monitoring, to increase dispatched driver-passenger pairings based on real-time data and real-world edge cases, to build tools to increase the efficiency of our customer service team, and more. You're an enthusiastic and experienced app developer looking to take your skills to the next level by joining our iOS team. You’re excited about scaling millions of lines of Swift code and hundreds of mobile developers. Our app is used by millions of people, and we take great pride in our work. This means excellent development practices, careful code architecture, and an organization built around rapid releases. Our codebase is written entirely in Swift with modern design patterns and coding standards. We rely on 3rd party libraries and contribute back to the community (including the Swift language itself). Responsibilities Help establish roadmap and architecture based on technology and our needs Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and distribute knowledge Share your knowledge by giving brown bags, tech talks, and promoting appropriate tech and engineering best practices Performing thoughtful code reviews for colleagues, and help others by conducting high-quality code reviews Participate in hiring activities: take part in technical interviews, live coding, share detailed feedback to hir
Other cities to consider
More places hiring for this role
Get new infrastructure and mlops engineer jobs in Toronto, Canada by email
Daily job updates · Unsubscribe anytime