Jobs in Canada

Infrastructure Engineer in Toronto

167 active opportunities · Updated October 2026

Explore current infrastructure engineer jobs in Toronto. Filter by work mode, employment type, experience, department, date posted and distance.

L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$108K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As an Observability team member, you are responsible for the operation and maintenance of our logging and metrics infrastructure. You ensure all teams at Lyft are aware of the operational health of their products by monitoring system availability and take a holistic view of our platform performance. You build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you find opportunities to improve our systems in order to push our platform forward. You provide our partners with the support they need to help them build robust large scale distributed systems. We count on the reliability of our infrastructure to empower Lyft teams to provide our customers rich experiences that are highly available with rock solid performance to ensure our transportation platform continues to connect people and places. As we grow our team, we are seeking experienced Infrastructure Engineer to ensure that as our Infrastructure continues to scale, our platform continues to provide an essential and dependable service that transports millions of people every day. Specifically we are searching for someone who brings fresh perspectives, enjoys collaborating with cross-functional teams in order to continually improve our products and services for our customers. Responsibilities: Maintain, improve, and develop tooling and systems that enhance the reliability, scalability, and efficiency of our platform. Assist engineering teams in defining service-level objectives (SLOs) and provide the necessary toolin

PythonAWSKubernetesAI
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$136K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As a Infrastructure Engineer at Lyft, you will run our Production Infrastructure by monitoring system availability and take a holistic view of our platform health. You will build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you will seek opportunities to optimize our systems in order to push our platform forward, anticipating our customers' needs in order to continually improve the platform. You will provide Lyft partner teams with operational support to help them build robust large scale distributed systems. About the Team Data Pipelines is at the heart of all critical data flowing through Lyft supporting hundreds of services that impact millions of drivers and passengers every day. Our team’s mission is to empower Lyft engineers to self-serve in building and maintaining data pipelines as needed to support products that deliver the world’s best transportation experience. We leverage a variety of technologies to store, stream and manage data making it available to our internal customers. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and

PythonAWSDockerKubernetes
F
📍 Toronto, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Forma.ai: Forma.ai is a Series B startup that's revolutionizing how sales compensation is designed, managed and optimized. We handle billions in annual managed commissions for market leaders like Edmentum, Stryker, and Autodesk. Our growth has been fuelled by our passion for fundamentally changing and shaping how companies use sales intelligence to drive business strategy. We’re welcoming equally driven individuals who are excited about creating something big! The Opportunity As a Staff Security Engineer, you will be a hands-on technical leader strengthening security across Forma's application, cloud infrastructure, development lifecycle, internal systems, and incident-response practices. Security today is shared across Engineering and DevOps. You'll work closely with both teams and have real room to shape how Forma approaches security as we grow. Depending on your interests and the needs of the business, the role could develop into a deeper individual-contributor position or help build a dedicated security team. You'll work directly with Engineering, DevOps, IT, Product, Legal, and Privacy to identify risks, design practical controls, automate security processes, and help teams ship secure and reliable software. What you'll do Cloud and infrastructure security Design and implement security controls across Forma's AWS environments, with a focus on IAM, least-privilege access, service identities, and account boundaries. Embed security requirements into Terraform and other Infrastructure as Code, and improve secrets, certificate, encryption-key, and credential management. Build automated checks for insecure configurations, excessive permissions, exposed resources, and configuration drift across Kubernetes, containers, serverless workloads, networking, and data services. Application, data, and AI security Run threat modelling and security architecture reviews for new products, services, APIs, data pipelines, and third-party integrations

PythonAWSKubernetesCI/CD
S
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -100%

C$162K – C$420K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Events Analytics Platform (EAP) team is responsible for the infrastructure that powers all of Sentry's time-series data and searching capabilities across billions of events with sub-second latency. We started this initiative by building Snuba, the primary storage and query service for Sentry's event data powered by ClickHouse, and we are now focused on unlocking deeper visibility and reporting across the terabytes of event data our users generate. As a Senior Software Engineer, you will lead efforts to push the boundaries of data visibility at Sentry. You will do this by expanding the capabilities of our search infrastructure, building new capabilities on top of our state-of-the-art storage layer and increasing the performance and integrity of Sentry’s core data services. You will also help shape Infrastructure's technical direction at Sentry and collaborate with Product and other Engineering teams to turn that vision into a reality. If you want to solve the hard problems that come with scaling event data into the petabyte range, this could be the job for you. In this role you will: Expand EAP's ability to deliver data at world-class speed and reliability. Architect and automate services and systems to scale reliably under growing demand. Make architectural trade-offs that balance product requirements with engineering constraints. Maintain and grow the team's code quality initiatives by regularly reviewing code and contributing to design decisions. Lead design and discussions around deliverables the team is working towards. Improve the maintainability and developer experience of the codebases EAP owns. Exa

PythonSQLPostgreSQLRedis
C
📍 Toronto, Ontario, Canada· Full-time
✓ Quality checkedCompany trend -91.5%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Security Clearance: Active Secret+ clearance strongly preferred; candidates eligible and willing to obtain clearance will also be considered. More information about Canadian Security Clearance is available here . As an Infrastructure Security Engineer, your key responsibilities include: Deploy, and manage infrastructure for Protected B classified environments, ensuring compliance with ITSG-33 and Canadian government standards Design and implement security controls for cloud (AWS, GCP, Azure) and hybrid/multi-cloud deployments Evaluate, implement, and manage security tools and technologies for training cluster and inference infrastructure hardening Implement security best practices including IAM, encryption, logging, and monitoring Participate in security incident response activities, including detection, analysis, containment, and remediation Conduct regular vulnerability assessments and penetration testing of infrastructure components Maintain comprehensive security documentation, procedures, and configurations for classified environments Maintain active Secret+ security clearance and adhere to all Canadian government security

AWSAzureGCPKubernetes
T-
📍 Toronto, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Role: The Machine Learning team at Tubi drives the innovation behind personalized user experiences. With the largest inventory in the industry and hundreds of millions of viewers, we tackle problems in the space of recommendations, search, content understanding, and ads optimization that shape the future of streaming. We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction and execution for scaling our machine learning capabilities while ensuring our distributed systems and infrastructure can support innovation at massive scale. You will combine technical depth with leadership excellence to guide teams that deliver both foundational ML systems and high-performance distributed services. This is a hybrid role for our Toronto office. What You'll Do: Lead and manage high-performing teams across ML engineering and ML infrastructure, fostering a culture of innovation, collaboration, and growth. Define and execute the strategic roadmap for ML systems, including recommendation, personalization, and ads optimization. Oversee the design, development, and deployment of scalable ML pipelines: data ingestion, feature engineering, model training, evaluation, and serving. Architect distributed systems to support ML workloads at scale, ensuring reliability, observability, and operational excellence. Partner closely with Product, Engineering, and Content teams to align on business goals and deliver impactful ML-driven experiences. Support best practices in experimentation, evaluation, and ML system monitoring. Ensure cost efficiency, scalability, and performance in ML infrastructure investments. Your Background: 10+ years of industry experience spanning machine learning engineering and distributed systems. 3+ years of leadership and management experience, with a proven ability to build and lead strong t

AWSMachine LearningAIGo
V
📍 Toronto, Ontario, Canada· Full-time
✓ Quality checkedCompany trend -88.9%

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta's Core Platform team provides the foundational infrastructure that powers all engineering at Vanta. We're expanding upmarket to support enterprise customers, which requires strategic investment in platform systems that ensure security, reliability, and developer productivity at scale. As we expand upmarket to support enterprise and regulated customers, we’re investing heavily in platform capabilities that scale securely while reducing cognitive load for product teams. As the Engineering Manager, Core Platform at Vanta, you'll own the foundational infrastructure that every engineer builds on, ensuring it scales with company growth while remaining fast, simple, and reliable. This team’s ownership spans shared services infrastructure, observability and monitoring, datastore management, and async work systems. Our Engineering Managers develop and grow high-performing teams that deliver significant value to our customers and enable our business to scale. This role sits at the intersection of technical architecture and team development, with real authority to set direction and grow a world-class platform team. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll do as an Engineering Manager at Vanta: Lead and grow high-performing platform engineering teams that deliver reliable, scalable infrastructure and operational excellence for Vanta’s products and customers Set technical direction and drive multi-quarter platform initiatives spanning infrastructure reliability, security, scalability, and developer experience across shared systems and services Partner closely with product engineerin

MongoDBAWSRestAI
S
📍 Toronto, Canada· Full-time
✓ Quality checkedCompany trend -91.4%

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe processes over $1T in payments volume per year, which is roughly 1% of the world’s GDP. The tremendous amount of data makes Stripe one of the best places to do machine learning. The ML Infra team builds services and tools that power every step in the ML lifecycle, including data exploration, feature generation, experimentation, training, deploying, serving ML models, and building LLM applications. With the phenomenal developments happening in the field of AI, we are positioned to accelerate the adoption of AI/ML across all parts of the company by building highly scalable and reliable foundational infrastructure. What you’ll do You will work closely with machine learning engineers, data scientists, and product engineering teams to enable seamless end-to-end experience in building solutions across data, analytics, and AI/ML platforms. You will build the next generation of ML Infra services and major new capabilities that substantially improve ML development velocity and MLOps maturity across the company. Responsibilities Designing and building scalable, reliable, and secure services for notebooks, ML model training, experimentation, serving, and LLM applications across multiple regions. Creating services and libraries that enable ML engineers at Stripe to seamlessly transition from experimentation to production across Stripe’s systems. Working directly with product teams and ML engineers to improve their day-to-day pr

RestMachine LearningAI
C
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -91.5%

From £215K/yr

Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the role. We’re building the next generation of agentic AI infrastructure at Cohere. This team sits at the intersection of ML systems, distributed infrastructure, and developer experience, creating the platform that powers autonomous AI agents at scale. You’ll work on hard, forward-looking problems with few established patterns, including secure code execution, agent state management, model routing, identity and authentication, and resource management for long-running agent workflows. This role is a strong fit for someone who combines systems depth with ML intuition. You should be comfortable building reliable infrastructure, thinking through distributed systems tradeoffs, and understanding how emerging agentic capabilities shape platform design. What you’ll work on. Secure execution environments for agent-generated code Identity, authentication, and trust boundaries for agents Model routing and orchestration across different model types and environments Rate limiting, quotas, and resource management for agent workflows State management, memory, and filesystem abstractions for agents. In this role you will: Turn emerging M

KubernetesGitRestAI
C
📍 Toronto, Ontario, Canada· Full-time
✓ Quality checkedCompany trend -91.5%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. As a Site Reliability Engineer you will: Build self-service systems that automate managing, deploying and operating services. This includes our custom Kubernetes operators that support language model deployments. Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems. Take steps required to ensure we hit defined SLOs, including pa

AWSAzureGCPKubernetes
C
📍 Toronto, Ontario, Canada· Full-time
✓ Quality checkedCompany trend -91.5%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About North: North is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications. Why This Role? This role offers a unique opportunity to shape how enterprises harness the power of AI in real-world applications. As a bridge between our core North product and our clients’ engineering teams, you’ll be at the forefront of solving complex problems and securely integrating AI into critical sectors such as finance, healthcare, and telecommunications. Our esteemed clients include industry leaders like RBC, Dell, and LG CNS. We are seeking engineers who deeply care about customers and want to work at the cutting edge of Agentic AI. In this role, you will: Lead end-to-end deployment of North in private cloud and on-premises environments, including planning, configuration, testing, and rollout. Partner with enterprise IT teams t

AWSAzureGCPKubernetes
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$108K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. We are building and maintaining a highly scalable asynchronous platform that empowers our organization to handle critical business cases. As a software engineering team, our mission is to create robust and innovative solutions that drive the success of our business and deliver unparalleled value to our customers. We adopt Infrastructure as Code practice to automate the provisioning and configuration of our resources, which helps reduce manual configuration and improve consistency. Our team culture is built on collaboration, open communication, and a supportive environment where each member's ideas are valued and contributions are recognized. We believe in the importance of fostering a positive workplace culture that inspires innovation and creativity. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and production readiness reviews, platform management and capacity planning ceremonies with cross-functional teams Document Infrastructure operations process and insights, identify repeatable actions and ruthlessly automate repetitive tasks Participate in our teams on-call rotations, respond to incidents and support other teams mitigate customer impacting events Experience: 5+ years experience working on teams responsible for software development, automation and systems engineering Experience building large-scale infrastructure, distributed systems or networks. Knowledge with SQS,

PythonAWSAzureGCP
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -72.4%

From C$172K/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. As the Engineering Manager for the Lakehouse Foundation team, you will lead a group of engineers responsible for the foundational data layer that all of Lyft's data systems and emerging AI workloads are built on. The team owns catalog and metadata management, table formats and storage, and the access patterns and gateways through which other engineering teams interact with Lyft's data. As Lyft converges on a unified lakehouse architecture, this team builds and operates the single source of truth that powers analytics, machine learning, experimentation, and every business decision made from data. You will play a key role in shaping the team's technical direction, partnering with peer Data Platform teams on a multi-year platform evolution, and developing engineers who operate with autonomy on systems of significant scale and complexity. Lyft's Infrastructure teams build the foundational systems that the rest of engineering depends on to move fast, ship reliably, and scale efficiently. These are high-leverage roles where the work you and your team do has a multiplicative effect across the company. We're looking for experienced leaders who can balance the discipline of operating critical infrastructure with the curiosity to keep evolving how Lyft builds. Engineering at Lyft is a place where managers and engineers operate with high ownership and strong technical judgment. Our engineers expect their managers to be honest, available, and focused on the work that matters: developing their teams, removing obstacles, and giving people the support they need to do their best work. We build teams that are inclusive, technically rigorous, and have a strong sense of ownership for what they build. Responsibilities: Lead a team responsible for Lyft's foundational data layer, including catalog and metadata management,

RestMachine LearningAIGo
O
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -63.6%

From C$110K/yr

Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Team The Auth0 Platform Tools team owns the incident management tooling, Slack-based tooling, StatusPage, and local development environments that Auth0 engineers rely on every day. That includes incident.io and the services we have built around it, Statuspage, custom Slack bot applications that automate our incident response and engineering operations workflows, the customer-facing web application behind status.auth0.com, Vivaldi, and Tilt - the tools engineers use to run Auth0 locally. We are seeking an engineer to help build new features across all of these tools. Our stack is primarily TypeScript and Node.js, with a React and Next.js front end, backed by Postgres and Redis, and deployed on Kubernetes on AWS. A significant portion of our incident and engineering operations automation is built on Tines, a no-code automation platform. Prior no-code experience is welcome, but we expect you to learn Tines here and become effective with it. Current initiatives include extending our incident tooling to meet FedRAMP requirements, taking full ownership of the status page, and improving how we communicate incident status to customers. There is real room to improve along the way, from test coverage to resilience to inherited technical debt. We build for two audiences: Auth0 engineers, who depend on our tooling every day, and Auth0's customers, who rely on the status page during incidents. We are looking for an engineer who cares about both and enjoys working wi

TypeScriptPythonReactNode.js
O
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -63.6%

From C$136K/yr

Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Datastores Engineer, Platform Infrastructure The Auth0 platform secures more than 100 million logins each day for customers all around the world - and we're growing fast! The Platform Infrastructure team enables Auth0 engineers to move faster by giving them tools to easily deploy and manage their services on AWS and Azure. This is a role with a huge impact. You will get to work with engineers throughout the organization and what you build will be a foundational piece of the infrastructure that allows Auth0 to scale for years to come. We are looking for Engineer who are passionate about distributed systems, availability, and delivering customer value to join our Platform Infrastructure Datastores team. Because we build and support the overall Auth0 platform, the ideal candidate is someone who is passionate about infrastructure, operations, databases and not intimidated by cross-organization coordination and collaboration. You will: Develop our large, distributed and highly-available infrastructure. Implement platform tools that allow feature teams to deploy and manage the datastores for their services. Research new technologies to accelerate new environment creation. Carry cross team initiatives from end to end: code reviews, design reviews, operational robustness, security hygiene, etc. Participate in the team's on-call rotation. You might be a good fit if you: Have 5-8 years of software development experience. Are proficient in or have a desire

SQLPostgreSQLMongoDBRedis
🔔

Get new infrastructure engineer jobs in Toronto, Canada by email

Daily job updates · Unsubscribe anytime