Jobs in Canada

Infrastructure Team Manager in Canada

470 active opportunities · Updated October 2026

Explore current infrastructure team manager jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.

S
📍 Canada· Full-time
✓ Quality checkedCompany trend -91.4%

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team The Product Design team is composed of several teams that work together to help define, create, and deliver all user-facing aspects of the Stripe brand and product. Product Designers are strategic partners to Engineering and Product Management within a dedicated product area. In partnership with the product team, they help define the user-facing experiences our users encounter on Stripe products, then translate that thinking into a complete experience that can be tested, shipped, and refined. They are responsible for creating well-functioning and beautiful products and experiences that users love and are eager to recommend to others. The Stripe Dashboard makes the simple things easy and the complex things possible for millions of businesses around the globe. We deliver products and experiences for all Stripe users including Startups, SMBs, Enterprises, and major platforms. Since our user base is so broad, it is essential that the experiences we build are configurable, customizable, and extensible. We take pride in crafting user-friendly and user-focused experiences. What you’ll do As part of the Money Management Design team, you will lead the design efforts for our global-by-default money management suite. We are building the rails for the next generation of commerce, where stablecoins move at the speed of the internet. Through Stripe Treasury, we enable businesses to hold dollar-denominated balances, settle cross-border transactions i

L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -73.4%

From C$1.4M/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Safety and Customer Care (SCC) team at Lyft manages over 1.7 million monthly human and AI interactions and serves as Lyft's primary direct touchpoint with riders and drivers. We handle critical infrastructure that powers both human associates and AI agents to make riders and drivers feel safe and comfortable while riding or driving with Lyft, transforming every support interaction into a moment of genuine connection. Agentic AI is at the center of how we scale that mission. We fine-tune and align open-source models, build AI-powered support agents, and develop end-to-end AI agents for safety case management, systems that reason over complex, high-stakes cases and drive them to resolution. SCC brings together ML, data, backend, and product engineers alongside data scientists and operations partners to transform these systems. As a Machine Learning Engineer on the SCC team, you will fine-tune and align models and build AI Agents that power how riders and drivers get help. Your work spans the full loop: post-training open-source models for our domain, composing them into multi-step agents, and building the evaluation that proves they are safe to ship in a customer-facing, safety-critical setting. Post-train and adapt open-source LLMs for SCC use cases using SFT, LoRA, and preference-tuning methods (RLHF, RLAIF, RLVR). Design and build AI-powered support agents and end-to-end agents for safety case management using LangGraph or equivalent agentic frameworks. Own the evaluation data flywheel, offline and online, that defines what "good" looks like and build benchmarks for the team to hill-climb. Turn interaction feedback into training data and learning signals, closing the data flywheel that continuously improves the models. Responsibilities: Conduct literature review and build post-training fra

PythonMachine LearningAIGo
L
📍 Toronto, Canada· Full-time
✓ High-confidence listingCompany trend -73.4%

From C$1.3M/yr

Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Safety and Customer Care (SCC) team at Lyft manages over 1.7 million monthly human and AI interactions and serves as Lyft's primary direct touchpoint with riders and drivers. We handle critical infrastructure that powers both human associates and AI agents to make riders and drivers feel safe and comfortable while riding or driving with Lyft, transforming every support interaction into a moment of genuine connection. As a Data Engineer on the SCC team, you will have ownership over the data modeling and pipelines that power SCC’s Associate and AI Agent Platform . Your efforts will be critical to the reliability of our pipelines, execution of third party data integrations, accurate reporting of agents performance, and efficiency improvements that can save millions of dollars / year. You will work cross-functionally to bridge Lyft's business goals with data engineering. Your efforts will allow access to business and user behavior insights, using huge amounts of Lyft data to fuel several teams such as Analytics, Data Science, Engineering, and many others. Responsibilities: Owner of the core data pipeline, responsible for scaling up data processing flow to meet the rapid data growth at Lyft Evolve data model and data schema based on business and engineering needs Implement systems tracking data quality and consistency Develop tools supporting self-service data pipeline management (ETL) SQL and MapReduce job tuning to improve data processing performance Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and distribute knowledge Collaborate cross-functionally with product, engineering, data science, and marketing teams to understand business problems and align on prioritization and solutions Experience: Bachelor's degree in Compute

PythonSQLAWSRest
T-
📍 Toronto, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Role: The Machine Learning team at Tubi drives the innovation behind personalized user experiences. With the largest inventory in the industry and hundreds of millions of viewers, we tackle problems in the space of recommendations, search, content understanding, and ads optimization that shape the future of streaming. We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction and execution for scaling our machine learning capabilities while ensuring our distributed systems and infrastructure can support innovation at massive scale. You will combine technical depth with leadership excellence to guide teams that deliver both foundational ML systems and high-performance distributed services. This is a hybrid role for our Toronto office. What You'll Do: Lead and manage high-performing teams across ML engineering and ML infrastructure, fostering a culture of innovation, collaboration, and growth. Define and execute the strategic roadmap for ML systems, including recommendation, personalization, and ads optimization. Oversee the design, development, and deployment of scalable ML pipelines: data ingestion, feature engineering, model training, evaluation, and serving. Architect distributed systems to support ML workloads at scale, ensuring reliability, observability, and operational excellence. Partner closely with Product, Engineering, and Content teams to align on business goals and deliver impactful ML-driven experiences. Support best practices in experimentation, evaluation, and ML system monitoring. Ensure cost efficiency, scalability, and performance in ML infrastructure investments. Your Background: 10+ years of industry experience spanning machine learning engineering and distributed systems. 3+ years of leadership and management experience, with a proven ability to build and lead strong t

AWSMachine LearningAIGo
C
📍 Toronto, Ontario, Canada· Full-time
✓ Quality checkedCompany trend -91.5%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Security Clearance: Active Secret+ clearance strongly preferred; candidates eligible and willing to obtain clearance will also be considered. More information about Canadian Security Clearance is available here . As an Infrastructure Security Engineer, your key responsibilities include: Deploy, and manage infrastructure for Protected B classified environments, ensuring compliance with ITSG-33 and Canadian government standards Design and implement security controls for cloud (AWS, GCP, Azure) and hybrid/multi-cloud deployments Evaluate, implement, and manage security tools and technologies for training cluster and inference infrastructure hardening Implement security best practices including IAM, encryption, logging, and monitoring Participate in security incident response activities, including detection, analysis, containment, and remediation Conduct regular vulnerability assessments and penetration testing of infrastructure components Maintain comprehensive security documentation, procedures, and configurations for classified environments Maintain active Secret+ security clearance and adhere to all Canadian government security

AWSAzureGCPKubernetes
G
📍 Canada· Full-time
✓ High-confidence listingCompany trend -100%

From C$107K/yr

Quick readStrong listing-quality and freshness signals

Location Details: Canada, Remote At GoDaddy, the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) , and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Contribute to the development of GoDaddy’s eCommerce and SSO infrastructure and Kubernetes systems on AWS. On a day-to-day basis you will be working on the team who designs, writes, tests and deploys the infrastructure and application management software for GoDaddy’s eCommerce applications. Expect to learn every day. What you'll get to do... Work as a polyglot engineer, writing and maintaining Infrastructure as code with frameworks/ ecosystems such as Java, Unix CLI, and NodeJS Build and operate infrastructure workflows and deployment pipelines using Kubernetes, Argo Workflows, Argo CD, and GitOps practices Design, build, and own services and APIs in Java, running on Kubernetes-based platforms across AWS and distributed systems Develop and support application and infrastructure delivery pipelines, enabling reliable releases of eComm, Auth and Infrastructure services Collaborate closely with other GoDaddy departments to help advance security and technical standards, maintain regulatory compliances while operating eComm & Auth platforms Your experience should include... 5+ years of strong backend software engineering experience in Java Hands-on experience with Kubernetes, including Helm, Kustomize, or equivalent tools to deploy and manage backend services Experience building and operating high-volume, mission-critical production systems on AWS with continuous deployment (CD) practices Strong experience with infrastructure as code, supporting backend applications and services Experience with observability a

JavaNode.jsSQLAWS
PE
📍 Palo Alto, CA· Full-time· Hybrid
✓ Quality checked

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role As a Platform Engineer on Palantir's Identity Platform team, you will design, build, and operate secure-by-design identity infrastructure and tooling. You will make identity governance and access management easier and more secure to implement for Palantirians and customers worldwide. As part of Palantir's best-in-class Information Security organization, you will research, implement, and scale innovative solutions that help Palantir stay ahead of a dynamic threat landscape. You will help the team build the next-generation identity platform that treats agents and workload identities as first-class principals: the access graph that makes entitlements legible, the policy engine that makes authorization enforceable and auditable, and the token-issuance layer that keeps access short-lived, scoped, and least-privilege by default. Your work will shape the arc of these components from design to production. You will join the high-performing Identity Platform team, engineers who are passionate about delivering identity outcomes at scale that reduce risk and friction. The team builds and operates the identity platforms that serve both corporate and production (customer-facing) infrastructure, and the paved-road tooling and secure baselines that teams across Palantir deploy on. Your goal will be to make the secure path the easy path. Your work will directly strengthen the identity substrate beneath Palantir's most critical deployments, from a globally distributed workforce to regulated and air-gapped environments.

A
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $187K/yr

Quick readStrong listing-quality and freshness signals

Airtable is the no-code app platform that empowers people closest to the work to accelerate their most critical business processes. More than 500,000 organizations, including 80% of the Fortune 100, rely on Airtable to transform how work gets done. Airtable’s infrastructure is evolving to meet the needs of our fast growing engineering org. We are looking for infrastructure engineers to join our team to help improve critical product infrastructure, with a focus on building systems that have a great developer experience and will scale as we grow. We currently have openings on: Asynchronous Serving: The Asynchronous Serving team is scaling critical systems used by Airtable’s most essential and up-and-coming product features, especially AI features. Upcoming projects include refactoring our background task queue to track its tasks in DynamoDB, adding quality of service to the job queue, and revamping a streaming service to handle 10x scale while being more resilient. Compute: The compute pod builds and manages our Kubernetes-based platform that supports every service at Airtable, including all new AI services such as vector databases, AI evals store, and document extraction and understanding services. We have a lot of exciting foundational work in our roadmap, such as Overhauling our network stack and service discovery, to simplify service setup and strengthen security Region level disaster recovery, and bringing up compute platform from 0->1 in a new region Building custom Kubernetes operators for reliably managing some of our most critical workloads Developer Platform : The Developer Platform team sits at the intersection of all engineering at Airtable, focusing on building the internal tooling, frameworks, and CI/CD systems that power our product teams. We strive to streamline developer workflows - from build and test cycles to production deployments—and foster a best-in-class developer experience. Join us if you’re passionate about creating high-lever

AWSKubernetesCI/CDRest
SA
📍 San Francisco, Canada· Full-time
✓ High-confidence listing

From $184K/yr

Quick readStrong listing-quality and freshness signals

Scale AI is seeking a highly skilled and motivated Software Engineer, ARC (Architecture, Reliability, & Compute) to join our dynamic Public Sector Engineering team. As a part of this team, you will define how the company ships software, establishing the patterns for deploying into complex government and high-security environments, rather than just running Terraform scripts. You will build and maintain internal CLIs/tools that standardize testing, deployment, environment management and are tools that engineering relies on to prevent downstream breakages. You will execute on automated deployment efforts to pay down tech debt, creating fully functional staging/testing environments, and defining the company's standard for safe deployments. You will: Design and implement secure scalable backend systems for Public Sector customers, leveraging Scale's modern and cloud-native AI infrastructure. Own services or systems and define their long-term health goals, while also improving the health of surrounding components. Re-architect the stack to run in compliant or restrictive environments. This requires designing swappable components (auth, storage, logging) to meet government/security mandates without breaking the product. Collaborate with cross-functional teams to define and execute the vision for backend solutions, ensuring they meet the unique needs of government agencies operating in secure environments. Participate actively in customer engagements, working closely with stakeholders to understand requirements and deliver innovative solutions. Contribute to the platform roadmap and product strategy for Scale AI's Public Sector business, playing a key role in shaping the future direction of our offerings. Must have: At least an active secret clearance and the ability & willingness to up level to TS/SCI with CI Poly. This is a requirement and candidates will not be considered who do not hold at least a secret clearance Ideally you'd have: Full Stack Development: Prof

AWSAzureGCPDocker
R
📍 Bellevue, WA; Menlo Park, CA· Full-time
✓ Quality checkedCompany trend -76.2%

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the Role + Team The Security Engineering team builds the identity and access control plane governing how employees, services, and agentic workloads access Robinhood’s critical infrastructure. We are a platform-building engineering team focused on creating secure, frictionless, and automated self-service access systems at scale. As a Staff Software Engineer on this team, you will shape Robinhood’s foundational identity architecture, define technical strategy, and build Tier-0 access infrastructure used across the entire company! From solving complex authentication and authorization challenges to setting secure guardrails for emerging AI and agentic workflows, your technical leadership will directly impact every product team at Robinhood! What You'll Do Access Control Plane Architecture: Define and build the core technical architecture for managing employee, service, non-human, and agentic access across company-wide resources Tier-0 Platform Engineering: Design, build, and operate highly available, resilient, and observable backend identity infrastructure using Go, Python, Rust, or comparable systems languages. AuthN & AuthZ Solutions: Implement modern identity protocols and access frameworks (OAuth 2.0, OpenID Connect, RBAC, ABAC, and policy-based access controls). AI & Agentic Access Governance: Develop secure access patterns, identity frameworks, governance, and observability guardrails for AI agents, LLM applications, and Model Context Protocol (MCP) systems. Self-Service Developer Experience: Create automated, developer-friendly platforms that allow teams to request, approve, provision, and audit access without ma

PythonVueAWSKubernetes
C
📍 Toronto, Ontario, Canada· Full-time
✓ High-confidence listingCompany trend -91.5%

From £215K/yr

Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the role. We’re building the next generation of agentic AI infrastructure at Cohere. This team sits at the intersection of ML systems, distributed infrastructure, and developer experience, creating the platform that powers autonomous AI agents at scale. You’ll work on hard, forward-looking problems with few established patterns, including secure code execution, agent state management, model routing, identity and authentication, and resource management for long-running agent workflows. This role is a strong fit for someone who combines systems depth with ML intuition. You should be comfortable building reliable infrastructure, thinking through distributed systems tradeoffs, and understanding how emerging agentic capabilities shape platform design. What you’ll work on. Secure execution environments for agent-generated code Identity, authentication, and trust boundaries for agents Model routing and orchestration across different model types and environments Rate limiting, quotas, and resource management for agent workflows State management, memory, and filesystem abstractions for agents. In this role you will: Turn emerging M

KubernetesGitRestAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team The Customer Experience team serves as a foundational operations pillar at DoorDash, dedicated to resolving friction within the last mile. We architect and oversee an expansive global network of support centers, obsessing over the user journey to ensure every interaction is seamless and reliable. Our mission is to scale a best-in-class support infrastructure that advocates for our community and delivers excellence at the lowest level of detail. About the Role DoorDash is looking for a Director of Customer Experience Strategy & Operations to define and scale the operating systems that shape how consumers, Dashers, and our broader marketplace experience DoorDash. This leader will own a broad portfolio across customer experience strategy, support operations, policies and processes, safety-critical contact handling, and cross-functional CX initiatives across business lines including New Verticals, International, Dasher, Marketplace, and DashPass. You will lead a high-performing S&O team, partner deeply with Product, Engineering, Analytics, Fraud, Operations, and senior business leaders, and translate ambiguous customer and operational problems into scalable solutions. We are especially excited about leaders who bring an AI-native operating mindset: people who have used AI, automation, lightweight tooling, or internal apps to solve real problems, improve team productivity, and redesign manual workflows. This role is an opportunity to apply that builder mindset at marketplace scale. You’re excited about this opportunity because you will… Define the vision, metrics, and operating principles for a best-in-class customer experience across consumers, Dashers, and emerging business lines. Lead a high-impact S&O team responsible for turning defining the customer experience using a multitude of inputs Redesign how CX work gets done by combining people, process, product, data, and AI-enabled workflows. Partner with Product, Engineering, Analytics, Fraud

AWSGitRestAI
DU
📍 San Francisco, Canada· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Team The Storage teams build and operate online stateful systems and abstractions that are reliable, efficient, secure and easy to use for DoorDash Engineering. The teams are responsible for understanding Product Engineering’s evolving needs and developing platform and infrastructure capabilities to serve them. The team currently supports CockroachDB, Cassandra, Kafka and Redis as well as data abstraction services to reduce the complexity of interacting with storage systems for Product Engineers. About the Role The Storage team is building and operating a high-performance, scalable, and reliable data abstraction layer that optimizes both efficiency and reliability. Our goal is to create a platform that manages itself and fades into the background—empowering engineers to focus on delivering product experiences our customers love. This role is available across two teams within Storage, each solving unique and high-impact challenges: One team is building the orchestration layer for DoorDash’s storage platform—unifying lifecycle management, operations, and self-serve APIs for databases and streaming systems, turning complex, stateful infrastructure into reliable, developer-friendly services used across the company. One team builds and operates the distributed data platform powering DoorDash's largest stateful workloads -- including Cassandra, which backs critical product surfaces across DoorDash, Wolt, and Roo. You'll design high-throughput data abstractions, smart clients, and platform services that make distributed data reliable and easy to work with at multi-petabyte, multi-million-QPS scale, with opportunities to go deep on distributed systems internals and contribute to the open-source Cassandra ecosystem. If you're passionate about distributed systems, developer experience, and building foundational infrastructure at scale, we'd love to hear from you. You must be located in San Francisco, Sunnyvale, Seattle, or the New York Metro Area for this hybrid pos

JavaSQLRedisAWS
CC
📍 Vancouver, British Columbia, Canada· Full-time
✓ High-confidence listing

From C$11K/mo

Quick readStrong listing-quality and freshness signals

Are you a self-starter who loves tackling challenging problems in a team-oriented, fast-paced environment? Do you love the idea of living in Vancouver? Our quantitative equity team is expanding, and we are looking for talented individuals whose skills and interests are aligned with the team’s mission to deliver superior investment performance to our clients through a disciplined, systematic investment approach. Jobs within our team lie at the exciting intersection of research, data science, finance, and technology. A culture of innovation, career growth and success through collaboration has enabled us to deliver superior investment performance to our clients for over two decades. Today, our clients entrust us with over $78+ billion USD in financial assets. What You Will Do This is an exciting opportunity for individuals who are passionate about learning the quantitative equity investment management business and excited to tackle a broad range of investment, mathematical, and technology challenges. You will start with a comprehensive training program, learning about various elements of quantitative equity investment management. You will then continue to a specialized group within our team based on a combination of your fit and preference. Your specialized role may involve building, scaling, managing, monitoring, and/or trading our quantitative model and portfolios. You will be supported with coaching and mentorship from senior members of the team. We are creating the conditions for you to grow your career steadily over time. Groups within our team where you can contribute and gain valuable experience include: Investment Process Management Our Investment Process Management (IPM) group plays a central role in building highly scalable, efficient, and robust process management systems. The processes and infrastructure we maintain are critical in the team’s ability to d

RestAIGoRust
AC
📍 Toronto, Canada· Full-time
✓ High-confidence listing

From C$75K/yr

Quick readStrong listing-quality and freshness signals

Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. About the Team Appian Customer Success is obsessed with delivering exceptional customer outcomes and high-velocity business impact. Our Public Sector team acts as a trusted strategic partner to government agencies and organizations, bringing their most vital ideas to life. By joining this high-performing team, you will strengthen and evolve your technical consulting skills while accelerating the adoption of our AI-Powered Process Automation platform across critical public infrastructure. As a Senior Consultant, you will step into a pivotal leadership role at the intersection of public sector strategy and modern technical innovation. This is your opportunity to move beyond isolated coding tasks and own the end-to-end delivery of highly visible, complex enterprise systems. You will lead small engineering teams, architect robust integrations, and serve as a trusted technical advisor to client stakeholders - enabling enterprise organizations to achieve true Enterprise-Grade Orchestration . What You’ll Do Lead Project Delivery: Oversee and drive the entire project lifecycle to define, design, and implement custom automation solutions using the Appian platform. Mentor & Develop Talent: Actively lead, coach, and mentor junior consultants through fast-paced software implementations, fostering a culture of technical excellence. Architect Integrations: Build seamless data pipelines and robust APIs to integrate multiple external enterprise systems using REST, SOAP, and JDBC connections. Drive Data Modeling: Design and launch sophisticated relational data

SQLAWSRestAgile
🔔

Get new infrastructure team manager jobs in Canada by email

Daily job updates · Unsubscribe anytime