A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We are a software engineering team with expertise in enabling ML models in production. We deploy AI models to run in variety of environments: air-gapped government networks, forward-deployed defense environments, edge nodes, and enterprises with strict data sovereignty requirements. Our customers rely on us for frontier AI capabilities running on hardware they control, often with constrained GPU resources and limited direct access. Rising to that challenge and meeting those expectations is what Palantir's excels at. We treat models like any other software: continuously tested, continually delivered, packaged for reproducible deployment, and built for long-term maintainability. You will own services end-to-end, and work across the full stack, from inference engines, GPU scheduling to deployment pipelines, observability, and integration with Palantir's platform. The goal is to deliver new models and capabilities quickly and continuously. Join us if you want to solve problems at the intersection of infrastructure and machine learning that directly enable critical customers.
Jobs in Canada
Infrastructure Team Manager in Canada
470 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure team manager jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.
From $184K/yr
Scale AI is seeking a highly skilled and motivated Software Engineer, Frontier AI Infrastructure to join our dynamic Public Sector Engineering team. As a part of this team, you will own the model inference layer - enabling state of the art models, debugging the latest AI tools, managing networking, debugging latency, and tracking pricing/usage metrics for AI models. You will lead technical discussions on the frontlines with cloud vendors and customers to deliver on critical contracts and to debug platform issues. You will also work upstream with Product to understand features before they break, moving us from "infra-only debugging" to proactive integration testing. You will: Design and implement secure scalable backend systems for Public Sector customers, leveraging Scale's modern and cloud-native AI infrastructure. Own services or systems and define their long-term health goals, while also improving the health of surrounding components Re-architect the stack to run in compliant or restrictive environments. This requires designing swappable components (auth, storage, logging) to meet government/security mandates without breaking the product. You will work with Product to build integration tests that catch issues early, shifting the focus from "infra-only debugging" to preventing failures upstream. Participate actively in customer engagements, working closely with stakeholders to understand requirements and deliver innovative solutions. Contribute to the platform roadmap and product strategy for Scale AI's Public Sector business, playing a key role in shaping the future direction of our offerings. Must have: At least an active secret clearance and the ability & willingness to up level to TS/SCI with CI Poly. This is a requirement and candidates will not be considered who do not hold at least a secret clearance Ideally you'd have: Full Stack Development: Proficiency in both front-end and back-end development, including experience with modern web develo
From C$160K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Staff FullStack Engineer, AI Acceleration Team The AI Acceleration Team The AI Acceleration Team builds AI features right where identity and agents meet. Companies are handing real access to AI agents faster than their controls can keep up, and two problems show up first: how do you authorize what an agent is allowed to do on someone's behalf, and how do you keep both people and agents from being phished? We take those on directly. Our North Star: to get AI right, you have to get identity right. We are a new team, and we shape the problems as much as we solve them. We bring a product mindset to the work: we dig into what customers actually need, define the outcome, and build toward it. We move fast, stay close to the real problem, and are honest about what works. The Staff FullStack Engineer Opportunity As a Staff FullStack Engineer on the AI Acceleration Team, you will own the technical build for solving agent authorization and phishing resistance. You see what needs doing and start it. This is 0-to-1 work, meaning you will move on problems before they are fully specified, always keeping customer success as the ultimate bar. You will help shape which problems we take on, make architectural calls that outlive any single project, and set the technical bar for how we work. You will enforce least privilege for what agents can do and make phishing-resistant access the default. We build with Vercel and AWS, and you will likely work heavily in JavaScript/TypeScri
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. As a Site Reliability Engineer you will: Build self-service systems that automate managing, deploying and operating services. This includes our custom Kubernetes operators that support language model deployments. Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems. Take steps required to ensure we hit defined SLOs, including pa
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About North: North is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications. Why This Role? This role offers a unique opportunity to shape how enterprises harness the power of AI in real-world applications. As a bridge between our core North product and our clients’ engineering teams, you’ll be at the forefront of solving complex problems and securely integrating AI into critical sectors such as finance, healthcare, and telecommunications. Our esteemed clients include industry leaders like RBC, Dell, and LG CNS. We are seeking engineers who deeply care about customers and want to work at the cutting edge of Agentic AI. In this role, you will: Lead end-to-end deployment of North in private cloud and on-premises environments, including planning, configuration, testing, and rollout. Partner with enterprise IT teams t
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About North: North is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications. Why This Role? Cohere’s team partners with Canadian public sector organisations to unlock transformative value through secure, ethical deployment of Generative AI (GenAI) solutions. We work collaboratively to address complex societal challenges while maintaining the highest standards of data security and compliance. You will work directly with public sector customers to quickly understand their greatest problems and design and implement solutions using Cohere's stack. This role offers a unique opportunity to shape how enterprises harness the power of AI in real-world applications. As a bridge between our core North product and our clients’ engineering teams, you’ll be at the forefront of solving complex problems and securely integrating A
$140K – $225K/yr
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As the Senior Software Engineer, Tooling and Development Infrastructure, you will play a critical role in shaping the developer productivity tools and automated testing strategy. You’ll collaborate closely with design, development, and quality teams to plan, design, and implement robust automated tools and services that ensure the quality and reliability of our AI software stack. You will be highly hands-on in your work and collaborate closely with stakeholders. This position offers a unique opportunity to influence the development of cutting-edge automation frameworks, foster a culture of quality, and contribute to the long-term success of the organization. What You Might Do Develop and implement automation frameworks and testing strategies that cover the entire software stack, from backend systems to user-facing features. Identify, evaluate, and integrate new tools that streamline development. This includes everything from code quality tools and to Infrastructure-as-Code (IaC) solutions. Lead continuous improvement efforts for our build, release, and test systems, ensuring a robust
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Developer Infrastructure org is the engine behind Robinhood's entire engineering organization — a collection of tightly integrated teams whose collective mission is to make every engineer at Robinhood faster, more reliable, and exponentially more productive. The org spans four major teams: DevX (Developer Experience), TestX (Test Infrastructure), Backend Platform, and Mobile Platform. DevX owns Robinhood's monorepo and Bazel-based build infrastructure — the critical layer between a developer writing code and that code being ready to ship — along with the company's full CI/CD pipeline and remote build execution cluster. TestX owns the infrastructure behind Robinhood's entire test experience: the integration test environments, and personal development environments that serve as miniature simulations of the full Robinhood system, giving engineers a safe, isolated space to test their code end-to-end before it ever touches production. Backend Platform and Mobile Platform own the core language runtimes, libraries, IDEs, and developer toolchains across Python, Go, TypeScript, Swift, and Android. Together, these teams share a single north star: leveraging AI and agentic systems
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold builders and sharp problem-solvers who are wired to deliver great outcomes. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. The DevX team’s mission is to build and operate the core developer infrastructure at Robinhood. Our team owns and scales the systems that thousands of engineers rely on daily, partnering with software developers across the company to make development fast, reliable, and cost-efficient! As a Staff Software Developer, you will act as a technical leader for our build and developer infrastructure, driving the strategy and execution of the systems thousands engineers depend on every day. Your work will span our build systems, CI pipelines, and remote development environments, ensuring engineers can code, test, and build with speed, safety, and reliability at scale. In this role, you will collaborate with teams across Robinhood to eliminate developer friction and raise the bar for engineering productivity. This is a high-visibility leadership opportunity to shape our developer ecosystem and set new standards of engineering efficiency! This role is based in our Toronto, ON office(s), with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. What you’ll do Architect the long-te
About the Team Come help us build and develop tools serving hundreds of engineers internally! We’re looking for a Fullstack Software Engineer to join our Developer Insights team. About the Role Our mission is to improve the developer experience of engineers at DoorDash by building various internal products, including our internal developer portal, Developer Insights. Our success as a platform team depends on the success of the product teams we serve. Because of this, we invest in building a strong community that encourages participation and promotes best practices. You’re excited about this opportunity because you will… Introduce cutting edge technologies to our engineering organization, including tools built on LLMs Build new features for Developer Insights (using Backstage.io) Improve the developer experience for all of our engineers Work and collaborate across team boundaries. Contribute features and bug fixes to upstream open-source projects. Mentor and educate your peers. Lead the team in a technical fashion and assist in roadmap planning and measurement of existing features. Represent the team at large in OKR and engineering all-hands presentations. Context switch from frontend to backend to data depending on the need that arises. We’re excited about you because… You have at least 2 years of experience in web technologies using Typescript with React on the frontend with Java, Kotlin, Python or Go backend experience. You have a product mindset and apply that to how you would build out platform services. You love systems and software, and you're proficient in both. You’re curious and dive deep into different system architectures. You are an organized and excellent written and verbal communicator. You have proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software Compensation The successful candidate’s starti
About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team We are a brand new team, working quickly to accelerate Stripes engineering productivity by effectively deploying LLM agents and tools to automate large swaths of the engineering workflow. We’re a cross-continent team, spanning North America and Europe - filled with high agency engineers with strong product perspectives. Some examples of what we work on: here and here . What you’ll do You will build the next generation of internal AI coding tools and platforms, to massively accelerate Stripe’s engineering productivity safely and with a clear focus on our users’ needs. Who you are We’re looking for someone who meets the minimum requirements to be considered for the role. If you meet these requirements, you are encouraged to apply. The preferred qualifications are a bonus, not a requirement. Minimum requirements We’re looking for someone who has: Strong software engineering skills and ability to write high quality code at high velocity 2-5 years of professional hands-on software development experience, able to write well-factored algorithms and have experience with commonly used data structure and algorithms Hands on experience building tools or platforms used by other engineers or internal users Strong collaboration skills, can work across workstreams within your team and contribute to your peers’ success Customer obsession, ability to articulate and represent customer experience in various forums to drive the right outcome Have t
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Product Education team designs, develops, and launches learning experiences for Lyft customers in the in-app and web Learning Center. We partner with stakeholders across Lyft to make sure customers get accurate, timely information — from in-app safety features to tax season tips to five-star guidance — reaching hundreds of thousands of customers across the US and Canada. Our ~200 tutorials support 25M+ enrollments a year, and we work closely with the Learning Platform product and engineering teams to build the features that support customer learning at scale. We're looking for a Senior Education Content Producer to own visual asset creation and accuracy for the Learning Center. Most of your time will be spent on video editing and motion graphics. You’ll design, create and edit new videos, animations and images, and keep our existing library accurate as the underlying product is constantly evolving. You'll also need solid producer skills: pitching and managing external vendors/agencies, and running the occasional small-scale shoot. You'll work alongside our other Content Producer, who also owns asset creation and updates, and our Learning Experience Designers, who write and structure tutorial content. Together, you'll make sure every customer-facing tutorial looks and feels like Lyft, and stays that way as the product changes. Responsibilities: Visual Asset Creation & Accuracy Using the Adobe Creative Suite and Figma, edit and animate new video, motion graphics, and image assets from concept through final delivery Own the day-to-day health of our visual asset library: auditing, updating, and troubleshooting images, live-action video, UI walkthroughs, and animations across dozens of tutorials Track versioning closely as the underlying product changes, proactively catching and
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are seeking a hands-on Data Center Technician, IT Contractor to support the installation, deployment, relocation, cabling, inventory, maintenance, and decommissioning of servers, network equipment, storage systems, and other IT infrastructure. The successful candidate will follow established procedures, maintain accurate documentation, and coordinate effectively with internal technical teams and data center personnel. This role will be a 6-month contract fully on-site, based out of the downtown Toronto, ON Beanfield data center, with travel to Tenstorrent office locations and other data center facilities as required. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You bring approximately 2 to 5 years of hands-on experience in data center operations, IT infrastructure, server or network hardware, or a similar technical role. You are comfortable working with servers, network equipment, storage systems, racks, rack layouts, and copper and fiber optic cabling. You are detail-oriented, organized, and able to follow documented procedures, cabling standards, installation instructions, and safety requirements. You communicate clearly, work well wi
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Software Platform team accelerates developer velocity and increases system reliability by building the foundational platforms and tools that power Robinhood engineering. Within this group, the Kubernetes Compute team focuses on building and operating a highly available, scalable Kubernetes-powered container platform. We ensure that our infrastructure seamlessly supports reliable application deployments, integrates core platform capabilities, and enables multi-region scalability. We are expanding our core container systems to support our next phase of technical growth! As a Senior Software Develope r, you will focus heavily on building, operating, and expanding our container provisioning platforms. You will be responsible for designing resilient container infrastructure and contributing to our technical migration to Amazon EKS to improve platform reliability. In this position, you will collaborate with engineering teams across Robinhood to deliver reliable platform integrations for core capabilities like networking and security. Your work will directly help our infrastructure scale efficiently while maintaining a high standard of safety and system uptime. This role is bas
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Software Platform team accelerates developer velocity and increases system reliability by building the foundational platforms and tools that power Robinhood engineering. Within this group, the Kubernetes Compute team focuses on building and operating a highly available, scalable Kubernetes-powered container platform. We ensure that our infrastructure seamlessly supports reliable application deployments, integrates core platform capabilities, and enables multi-region scalability. We are expanding our core container systems to support our next phase of technical growth! As a Software Developer, you will focus on building, maintaining, and scaling our container provisioning platforms. Working alongside senior engineers, you will write code to improve our infrastructure capabilities and actively participate in our technical transition to Amazon EKS. In this role, you will collaborate with teams across the organization to ensure robust platform integrations for everyday application needs like security and networking. Your efforts will directly improve system visibility, automation, and reliability across the platform. This role is based in our Toronto office(s), with in-perso
Other cities to consider
More places hiring for this role
Get new infrastructure team manager jobs in Canada by email
Daily job updates · Unsubscribe anytime