About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is our evaluation platform — the unified evals backbone that lets teams measure, trace, and trust the quality of LLM and agent systems across the company, powering trace/score ingestion, LLM-as-judge workflows, agent simulations, and LLM observability for the tens of millions of daily requests flowing through our LLM Gateway. We also own core platform surfaces including the Agent Gateway, open-weights model serving and batch inference, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, with a primary focus on our evals and LLM observability platform: the systems that let teams evaluate, trace, and continuously improve the quality of LLM and agent products. You’ll work across evaluation frameworks and SDKs, OpenTelemetry-based trace/score ingestion, LLM-as-judge and offline/online eval pipelines, agent simulations, data pipelines, backend services, and observability. This role is ideal for an engineer who enjoys building reliable measurement and quality primitives in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and evaluation methodologies are evolving quickly. You’re excited about this opportunity because you will… Build the infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Work on our unified evals platform — evaluation SDKs, OpenTelemetry trace/score ingestion, LLM-as-judge, offline and online eval pipelines, and agent simulations — alongside the LLM Gatew
Jobs in Canada
Infrastructure Engineer in Canada
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current infrastructure engineer jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, leading the design and architecture of our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly. You’re excited about this opportunity because you will… Lead the design of infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Own and evolve our open-weights serving stack — real-time GPU endpoints, high-thr
From $187K/yr
Airtable is the no-code app platform that empowers people closest to the work to accelerate their most critical business processes. More than 500,000 organizations, including 80% of the Fortune 100, rely on Airtable to transform how work gets done. Airtable’s infrastructure is evolving to meet the needs of our fast growing engineering org. We are looking for infrastructure engineers to join our team to help improve critical product infrastructure, with a focus on building systems that have a great developer experience and will scale as we grow. We currently have openings on: Asynchronous Serving: The Asynchronous Serving team is scaling critical systems used by Airtable’s most essential and up-and-coming product features, especially AI features. Upcoming projects include refactoring our background task queue to track its tasks in DynamoDB, adding quality of service to the job queue, and revamping a streaming service to handle 10x scale while being more resilient. Compute: The compute pod builds and manages our Kubernetes-based platform that supports every service at Airtable, including all new AI services such as vector databases, AI evals store, and document extraction and understanding services. We have a lot of exciting foundational work in our roadmap, such as Overhauling our network stack and service discovery, to simplify service setup and strengthen security Region level disaster recovery, and bringing up compute platform from 0->1 in a new region Building custom Kubernetes operators for reliably managing some of our most critical workloads Developer Platform : The Developer Platform team sits at the intersection of all engineering at Airtable, focusing on building the internal tooling, frameworks, and CI/CD systems that power our product teams. We strive to streamline developer workflows - from build and test cycles to production deployments—and foster a best-in-class developer experience. Join us if you’re passionate about creating high-lever
About the Team DoorDash Labs is an independent team within DoorDash. We explore robotics and automation to transform last mile logistics in the long term. If you have a passion for applying robotics solutions in a service used by millions of people, then we want to talk to you! About the Role We're hiring a Robotics Infrastructure Engineer in our Autonomy Software team. In this role, you'll own, build, and manage the infrastructure that makes aerial autonomy development possible. You'll work on the onboard systems that keep a drone alive (process management, health monitoring, parameterization) and the development environment that makes the team fast (build systems, CI/CD, logging, debugging, regression testing). This is not cloud infrastructure. This is real-time, fault-tolerant, onboard software for vehicles that cannot gracefully restart at 50 meters altitude. You're excited about this opportunity because you will… Play an integral role on a small and focused team Develop and own critical onboard components: process management, health monitoring, configuration management, and message passing Own the build system (C++, Python) and middleware layer (ROS2), including cross-compilation for Jetson targets Design and maintain CI/CD pipelines and regression testing infrastructure Build and manage the parameterization system, including schema definition, validation, migration, and deployment Build robotics logging, plotting, and debugging tools that make the entire team more productive Work closely with the simulation team to support SIL/HIL development workflows Define reliability standards for onboard software: watchdogs, failover, and graceful degradation We're excited about you because… You have prior experience at a robotics company in a similar infrastructure role You have experience with robotics middleware (ROS2, LCM, eCal, Apex.AI) You have experience with build systems and package managers (CMake, Bazel, Nix, Conan) You have experience with NVidia Jetson and Je
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Forward Deployed Infrastructure Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for Palantir platforms, products, and deployments. You'll get to use your creativity to develop novel solutions to evolving challenges and automate processes wherever possible, using whichever tools are best for the job including industry-leading LLM and AI technology! As a Forward Deployed Infrastructure Engineer, every day is different! You will be developing software and providing high-quality support for software systems that are critical to solving our government’s greatest challenges. We strongly believe in engineering teams being responsible for the operations of their services in production. As such, you’ll work closely with forward deployed teams and product teams to participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.
From C$108K/yr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As an Observability team member, you are responsible for the operation and maintenance of our logging and metrics infrastructure. You ensure all teams at Lyft are aware of the operational health of their products by monitoring system availability and take a holistic view of our platform performance. You build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you find opportunities to improve our systems in order to push our platform forward. You provide our partners with the support they need to help them build robust large scale distributed systems. We count on the reliability of our infrastructure to empower Lyft teams to provide our customers rich experiences that are highly available with rock solid performance to ensure our transportation platform continues to connect people and places. As we grow our team, we are seeking experienced Infrastructure Engineer to ensure that as our Infrastructure continues to scale, our platform continues to provide an essential and dependable service that transports millions of people every day. Specifically we are searching for someone who brings fresh perspectives, enjoys collaborating with cross-functional teams in order to continually improve our products and services for our customers. Responsibilities: Maintain, improve, and develop tooling and systems that enhance the reliability, scalability, and efficiency of our platform. Assist engineering teams in defining service-level objectives (SLOs) and provide the necessary toolin
C$113.4K – C$162K/yr
We believe communication belongs to everyone. We exist to democratize phone service. TextNow is evolving the way the world connects and that's because we're made up of people with curious minds who bring an optimistic, yet critical lens into the work we do. We're the largest provider of free phone service in the nation. And we're just getting started. Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between! This role is about impact at scale. You’ll shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded into everyday engineering practice. Using AI is not optional, it’s expected. From design and architecture to implementation, testing, debugging, documentation, and operational analysis, you will actively leverage AI tools to increase velocity, improve code quality, and make better technical decisions. We provide a robust suite of AI-powered development tools and workflows to support you, and we expect you to continuously evolve how you use them to raise the bar for efficiency, clarity, and product excellence across the organization. What You'll Do Ensure System Reliability: Design, build, and maintain scalable, resilient, and highly available systems to support TextNow’s infrastructure and services. Automation & Infrastructure as Code: Develop and maintain automation using Terraform, Ansible, and other tools to enable efficient deployment, scaling, and operations of cloud-based systems (AWS preferred). Incident Response & On-Call Support: Participate in an on-call rotation, troubleshoot issues, and drive incident resolution to minimize downtime and improve syste
About the Team Come help us build and develop tools serving hundreds of engineers internally! We’re looking for a Fullstack Software Engineer to join our Developer Insights team. About the Role Our mission is to improve the developer experience of engineers at DoorDash by building various internal products, including our internal developer portal, Developer Insights. Our success as a platform team depends on the success of the product teams we serve. Because of this, we invest in building a strong community that encourages participation and promotes best practices. You’re excited about this opportunity because you will… Introduce cutting edge technologies to our engineering organization, including tools built on LLMs Build new features for Developer Insights (using Backstage.io) Improve the developer experience for all of our engineers Work and collaborate across team boundaries. Contribute features and bug fixes to upstream open-source projects. Mentor and educate your peers. Lead the team in a technical fashion and assist in roadmap planning and measurement of existing features. Represent the team at large in OKR and engineering all-hands presentations. Context switch from frontend to backend to data depending on the need that arises. We’re excited about you because… You have at least 2 years of experience in web technologies using Typescript with React on the frontend with Java, Kotlin, Python or Go backend experience. You have a product mindset and apply that to how you would build out platform services. You love systems and software, and you're proficient in both. You’re curious and dive deep into different system architectures. You are an organized and excellent written and verbal communicator. You have proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software Compensation The successful candidate’s starti
About the Team The Storage organization builds and operates the online stateful systems and abstractions that DoorDash Engineering depends on: reliable, efficient, secure, and easy to use. Within Storage, the Distributed Caching team owns every caching offering at DoorDash end to end, including ElastiCache (Redis/Valkey), Boulder (our KVRocks-based key-value store for high-QPS feature serving), Entity Cache (read Bill Shen’s engineering blog post, “ High-Performance Proxy Cache for DoorDash Services ”), and the Distributed Lock Service, plus the smart clients (asgard-redis, valkey-go) that sit in front of them. These systems back critical product surfaces across DoorDash, Wolt, and Deliveroo: the team runs roughly 400 ElastiCache clusters serving hundreds of millions of GET requests per second in aggregate, and Boulder, our offline-to-online feature store, serves billions of feature lookups per second at peak. About the Role The team owns provisioning of clusters and the smart clients that sit in front of them, baking in sensible defaults so that other engineering teams get a turnkey caching solution instead of having to run their own. You'll help drive Boulder's evolution to scale further, improve cost efficiency, enhance performance, and support real-time updates; re-platform the Distributed Lock Service onto a strongly consistent backend; and build the self-serve tooling and recommendation engine that let customers describe a workload (QPS, TTL, payload size, latency profile) and get the right backend without talking to a human. You'll go deep on cache invalidation, replication, sharding, compaction, and failover, while shipping the guardrails, automation, and observability that keep this scale operable by a small team. You must be located in San Francisco, Seattle, or the New York Metro Area for this hybrid position. You will report to the Engineering Manager on the Distributed Caching team within the Storage organization. You’re excited about this opportunity b
About the Team The Code Quality team sits within the Developer Platform organization and owns the systems that keep DoorDash's codebase healthy and secure as it scales: static analysis, quality gates, test frameworks, regression infrastructure, and tooling. Our job is to make sure the signals engineers rely on before shipping — test results, coverage, performance feedback etc — are fast and trustworthy. The decisions we make about tooling and standards directly shape how confidently and quickly engineering teams at DoorDash can ship to production. About the Role We're looking for Software Engineers to help build and maintain the systems that validate code quality across DoorDash's engineering org, treating our tooling as a critical product for the engineers who rely on it every day: static analysis and quality gates, test frameworks and regression infrastructure. You’ll design the tooling and automation that will help derive trustworthy quality signals, integrate them into the development lifecycle, and make it easy for engineers to execute reliable, repeatable workflows. You will collaborate across the engineering org, partnering directly with the teams who use what you build to understand the accuracy, reliability and performance of their functionality. You will report into the Engineering Manager on our Code Quality team in our Developer Platform organization. You must be located in either San Francisco, CA, Sunnyvale, CA, Los Angeles, CA, Seattle, WA, or New York, NY. You're excited about this opportunity because you will… Build and maintain quality tooling — static analysis, quality gates, coverage reporting, test frameworks, regression infrastructure — and integrate it directly into our developer workflows and CI/CD pipelines Define and derive quality signals - flakiness, pass rate, coverage, performance, scale readiness etc - Build tooling that improves everyday engineering workflows, including local development, CI/CD, debugging, and rollou
About the Team Our mission is to provide a world-class development experience that makes DoorDash's web engineers among the most productive in the industry. We achieve this by creating the tools that enable all teams at the company to ship features quickly and reliably. Because their success is our success, we are deeply invested in building a strong, collaborative web community that champions best practices and welcomes participation. About the Role As a Software Engineer on the Developer Experience team, you will build the foundational pieces for all DoorDash, Wolt and Deliveroo Web applications. These include monorepos, build & CI systems and agent-first development tooling. You will work closely with engineers and other internal stakeholders to deliver large and impactful initiatives. Additionally, you will be a culture carrier for our Web engineers through mentorship, education, and engagement of your peers. You will report into the Engineering Manager of our Web Developer Experience team in our Developer Platform organization. You must be located in either San Francisco, CA, Sunnyvale, CA, Los Angeles, CA, Seattle, WA, or New York, NY. You’re excited about this opportunity because you will… Shape the future of Web Development. You will have a direct and meaningful impact on the daily workflows of every web engineer at the company, enhancing their productivity and overall developer experience. Build from the ground up. You will architect and implement foundational libraries, cutting-edge build systems, and innovative development tools that serve as the bedrock for all of our web applications. Solve complex, high-impact challenges. You will tackle some of the most significant technical hurdles in web engineering, and the solutions you deliver will be leveraged by hundreds of employees across numerous product teams. Act as a force multiplier. Your work will directly empower product teams to build, test, and release new features to our customers fa
About the Team The Storage teams build and operate online stateful systems and abstractions that are reliable, efficient, secure and easy to use for DoorDash Engineering. The teams are responsible for understanding Product Engineering’s evolving needs and developing platform and infrastructure capabilities to serve them. The team currently supports CockroachDB, Cassandra, Kafka and Redis as well as data abstraction services to reduce the complexity of interacting with storage systems for Product Engineers. About the Role The Storage team is building and operating a high-performance, scalable, and reliable data abstraction layer that optimizes both efficiency and reliability. Our goal is to create a platform that manages itself and fades into the background—empowering engineers to focus on delivering product experiences our customers love. This role is available across two teams within Storage, each solving unique and high-impact challenges: One team is building the orchestration layer for DoorDash’s storage platform—unifying lifecycle management, operations, and self-serve APIs for databases and streaming systems, turning complex, stateful infrastructure into reliable, developer-friendly services used across the company. One team builds and operates the distributed data platform powering DoorDash's largest stateful workloads -- including Cassandra, which backs critical product surfaces across DoorDash, Wolt, and Roo. You'll design high-throughput data abstractions, smart clients, and platform services that make distributed data reliable and easy to work with at multi-petabyte, multi-million-QPS scale, with opportunities to go deep on distributed systems internals and contribute to the open-source Cassandra ecosystem. If you're passionate about distributed systems, developer experience, and building foundational infrastructure at scale, we'd love to hear from you. You must be located in San Francisco, Sunnyvale, Seattle, or the New York Metro Area for this hybrid pos
From C$136K/yr
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As a Infrastructure Engineer at Lyft, you will run our Production Infrastructure by monitoring system availability and take a holistic view of our platform health. You will build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you will seek opportunities to optimize our systems in order to push our platform forward, anticipating our customers' needs in order to continually improve the platform. You will provide Lyft partner teams with operational support to help them build robust large scale distributed systems. About the Team Data Pipelines is at the heart of all critical data flowing through Lyft supporting hundreds of services that impact millions of drivers and passengers every day. Our team’s mission is to empower Lyft engineers to self-serve in building and maintaining data pipelines as needed to support products that deliver the world’s best transportation experience. We leverage a variety of technologies to store, stream and manage data making it available to our internal customers. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and
About the Team Our team exists to empower DoorDash Mobile engineers. We are the champions of three core pillars: Quality, Velocity, and Efficiency. Ultimately, our goal is to build the supportive infrastructure that allows our fellow engineers to build, ship, and operate apps at massive scale—with confidence and ease. We believe that a great developer experience leads to a great customer experience. About the Role We are looking for an Engineering Manager to lead our Mobile Developer Experience team (also known internally as the Mobile Foundations team). In this role, you’ll not only set the technical vision but also nurture the culture required to build world-class mobile applications. Your team will concentrate on one vertical—ZeroKit, which enables rapid mobile prototyping via agents—and two horizontal infrastructure pillars: the iOS and Android monorepos. You will partner with senior engineers and stakeholders across the company to design systems that make our platform faster, more reliable, and more efficient. You will be a partner to your customers—your fellow engineers—working side-by-side to understand their hurdles and solve their immediate challenges. You’ll also look to the future, anticipating needs so we can deliver solutions before they become blockers. Crucially, this role will support our growing international engineering footprint and unified technology stack, meaning you and your team will work across our DoorDash, Wolt, and Deliveroo brands. This is a unique opportunity to lead a team through a mix of exciting greenfield initiatives and the refinement of established, successful tools. You must be located in the following locations for this opportunity: San Francisco, CA; Sunnyvale, CA; Seattle, WA; Los Angeles, CA; New York, New York. You’re excited about this opportunity because you will… Develop and maintain foundational components to enable DoorDash Engineers to excel at mobile engineering. Lead the development and strategy for ZeroKit to enabl
We take play seriously. We’re looking for curious adventurers ready to find their party, fueled by imagination and drive to build what’s never been built before. At Hasbro and Wizards of the Coast, you’ll collaborate with passionate teams to reimagine our iconic brands and create experiences that spark joy, connection, and community through the magic of play. This is your chance to shape legendary play that lasts a lifetime. At Wizards of the Coast, we connect people around the world through play and imagination. From our genre-defining games like Magic: The Gathering® and Dungeons & Dragons® to our growing multiverse, we continue to innovate and build new ways to foster friendship and connection. That’s where you come in! The cloud platform underpins how services are built, deployed, secured, and operated across the organization. As a Principal Cloud Infrastructure Engineer, you will define and drive the governance model and automation strategy for cloud infrastructure. This includes setting policy-as-code standards, automating compliance and cost controls, and ensuring infrastructure is provisioned and operated through consistent, auditable, self-service pathways rather than tailored or manual processes. You will operate at both a strategic and hands-on level and will set direction for how cloud resources are governed and automated while also building the tooling and guardrails that make that direction real. Success in this role means engineering teams can self-service the build and operation of infrastructure with confidence that it is secure, cost-aware, and consistent by default, without slowing delivery down. What you'll do Cloud Governance Define and enforce policy-as-code standards (tagging, naming, encryption, network segmentation, access boundaries) across cloud accounts/subscriptions Establish guardrails using tools such as OPA, Sentinel, AWS Config/Service Control Policies to report on and prevent drift from
Other cities to consider
More places hiring for this role
Get new infrastructure engineer jobs in Canada by email
Daily job updates · Unsubscribe anytime