Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Observability team's mission is to build and own Robinhood's full-stack observability platform — the foundation that keeps every product, service, and customer experience running reliably at scale. We design and operate the systems that give engineers deep visibility into how Robinhood's infrastructure behaves, ensuring that when something goes wrong, the right people know immediately and can act fast. Our work spans metrics, logs, distributed tracing, and alerting pipelines, and we partner closely with the Robinhood Command Center to ensure our observability systems meet or exceed 99.9% uptime. We believe in monitoring our own monitors — the observability infrastructure is a product, not just a tool. As a Staff Software Engineer on the Observability team, you will be the technical lead shaping the roadmap for how Robinhood observes itself at scale. You will own the observability control plane end-to-end, making architectural decisions that directly impact the reliability and operational health of Robinhood's entire product surface. You'll lead a team of six engineers, drive the strategy for cost-efficient telemetry ingestion, build in-house solutions where off-the-shelf
Jobs in Canada
Staff Platform Engineer Observability in Canada
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current staff platform engineer observability jobs across Canada. Filter by work mode, employment type, experience, department, date posted and distance.
From $250K/yr
About Scale Scale’s mission is to develop reliable AI systems for the world’s most important decisions. As the leading AI data foundry, we provide the high-quality data and full-stack technologies that power the world’s most advanced models — fueling breakthroughs in generative AI, defense, and autonomous vehicles. We partner with leading enterprises and governments to bring AI into production that performs when it matters most, combining rigorous evaluation with full-stack deployment so our customers can build AI they can trust. About the Team Applied Intelligence Systems (AIS) is part of the Scale Generative AI Platform (SGP), focused on pushing the frontier of what agentic applications can do across diverse enterprise and government use cases. We build the infrastructure and tooling that power agentic AI in production, paired with applied ML research, design, and evaluation to ensure these systems perform reliably at the scale our customers demand. AIS spans multiple workstreams — agent evaluation and oversight, orchestration and tool-use infrastructure, model and systems optimization, and applied research on new agent capabilities — and this role is not scoped to any single one of them. We’re growing fast, with increasing traction across both commercial and public sector customers, and we’re just getting started — this team will define what dependable, production-grade agentic AI looks like. About the Role As a Staff Machine Learning Research Engineer, you will operate across the full breadth of AIS’s technical needs — wherever the hardest ML problem in agentic AI happens to be that quarter. This could mean training and fine-tuning models, designing evaluation and observability systems, building improvement loops from production data, prototyping novel agent architectures, or designing internal systems and tooling that boost productivity across teams. You’re not tied to one team’s roadmap; you’re expected to move to where the technical leverage is highest, and t
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the Role + Team The Security Engineering team builds the identity and access control plane governing how employees, services, and agentic workloads access Robinhood’s critical infrastructure. We are a platform-building engineering team focused on creating secure, frictionless, and automated self-service access systems at scale. As a Staff Software Engineer on this team, you will shape Robinhood’s foundational identity architecture, define technical strategy, and build Tier-0 access infrastructure used across the entire company! From solving complex authentication and authorization challenges to setting secure guardrails for emerging AI and agentic workflows, your technical leadership will directly impact every product team at Robinhood! What You'll Do Access Control Plane Architecture: Define and build the core technical architecture for managing employee, service, non-human, and agentic access across company-wide resources Tier-0 Platform Engineering: Design, build, and operate highly available, resilient, and observable backend identity infrastructure using Go, Python, Rust, or comparable systems languages. AuthN & AuthZ Solutions: Implement modern identity protocols and access frameworks (OAuth 2.0, OpenID Connect, RBAC, ABAC, and policy-based access controls). AI & Agentic Access Governance: Develop secure access patterns, identity frameworks, governance, and observability guardrails for AI agents, LLM applications, and Model Context Protocol (MCP) systems. Self-Service Developer Experience: Create automated, developer-friendly platforms that allow teams to request, approve, provision, and audit access without ma
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Lyft Business organization builds the platform that lets organizations move the people they depend on employees, clients, and customers. From the foundational mechanics of a company funding and provisioning a ride to the tailored workflows that get an employee to a client meeting or home after a late shift, we sweat the small stuff to make Lyft the transportation solution businesses build their programs on. As a Staff Software Engineer on the Lyft Business Team, you will act as a critical technical leader, taking holistic ownership of complex systems that span the rider experience, the administrator experience, and the enterprise integrations behind them — defining strategic roadmaps, driving cross-functional alignment, and driving engineering excellence to improve how organizations move their people. Responsibilities: Shape the long-term architecture for systems, taking accountability for both short-term functionality and long-term health Translate high-level business goals into actionable engineering projects. Own the technical roadmap from conception to delivery, managing cross-team dependencies and mitigating risks Champion improvements in system, observability, performance, and tech debt reduction, extending your influence beyond your immediate team Establish best practices for deployment, alerting, and on-call health. Take holistic ownership of the platform's stability, tracking down issues root causes and building preventive safeguards for your immediate scope and beyond. Drive alignment across product, design, and operations. Proactively resolve bottlenecks and make decisive trade-offs to protect system architecture from competing organizational priorities Mentor and level up the engineers around you. Delegate stretch opportunities, lead cross-team reviews, and foster a culture of s
From C$160K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Streaming Foundations team builds services and operates data pipeline infrastructure to support event streaming, messaging, and analytics use cases. We are looking for a Software Engineer who is passionate about distributed systems, platform engineering, and solving data-intensive problems at scale. In this high-impact role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Auth0 to scale for years to come. What you’ll be doing Help set the technical direction for the team and influence the engineering roadmap for the Platform’s streaming capabilities Design and lead the implementation of our most complex and critical systems for data-intensive use cases. Research and champion new technologies and architectural patterns to solve strategic challenges and scale the platform. Lead and influence cross-functional initiatives, ensuring technical alignment and successful execution across multiple teams. Improve the operational posture of our systems by designing for observability, reliability, and scalability, and by mentoring others in operational best practices. Coach and mentor senior engineers and act as a technical leader across the engineering organization. Collaborate with different stakeholders like product teams whenever needed. What you’ll bring to our teams 7+ years of software development experience in a fast-paced, agile environment Experience working with Golang or Java is preferred H
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The internal infrastructure team is responsible for building world-class infrastructure and tools used to train, evaluate and serve Cohere's foundational models. By joining our team, you will work in close collaboration with AI researchers to support their AI workload needs on the cutting edge, with a strong focus on stability, scalability, and observability. You will be responsible for building and operating superclusters across multiple clouds. Your work will directly accelerate the development of industry-leading AI models that power Cohere's platform North. Please Note: All of our infrastructure roles require participating in a 24x7 on-call rotation, where you are compensated for your on-call schedule. As a Staff Software Engineer, you will: Build and scale ML-optimized HPC infrastructure : Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads. Optimize for AI/ML training : Collaborate with cloud providers to fine-tune infrastructure for cost efficiency, reliability, and performance , leveraging technologies like R
From C$90K/yr
Platform administration isn't the part of the product anyone screenshots for a demo. Nobody's writing a blog post about your identity provisioning flow. But it's the thing every other team at Diligent quietly depends on more than any other dev team in the company and when it breaks, everyone notices immediately. If that kind of quiet, high-stakes ownership sounds appealing rather than thankless, keep reading. This role is for someone who wants real skin in the game: you build it, you ship it, you support it. We're a high-initiative team that improves things we see first and asks permission later, building secure, event-driven microservices in TypeScript on AWS, and treating infrastructure as code the same way we treat application code: with rigor, not as an afterthought. Here's a breakdown of what you'll do (not all of it, just the important stuff) Design and build secure, scalable full-stack services using AWS serverless tech (Lambda, SQS, API Gateway) — with real attention to event-driven patterns and observability, not just "does it work on my machine." Own your services in production. That means building good observability, keeping an eye on alerts, and responding to them before they become bigger problems. Build infrastructure as code with AWS CDK and push for CI/CD that ships safely and often — not "big bang" releases you have to pray over. Design RESTful APIs other teams will actually want to consume: clear contracts, sane versioning, no surprises. Write tests — unit, integration, end-to-end — as part of how you build, not a chore you do after. Show up to architecture discussions with opinions and documentation, not just vibes. Use AI tools to move faster on coding, debugging, testing, and research — but you're still the one who validates the output. These are the essentials you'll need to get an interview 2-3 years of professional software engineering experience in an agile, full-stack-focused environment. Solid full-stack fundamentals: request lifecycles, d
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are seeking an experienced Staff Software Engineer to join Okta's Universal Directory Platform team within the Product Platform Pillar. The team serves as the intelligent core of the enterprise security fabric, maintaining the source of truth for all identity assets and their associated relationships. Opportunity This position will be involved into development, design, and maintenance of our highly performant, reliable, and scalable platform, which is critical for managing user lifecycles, groups, and memberships. The successful candidate will possess experience in building and deploying scalable, reliable software and infrastructure within a cloud environment. What you’ll be doing Understand our identity management group codebase and development process: Jira, Technical Designs, Code Review, Testing, and Deployment. Develop and implement frameworks and toolings for our Universal Directory Service platform. Design and implement high-performance distributed scalable and fault-tolerant software components. Quickly deliver high-quality bug fixes and handle customer-reported issues. Conduct quality code reviews and automated testings. Partner with our Product Development, QA, and Site Reliability Engineering teams for scoping the development and deployment work. What you’ll bring to the role The ideal candidate is someone who is experienced building software systems to manage and deploy reliable and performant infrastructure and prod
About the Role We are a small team of AI builders in Paytm Labs. As a Staff AI Platform Engineer, you will work across inference and agentic systems. You will contribute to Paytm's AI inference platform (Pi), serving internal teams and enterprise customers - running our own coding and domain-specific models (voice, vision, risk, fintech workflows) as well as third-party models. You will also architect and build the platform that enables autonomous AI agents to operate safely and reliably in production - the runtime, orchestration, and developer tooling for agents to reason, plan, use tools, and execute complex multi-step workflows, automating both software development and business processes. You will work at the intersection of LLMs, distributed systems, and production fintech infrastructure, helping define how inference and agentic AI are built and deployed across payments, risk, fraud, collections, support, and developer experience.
From C$1.4M/yr
About the Role: We're hiring Senior and Staff Data Platform Engineers to join the Data Infrastructure teams in Toronto. Together these teams own the infrastructure that processes billions of events per day: Spark-on-Kubernetes, Flink and Kinesis pipelines, a multi-petabyte Delta Lake, a large-scale MemoryDB feature store, Databricks multi-environment operations, and the catalog and lifecycle systems that govern it. The team is small and senior. Each engineer owns major platform components: you design it, build it, and support it in production. This is a hybrid-role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: Spark-on-Kubernetes — EKS-based compute platform for Spark workloads: cluster configuration, Pod Identity IAM, job environment setup, Kustomize overlays, and shadow canary validation Event ingestion — Rust services and Flink jobs processing billions of events per day over Kinesis; throughput, reliability, on-call response, and AI-assisted operational tooling to reduce toil Platform infrastructure — Terraform modules for environment provisioning, cross-account AWS IAM, ARC runner infrastructure, and CI/CD for data platform changes Feature store and ML compute — Flink-based real-time feature pipelines feeding a large-scale MemoryDB cluster; GPU capacity governance and Databricks multi-environment operations for ML training workloads Workflow orchestration and CDC — Airflow-based DAG deployment, change data capture pipeline operations, and data quality monitoring Your Background: 3+ years building and operating production data platform infrastructure at the cluster or platform level, across Spark, Flink, Kinesis, Kubernetes, or equivalent Deep experience in at least one of: Spark-on-K8s cluster operations, Rust-based data or systems engineering, Kubernetes platform engineering and IaC, or data catalog and governance tooling Production AWS experience or equivalent: EKS, S3, Kinesis, and mu
$165K – $247K/yr
Amplitude is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Winter 2026 Report, Amplitude is the best-in-class solution for product, data, and marketing teams. Learn more at amplitude.com . As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do. Amplitude’s Commitment to Diversity Equity & Inclusion (DEI): Amplitude believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive. About the Role Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Senior Platform Engineer, you'll own medium-to-high-complexity platform projects end-to-end and help shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll partner with Staff engineers and product teams to make Kubernetes effortless across the engineering org, building self-service automation and scalable AWS infrastructure that lets product teams ship faster, safer, and with less cognitive load. If you're excited about building the systems that other engineers will rely on every day, this role is for you. Key Resp
From C$100K/yr
This position is based in Vancouver, BC , within Diligent’s Technical Center of Excellence. We are currently hiring candidates who are based in or able to work from Vancouver . Software Engineer — Platform AI Service Levels: Software Engineer II Senior Software Engineer Staff Software Engineer Location: Vancouver Position Overview As a Software Engineer on Diligent's Platform AI team, you'll help design, build, and operate the core services that power AI-driven capabilities across Diligent's global product suite. You'll build secure, scalable, serverless services on AWS that translate AI research and models into commercial-quality, production-ready solutions — enabling customers to derive insights from their governance data. You'll work closely with AI researchers, product managers, and other engineering teams, owning your services end-to-end: architecture, implementation, deployment, and monitoring. The team operates with a strong AI-augmented engineering culture — using AI tools to accelerate coding, testing, debugging, and delivery — while applying sound judgment about when and how to apply them. Key Responsibilities Design and implement secure, scalable, fault-tolerant, high-performing solutions using AWS serverless technology — event-driven, highly observable, and built with infrastructure as code. Collaborate with AI researchers/engineers to translate AI and LLM capabilities into robust, production-grade services, and help other teams integrate them. Build and maintain the pipelines needed to deploy, monitor, and manage AI services at scale — observable, resilient, and cost-effective. Use AI-powered development tools (code assistants, test generation, architecture exploration) responsibly to accelerate delivery and improve quality, always validating outputs. Participate in architecture discussions and design reviews, and contribute to product design by understanding customer problems — especially where AI can offer a breakthrough solution. Work in
From C$160K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Overview At Okta, we are building a world where anyone can safely use any technology, and within Okta Platform R&D, we're specifically focused on protecting the future of work. We are a globally-minded company dedicated to delighting our customers across the world. We are looking for an experienced full stack Staff Software Engineer (with a focus on UI) who is passionate about building large-scale, mission-critical infrastructure in a fast-paced agile environment. The ideal candidate will be excited about joining a platform team that is driving several cross functional initiatives within the company. You will be responsible for defining the forward-looking vision for Okta’s UI platform infrastructure. You will work with a globally distributed engineering team to build the frameworks and systems that empower all feature teams to effectively deliver world-ready software. You will play a crucial role in developing and maintaining shared libraries and tools used across the entire organization, ensuring accessibility, scalability, and adherence to Okta’s coding standards. If you are passionate about technical leadership and have a deep expertise in transforming complex products for a global audience, then we want to talk to you. In this role, you’ll get to… Define, own, and drive the architectural vision and technical roadmap for a major UI platform subsystem (e.g., internationalization and localization infrastructure), securing buy-in across dependen
About Forma.ai: Forma.ai is a Series B startup that's revolutionizing how sales compensation is designed, managed and optimized. We handle billions in annual managed commissions for market leaders like Edmentum, Stryker, and Autodesk. Our growth has been fuelled by our passion for fundamentally changing and shaping how companies use sales intelligence to drive business strategy. We’re welcoming equally driven individuals who are excited about creating something big! About the Team Engineers on this team construct our rules-based calculating engine for processing sales commissions. This might sound simple if you have never been exposed to sales comp plans, it is not! We are low on meetings, high on accountability. Most of the team are in EST time zone but we have a few located in PST and Central as well. We are far from maintenance / progressive evolution in many areas, there is a lot of room to make a big impact in the overall design. What you’ll be doing Reporting to the Manager of Data Platform, you will play a critical role in the evolution of our Spark based data platform. You'll lead development efforts for our complex, data-rich platform features while being an example to the team of code quality and thoughtful software design. You will be working on the most challenging code at Forma. As a Staff Engineer, you are expected to operate with a high degree of ownership and trust. This includes proactively identifying architectural risks, surfacing edge cases or constraints others may not see, and advocating for improvements that strengthen the long-term integrity of the system. We value engineers who bring forward thoughtful perspectives - even when they challenge assumptions - and who help the team see around corners. You will: Design and evolve backend services that power product workflows. Architect data models representing hierarchical & graph structures, relationships, and large-scale enterprise datasets. Build
From C$160K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Staff Software Reliability Engineer - Data Platform About the Team The Data Platform team is responsible for the foundational data services, systems, and data products for Okta that benefit our users. Today, the Data Platform team solves challenges and enables: Streaming analytics Interactive end-user reporting Data and ML platform for Okta to scale Telemetry of our products and data Our elite team is fast, creative and flexible. We encourage ownership. We expect great things from our engineers and reward them with stimulating new projects, new technologies and the chance to have significant equity in a company. Okta is about to change the cloud computing landscape forever. About the Position This is an opportunity for experienced Software Reliability Engineers to join our fast growing Data Platform organization that is passionate about scaling high volume, low-latency, distributed data-platform services & data products. In this role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Okta to scale for years to come. As a member of the Data Platform team, you will be responsible for designing, building, and deploying the systems that power our data analytics and ML. Our analytics infrastructure stack sits on top of many modern technologies, including Kinesis, Flink, ElasticSearch, and Snowflake. We are looking for experienced Software Engineers who can help desi
Other cities to consider
More places hiring for this role
Get new staff platform engineer observability jobs in Canada by email
Daily job updates · Unsubscribe anytime