Must be based in Vancouver The role We're hiring a dedicated data engineer to own the production data platform that our delivery, product, and engineering teams run on; designing integrated, governed data pipelines and delivering automated reporting, AI-assisted workflows, and predictive signals on top of them. You'll write production code, design systems, own CI/CD, and be accountable for the correctness of data that leaders make decisions on. What you'll do Design and operate our cloud data platform: ingestion, transformation, orchestration and serving. Integrate data from across the business (delivery tooling, CRM, product telemetry, finance, support and customer feedback systems) with shared identifiers, data contracts and lineage. Build automated and continuously refreshed reporting so teams manage by exception rather than chasing status. Connect approved AI agents to governed data with structured outputs, provenance, guardrails and human approval in the loop. Build feature pipelines and the MLOps controls behind predictive use cases: tests, versioning, promotion gates and drift monitoring. Own the engineering standards for data: testing, observability, environment promotion, PII classification and access control. What you'll bring Strong software engineering fundamentals: production-quality code, API and interface design, testing discipline, systems design. Real experience building and operating production data platforms on a cloud warehouse or lakehouse (Snowflake and AWS preferred) with dbt and a modern orchestrator. Practical AI tooling experience: something shipped, not prototyped. LLM-backed classification, extraction or structured-output pipelines; agent and tool-calling workflows; retrieval; evals. You can reaso
Jobiba hiring network
Senior Software Reliability Engineer Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Here's a summary of the role: Do you love building scalable cloud platforms and solving complex engineering problems with modern technologies? As a Senior Software Engineer at Diligent, you'll design and deliver high-performing , serverless applications that power our global SaaS platform. You'll work extensively with TypeScript, Node.js, AWS, and event-driven microservices, owning services from design to deployment and production monitoring. This is an opportunity to influence technical decisions, mentor engineers, and explore how AI can transform software development and engineering productivity. If you're passionate about cloud-native architectures, distributed systems, and building software that scales to millions of users, we'd love to meet you. Here's a breakdown of what you'll do (not all of it, just the important stuff): Design and build scalable backend services and event-driven microservices using TypeScript and AWS. Develop secure APIs and integrations that power reporting, analytics, and dashboard experiences. Build and maintain serverless solutions using AWS services such as Lambda, EventBridge , SQS, and DynamoDB. Drive engineering excellence through testing, observability, automation, and production readiness practices. Contribute to infrastructure-as-code and CI/CD pipelines using AWS CDK and modern DevOps practices. Mentor engineers, participate in architecture discussions, and champion the use of AI tools to improve development efficiency. These are the essentials you'll need to get an interview: 6-8 years of professional software engineering experience. Strong experience with TypeScript, Node.js, and modern backend development patterns. Hands-on experience building cloud-native applications on AWS. Strong understanding of serverless architectures and event-driven microserv
Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Platform Engineering team. In this role, you will support the design and development of shared platforms used across Scale. This includes designing our foundational data platforms and lifecycle, architecting Scale’s core cloud infrastructure and orchestration stack, and redefining how engineers develop, build, test, and deploy software at Scale. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the design, and implementation of our foundational platforms and systems, working closely with stakeholders and internal customers to understand and refine requirements. Collaborating with cross-functional teams to define, design, and deliver new features. Proactively identifying opportunities for, and driving improvements to, current p
The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Senior Software Engineer, you will lead the development of a vertical feature or a horizontal capability to include defining requirements with stakeholders and implementation until it is accepted by the stakeholders. You will: Lead the design and implementation of scalable backend systems and distributed architectures for Federal customers. Manage the full lifecycle of feature development from requirement definition to deployment on classified networks. Direct the orchestration of asynchronous agent fleets to meet mission requirements. Lead customer engagements to translate mission needs into technical requirements. Own the communication with stakeholders to ensure implementation meets defined acceptance criteria. Conduct technical reviews and identify risks within machine learning infrastructure and model serving. Drive the platform roadmap by providing technical specifications for Federal product offerings. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in developing and deploying applications in a cloud-native environment. Understanding of containerization (e.g., Docker) and contai
About Scale At Scale AI, our mission is to accelerate the development of AI applications. For 8 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including: generative AI, defense applications, and autonomous vehicles. With our recent Series F round, we’re accelerating the abundance of frontier data to pave the road to Artificial General Intelligence (AGI), and building upon our prior model evaluation work with enterprise customers and governments, to deepen our capabilities and offerings for both public and private evaluations. About Data Engine Our Generative AI Data Engine powers the world’s most advanced LLMs and generative models through world-class RLHF (Reinforcement Learning with Human Feedback), human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. Our Approach As part of the interview process, you’ll be considered for opportunities across several teams within the GenAI Engineering organization, based on your interests, expertise, and business needs. Potential team placements include Allocation, Growth, Frontier Data, Trust & Safety, Pay, Operator, or Tasking Experience. Together, these teams power Scale’s AI data operations - from building high-impact datasets that push the boundaries of LLM capabilities, to optimizing contributor onboarding and incentives, to safeguarding data integrity through advanced trust, safety, and security measures. They work at the intersection of ML, operations, and analytics to ensure we deliver the highest-quality data at scale. Responsibilities: Design, build, and maintain robust, scalable systems across the full stack, including front-end, back-end, and infrastructure layers Implement high-impact features using modern technologies such as TypeScript, React, Node.js, MongoDB, Elasticsearch, and Temporal Collaborate closely with internal operators (your use
Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Identity Engineering team. In this role, you will help support the design and development of core software systems specifically focused on identity, access management, authorization, and authentication. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the design, and implementation of our identity infrastructure to ensure secure authentication and authorization across enterprise systems. Build software for authentication mechanisms such as Single Sign-On (SSO), Multi-Factor Authentication (MFA), and federated identity solutions (SAML, OAuth, OpenID Connect). Build software for authorization mechanisms such as Relation-based access control (ReBAC), Attribute-based access control (ABAC), Role-based access cont
OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leader in full stack automation fabric solutions for mission-critical business processes. With the first SaaS-based composable automation platform specifically built for ERP, we believe in the transformative power of automation. Our unparalleled solutions empower you to orchestrate, manage and monitor your workflows across any application, service or server — in the cloud or on premises — with confidence and control. Redwood’s global team of automation experts and customer success engineers provide solutions and world-class support designed to give you the freedom and time to imagine and define your future. Get out of the weeds and see the forest, with Redwood Software. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT We are looking for a Senior Software Engineer to join our global engineering organization and help lead the design, development, and enhancement of our solutions and products. As part of our engineering team, you’ll build high quality, scalable and secure software that powers critical file exchange capabilities for over 5000 global enterprise customers. This role is ideal for an experienced backend engineer who thrives in a collaborative environment, takes ownership of complex systems, and enjoys working across the full stack—from APIs to UI to cloud integration. Design and develop robust, high-performance, highly secure file transfer products and solutions in Modern C++ (version 11 and later) and Java Build and maintain RESTful integrations using JSON and XML Contribute to and enhance web-based UIs using JavaScript, JQuery, and related frameworks. Implement and optimize support for file transfer protocols (SFTP, FTPS, HTTP/S, FTP/S)&
About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems. The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible. Responsibilities and Duties Develop, own, and maintain tools and services to support the software build and release process Deploy and maintain services with Kub
About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems. The Team An exciting opportunity to join a new team within the Software Operations group. The Build Engineering team is a new function within Software Infrastructure, which focuses on the overall process of building and integration of the Machine Learn ing S oftware S tack. You will work closely with the QA and development teams to get an understanding of how our ML SW stack is built, helping to ensure good build practices, and proving that the stack works together and is reproducible in secure, sandboxed environments. Responsibilities and Duties Developing our internal t
About the job Build the debugger that helps developers unlock more from Graphcore AI processors. As a Senior Software Engineer in our Debugger team, you will help define and implement Graphcore’s next-generation debugging capability. Your work will support developers building and optimising workloads on our advanced AI processors. You will adapt and expand debugger functionality, resolving complex issues across software and hardware boundaries. The tools you build will help internal and external users understand behaviour, improve performance and move faster. You will work closely with software, firmware, hardware, partner and customer teams. This role offers rare depth across processor architecture, toolchains and real developer workflows. The team and culture Work happens close to the technology, with engineers expected to investigate deeply, speak up and take ownership. The team uses Agile ways of working to keep progress visible and decisions moving. You will collaborate across software, firmware and hardware teams to identify debug feature opportunities. Decisions are shaped by technical evidence, user needs and the judgement of engineers closest to the problem. What we’re looking for · Experience using debuggers to resolve complex program issues. · Strong low-level programming skills in C, C++ or Rust. · Strong understanding of processor architectures. · Ability to communicate clearly across software, firmware and hardware teams. · A proactive, self-driven approach to improving product quality and functionality. · Familiarity with compiler toolchains, debugging protocols, Python, IDE development or PyTorch. While we have outlined a set of requirements, we value transferable skills and diverse experiences. We also welcome engineers returning to the profession after a career break, including through returnship routes. Benefits · Flexible working: Balance your work and personal life with greater flexibility · Generous leave: Take time to rest, recharge and enjoy
About Clutch Clutch is Canada’s largest online used car retailer, delivering a seamless, hassle-free car-buying experience to drivers everywhere. Customers can browse hundreds of cars from the comfort of their home, get the right one delivered to their door, and enjoy peace of mind with our 10-Day Money-Back Guarantee. Named one of Canada’s Top Growing Companies two years in a row and awarded a spot on LinkedIn’s Top Canadian Startups list, we’re looking to add curious, hard-working, and driven individuals to our growing team. Headquartered in Toronto, Clutch was founded in 2017. Clutch is backed by world-class investors including Canaan, BrandProject, Real Ventures, D1 Capital, and Upper90. To learn more, visit clutch.ca About the role At Clutch, we believe advances in AI are fundamentally changing how great products are built. Traditional handoffs between product, design, and engineering are giving way to smaller, faster teams where builders can take ideas from concept to production in days instead of months. Our engineers don't just write code, they own outcomes. That means working directly with stakeholders to understand the problem, shaping the solution, thinking through design and user experience, and building scalable products end to end, using strong engineering principles and AI-powered development tools. We're looking for people who see a broken process or an unanswered question and feel compelled to fix it, not wait for it to land in their queue. This is a role for people who thrive on ownership of the problem, the solution, and the outcome. You'll have the autonomy to take initiatives from idea to launch, using whatever tools the job requires (code, AI, prototypes, judgment) to get there. If you're energized by ambiguity, moving fast, and building products with real business impact, you'll thrive at Clutch. Technology at Clutch We believe great engineers can learn great technology. Our team builds with technologies like TypeScript, React, Express, Postgr
About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, leading the design and architecture of our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly. You’re excited about this opportunity because you will… Lead the design of infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Own and evolve our open-weights serving stack — real-time GPU endpoints, high-thr
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . Responsibilities: Design and develop control‑plane and system services for a firewall platform, operating above the networking dataplane (OSI Layer 3+). Build and maintain management interfaces and APIs , including configuration frameworks, CLI, secure access (SSH), logging, and diagnostics. Implement security services such as authentication (SAML), guest access, SSL/TLS handling, and certificate lifecycle management . Develop embedded services for content filtering , including URL categorization, reputation, and rating systems, along with SNMP/telemetry support. Lead efforts in licensing systems , vulnerability remediation, and secure architecture in collaboration with cross‑functional teams. Requirements: 8+ years of experience in C/C++ systems development for embedded platforms, security appliances, or network operating systems. Strong hands‑on expertise with configuration frameworks, service APIs, CLI development, logging systems, and secure remote access (SSH) . Solid understanding of L3–L7 networking and security concepts , including SSL/TLS, au
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . About the Role SonicWall is seeking a Senior Software Engineer with deep IPv6 expertise to lead the design, implementation, and hardening of IPv6 capabilities across the Sonicwall firewall firmware stack. As a Principal Engineer, you will be the technical authority for IPv6 feature parity, protocol correctness, and interoperability across SonicWall Next-Generation Firewall (NGFW) platforms. This role is ideal for engineers with hands-on experience delivering IPv6 in high-performance router, switch, or carrier-grade networking products who are looking to apply that expertise in the embedded network security domain. Responsibilities Lead IPv6 Architecture & Feature Delivery — Own end-to-end design and implementation of IPv6 features across the Sonicwall frewall networking stack, including IPv6 routing, interface addressing, DHCPv6, NDP, MLDv2, and IPv6 Policy-Based Routing (PBR). Drive IPv6 Feature Parity — Identify and close gaps between IPv4 and IPv6 feature sets across subsystems — including VPN Tunnel Interfaces, IPsec/IKEv2, SSL-VPN, Syslog, SNMP, and NAT64 — ensuring consistent dual-stack and
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . About the Role SonicWall is seeking a Software Dev Senior Engineer (Dataplane) with deep expertise in High Availability, Fast-Path Packet Processing, and Distributed Systems to architect and optimize the dataplane HA subsystem across SonicWall NGFW hardware and virtual firewalls — driving sub-millisecond session state synchronization, zero-packet-loss failover, and line-rate forwarding for IPsec, SSL-VPN, and TCP/UDP flows at 10G–100G+. Responsibilities Dataplane HA Architecture — Own fast-path HA engine design covering Active/Standby mirroring, Active/Active flow redistribution, and zero-drop VMAC/VIP failover. Stateful Session Mirroring — Build low-overhead dataplane sync for TCP/UDP 5-tuple tables, IPsec SA fast-path, SSL-VPN contexts, and L7 states across HA peers at line rate. Lock-Free Session Tables — Design NUMA-aware, lock-free (RCU/CAS/seqlock) session tables for multi-core pipelines sustaining tens of millions of concurrent sessions at sub-microsecond latency. High-Performance Sync Fabric — Engineer zero-copy ring buffer / DPDK mempool-based sync with prioritized queuing for IPsec/SSL-VPN state
Get new senior software reliability engineer jobs by email
Daily job updates · Unsubscribe anytime