Jobiba hiring network

Staff Software Reliability Engineer Data Platform Jobs

3,518 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current staff software reliability engineer data platform jobs. Use filters to narrow by work mode, employment type, experience and date posted.

F
Forma.ai
📍 Toronto• Full-time• From C$190K/yr
15 days ago

About Forma.ai: Forma.ai is a Series B startup that's revolutionizing how sales compensation is designed, managed and optimized. We handle billions in annual managed commissions for market leaders like Edmentum, Stryker, and Autodesk. Our growth has been fuelled by our passion for fundamentally changing and shaping how companies use sales intelligence to drive business strategy. We’re welcoming equally driven individuals who are excited about creating something big! About the Team We build enterprise software that helps organizations optimize sales performance and improve go-to-market agility. Our engineering organization includes multiple product application teams responsible for delivering core customer-facing capabilities. We are seeking Senior Backend Engineers to join our application teams. You’ll work alongside staff, senior, and early-career engineers to design, build, and scale backend systems that power enterprise-grade product workflows. This is an opportunity to work on complex product and data problems while contributing meaningfully to technical decisions, system quality, and team delivery. We are low on meetings and high on accountability. Most of the team is in the EST time zone, with a few located in AST, PST, and Central as well. What you’ll be doing You will play an important role in the continued evolution of our application stack. You will design and build backend capabilities for complex product workflows, contribute to system design discussions, and help ensure our systems remain maintainable, reliable, and scalable as we grow. As a Senior Backend Engineer, you are expected to operate with strong ownership and sound technical judgment. This includes identifying risks in the work you own, surfacing edge cases, asking thoughtful questions, and proposing improvements that strengthen the quality and reliability of the system. You will: Design and build backend services that power complex product workflows. Contribute to d

javascripttypescriptpython
View job →
F
15 days ago

About Forma.ai: Forma.ai is a Series B startup that's revolutionizing how sales compensation is designed, managed and optimized. We handle billions in annual managed commissions for market leaders like Edmentum, Stryker, and Autodesk. Our growth has been fuelled by our passion for fundamentally changing and shaping how companies use sales intelligence to drive business strategy. We’re welcoming equally driven individuals who are excited about creating something big! Senior Staff Backend Engineer About the Team We build enterprise software that helps organizations optimize sales performance, enabling go-to-market agility. Our engineering organization includes multiple product application teams responsible for delivering core customer-facing capabilities. We are seeking a Senior Staff Backend Engineer to join our application teams and help set technical direction across multiple domains within engineering. You'll work alongside staff, senior, and early-career engineers, and partner closely with engineering leadership to define, evolve, and scale the systems that power enterprise-grade product workflows. This is an opportunity to own complex, multi-domain technical problems and shape product direction beyond a single team. We are low on meetings, high on accountability. Most of the teams are in the EST time zone, but we have a few located in AST, PST, and Central as well. What you'll be doing You will play a pivotal role in shaping the technical direction of our application stack across multiple domains. You will lead development efforts for our most complex initiatives, the kind that span two or more teams or product areas, and serve as a technical benchmark for system design, code quality, and long-term maintainability. You'll operate at the intersection of data modelling, business logic, and enterprise-scale reliability, and your work will often set standards that neighboring teams adopt. This remains a hands-on

javascripttypescriptpython
View job →
M
Mongodb
📍 United States• Full-time• From $137K/yr
1mo ago

The Code Gen team is tasked with building AI-powered code transformation tools that transform rigid, legacy applications that suffer from poor scalability and high operating costs into modern, microservices-based architectures that are built on top of MongoDB. Join our team and be at the forefront of innovation and creativity. We are looking for a Staff Engineer with domain expertise and years of experience in modernizing legacy applications that are based on traditional database systems. A significant advantage is profound prior experience in leveraging AI, particularly LLMs and GenAI capabilities, to enable reliable, self-driving automation of the code transformation, iterative build, and test processes. In this role, you will be instrumental in initiating technical strategies and ideas, lead the Code Gen team in designing, building, and optimizing our code transformation workflow and tools. You will work on critical components that ensure the scalability, efficiency, and reliability of our services. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation, build and test. This role will be based remotely in North America. A strong candidate for this position will have Extensive experience (8+ years) in software development and operations, with a proven track record of delivering high performance, correctness, and architectural excellence in fast-paced environments Experience using Relational Databases such as Oracle, MySQL, Microsoft SQL Server or PostgreSQL Experience with tools and methodologies for code analysis, refactoring, and automated testing Experience in designing and implementing complex software systems, collaborating effectively with engineers of all experience levels to achieve high reliability and performance Practical knowledge of integrating GenAI into large-scale, complex systems, including a clear unde

sqlpostgresqlmysql
View job →
M
Mongodb
📍 Alberta• Full-time• From C$159K/yr
1mo ago

The Code Gen team is tasked with building AI-powered code transformation tools that transform rigid, legacy applications that suffer from poor scalability and high operating costs into modern, microservices-based architectures that are built on top of MongoDB. Join our team and be at the forefront of innovation and creativity. We are looking for a Staff Engineer with domain expertise and years of experience in modernizing legacy applications that are based on traditional database systems. A significant advantage is profound prior experience in leveraging AI, particularly LLMs and GenAI capabilities, to enable reliable, self-driving automation of the code transformation, iterative build, and test processes. In this role, you will be instrumental in initiating technical strategies and ideas, lead the Code Gen team in designing, building, and optimizing our code transformation workflow and tools. You will work on critical components that ensure the scalability, efficiency, and reliability of our services. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation, build and test. This role will be based remotely in North America. A strong candidate for this position will have Extensive experience (8+ years) in software development and operations, with a proven track record of delivering high performance, correctness, and architectural excellence in fast-paced environments Experience using Relational Databases such as Oracle, MySQL, Microsoft SQL Server or PostgreSQL Experience with tools and methodologies for code analysis, refactoring, and automated testing Experience in designing and implementing complex software systems, collaborating effectively with engineers of all experience levels to achieve high reliability and performance Practical knowledge of integrating GenAI into large-scale, complex systems, including a clear unde

sqlpostgresqlmysql
View job →
SL
15 days ago

Title: Staff Site Reliability Engineer, Product Area Focus Location: Noida/ Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering team. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedite

pythonjavareact
View job →
SL
15 days ago

Title: Staff Site Reliability Engineer, Product Area Focus Location: Noida / Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering team. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedite

pythonjavareact
View job →
R
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering (SRE) team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust observability solutions, lead incident response, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. You Will: Architect and Implement Observability: Design, build, and lead the implementation of comprehensive monitoring, logging, and tracing solutions. Create dashboards and metrics that provide real-time visibility into system health and performance, enabling proactive issue detection. Define and Drive Reliability Standards: Work with product and engineering teams to define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to monitor and report on these metrics, holding teams accountable and ensuring we maintain high reliability standards while balancing innovation speed. Lead Incident Management and Response: Act as a senior leader during high-impact incidents, guiding the team to rapid resolution

pythongcpdocker
View job →
M
Mongodb
📍 Bengaluru• Full-time
1mo ago

We are seeking a Staff Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Staff Site Reliability Engineer on this new team, you will be responsible for providing technical leadership for the operational foundations that enable deployment at scale of AI applications. You will own the reliability architecture of the platform as it expands across regions and cloud providers, and set the technical direction for how the platform is operated, including capacity planning, multi-cloud expansion, incident response, and SLO discipline. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Bengaluru for our hybrid working model. Position Expectations Own the reliability architecture of the platform across regions and cloud providers Collaborate with the teams building the platform, providing internal support and guidance on operability, capacity, and best practices Set operational standards for the team: on-call quality, incident response, SLO discipline Mentor and technically develop the SRE team Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure Qualifications 10+ years of experience working on software and operating distributed systems, with deep Kubernetes expertise, including designing or evolving multi-cluster platforms Proficiency in Python, Go, or a similar programming language Understand workload isolati

pythonmongodbaws
View job →

You’ll shape the future of a business‑critical platform as the technical lead across both product engineering and cloud infrastructure. You’ll modernize a mature .NET application running on AWS today, while steering its evolution toward a cloud‑native, React/Node.js, AI‑enabled architecture. If you enjoy owning architecture end‑to‑end, from backend and frontend through CI/CD, DevOps, and AWS infrastructure, this role gives you real influence at Staff Engineer level and the opportunity to set engineering standards that others follow. You’ll spend your time leading complex .NET and React features, designing scalable AWS infrastructure with Infrastructure as Code, and building automation that makes releases fast, safe, and repeatable. You’ll work on performance, reliability, and modernization in equal measure—fixing what’s slowing the platform down today and designing what it will look like in the next generation. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture and development of enterprise .NET services and APIs that power a business‑critical platform. Design and operate AWS infrastructure (using AWS CDK in TypeScript) to support secure, scalable, multi‑environment deployments. Build and optimize CI/CD pipelines (AWS CodePipeline, CodeBuild, Windows build agents) to make shipping .NET and React changes fast and reliable. Drive modernization initiatives across the stack, including clean architecture, refactoring legacy components, and reducing technical debt. Design and tune PostgreSQL and MSSQL database solutions for performance, scalability, and reliability. Mentor engineers and influence engineering practices across teams, raising the bar on cloud, DevOps, and software design. These are the essentials you’ll need to get an interview Significant experience (typically 8+ years) delivering and operating scalable enterprise software, owning both application code and cloud infrastructure. Deep hands‑on expertise with C

typescriptreactnode.js
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $252K/yr
15 days ago

Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale our product in a fast-paced environment. The ideal candidate will have a strong understanding of software engineering principles and practices, as well as experience with large-scale distributed systems. You will be responsible for owning large new areas within our product, working across backend, frontend, and interacting with LLMs and ML models. You will solve hard engineering problems in scalability and reliability. You will: Own large new areas within our product Work across backend, frontend, and interacting with LLMs and ML models Deliver experiments at a high velocity and level of quality to engage our customers Work across the entire product lifecycle from conceptualization through production Be able, and willing, to multi-task and learn new technologies quickly Ideally you'd have: 7+ years of full-time engineering experience, post-graduation Experience scaling products at hyper growth startups Experience tinkering with or productizing LLMs, vector databases, and the other latest AI technologies Proficient in Python or Javascript/Typescript, and SQL Experience with Kubernetes Experience with major cloud providers (AWS, Azure, GCP) Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval

javascripttypescriptpython
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. The Staff SRE, Classified Opportunity This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a technical leader within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. Security Clearance: Active U.S. TS/SCI clearance with Full Scope Poly Compliance Expertise: Proven experience navigating Federal and DoD compliance frameworks, specifically FedRAMP and Impact Level 6 (IL6) What You’ll Do Work with various teams to design and implement scalable, and reliable network solutions Maintain a highly available cloud infrastructure edge for the Okta identity platform C

pythonawsdocker
View job →
M
Mongodb
📍 Boston; Miami; New York City; Pittsburgh; Raleigh; United States• Full-time• From $126K/yr
1mo ago

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of our Boston, New York City, Raleigh, Miami, Pittsburgh or remotely in the United States while physically based in an Eastern or Central time zone location. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure g

pythonmongodbaws
View job →

The Team MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of either our Dublin or Cork office or remotely in Ireland. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Build for reliability, making services and infrastructure avail

pythonmongodbaws
View job →

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of our Toronto or Montreal office or remotely in the Canada while physically based in an Eastern or Central time zone location. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineerin

pythonmongodbaws
View job →
O
Okta
📍 Washington• Full-time• From $174K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. The Staff SRE, Classified Opportunity This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a technical leader within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. Security Clearance: Active U.S. TS/SCI clearance with Full Scope Poly Compliance Expertise: Proven experience navigating Federal and DoD compliance frameworks, specifically FedRAMP and Impact Level 6 (IL6) What you’ll be doing: Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting hi

pythonsqlpostgresql
View job →
🔔

Get new staff software reliability engineer data platform jobs by email

Daily job updates · Unsubscribe anytime