Jobiba hiring network

Staff Software Reliability Engineer Data Platform Jobs

3,518 active opportunities · Updated for October 2026

Fresh results

14 shown

Explore current staff software reliability engineer data platform jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As the Team Lead for Initiator & Protocol Engineering, you will spearhead the critical bridge between our industry-leading FlashArray and the Linux/VMWare ecosystems. You will drive the performance and reliability of our storage protocol stacks—spanning NVMe over Fabrics and Fibre Channel—ensuring Pure Storage remains the gold standard for enterprise connectivity. Collaborating closely with cross-functional hardware and software teams, you’ll mentor a high-caliber engineering squad to solve complex kernel-level challenges and influence the global Linux upstream community. WHAT YOU'LL DO Own the Protocol Lifecycle: Lead the development, maintenance, and optimization of Linux and VMWare initiator stacks (NVMeoF, FC-SCSI, iSCSI) and target drivers to ensure seamless, high-performance integration with Pure FlashArray. Drive System Resilience: Architect enhancements for Fibre Channel and NIC driver stacks that improve RAS (Reliability, Availability, and Serviceability), specifically focusing on multipathing logic and link health monitoring. Technical Leadership & Mentorship: Guide a team of senior and junior engineers through complex project deliveries, conducting deep-dive code reviews and setting the technical bar for C/C++ and Python development within the kernel space. Solve the Impossible: Act as the final escalation point for the most challenging system-level bugs found in the field or internal testing, u

pythonawslinux
View job →
E
14 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the FlashArray team to build high-performance, resilient storage software that powers mission-critical applications worldwide. As a core member of an agile engineering group, you will design and deliver zero-downtime algorithms that directly shape our enterprise storage and public cloud offerings. In this role, you will partner closely with global systems engineers and product teams to translate complex technical challenges into scalable, real-world solutions. Your work will directly impact how thousands of global enterprises manage, protect, and scale their data seamlessly. WHAT YOU’LL DO Design & Deliver Resilient Systems: Architect and implement high-performance algorithms for enterprise storage products, ensuring platform reliability, end-to-end delivery from concept to release, and six-nines availability. Expand Cloud Architecture: Extend core platform capabilities into public cloud environments (such as AWS), driving performance and agility for both traditional IT and cloud-native applications. Drive Problem Solving & Quality: Analyze and resolve complex systems software challenges, optimizing storage internals and contributing to continuous continuous platform upgrades. Collaborate & Mentor: Partner with multidisciplinary engineering peers to review code, refine architecture, and maintain high engineering standards across distributed software projects. WHAT YOU BRING Systems Programming Exp

pythonjavaaws
View job →
DC
14 days ago

Role Overview Build the software services that power products, platforms, and better business decisions. As a Software Engineer II, you’ll develop scalable backend applications, APIs, integrations, and AI-enabled features using Python and cloud technologies. You’ll contribute to solutions from design through production, helping improve reliability, performance, security, and developer productivity. This is an opportunity to solve meaningful engineering challenges, grow your technical ownership, and collaborate with experienced engineers across the development lifecycle. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design and build scalable backend services, REST APIs, integrations, and reusable software components using Python and, where relevant TypeScript. Develop data ingestion, transformation, service-to-service communication, and automation capabilities that support reliable product experiences. Contribute to AI-enabled features and use AI development tools responsibly to improve coding, testing, research, documentation, and delivery. Apply sound engineering practices across architecture, performance, security, testing, debugging, and maintainability. Deploy and operate services using AWS and CI/CD workflows, contributing to monitoring, troubleshooting, documentation, and continuous improvement. Partner with engineers and cross-functional colleagues through design discussions, code reviews, technical problem-solving, and knowledge sharing. These are the essentials you’ll need to get an interview 3–5 years of professional experience building and delivering production software in an agile environment. Strong hands-on experience with Python and backend development, including APIs, integrations, or service-oriented applications. Experience working with cloud platforms, preferably AWS, and familiarity with deployment or CI/CD practices. Working knowledge of software design principles, testing, debugging, performance optimization, an

typescriptpythonreact
View job →

As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams. Success in this role means smoother launches, clearer roadmaps, stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale. This role can be based remotely on the East Coast What You'll Do Drive Program Planning & Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep work clearly tracked in Jira, and deliver on time Strengthen Production Reliability – Lead change management and launch readiness programs. Partner with SREs and product teams to define and operationalize SLOs/SLIs, and use incident data, metrics, and capacity signals to drive prioritization and continuous improvement Lead Cross-Functional Coordination – Align SRE with Security, Compliance, Cloud platform, and other engineering teams. Coordinate cross-team incident response, ensure clear follow-through, and build trust as the go-to driver of complex, multi-team efforts Build Scalable Systems & Processes – Design lightweight frameworks and communication patterns that help SRE deliver reliably at scale. Work yourself out of the "hero" role by leaving teams better-equipped to execute independently Requirements 8+ years in technical program management, engineering management, or a comparable technical role partnering with software engineering teams Proven track record leading large-scale, cross-team platform initiatives through ambiguity and change Strong knowledge of production change management, software development lifecycle, and reliability metrics (SLOs, SLIs) Skilled at shaping roadmaps and managing dependencies Able to query and interpret metrics, logs, or other data s

mongodbawsazure
View job →
M
1mo ago

As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams. Success in this role means smoother launches, clearer roadmaps, stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale. This role can be based out of our Dublin or Cork office or remotely in Ireland. What You'll Do Drive Program Planning & Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep work clearly tracked in Jira, and deliver on time Strengthen Production Reliability – Lead change management and launch readiness programs. Partner with SREs and product teams to define and operationalize SLOs/SLIs, and use incident data, metrics, and capacity signals to drive prioritization and continuous improvement Lead Cross-Functional Coordination – Align SRE with Security, Compliance, Cloud platform, and other engineering teams. Coordinate cross-team incident response, ensure clear follow-through, and build trust as the go-to driver of complex, multi-team efforts Build Scalable Systems & Processes – Design lightweight frameworks and communication patterns that help SRE deliver reliably at scale. Work yourself out of the "hero" role by leaving teams better-equipped to execute independently Requirements 8+ years in technical program management, engineering management, or a comparable technical role partnering with software engineering teams Proven track record leading large-scale, cross-team platform initiatives through ambiguity and change Strong knowledge of production change management, software development lifecycle, and reliability metrics (SLOs, SLIs) Skilled at shaping roadmaps and managing dependencies Able to query and interpret

mongodbawsazure
View job →
DC
Diligent Corporation
📍 Netherlands• Full-time
14 days ago

Role Overview You’re a hands-on backend engineer who enjoys owning features end to end and working on real products that customers rely on every day. In this Software Engineer II role, you’ll help build and evolve a Third Party Risk Management SaaS platform using Laravel and PHP, designing scalable APIs and services that keep performance and reliability front and center. You’ll work in a product-focused team that owns its services from architecture and implementation through deployment, monitoring, and continuous improvement. You’ll mentor junior engineers, influence technical decisions, and use modern AI-powered tools thoughtfully to ship better code faster. If you’re looking for a mid-level role with real ownership, modern tooling, and the chance to grow your impact, this is for you. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design, build, and maintain backend features and RESTful APIs in Laravel within a modern TALL stack environment. Own well-defined stories from implementation through deployment, monitoring, and iteration, ensuring performance and reliability. Contribute to architectural discussions and technical decisions that shape the Third Party Risk Management platform. Review code, improve test coverage, and strengthen CI/CD and engineering standards across the team. Mentor Software Engineer I colleagues through code reviews, pairing, and knowledge sharing. Use AI tools (e.g. GitHub Copilot, ChatGPT) to accelerate coding, debugging, testing, and documentation—while critically validating outputs and ensuring safe, responsible use. These are the essentials you’ll need to get an interview 3–5 years of professional software engineering experience in an agile, fast-paced environment. Strong experience with PHP and Laravel, ideally within the TALL stack (Tailwind, Alpine.js, Laravel, Livewire). Solid understanding of relational databases (MySQL or MariaDB), including data modelling and query optimisation. Experience designin

reactvuesql
View job →

Senior Product Manager, Robotics & Autonomy What we're doing isn't easy, but nothing worth doing ever is. At Diligent Robotics, we envision a future powered by robots that work seamlessly with human teams. We build artificial intelligence that enables service robots to collaborate with people and adapt to dynamic human environments. Our robots operate every day in hospitals, helping healthcare staff spend less time on routine work and more time caring for patients. Operating a real-world fleet gives us something few robotics companies have: continuous customer feedback and operational data that directly shapes the next generation of Physical AI. We're looking for a Senior Product Manager, Robotics & Autonomy to define and execute the product strategy for some of the most critical capabilities in our robotics platform. You'll work at the intersection of robotics, autonomy, AI, and software engineering to translate business priorities, customer needs, and technical opportunities into a clear product roadmap that drives measurable outcomes. This role is ideal for someone who understands complex autonomous systems and enjoys working alongside world-class engineers to bring ambitious technology from concept into production. Responsibilities Own the product strategy and roadmap for key Robotics and Autonomy initiatives, balancing customer impact, technical feasibility, and long-term platform investments. Define product requirements for autonomy, navigation, perception, fleet intelligence, simulation, and robotics platform capabilities. Partner closely with Engineering, AI, Robotics, Customer Success, Operations, and Leadership to align priorities across the organization. Translate customer feedback, fleet telemetry, and operational insights into product decisions that improve robot performance, reliability, and user experience. Prioritize investments using data, customer value, technical complexity, and business impact. Drive cross-functional execution from concep

agilemachine learningai
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
14 days ago

Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is experiencing rapid growth. You’ll play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features, while also balancing scalability, reusability, and performance. As a Staff Engineer on the Application Engineering team, you’ll serve as a senior technical leader - responsible for designing and delivering high-quality, scalable software systems that power Cohere Health’s core platform. You’ll act as a multiplier, elevating the technical bar for the team, mentoring engineers, and partnering with product, data, and clinical teams to deliver solutions that meet compliance, quality, and performance standards. This role is ideal for engineers who thrive on solving complex problems in healthcare, have deep expertise in building distributed systems, and want to influence architecture and engineering practices at scale. What you’ll do: Technical Leadership & Architecture Define and drive the architecture of large-scale, distributed application systems across the Cohere platform. Ensure solutions are secure, performant, maintainable, and compliant with NCQA, CMS, and payer requirements. Champion engineering best practices in CI/CD, testing, release management, and observability. Hands-On Engineering Write clean, maintainable, and well-tested code, primarily in modern frameworks (e.g., Python, TypeScript/React, Java/Kotlin). Lead the development of core features and APIs that directly impact providers, payers, and patients. Partner with DevOps and Data teams to ensure seamless integration, scalability, and operational readiness. Quality & Compliance Focus Embed automated testing, monitoring, and release safeguards into the development lifecycle. Proactively address compliance and audit-readiness requirements in application

typescriptpythonjava
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
14 days ago

Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is experiencing rapid growth. You’ll play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features while also balancing scalability, reusability, and performance. As a Staff Engineer on the Application Engineering team, you’ll serve as a senior technical leader - responsible for designing and delivering high-quality, scalable software systems that power Cohere Health’s core platform. You’ll act as a multiplier, elevating the technical bar for the team, mentoring engineers, and partnering with product, data, design , clinical and payment teams to deliver solutions that meet compliance, quality, and performance standards. This role is ideal for engineers who thrive on solving complex problems in healthcare, have deep expertise in building distributed systems including data solutions, and want to influence architectural decisions for security and scale, drive cross-collaborations for alignment, establish technical standards for consistency and evolve both application and data engineering best practices at scale. What you’ll do: Technical Leadership & Architecture Define and drive the architecture of large-scale, distributed application systems across the Cohere platform. Ensure solutions are secure, performant, maintainable, and compliant with NCQA, CMS, and payer requirements. Champion platform engineering best practices in CI/CD, testing, release management, and observability. Hands-On Engineering Write clean, maintainable, and well-tested code, primarily in modern frameworks (e.g., Python, TypeScript/React, Java/Kotlin). Lead the development of core features and APIs that directly impact providers, payers, and patients. Partner with DevOps and Data teams to ensure seamless integration, scalability, and operati

typescriptpythonjava
View job →

About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We’re entering a world where AI agents don’t just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents’ ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova’s throughput, correctness, and operational rigor grows dramatically. We’re looking for a Staff Software Engineer who wants to go deep on both the engine internals and the infrastructure underneath it. You’ll work across the full stack of a modern OLAP system — query planning and execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — while driving meaningful improvements to performance, cost-efficiency, and reliability at scale. You’ll influence technical direction through your work, your design reviews, and your mentorship of other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of customers. What You’ll Do Build and evolve core query engine infrastructure Work across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Design and implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Design for high-throughput automated quer

pythonjavaredis
View job →
D
Datadog
📍 New York• Full-time• From $244K/yr
1mo ago

Role Summary: Datadog is seeking a Staff Software Engineer to help shape the future of our Bring Your Own Cloud (BYOC) Logs offering by unifying observability pipelines with log management software that customers deploy and manage in their own infrastructure. This role will focus on building and scaling systems that process, route, and store high-volume observability data within customer-managed infrastructure. You will operate as a hands-on technical leader, driving architecture, cross-team delivery, and product direction across a complex and evolving space. This is a high-impact opportunity to influence product strategy, mentor engineers, and solve deeply technical challenges at scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Make customer-controlled deployments feel like a managed Datadog product: deployment, upgrades, configuration, observability, diagnostics, reliability, and secure operation across diverse customer cloud environments Build and scale high-throughput systems for log processing, routing, and transformation across distributed environments Lead cross-team initiatives, aligning engineers, product managers, and stakeholders to deliver complex, multi-team projects Design and implement software that runs reliably that customers deploy and operate within their own cloud infrastructure. Improve system performance, scalability, and cost efficiency through thoughtful trade-off analysis and capacity planning Contribute hands-on to critical code paths, debugging, and deployment challenges in customer environments Who You Are: You have significant experience building software that is installed, deployed, and operated in customer environments rather than only as a fully managed SaaS service. You have strong expertise in distributed systems,

awsazuregcp
View job →
O
20 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know Okta Okta is The World’s Identity Company. We free everyone to safely use any technology anywhere, on any device or app. Our Workforce and Customer Identity Clouds enable secure yet flexible access, authentication, and automation that transforms how people move through the digital world, putting Identity at the heart of business security and growth. At Okta, we celebrate a variety of perspectives and experiences. We are not looking for someone who checks every single box we’re looking for lifelong learners and people who can make us better with their unique experiences. Join our team! We’re building a world where Identity belongs to you. About Technology Data and Intelligence at Okta At Okta, the Technology Data and Intelligence (TDI) team drives internal efficiency through secure, scalable, and innovative systems. TDI partners with teams across the company to build and support the infrastructure, automation, and enterprise applications that keep operations running smoothly. Focused on enabling productivity and aligning technology with business goals, TDI plays a vital role in both day-to-day operations and long-term strategic growth. The Staff Software Engineer Opportunity We are looking for a Staff Software Engineer to join our growing team in TDI and help scale our internal business solutions with a sharp focus on security, reliability, scalability, and intelligent automation. You will be responsible for designing and developing customization

javascriptpythonjava
View job →
O
Okta
📍 India• Full-time
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know Okta Okta is The World’s Identity Company. We free everyone to safely use any technology anywhere, on any device or app. Our Workforce and Customer Identity Clouds enable secure yet flexible access, authentication, and automation that transforms how people move through the digital world, putting Identity at the heart of business security and growth. At Okta, we celebrate a variety of perspectives and experiences. We are not looking for someone who checks every single box we’re looking for lifelong learners and people who can make us better with their unique experiences. Join our team! We’re building a world where Identity belongs to you. About Technology Data and Intelligence at Okta At Okta, the Technology Data and Intelligence (TDI) team drives internal efficiency through secure, scalable, and innovative systems. TDI partners with teams across the company to build and support the infrastructure, automation, and enterprise applications that keep operations running smoothly. Focused on enabling productivity and aligning technology with business goals, TDI plays a vital role in both day-to-day operations and long-term strategic growth. The Staff Software Engineer Opportunity We are looking for a Staff Software Engineer to join our growing team in TDI and help scale our internal business solutions with a sharp focus on security, reliability, scalability, and intelligent automation. You will be responsible for designing and developing customization

javascriptpythonjava
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Cortex CoWork team is defining the future of AI for enterprise data. Our mission is to transform how the world’s largest enterprises interact with their data through flagship products like Snowflake (CoWork) Intelligence . As a Principal AI Engineer , you will be a technical North Star for our AI initiatives. You won't just execute on a roadmap; you will help define it. You will tackle the most complex, "frontier" problems in agentic reasoning, NL-to-SQL, and enterprise-scale RAG, ensuring our AI products are not only innovative but fundamentally reliable and scalable for the Fortune 500. What you will do in this role: Technical Strategy & Architecture: Define the long-term technical vision for Snowflake Intelligence. Lead the architectural design of multi-agent systems, complex tool-use frameworks, and self-correcting NL-to-SQL engines. Drive Industry-Leading Reliability: Move beyond simple evals to build world-class, automated "hill-climbing" infrastructure. You will establish the methodology for how Snowflake measures and guarantees LLM performance across diverse customer schemas. Cross-Functional Influence: Partner with Product and Engineering leadership to align AI capabilities with business goals. You will bridge the gap between Research (modeling) and Product

🔔

Get new staff software reliability engineer data platform jobs by email

Daily job updates · Unsubscribe anytime