Jobiba hiring network

Staff Software Reliability Engineer Data Platform Jobs

3,518 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current staff software reliability engineer data platform jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
Okta
📍 Toronto• Full-time• From C$160K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Staff Software Reliability Engineer - Data Platform About the Team The Data Platform team is responsible for the foundational data services, systems, and data products for Okta that benefit our users. Today, the Data Platform team solves challenges and enables: Streaming analytics Interactive end-user reporting Data and ML platform for Okta to scale Telemetry of our products and data Our elite team is fast, creative and flexible. We encourage ownership. We expect great things from our engineers and reward them with stimulating new projects, new technologies and the chance to have significant equity in a company. Okta is about to change the cloud computing landscape forever. About the Position This is an opportunity for experienced Software Reliability Engineers to join our fast growing Data Platform organization that is passionate about scaling high volume, low-latency, distributed data-platform services & data products. In this role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Okta to scale for years to come. As a member of the Data Platform team, you will be responsible for designing, building, and deploying the systems that power our data analytics and ML. Our analytics infrastructure stack sits on top of many modern technologies, including Kinesis, Flink, ElasticSearch, and Snowflake. We are looking for experienced Software Engineers who can help desi

javaawskubernetes
View job →
DU
13 days ago

About the Team Data is at the foundation of DoorDash success. The Data Engineering team builds database solutions for various use cases including reporting, product analytics, marketing optimization and fi nancial reporting. Team serves as the foundation for decision-making at DoorDash. About the Role DoorDash is looking for a Sta ff Software Engineer,Data to be a technical lead and help architect and scale our data reliability, data infrastructure, automation and tools to meet growing business needs. You’re excited about this opportunity because you will... Own critical data systems that support multiple products/teams Develop, implement and enforce best practices for data infrastructure and automation Design, develop and implement large scale, high volume, high performance data models and pipelines for Data Lake and Data Warehouse Improve the reliability and scalability of our Ingestion, data processing, ETLs, Reporting tools and data ecosystem services Manage a portfolio of data products that deliver high-quality, trustworthy data Help onboard and support other engineers as they join the team We’re excited about you because... 8+ years of professional experience as a hands-on engineer and technical leader leading multiple projects 6+ years experience working in data platform and data engineering or a similar role You have proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software Pro fi ciency in programming languages such as Python/Kotlin/Scala 4+ years of experience in ETL orchestration and work fl ow management tools like Air fl ow Expert in database fundamentals, SQL, data reliability practices and distributed computing 4+ years of experience with the Distributed data/similar ecosystem (Spark, Presto) and streaming technologies such as Kaa/Flink/Spark Streaming Excellent communication skills and experience working

pythonsqlaws
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $252K/yr
13 days ago

Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human evaluation and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Platform Engineering team. In this role, you will lead the design and development of core data storage, streaming, caching, and indexing platforms and underlying systems. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the architecture, design, implementation, and reliability of our foundational data platforms and systems, working closely with stakeholders and internal customers to understand and refine requirements. Collaborate with cross-functional teams to define, design, and deliver new features. Proactively identify opportunities for, and driving improvements to, current programming practices, including process enhancements and tool upgrades. Present technical information to teams and stakeholders, providing

mongodbredisaws
View job →
T
Twilio
📍 - India• Full-time• Remote
1mo ago

Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as our next Staff engineer (L4), Twilio’s Segment team. About the job As a Staff Engineer on the Twilio Segment Data platform/ pipelines team, you’ll build and scale systems that process several hundred thousands of data points per second. You will lead the development of high-scale ingestion and data processing systems You'll be designing, operating and maintaining complex distributed systems, ensuring reliability, performance, and cost-efficiency while querying petabytes of data for our customer data platform (CDP). Responsibilities In this role, you’ll: Design and deliver robust, high-scale routing experiences for the Data platform/ pipelines team for Twilio Segment. Ship features that opt for high availability and throughput with eventual consistency Collaborate with engineering and product leads, as well as teams across Twilio Segment Support the reliability and security of the platform Build and optimize globally available and high

REMOTEjavaawsgcp
View job →
A
13 days ago

Who We Are Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. Addepar unifies portfolio, market and client data in a total portfolio view and delivers AI-powered insights within investment and client workflows. More than 1,400 firms in nearly 60 countries use Addepar to manage and advise on nearly $9 trillion in assets. Its open platform integrates with nearly 650 software, data and consulting partners to power end-to-end investment operations across firms of all sizes and complexity. Addepar supports clients worldwide with offices in New York City, Salt Lake City, London, Edinburgh, Pune, Dubai, Geneva, Singapore and São Paulo. The Role We are currently seeking a Staff Software Engineer, Infrastructure to join the AI Platform team that powers seamless insights and interaction through natural language and data intelligence across our AI products. As a Staff Software Engineer, you’ll architect, build, and operate the backend and platform systems that power AI Platform. You’ll work across service design, distributed systems, cloud infrastructure, event-driven processing, observability, CI/CD, and production reliability, helping shape the technical direction of a platform that supports scalable, client-facing AI experiences. This role requires a strong software engineering foundation combined with deep infrastructure and systems thinking. We are looking for an engineer who can write high-quality production code, make sound architectural tradeoffs, and own platform capabilities end-to-end — not someone focused only on scripting, cloud configuration, or infrastructure tooling in isolation. You will collaborate closely with frontend, product, and AI/ML engineers to deliver reliable, secure, and scalable systems that align with Addepar’s standards of performance, resilience, and trust. Applicants must have legal authorization to work in the country where this role is based o

pythonjavasql
View job →

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The DevX team's mission is to build and operate the core developer infrastructure at Robinhood. Our team owns and scales the systems that thousands of engineers rely on daily, partnering with software developers across the company to make development fast, reliable, and cost-efficient. We are a team that takes pride in technical excellence, platform reliability, and building tools that have an outsized impact on engineering velocity across the entire organization! As a Staff Software Engineer on the DevX team, you will serve as a technical leader for our build and developer infrastructure, driving the strategy and execution of the systems thousands of engineers depend on every day. Your work will span our build systems, CI pipelines, and remote development environments, ensuring engineers can code, test, and build with speed, safety, and reliability at scale. You will collaborate with engineering teams across Robinhood to eliminate developer friction, raise the bar for engineering productivity, and shape the long-term direction of our developer ecosystem. This is a high-visibility opportunity to set new standards of engineering efficiency and mentor the next generation of in

pythonvueaws
View job →
I
Instacart
📍 United States - Remote• Full-time• Remote• From $265K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Instacarts Data Infrastructure organization builds and operates the systems that power our company’s data ecosystem, including a modern open data lakehouse on Apache Iceberg, a multi-engine compute platform for stream and analytical workloads, and self-serve tooling that helps Product, Data Science, ML, Ads, Finance, and engineering teams move fast with data. We’re looking for a Staff Software Engineer, Data Infrastructure to join our Data Governance and Foundations Team. In this role, you’ll serve as a senior technical leader owning the architecture and delivery of our open lakehouse foundation, governance and access patterns, and multi-engine compute strategy—balancing today’s reliability with the next three to five years of scale, maturity, and cost efficiency. You’ll collaborate closely with engineering leadership and stakeholders across Data Science,

REMOTEpythonsqlaws
View job →

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Unified Data Store (UDS) team is the architect of Airbnb’s global system-of-record. We design, build, and operate the mission-critical storage platform that powers every user profile, listing, reservation, and financial transaction on the platform. Supporting over 150 million users worldwide, our work is the bedrock of Airbnb’s reliability and efficiency. As a Staff Engineer in our Brazil Engineering Hub , you will join a high-impact group of technical leaders who value craft and operational excellence. You won't just be managing data; you will be building a modern distributed infrastructure service that enables hundreds of product teams to ship features with total confidence. The Difference You Will Make: We are looking for a hands-on technical leader who thrives on solving deep architectural challenges and leading through ambiguity. As a Staff Engineer, you will serve as the Technical North Star for the UDS Client Stack, ensuring our data access layer is seamless, high-performance, and future-proof. A Typical Day: Define Technical Strategy: Lead the multi-year roadmap and long-term architecture for the UDS client stack, balancing immediate execution with systemic platform evolution. Architect for Scale: Design and operate a high-performance data access layer that abstracts complexities like indexing, replication, and global consistency models. Drive Engineering Excellence: Lead deep-dive design reviews and establish best practices for building fault-tolerant distributed systems across the organization. Empower Developers: Act as a bridge between infrastructure and

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale is seeking a Staff Software Engineer to lead the technical vision for our Infrastructure team. As a Staff Engineer, you will be responsible for the architectural evolution of our control plane and data plane, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. You will act as a force multiplier, setting the standards for Kubernetes-based cloud-native infrastructure while mentoring engineers and driving cross-functional alignment across the Ray open-source community and our proprietary product teams. Key Responsibilities Architectural Leadership: Define and drive the multi-year technical roadmap for services that orchestrate Ray clusters across diverse cloud and on-premises environments. Systemic Optimization: Lead the design and optimization of high-performance control plane components specifically tailored for large-scale, heterogeneous AI/ML workloads. Platform Reliability: Establish the organization-wide standards for the reliability, scalability, and observability of Anyscale-managed infrastructure. Strategic Integration: Direct the long-term strategy for accelerator integration (GPUs, TPUs) and container management to ens

pythonawsazure
View job →
R
1mo ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Storage Platform team builds and operates the platform that powers database access across Robinhood. We own relational (Postgres/Aurora), key-value (DynamoDB), and caching systems, along with the SDKs, control plane automation, and data plane services that enable safe and reliable access at scale. Our mission is to standardize and strengthen how services connect to storage, improve reliability and performance, and reduce operational overhead through automation. We manage thousands of databases and hundreds of caching clusters supporting millions of users and critical brokerage workloads. Availability is our highest priority — our systems are designed to meet strict uptime targets, including no downtime during market hours. As a Staff Software Engineer , you will design and evolve the core infrastructure that underpins Robinhood’s storage systems. You’ll lead complex distributed systems initiatives such as horizontal sharding, proxy-based query routing, connection pooling, and cross-shard transactions. You’ll work on improving database reliability, performance, and cost efficiency across multi-region deployments. This role has a direct impact on system availability, laten

vuesqlpostgresql
View job →
G
Gitlab
📍 United Kingdom; Remote, United States• Full-time• Remote• From $126.4K/yr
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve our production infrastructure. This is a single application for Site Reliability Engineering opportunities across our Infrastructure Platforms department. Rather than asking you to choose the right team or level upfront, we evaluate your skills holistically and match you to the opportunity that best aligns with your experience

REMOTEawsgcpkubernetes
View job →

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. As a Staff Site Reliability Engineer (SRE) at GitLab, you’ll help keep all user-facing services and production systems reliable, scalable, and efficient. Our SREs combine a pragmatic operations mindset with strong software engineering practices to drive automation, reduce toil, and improve resilience across our platform. In the Environment Automation specialization, your focus is on operating and automating hundreds of GitLab environments—from initial provisioning to day-to-day maintenance tasks. Unlike other SRE roles, this position centers on automating the lifecycle of many tenant environments, ensuring they remain secur

awsgcpkubernetes
View job →
P
Prophecy
📍 Bengaluru• Full-time
13 days ago

About Prophecy The leader in AI-native data preparation and analysis, Prophecy is revolutionizing how the world's top enterprises turn data chaos into reliable insights. We introduce the AI-native data lifecycle (generate, refine, deploy) where our industry leading AI agents and humans work hand-in-hand in visual and document interfaces to analyze, transform and prepare data, to ship trusted insights at enterprise scale. To learn more, visit us on LinkedIn. Position Summary This is an exceptional opportunity to be a technical leader on the team that powers the heart of Prophecy: the Runtime, our orchestration and execution service. Every data workflow , whether generated by our AI data agents or built by a human in the visual editor , is planned, scheduled, and run at enterprise scale by Runtime, directly against the customer's own platform (Databricks, Snowflake, BigQuery, and more). When a business user asks an agent a question and sees trusted results seconds later, Runtime is what made it happen. As a Staff Engineer, you will set the technical direction for this service: the distributed execution engine, the task coordinator, the multi-engine connectors, and the fast interactive previews that make the AI-native experience feel instant. You will own hard, ambiguous, high-leverage problems end-to-end , from architecture to production reliability , and raise the bar for the engineers around you. We're seeking A players to work with A players. We're a high-growth company on a once-in-a-lifetime journey to revolutionize the way data is prepared and analyzed. Be a part of it. The impact you will have Own the engine behind the product. Architect and evolve the runtime service that turns agent- and human-authored data workflows into high-performance jobs that run reliably in production and interactively during authoring. This is deep-tech at the center of the platform. Make it fast and make it scale. Design the distributed execution and task-coordination layer , schedul

pythonjavarest
View job →

Aliases: Staff Engineer, Senior Staff Engineer, Lead Engineer, Senior Technical Lead, Senior Architect, Platform Architect, Data Architect, Solutions Architect (Engineering), Principal Engineer at a smaller company About Truveta Truveta is the world's first health provider-led data platform with a vision of Saving Lives with Data. Our mission is to enable researchers to find cures faster, empower every clinician to be an expert, and help families make the most informed decisions about their care. Achieving Truveta's ambitious vision requires an incredible team of talented and inspired people with a unique combination of health, software, and big data expertise who share our company values. This opportunity Join the founding team of the Truveta India Development Center and play a pivotal role in shaping its future. As one of the early members, you will help build and scale a high-impact organization while contributing to the products and platforms that advance healthcare through data and AI. This is an unusual chance to influence both the technical direction and the culture of a growing global engineering hub. The Data Platform group in India owns parts of the pipeline that ingests, processes, stores, and serves healthcare data at a scale very few organizations operate at. Patients, doctors, and medical researchers deserve the same engineering rigor that has transformed other industries, and that is the work: correctness under volume, reliability under failure, and speed without cutting the corners that regulated health data does not allow anyone to cut. Who we need We are seeking engineers who think in platforms. You will own a processing domain or a multi-team data capability, define its architecture, and make that architecture legible enough that the teams building inside it can move quickly without asking permission. The problems at this level are the recurring ones: the class of data-quality defect that keeps retu

B
Brex
📍 Vancouver• Full-time• $240K – $285K/yr
13 days ago

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Staff Software Engineer in Banking, you will help shape the technical direction of one of Brex’s most strategic and complex product areas. The Banking org is both a product and platform org, it owns the Brex Business Account product and AP offerings like Bill Pay and Vendors, and the underlying money movement platform and partner integrations that power those experiences. In this role, you’ll work across customer-facing product surfaces and core financial infrastructure, driving architecture, reliability, and

aiexcelaccounting
View job →
🔔

Get new staff software reliability engineer data platform jobs by email

Daily job updates · Unsubscribe anytime