Jobiba hiring network

Distributed Systems Engineer Data Platform Delivery Database Retrieval Jobs

1,301 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current distributed systems engineer data platform delivery database retrieval jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
Okta
📍 Washington• Full-time• From $194K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Position Overview: We are seeking a highly technical Staff Observability Site Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to delivering a world class, comprehensive, scalable Observability Platform that enables our SRE teams and business partners. You will treat infrastructure as code —utilizing Terraform and strong coding proficiency in Go, Python, or Ruby —to automate the deployment of agents and collectors across complex distributed systems. Key Responsibilities Automated Infrastructure: Design, build, and maintain scalable observability infrastructure using tools like Terraform. Splunk Engineering: Optimize the collection, processing, and storage of log data to ensure high reliability and low latency of our Splunk services Incident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "observability-driven development." Automation: Eliminate "toil" by automating the deployment and scaling of observability agents and collectors. Required Skills & Experience (The Essentials) Log Management: Minimum 5+ Experience scaling and managing Splunk Cloud at scale (1000+ SVCs), including Workload Management (WLM) and HEC optimization. Visualization: Expertise in creating intuitive, actionable Splunk dashboards that correlate data across multiple sources. SRE Mindset: Minimum 5+ years of experience in an SRE, Dev

pythonawsgcp
View job →
L
Lyft
📍 San Francisco• Full-time• $1.3M – $1.6M/yr
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. With over half a billion rides and counting, Lyft is solving hard problems in a rapidly growing domain with a lot of data and creative solutions in Marketplace, Mapping, Fraud, Growth and beyond. Building a next-generation platform for low-cost, ultra-immersive transportation to improve people's lives requires robust, scalable software systems operating at massive scale. Our highly motivated Software Engineers work on these challenging problems and build the systems that directly impact various aspects of our core business. If you are a critical thinker with experience building software systems, passionate about solving business problems through well-crafted code and working in a dynamic, creative, and collaborative environment, we are searching for you. As an associate software engineer, you will be designing, building, and launching the services that power the platform's core products. Compared to similarly-sized technology companies, the set of problems that we tackle is incredibly diverse. They cut across transportation, distributed systems, backend services, mapping, personalization, and real-time infrastructure. We are hiring motivated engineers across each of these areas. We're looking for someone who is passionate about solving problems with code, building reliable and maintainable systems, and is excited about working in a fast-paced, innovative, and collegial environment. You will report to a Software Engineering Manager. Responsibilities: Partner with Engineers, Data Scientists, Product Managers, and Business Partners to build software for business and user impact Perform technical analysis and build proof-of-concept prototypes to explore and propose solutions to both new and existing problems Design and develop software components, services, and APIs Write production quality code to launc

pythonaigo
View job →

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op

pythonawsrest
View job →

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Sr. Staff Software Development Engineer-AI Security to join our team. This is a Hybrid (based in San Jose, CA or Bellevue, WA with a 3 days in office requirement) role, reporting to the Director of Software Engineering in the Emerging Tech department. You will be responsible for designing and implementing core infrastructure components and distributed systems, serving as a foundational architect for our AI security solution. This high-impact role focuses on scaling security infrastructure to support hundreds of millions of users, collaborating with stakeholders across the development lifecycle to drive innovation and technical excellence. What you’ll do (Role Expectations) Architect, develop, and optimize a low-latency, high-throughput AI Security plane utilizing Rust, specifically leveraging its async/await model for highly efficient I/O and service-oriented architecture Build resi

vueawskubernetes
View job →
L
Lyft
📍 Toronto• Full-time• From C$108K/yr
26 days ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. We are looking for experienced backend software engineers to join our claims tech engineering team. Our vision is to tangibly reduce risk on the Lyft platform, and by extension reduce insurance cost. Our team is dedicated to centralizing the entire claims operation onto a unified risk platform. This consolidation of data, workflows, and communications aims to foster proactive measures, enhance efficiency, ensure consistency, and provide valuable insights. These efforts are designed to effectively reduce claims costs as Lyft's operations expand. Additionally, our team is responsible for maintaining robust relationships with our third-party insurance partners, guaranteeing timely, proactive, and precise sharing of claim data. Responsibilities: Write well-crafted, well-tested, readable, maintainable code Own feature from product spec to successful high quality development, deployment and maintenance Participate in code reviews to ensure code quality and distribute knowledge Respond to external questions and requests. Unblock, support and communicate with stakeholders to achieve results Experience: 3+ years of relevant professional experience Experience with object-oriented programming Experience in distributed systems Experience working with databases, relational or NoSQL Write clear, scalable and clear design documentation Design, build and improve a set of team owned components Benefits: Extended health and dental coverage options, along with life insurance and disability benefits Mental health benefits Family building benefits Child care and pet benefits Access to a Lyft funded Health Care Savings Account RRSP plan with company match to help save for your future In addition to provincial observed holidays, salaried team members are covered under Lyft's flexible paid time off policy. The policy allows

R
Roblox
📍 San Mateo• Full-time• From $243.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Backend Engineer on the User Communities team, you’ll own the backend systems that power social engagement features (Announcements, Forums, Polls, Roles & Permissions systems) for millions of players and creators on Roblox. Your day-to-day will focus on building and scaling the infrastructure that makes both in-experience and on-platform communication and player engagement fast, safe, and reliable. That means driving architectural decisions, reasoning through trade-offs, influencing stakeholders, championing user-first safety standards, and defining developer-facing APIs so creators can build richer experiences. You'll work closely with frontend and backend engineers, product, design, and data scientists across teams, and you'll have real influence over our technical and product direction. If you're an experienced engineer who gets excited about large-scale distributed systems and wants to shape the way millions of people connect inside virtual worlds, we'd love to talk. You Will: Shape the backend engineering culture across the Communities team. Your influence extends beyond your immediate pod to improve how we build, review, and ship software across the entire org. Act as

javaawsgit
View job →
B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab

kubernetesmachine learningai
View job →
P
1mo ago

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid Europe is building Plaid in Europe—the open banking platform that makes international expansion one click away for fintechs who want to scale globally. We are the only transatlantic open banking provider, uniquely positioned to help fintechs scale globally with a single Plaid integration. Our team localises, scales, and operates the bank connectivity layer that powers payments, underwriting, onboarding, and identity—abstracting fragmented European rails and regulations into a single, reliable platform. The team owns four core pillars: - Account-to-Account Payments - Cash Flow & Income Insights for Credit - Bank Connectivity Across Europe - Consumer Experience This team operates at the intersection of distributed systems, financial infrastructure, and consumer UX. We solve hard problems in payments reliability, data quality, regulatory complexity, and cross-market expansion—enabling fintechs to launch in new countries with minimal additional engineering effort. We are a high-impact, cross-functional engineering team working closely with US platform teams, GTM, compliance, and operations to make Plaid the default open banking layer for global fintechs operating in Europe. You will serve as t

pythonawsai
View job →

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We are the Data Foundation & AI team within Plaid’s Data organization. Our mission is to build the shared ML and AI infrastructure that powers intelligent capabilities across Plaid’s product suite. We develop the foundational systems, models, and data assets that transform Plaid’s unique financial network data into scalable, general-purpose representations that teams across the company can leverage. Our work spans the full ML lifecycle — from large-scale data curation and model pretraining to production serving, evaluation, and monitoring. As part of the team, you’ll work at the intersection of machine learning infrastructure, applied AI, and distributed systems, helping establish the core AI platform that enables innovation across Plaid. As a Staff Machine Learning Engineer, you will lead the technical strategy and development of Plaid’s foundation models, driving key decisions across pretraining objectives, model architecture, and fine-tuning approaches that power a wide range of downstream product applications. You will serve as the technical lead for the full machine learning lifecycle, overseeing everything from data curation and experimentation to production deployment, feature management,

pythonawsmachine learning
View job →
R
Roblox
📍 San Mateo• Full-time• From $243.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Orchestration pod within Engine Productivity, you'll design and run the platform that executes large-scale end-to-end and integration tests, running the real, shipping client on real hardware, across Roblox's data centers, cloud, and our own device labs, so our engineering teams can ship the engine, clients, Studio, and more with speed and confidence. Every Roblox engine, client, and Studio change, along with the experiences built on top of them, should ship with confidence, and the Orchestration team is the layer that makes that possible. We build large-scale distributed services that turn thousands of test suites into a reliable, push-button pipeline: fanning work out across fleets of machines and real devices, moving artifacts to where they're needed, managing single- and multi-client test state, and giving test owners and maintainers a system to validate their own runs. It looks a lot like building a specialized cloud platform, with capacity-aware scheduling, isolation and sandboxing, and smart retry and backoff, plus the classic distributed systems problems (fairness, efficiency, failure handling, and reliability) at Roblox scale. You Will: Design a

pythonawsgit
View job →
O
Okta
📍 Bengaluru• Full-time
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Required Skills: 9+ years in Python and its libraries (e.g., Pandas, Boto3) for data manipulation, ETL processes, and developing serverless functions to interact with foundational AI services on platforms like AWS Bedrock . 6+ years with modern front-end frameworks (e.g., React, Next.js, TypeScript), and an ability to collaborate effectively with Product & Design. Deep hands-on experience with AWS solutioning , including designing and deploying production applications using core services such as AWS EC2, AWS Lambda, Amazon S3, Amazon RDS, and API Gateway . Familiarity with Infrastructure as Code (IaC) tools like CloudFormation or Terraform is essential. Proven experience building and deploying production-grade AI applications and services. Experience with agentic frameworks like LangChain, LangGraph, and LangSmith is a strong plus. Strong background in building distributed systems, microservices, and resilient APIs (REST/GraphQL) on cloud platforms (AWS/GCP/Azure). Demonstrated ability to influence technical strategy and lead architecture decisions across global teams. Excellent communication skills with experience working effectively across time zones and cultures. Working experience with foundational AI services, with a desire to use a unified platform like AWS Bedrock for Generative AI development and deployment Education and Certifications A Bachelor’s degree in Computer Science, Information Systems, or equivalent years of industry experience AWS cl

typescriptpythonreact
View job →
G
Godaddy
📍 United States• Full-time• From $154K/yr
1mo ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the world, delivering the object, block, and file storage platforms that power GoDaddy's hosting infrastructure, internal services, OpenStack environments, and next-generation AI/HPC workloads. If you're passionate about distributed systems, storage architecture, and solving failure scenarios at massive scale, this is an opportunity to work on infrastructure few engineers will experience in their careers. Ceph is a strategic platform at GoDaddy — not an ancillary service. Our global footprint includes 80+ production clusters, 20,000+ OSDs, 1,830 storage nodes, 300 PB of raw capacity, and 69 billion objects spanning five datacenters across three continents. The platform supports RBD, RGW (S3/Swift), and CephFS workloads through more than 1,550 pools, 574,000 placement groups, and 900+ MDS daemons, creating engineering challenges that demand deep expertise in storage architecture, data durability, performance optimization, automation, and observability. As a Lead Senior Site Reliability Engineer, you'll serve as one of the principal technical leaders for GoDaddy's Ceph platform. You'll design the next generation of storage clusters, lead major platform upgrades, drive capacity and hardware strategy, and establish the standards that govern how the platform scales. You'll be the engineer the team turns to for the most complex s

pythonkubernetesai
View job →
R
Roblox
📍 San Mateo• Full-time• From $216.7K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Privacy Software Engineer on the Privacy Infrastructure team, you will design and build the foundational platforms, services, and controls that enable Roblox to protect user data and meet global privacy obligations at scale. You will develop privacy-by-design solutions that support data governance, user privacy rights, data discovery, retention, auditing, and regulatory compliance across a rapidly growing ecosystem of products and services. This role sits at the intersection of distributed systems, data platforms, security, and privacy engineering. You will partner closely with engineers, product teams, security, legal, and policy stakeholders to build scalable privacy infrastructure that is deeply integrated into Roblox's development workflows. Your work will directly influence how we responsibly manage data for millions of users while enabling innovation across the platform. You Will Design and build scalable privacy infrastructure that enables Roblox to discover, govern, protect, and manage personal data at scale. Develop backend services, APIs, and data pipelines that embed privacy controls into engineering workflows and enable fulfillment of user privacy rights at global sc

pythonawsgit
View job →
S
Stripe
📍 Nyc Privy• Full-time
1mo ago

Who we are About Privy Our mission is to make privacy and user ownership the default online. To do so, we build simple, flexible APIs and tools for developers that make it easy to build new products on crypto rails. Privy owns the abstractions and infrastructure layer above wallets, integrating across chains, third-party providers, and Stripe products like Treasury and Link. We get to solve hard technical problems while leveraging Stripe's distribution to reach customers like Ramp, Klarna, Deel, Kraken, Hyperliquid, and Fomo — powering experiences for both mainstream users and crypto natives. Learn more about Privy: Privy and Stripe: Bringing crypto to everyone About the team Engineering at Privy is distinguished by: High urgency: Shipping very small iterations, very fast, to learn very quickly. Product taste: Our customers are developers, and to build effective products for them requires technical knowledge - you will often be "the PM". Security mindset: A great portion of our product is trust. While we have a dedicated security team, every engineer brings security to their designs from the start. In practice, we use boring technology like Node, React, and AWS so we can focus our engineering energy entirely on pushing the boundaries of Privy's core product, e.g. through hardware enclaves, multi-region low latency APIs, and blockchain abstractions that are accessible to mainstream developers. What you'll do Design and build the backend systems that power wallets, identity, and onchain infrastructure at scale Create platform primitives and APIs that enable teams across Privy and Stripe to build faster Lead complex technical initiatives across architecture, data, and distributed systems Improve the scalability, reliability, and performance of our core platform Help shape our technical direction through high-leverage engineering work Who you are Minimum requirements 8+ years of experience building and maintaining a production system at scale An understandin

reactawsai
View job →
M
1mo ago

The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our Toronto or Vancouver offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design primitives of at least one of AWS, Azur

mongodbawsazure
View job →
🔔

Get new distributed systems engineer data platform delivery database retrieval jobs by email

Daily job updates · Unsubscribe anytime