Jobiba hiring network

Distributed Systems Engineer Data Platform Delivery Database Retrieval Jobs

1,301 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current distributed systems engineer data platform delivery database retrieval jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We’re entering a world where AI agents don’t just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents’ ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova’s throughput, correctness, and operational rigor grows dramatically. We’re looking for a Staff Software Engineer who wants to go deep on both the engine internals and the infrastructure underneath it. You’ll work across the full stack of a modern OLAP system — query planning and execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — while driving meaningful improvements to performance, cost-efficiency, and reliability at scale. You’ll influence technical direction through your work, your design reviews, and your mentorship of other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of customers. What You’ll Do Build and evolve core query engine infrastructure Work across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Design and implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Design for high-throughput automated quer

pythonjavaredis
View job →

About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We're entering a world where AI agents don't just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents' ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova's throughput, correctness, and operational rigor grows dramatically. We're looking for a Senior Software Engineer who wants to go deep on the engine internals and the infrastructure underneath. You'll own significant components of a modern OLAP system — across query execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — and drive meaningful improvements to performance, cost-efficiency, and reliability. You'll grow your technical influence through the quality of your code, your design contributions, and your collaboration with other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of enterprise customers. What You'll Do Build and improve core query engine components Contribute across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Help ensure Nova's components support

pythonjavaredis
View job →
M
Mongodb
📍 Bengaluru• Full-time
1mo ago

MongoDB Technical Services Engineers use their exceptional problem solving and customer service skills, along with their deep technical experience, to advise customers and to solve their complex MongoDB problems. Technical Service Engineers are experts in the entire MongoDB ecosystem - database server, drivers, cloud and infrastructure. This also includes services such as Atlas (database as a service), or Cloud Manager (which helps customers with automation, backup and monitoring of their MongoDB systems). Our engineers combine their MongoDB expertise with passion, initiative, teamwork and a great sense of humor to help our customers to be successful with MongoDB. We are looking to speak to candidates who are based in Bengaluru for our hybrid working model. Cool things you’ll do You'll be working alongside our largest customers, solving their complex challenges - resolving questions on architecture, performance, recovery, security, and everything in between. You'll be an expert resource on best practices in running MongoDB at scale, whatever that scale may be. You'll be an advocate for customers' needs - interfacing with our product management and development teams on their behalf. And you'll contribute to internal projects, including software development of support tools for performance, benchmarking, and diagnostics. What you need We consider all candidates with an eye for those who are self-taught, insatiably curious, and multi-faceted. The ideal candidates should have strong technical experience in one (or more) of the following areas Systems administration Distributed systems Network Administration Database architecture and administration Application Architecture Data architecture and design Performance tuning and benchmarking Extra bonus points if you have experience in one or more of Java, Python, Ruby, C, C++, C#, Javascript, node.js, Go, PHP, or Perl If you have an operations background, we prefer experience administering large-scale production environments

javascriptpythonjava
View job →
M
Mongodb
📍 Gurugram• Full-time
1mo ago

MongoDB Technical Services Engineers use their exceptional problem solving and customer service skills, along with their deep technical experience, to advise customers and to solve their complex MongoDB problems. Technical Service Engineers are experts in the entire MongoDB ecosystem - database server, drivers, cloud and infrastructure. This also includes services such as Atlas (database as a service), or Cloud Manager (which helps customers with automation, backup and monitoring of their MongoDB systems). Our engineers combine their MongoDB expertise with passion, initiative, teamwork and a great sense of humor to help our customers to be successful with MongoDB. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Following the successful completion of the probationary period, the candidate ​may be required to work a Tuesday-to-Saturday schedule, with Sundays and Mondays designated as weekly days off​. Cool things you’ll do You'll be working alongside our largest customers, solving their complex challenges - resolving questions on architecture, performance, recovery, security, and everything in between. You'll be an expert resource on best practices in running MongoDB at scale, whatever that scale may be. You'll be an advocate for customers' needs - interfacing with our product management and development teams on their behalf. And you'll contribute to internal projects, including software development of support tools for performance, benchmarking, and diagnostics. As an ideal candidate, you will have We consider all candidates with an eye for those who are self taught, curious, and multi-faceted. Our ideal ATSE II candidate should have: 4+ years of relevant experience Strong understanding and grasp of the following areas Systems administration Troubleshooting Systems Scalable and Highly available distributed systems Network Administration Application Architecture Data architecture and design Understanding of database archit

javascriptpythonjava
View job →
M
Mongodb
📍 Dublin• Full-time
1mo ago

MongoDB Technical Services Engineers use their exceptional problem solving and customer service skills, along with their deep technical experience, to advise customers and to solve their complex MongoDB problems. Technical Service Engineers are experts in the entire MongoDB ecosystem - database server, drivers, cloud and infrastructure. This also includes services such as Atlas (database as a service), or Cloud Manager (which helps customers with automation, backup and monitoring of their MongoDB systems). Our engineers combine their MongoDB expertise with passion, initiative, teamwork and a great sense of humor to help our customers to be successful with MongoDB. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Cool things you’ll do You'll be working alongside our largest customers, solving their complex challenges - resolving questions on architecture, performance, recovery, security, and everything in between. You'll be an expert resource on best practices in running MongoDB at scale, whatever that scale may be. You'll be an advocate for customers' needs - interfacing with our product management and development teams on their behalf. And you'll contribute to internal projects, including software development of support tools for performance, benchmarking, and diagnostics. What you need We consider all candidates with an eye for those who are self-taught, insatiably curious, and multi-faceted. The ideal candidates should have strong technical experience in one (or more) of the following areas Systems administration Distributed systems Network Administration Database architecture and administration Application Architecture Data architecture and design Performance tuning and benchmarking Extra bonus points if you have experience in one or more of Java, Python, Ruby, C, C++, C#, Javascript, node.js, Go, PHP, or Perl If you have an operations background, we prefer experience administering large-scale production environments, i

javascriptpythonjava
View job →
M
Mongodb
📍 Mexico City• Full-time
1mo ago

MongoDB Technical Services Engineers use their exceptional problem solving and customer service skills, along with their deep technical experience, to advise customers and to solve their complex MongoDB problems. Technical Service Engineers are experts in the entire MongoDB ecosystem - database server, drivers, cloud and infrastructure. This also includes services such as Atlas (database as a service), or Cloud Manager (which helps customers with automation, backup and monitoring of their MongoDB systems). Our engineers combine their MongoDB expertise with passion, initiative, teamwork and a great sense of humor to help our customers to be successful with MongoDB. We are looking to speak to candidates who are based in Mexico City for our hybrid working model. Cool things you’ll do You'll be working alongside our largest customers, solving their complex challenges - resolving questions on architecture, performance, recovery, security, and everything in between. You'll be an expert resource on best practices in running MongoDB at scale, whatever that scale may be. You'll be an advocate for customers' needs - interfacing with our product management and development teams on their behalf. And you'll contribute to internal projects, including software development of support tools for performance, benchmarking, and diagnostics. What you need We consider all candidates with an eye for those who are self-taught, insatiably curious, and multi-faceted. The ideal candidates should have strong technical experience in one (or more) of the following areas Systems administration Distributed systems Network Administration Database architecture and administration Application Architecture Data architecture and design Performance tuning and benchmarking Extra bonus points if you have experience in one or more of Java, Python, Ruby, C, C++, C#, Javascript, node.js, Go, PHP, or Perl If you have an operations background, we prefer experience administering large-scale production environmen

javascriptpythonjava
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Research Program Management team partners with researchers and engineers to advance the development of increasingly capable, safe, and beneficial AI systems. We work alongside teams developing our core models, helping turn ambitious research goals into coordinated execution across model training, alignment and safety, and research infrastructure. The team also regularly collaborates with our closest cross-functional partners such as Security, Applied product and engineering, Strategy, and Scaling. About the Role As a Research Program Manager, you will embed with research teams and help drive some of the most technically complex and consequential work behind OpenAI’s model development. Depending on your focus, your work may span training, reasoning, evaluations, compute, research infrastructure, safety, model launch readiness, and governance. You will translate evolving research priorities into actionable programs, help teams navigate technical and operational tradeoffs, and keep important work moving as new issues emerge. This is a hands-on technical role: you will engage directly with research workflows, experimental results, technical systems, and engineering constraints; not simply coordinate from the sidelines. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. You might thrive in this role if you: Have 5+ years of experience in research program management, technical program management, or related roles in fast-moving environments. Can engage substantively with researchers and engineers on topics such as model training, experimental design, model safety, evaluation methods, data workflows, compute infrastructure, or distributed systems. Are comfortable working directly with technical tools, research data, experimental results, or operational workflows to understand problems and develop practical solutions. Have a strong track record of movi

awsrestai
View job →

As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental research problems and collaborating with Datadog's product and engineering teams to translate research advances into products. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Conduct research in generative AI and machine learning, building specialized foundation models and trained agents for observability Train multimodal models on large-scale, diverse telemetry data (metrics, logs, traces, topology, events) using distributed training infrastructure Design and build simulated environments and RL training loops for on-policy agent training and evaluation Collaborate with cross-functional teams (Product, Engineering) to integrate capabilities like multimodal world modeling and autonomous agents into Datadog's products Stay at the forefront of foundation models, world models, and RL-based agent research Contribute to r

gitmachine learningai
View job →
D
Datadog
📍 New York• Full-time• From $320K/yr
1mo ago

As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental research problems and collaborating with Datadog's product and engineering teams to translate research advances into products. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Conduct research in generative AI and machine learning, building specialized foundation models and trained agents for observability Train multimodal models on large-scale, diverse telemetry data (metrics, logs, traces, topology, events) using distributed training infrastructure Design and build simulated environments and RL training loops for on-policy agent training and evaluation Collaborate with cross-functional teams (Product, Engineering) to integrate capabilities like multimodal world modeling and autonomous agents into Datadog's products Stay at the forefront of foundation models, world models, and RL-based agent research Contribute to r

gitmachine learningai
View job →

Observability Pipelines (OP) is Datadog's on-premise, vendor-agnostic telemetry pipeline product. As an Engineering Manager on the team, you'll own people management and engineering execution for one of OP's core missions, spanning areas like Integrations (ingesting from and routing to the many source and destination systems customers rely on), streaming insights, cost control, or pipeline capabilities, reliability and scalability. You'll partner directly with Product to help shape the roadmap, and work closely with your peer EMs and senior ICs to define how OP operates and grows. This is an opportunity to build your management craft while having real influence over the technical direction of a fast-growing product area. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Own people management and engineering execution Establish a strong operating rhythm for the team Drive high standards for on-call rotations and incident response Partner with Product on the roadmap, balancing product priorities with technical realities Lead, coach, and grow the careers of engineers on your team Who You Are: Experienced managing engineers directly, comfortable owning a team’s operating rhythm end-to-end, from planning through execution and stakeholder communication to incident and on-call ownership Have a technical background in distributed systems and data infrastructure Have experience with on-premises or customer-installed software concepts A product-minded partner to have on the team — you enjoy working with Product on strategy Experience with high-performance or Rust-based data pipeline systems Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications o

aigorust
View job →

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<

pythonawsazure
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the team The Applied team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the role: We're seeking a Data Engineer to take the lead in building our data pipelines and core tables for OpenAI. These pipelines are crucial for powering analyses, safety systems that guide business decisions, product growth, and prevent bad actors. If you're passionate about working with data and are eager to create solutions with significant impact, we'd love to hear from you. This role also provides the opportunity to collaborate closely with the researchers behind ChatGPT and help them train new models to deliver to users. As we continue our rapid growth, we value data-driven insights, and your contributions will play a pivotal role in our trajectory. Join us in shaping the future of OpenAI! In this role, you will: Design, build and manage our data pipelines, ensuring all user event data is seamlessly integrated into our data warehouse. Develop canonical datasets to track key product metrics including user growth, engagement, and revenue. Work collaboratively with various teams, including, Infrastructure, Data Science, Product, Marketing, Finance, and Research to understand their data needs and provide solutions. Implement robust and fault-tolerant systems for data ingestion and processing. Participate in data architecture and engineering decisions, bringing your strong experience and knowledge to bear. Ensure the security, integrity, and compliance of data according to industry and company standards. You might thrive in this role if you: Have 3+ years of experience as a data engineer and 8+ years of any software engineering experience(including data engineering). Proficiency in at least one programming language commonl

pythonjavaaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world. Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure. About the Role We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology. You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more. There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future. In this role, you will: Build, lead, and grow high-performing infrastructure engineering teams. Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search. Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach. Champion pragmatic use of agent technology to amplify execution velocity. Reduce operational toil and incident frequency through better abstractions, gua

awsrestai
View job →
M
Mongodb
📍 United States• Full-time• From $151K/yr
1mo ago

Join the MongoDB Server Query Execution team, and help us build a world-class distributed open-source database. Our team plays a crucial role in the performance and efficiency of MongoDB's data processing. We are responsible for building and improving the core execution engine that powers all queries, taking a logical query plan produced by the optimizer and turning it into reality. This includes developing the physical operators for data retrieval and manipulation, improving the runtime for complex analytical and transactional workloads, and owning critical components such as our new execution engine. In addition to the core server, we support the query execution needs of other major products like Atlas Streams, Atlas Search and Vector Search, and mongosync, making our work vital to the entire MongoDB ecosystem. You will be joining a globally distributed team with a significant presence in both North America and Europe. While this role is based in the NAMER region, you will regularly collaborate closely with colleagues across different time zones. We support both office-based work in our North America hubs like New York, as well as remote work. We have tons of interesting problems to solve with a direct impact on users for transactional, time-series, and analytical workloads. To meet the ever-increasing data demands of modern applications, we are actively evolving our query system; this includes strategically re-architecting and improving key components of our query execution engine. We need your help to design and build the core of a distributed, flexible schema document database. This role can be based out of one of our North America offices, such as NYC or Palo Alto, or remotely across North America. Candidate Profile 10+ years of hands-on, professional experience in query engine development or database internals Experience with building production-level code with a large user base, robust design structure and rigorous code quality Degree in Computer Science or

mongodbawsazure
View job →
E
15 days ago

We build and operate the compute infrastructure our researchers run on, supporting large-scale processing of historical market data and model training on our own hardware across multiple data centers. Our environment includes bare-metal Linux, virtualization, storage, and GPU clusters, where performance, reliability, and predictable system behavior are critical. Our Infrastructure team covers monitoring and automation, distributed storage, hardware and OS provisioning, GPU clusters and workload scheduling, high-speed networking, L2/L3 Linux support, and security engineering. Engineers here own their tasks end to end, so there's room to go deeper in your area and pick up the parts you haven't touched yet. We’re looking for a Linux Infrastructure Engineer who can work hands-on with server and cluster environments, from deployment and configuration to performance tuning, troubleshooting, and ongoing improvement What You’ll Be Doing: Deploying, configuring, and maintaining Linux-based bare-metal servers across our data centers Building and operating clustered environments, including virtualization, storage, GPU compute, and database clusters Troubleshooting complex Linux, hardware, networking, and cluster-level issues Performance tuning for throughput, latency, stability, and resource utilization Monitoring infrastructure health and performance, identifying bottlenecks, and preventing recurring issues Supporting the full server lifecycle: provisioning, setup, upgrades, and maintenance Improving reliability and predictability during failures, maintenance, and scaling Automating provisioning, configuration, and operational tasks, primarily using Ansible and scripting What We Look For In You: Strong hands-on Linux administration and troubleshooting experience Production experience with on-premise, bare-metal infrastructure Good understanding of Linux performance and bottleneck analysis Experience with: infrastructure monitoring and troubleshooting production issues,

REMOTElinuxansible
View job →
🔔

Get new distributed systems engineer data platform delivery database retrieval jobs by email

Daily job updates · Unsubscribe anytime