Jobs in United States

Distributed Systems Engineer Data Platform Delivery Database Retrieval in United States

432 active opportunities · Updated October 2026

Explore current distributed systems engineer data platform delivery database retrieval jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 New York, NY, United States· Full-time
✓ Quality checkedCompany trend -99.2%

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role The Credit Engineering team builds the systems that enable Ramp to manage credit lifecycles for businesses end-to-end, from initial underwriting through ongoing portfolio strategy optimization. As an engineer on Credit, you’ll own systems that influence $100+ billion in annual payment volume through millions of decisions across our product stack in support of the most ambitious FinTech portfolio in the United States. We are looking for a backend engineer to own the technical roadmap for underwriting, limit management, portfolio optimization, and agentic workflows. You will build robust systems, automate complex operational tasks, and maintain the high-frequency controls that power all of Ramp’s products. If you are obsessed with correctness, excited by ambiguity, and passionate about problems at the intersection of distributed systems and financial strategy, this role is for you. What You'll Do Architect scalable, stateful systems for automated limit management to optimize Ramp’s charge card portfolio. Envision and build the next gene

PythonRestAIGo
R
📍 New York, NY, United States· Full-time
✓ High-confidence listingCompany trend -99.2%

From $10K/yr

Quick readStrong listing-quality and freshness signals

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role Ramp is, at its core, a fintech company. Our financial infrastructure enables us to issue cards, move money globally, and manage treasury flows for our customers. Stablecoins are emerging as a critical lever to make these flows faster, cheaper, and more accessible. We are looking for a Software Engineer to join our engineering team with a dedicated focus on stablecoin-based payment and treasury systems. In this role, you will act as both a technical owner and multiplier: working on core stablecoin services, ensuring secure and reliable integrations with external partners, and guiding Ramp’s evolution toward next-generation fintech solutions. Our ideal candidate combines deep curiosity about eventually consistent distributed systems with strong financial infrastructure experience, thrives in highly regulated and mission-critical environments, and has a passion for diving into new problem spaces. What You'll Do Help build & scale Ramp’s stablecoin-based payment and treasury infrastructure Partner with external providers and internal

PythonJavaAWSAzure
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Monetization team is a cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products, including next-generation ads experiences that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to build measurement systems that connect ad interactions to meaningful advertiser outcomes while protecting user privacy. In this foundational role, you’ll design infrastructure for conversion signals, attribution, reporting, and feedback loops across OpenAI’s ads products. This role is ideal for engineers who have built large-scale ads measurement, data, experimentation, marketplace, or distributed systems and want to apply that experience in a highly ambiguous 0→1 environment. You’ll work across event collection and normalization, deduplication and matching, attribution and modeled measurement, privacy-safe aggregation, reporting, and high-quality labels for ads optimization. We are hiring engineers who can independently own complex systems, make sound technical tradeoffs, and help define what should be built. You’ll work closely with Ads Delivery, Ads ML, Product, Research, Privacy, Data Sc

AWSRestAIGo
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Software Engineering team is responsible for designing and building the scalable, performant, and secure backend systems that power our products—from early prototypes to large-scale deployments. We collaborate closely with product, hardware, and full-stack teams to ensure our infrastructure enables fast iteration while setting a strong foundation for long-term growth. About the Role As a Backend Engineer , you will design and build services, APIs, and infrastructure that support evolving product needs. You’ll apply a deep understanding of backend systems and maintain enough end-to-end context—from hardware to cloud—to guide technical decisions that best serve the product and team. We’re looking for engineers who thrive in fast-paced, collaborative environments and care deeply about building robust systems that scale. This role is based in San Francisco, CA . We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Architect, build, and maintain high-performance, secure backend systems. Design APIs, data models, and infrastructure to support evolving product needs. Balance near-term development velocity with long-term maintainability and scalability. Collaborate with cross-functional teams to ensure cohesive, end-to-end solutions. You might thrive in this role if you: Have 7+ years of professional software engineering experience, with a focus on backend systems. Have a proven track record of building and scaling systems from early stage to large scale. Are proficient with Python and Go, and familiar with a range of server-side technologies. Have a strong grasp of system design, performance optimization, and security best practices. Can reason about full-stack tradeoffs from hardware through cloud infrastructure. (Nice to have) Have experience with distributed systems and cloud architectures. (Nice to have) Bring a background in instrumentation, analytics, and performanc

PythonAWSRestAI
S
📍 Bellevue, Washington, United States· Full-time
✓ High-confidence listingCompany trend -93.3%
Quick readStrong listing-quality and freshness signals

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake is now building a world class OLTP service on Postgres. We’re hiring Senior Postgres Engineers to help build the new Snowflake Postgres service. This is an exciting opportunity to help build a large scale, multi-cloud Postgres offering with access to an existing customer base. AS A SENIOR POSTGRES ENGINEER YOU WILL: Develop orchestration layer/ control-plane of large scale databases Work with AWS, Azure, and GCP APIs Build High Availability and Disaster Recovery solutions Tune Postgres to operate at scale for some of the largest datasets in the world Secure and ensure customer data is protected Work alongside some of the brightest minds in the industry and help redefine the space by inventing at various layers of the software stack — from broad distributed systems to low-level optimizations that fundamentally change the performance characteristics of the system to cater to our customer needs. OUR IDEAL SENIOR POSTGRES ENGINEER WILL HAVE: 7+ years industry experience designing, building and developing large scale systems in production Experience building and maintaining distributed, highly available, fault tolerant services Excellent understanding of low level operating systems concepts including multi-threading, memory management, networking, storage, performance,

PythonJavaPostgreSQLAWS
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? The Data Infrastructure team at Cohere is responsible for the storage and data movement layer underlying every model training run. We're building the unified storage layer that feeds our training workloads. It needs to serve petabytes of training data and model checkpoints fast enough to keep thousands of GPUs busy across several training clusters. In this role, you’d have an opportunity to build this system from the ground up. You’d be a key contributor, working on a problem few teams have had to solve at this scale. In this role, you will: Design, build, and operate the distributed storage system that feeds model training and evaluation. Run this system multiple on Kubernetes clusters at petabyte scale. Work with researchers and training-infra teams on how jobs actually read and write data, and turn that into throughput, latency, and durability requirements Work through the networking, I/O, and consistency problems of moving large datasets and checkpoints across regions and backends, with GPU idle time and time-to-insight as the measures of success You may be a good fit if you have: Strong storage fundamentals,

PythonKubernetesGitRest
H
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Advanced Engineer, Power System Controls plays a key role in the design, development, and implementation of control software for Hyliion’s Karno Power Module. This position focuses on systems involving combustion, thermal management, pressure regulation, high-voltage, and power management. The engineer will be responsible for developing and validating control algorithms, tuning system parameters, and analyzing data to ensure performance meets engineering specifications. Additional responsibilities include preparing technical documentation, supporting root cause analysis, and ensuring timely, high-quality software delivery. The role requires cross-functional collaboration and occasional travel to support system testing and troubleshooting. Duties and Responsibilities Design, develop, and implement high-quality control software for Hyliion’s Karno Power Module, which includes combustion, thermal, pressure, high-voltage and power management systems. Define and conduct tests to verify software and tune control parameters to meet key performance indicators. Prepare reports and technical documentation related to system performance, control strategies, and compliance. Process and analyze data to verify software against engineering specifications, support root cause analysis and for optimizing performance. Ensure on time delivery with quality. Assist product team in defining customer requirements and generate corresponding engineering specifications. Qualifications Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions. Qualifications include: Education, Experience and Cert

PythonAIC++Go
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

OpenAI’s charter calls on us to ensure the benefits of AI are distributed broadly and safely. Our Health AI team focuses on expanding access to high-quality medical expertise and aims to set a high standard for deploying AI responsibly in high-stakes domains. Improving health will be one of the defining impacts of AGI. Today, millions of people lack access to reliable medical information, and clinicians around the world face increasing time and resource constraints. We are building AI systems that support patients, clinicians, and health workers, while meeting the highest standards for safety, reliability, and privacy. We are seeking full stack software engineers to help build and scale products used by consumers and care providers globally. You will work closely with product, design, and research teams to ship real systems in a fast-moving, high-impact environment. In this role, you will: Design and build scalable fullstack systems for consumer and enterprise health. Own end-to-end feature development—from early design and implementation through deployment, monitoring, and iteration. Build and maintain data pipelines and services that meet strict privacy, security, and compliance requirements (e.g., HIPAA). Collaborate closely with researchers and safety teams to integrate reliability, evaluation, and guardrails into production systems. Debug, optimize, and harden systems to support high availability, performance, and global scale. Take ownership of ambiguous problems and drive them to practical, high-quality solutions. You might thrive in this role if you: Are deeply motivated by improving health outcomes and expanding access to medical expertise. Are a strong engineer who enjoys building durable, well-designed systems. Have 5+ years of experience writing maintainable, production-quality code. Can operate with high agency—owning problems end-to-end with minimal supervision. Enjoy working in fast-moving, cross-functional teams with engineers, product managers, desi

AWSGitRestAI
S
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.6%

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role At Sentry, Support is an engineering discipline. Our customers are the greatest technical minds in the world—developers at elite enterprises building the future of software—and they deserve answers that go deeper than a knowledge base link. We're looking for an APAC Technical Support Engineer based in San Francisco to join our global Support Engineering team. This role is designed to provide APAC coverage to our users; with the shift being Sunday through Thursday 4PM-12AM PST. We are architecting the Technical Support engine . We’re looking for an experienced engineer to help us redefine the standard of technical support by combining deep human expertise with autonomous agentic systems. You are a debugger of both code and systems. You will treat support volume as a data signal to build automated resolution paths, ensuring our human engineers only touch the most complex, high-impact architectural puzzles. Sentry Support Engineers aren't just clearing queues; they are Orchestrators . You will engage with our users across GitHub, Discord, and our internal systems, while acting as the Technical Lead for our Agentic Ops. You ensure that when a developer asks a complex question, our systems have the right context and a seamless "Human-in-the-Loop" path to you when deep, nuanced expertise is required. In this role you will Master the Sentry Ecosystem & Support Elite Developers Deep-Dive Debugging: Perform root-cause analysis on complex issues and distributed tracing gaps across polyglot environments. Support the Great Minds: Act as a strategic consultant for senior engineers at our largest enterprise customers, s

JavaScriptPythonJavaReact
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

PythonAWSLinuxRest
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i

PythonReactNode.jsAngular
H
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Manager, Electrical Engineering is responsible for the electrical systems of the KARNO generator, including high-voltage power electronics, battery systems, low- and high-voltage architecture, wiring harnesses, and the hardware that converts linear motion into electrical output. This is a working manager role: the position leads and develops a team of electrical engineers while remaining directly involved in technical execution, including circuit architecture, schematic review, and hardware bring-up in the lab. The Manager is accountable for the technical excellence, safety, and reliability of the electrical engineering function, and for establishing the design standards and review practices the team works to. The position plans team capacity, owns hiring and development for the electrical engineering staff, and partners with mechanical, controls, supply chain, and program management on system integration. KARNO systems are deployed in data center, military, and industrial applications. AI at Hyliion At Hyliion, AI is core to how we work. We equip every team member with leading AI tools and count on you to use them — to move faster, solve harder problems, and help us realize the full potential of KARNO technology for the world. Duties and Responsibilities Own critical electrical designs personally, including regular time at the bench and in the test cell, while leading the team as a practicing engineer. Lead the electrical engineering team in the design and development of KARNO generator electrical systems, including high-voltage power electronics, battery systems, linear generator power stages, and low-voltage controls hardwar

AIGoExcelSEM
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI’s Research Program Management team partners with researchers and engineers to advance the development of increasingly capable, safe, and beneficial AI systems. We work alongside teams developing our core models, helping turn ambitious research goals into coordinated execution across model training, alignment and safety, and research infrastructure. The team also regularly collaborates with our closest cross-functional partners such as Security, Applied product and engineering, Strategy, and Scaling. About the Role As a Research Program Manager, you will embed with research teams and help drive some of the most technically complex and consequential work behind OpenAI’s model development. Depending on your focus, your work may span training, reasoning, evaluations, compute, research infrastructure, safety, model launch readiness, and governance. You will translate evolving research priorities into actionable programs, help teams navigate technical and operational tradeoffs, and keep important work moving as new issues emerge. This is a hands-on technical role: you will engage directly with research workflows, experimental results, technical systems, and engineering constraints; not simply coordinate from the sidelines. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. You might thrive in this role if you: Have 5+ years of experience in research program management, technical program management, or related roles in fast-moving environments. Can engage substantively with researchers and engineers on topics such as model training, experimental design, model safety, evaluation methods, data workflows, compute infrastructure, or distributed systems. Are comfortable working directly with technical tools, research data, experimental results, or operational workflows to understand problems and develop practical solutions. Have a strong track record of movi

AWSRestAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $320K/yr

Quick readStrong listing-quality and freshness signals

As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental research problems and collaborating with Datadog's product and engineering teams to translate research advances into products. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Conduct research in generative AI and machine learning, building specialized foundation models and trained agents for observability Train multimodal models on large-scale, diverse telemetry data (metrics, logs, traces, topology, events) using distributed training infrastructure Design and build simulated environments and RL training loops for on-policy agent training and evaluation Collaborate with cross-functional teams (Product, Engineering) to integrate capabilities like multimodal world modeling and autonomous agents into Datadog's products Stay at the forefront of foundation models, world models, and RL-based agent research Contribute to r

GitMachine LearningAIGo
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<

PythonAWSAzureGCP
🔔

Get new distributed systems engineer data platform delivery database retrieval jobs in United States by email

Daily job updates · Unsubscribe anytime