Jobs in United States

Distributed Systems Engineer Data Platform Delivery Database Retrieval in United States

432 active opportunities · Updated October 2026

Explore current distributed systems engineer data platform delivery database retrieval jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $153.1K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Agent Infra team builds the infrastructure that runs autonomous AI agents at Roblox for millions of creators building experiences on our platform, and for thousands of Roblox engineers shipping production code every day. Agent workloads break the assumptions most infrastructure is built on. They run for hours instead of milliseconds, they're non-deterministic, they execute code nobody wrote, and they need real credentials against real systems to be useful at all. Making that safe, durable, and cost effective at Roblox’s scale is the hard problem our team is solving. As an early member of the Agent Infra team you’ll work alongside an experienced engineering team on systems that are already in front of real users, at scale, in a field that didn’t exist two years ago. You Are Early in your career : 1-3 years of professional experience or a recent grad; with strong CS fundamentals and an interest in learning infrastructure, distributed, and agentic systems. Someone who builds things: You have projects, internships, or open-source work you'd genuinely enjoy walking us through. Into AI agents. You've built something with them, even something small, and you can tell us what broke. Comfortable

PythonAWSGitRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Post-Training Frontiers team is responsible for training the frontier agents OpenAI ships to the world (GPT-Next). We train the flagship agentic models behind Codex, ChatGPT, and the API through large-scale reinforcement learning. The team’s work spans four areas. First, execution and science: working with teams across OpenAI to decide what can go into the final model and how, using scientific experiments and evals that are representative of the final pipeline so issues can be recognized early. Second, RL scaling: executing the final large-scale reinforcement learning run, making sure GPUs are used efficiently and training stays healthy. Third, research: improving horizontal capabilities like instruction following, factuality, memory, and multi-agent behavior, where the team’s broad visibility helps identify cross-cutting improvements across teams and domains. Fourth, engineering: maintaining the infrastructure stack and internal tools to ensure that both the final run and all integrations go as smoothly as possible and that the systems are easy to work with. About the Role This role focuses on keeping our frontier RL training runs fast, reliable, and unblocked. You will work across engineering and infrastructure problems as they emerge, from scaling and orchestration issues to inference bottlenecks, numerical problems, and hardware failures, as well as supporting large horizontal integrations in the big run, like multi-agent capabilities or memory. This is a role for a strong generalist who quickly learns anything needed for the task, has high attention to detail, debugs deeply, and is motivated by fixing the highest-impact problem in front of the team. In this role, you will: Keep large-scale async RL training runs moving by jumping into the most urgent engineering and infrastructure problems. Debug issues across training systems, inference, orchestration, scaling, and distributed infrastructure. Improve the reliability and efficiency of RL trai

AWSRestAIGo
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Workload Networking team is responsible for the collective communication stack used in our largest training jobs. Using a combination of C++ and CUDA we work on novel collective communication techniques that enable efficient training of our flagship models on our largest custom built supercomputers. The models we train are key ingredients to the AI research progress at OpenAI and the field as a whole, and we continually incorporate learnings from our entire research org into our training platform. About the Role As a Software Engineer, Networking you will design and implement custom networking collectives that are tightly integrated into our training stack. We’re looking for people who have a background in low level performance critical software. Experience with collective communication is a bonus. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Collaborate closely with ML researchers to design and implement efficient collective operations in C++ and CUDA. Ensure that our largest training jobs take full advantage of the different network transports used in our supercomputers. Work on simulations to inform our future supercomputer network designs. You might thrive in this role if you: Have written distributed algorithms using RDMA in the past. Are comfortable writing low level performance sensitive CPU and/or GPU code. Are familiar with network simulation techniques. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voic

AWSRestAIC++
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $192K/yr

Quick readStrong listing-quality and freshness signals

Coordination Systems provides foundational distributed systems building blocks for internal Datadog platforms. Our services cover sharding, consensus, resource protection, configuration distribution, and much more. We are looking for a manager to lead the Coordination Systems - Storage team. This team provides essential configuration storage and distribution systems that are depended upon by almost every service and pod at Datadog. We power critical runtime configuration (e.g. feature flags), complex control planes (e.g. dynamic sharding configuration), and much more. Storage is one of four subteams within Coordination Systems. If successful, the candidate will have opportunities to lead other growing and impactful areas such as Resource Protection. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: (Describe role responsibilities here/max 6 bullets) Lead a core team of 5 engineers (distributed, with majority in NYC) Lead ceremonies, prioritize and delegate project Stay hands-on with the code, e.g. isolated features, small remediations, investigation follow ups Stay actively involved in operations, incidents, root cause analysis, etc. Constantly promote a culture of operational excellence, organizing gamedays, conducting operational reviews, staying proactive with reliability Who You Are: (Describe role qualifications here/max 6 bullets) Strong distributed systems skills, able to understand and account for a variety of failure modes, well-versed in end-to-end o11y, validation testing, simulation setup, etc. Worked on platform teams before, providing critical infrastructure to internal stakeholders Experienced in handling significant incidents, both as a responder and follow-up ow

AIRustExcelSEM
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and inform next-generation infrastructure design. About the Role We are seeking Performance Modeling Engineers to develop and apply modeling tools that evaluate AI system performance and inform architectural decisions. In this role, you will work closely with the Performance Modeling Lead and partner teams to analyze system behavior, run simulations or analytical models, and help quantify tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks and applying them to real-world questions that impact system design and vendor decisions. This role is well-suited for engineers with strong software or modeling backgrounds who are interested in developing deeper expertise in system architecture and AI infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Develop and maintain performance modeling tools and frameworks. Build models to evaluate system behavior across: compute, memory, and interconnect subsystems distributed system scaling and bottlenecks. Run simulations and analytical models to support architectural tradeoff analysis. Collaborate with performance modeling lead and system architects to answer forward-looking design questions. Analyze and interpret modeling outputs, translating results into actionable insights. Validate models against real system measurements and workload behavior. Contribute to improving modeling fidelity, usability, and scalability. Qualifications Strong software engineeri

AWSRestAIRust
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Requirements: 5+ years of experience writing high-quality production code Experience building high-performance distributed systems at a large scale (the more battle scars, the better) Strong cloud skills Strong knowledge of low-level operating system foundations (Linux kernel, file systems, containers, etc.) Experience with performance engineering (tell us a story of when you shaved off a few milliseconds!) Ability to work in-person in our NYC or SF office. Prior experience with Rust is nice to have, but not required. Ability to participate in on-call rotation and respond to production incidents.

LinuxRestAIGo
C
📍 United States· Full-time· Remote
✓ Quality checkedCompany trend -100%

As an Engineering Manager on Coder’s Core Workspaces team, you’ll lead engineers building and evolving the systems behind our agentic development experience. You’ll help make agents more capable, reliable, and useful across real development environments. You’ll guide technical direction while growing the team and keeping execution sharp. You’ll work closely with Engineering, Product, and Design across the agent harness, integrations, and developer workflows. What you’ll do here Lead and grow a team within our Workspaces organization. Set technical direction across the agent harness, integrations, and workflows. Stay close to the code and contribute to architecture and implementation decisions. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve reliability, performance, and operability across agentic systems. Coach engineers, raise the technical bar, and create clarity around priorities and tradeoffs. What we’re looking for Experience managing and growing software engineering teams. Strong hands-on engineering experience with React and TypeScript . Experience with Go . Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS . Strong technical judgment and comfort working through ambiguity. A track record of helping engineers grow while maintaining a high execution bar. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP , agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute. Experience building abstractions across multiple model providers. Deep experience with AWS, Kube

TypeScriptReactAWSDocker
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an

KubernetesGitAIGo
S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.6%

From $220K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. As an Engineering Manager, you’ll lead a team of engineers owning critical workflows such as checkout and invoicing, while also developing new features to help customers manage their spend growth. In this role, you’ll partner across the organization to ensure our customers redeem everything Sentry has to offer and budget for future expansion. In this role you will Strategic Planning & Roadmap: Define and drive the team's roadmap. Align team goals with organizational objectives and contribute to the overall platform strategy. Technical Guidance & Operational Excellence: Provide technical leadership and guidance on complex distributed systems and design. Ensure the team is proactively identifying areas for improvement. Cross-functional Collaboration: Partner closely with business and technical teams to translate business goals into actionable objectives and scalable solutions. Team Leadership & Development: Lead, mentor, and grow a team of talented engineers, including Staff-level engineers. Build a culture of technical excellence, collaboration, continuous

JavaScriptTypeScriptPythonJava
D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%
Quick readStrong listing-quality and freshness signals

As a Senior Product Manager for Serverless at Datadog, you will define and deliver products that help developers monitor and operate serverless applications at scale. You’ll own the strategy and execution for Datadog’s AWS Serverless observability offering, building experiences that provide visibility into distributed systems and simplify debugging and operations. This role sits at the intersection of cloud infrastructure, developer experience, and AI-powered workflows, and is ideal for a PM who thrives in highly technical product areas. You will work cross-functionally and with external partners to shape how customers build and run modern serverless applications. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own and drive the roadmap for Datadog’s AWS Serverless observability products, including Lambda, Fargate, and Step Functions Define how developers monitor, debug, and operate serverless systems across distributed environments Partner with engineering and design to deliver end-to-end product capabilities from concept through launch and iteration Collaborate with AWS product teams to align roadmaps and deliver joint solutions for shared customers Engage with customers to understand serverless adoption patterns and validate product direction Define and track success metrics such as adoption, usage, and impact on developer workflows Who You Are: 5+ years of product management experience building technical products in areas such as cloud infrastructure, developer platforms, or observability Strong understanding of distributed systems, cloud-native architectures, and modern application development practices Familiarity with serverless technologies, containers, Kubernetes, or microservices environments Comfortable working closely with engineers and discu

AWSKubernetesMicroservicesAI
M
📍 San Francisco, California, United States
✓ Quality checkedCompany trend +212.5%

Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Vice President, Software Engineering Overview Decision Stream is Mastercard's next-generation AI-native decisioning platform, designed to power intelligent, real-time decisions across fraud, authentication, payments, and risk. Built with a startup mindset and enterprise-scale ambition, the platform combines innovation, speed, and engineering excellence to redefine decisioning across Mastercard's global ecosystem. We are seeking a visionary and hands-on VP of Software Engineering to help build and scale the platform. This leader will partner closely with Product, Architecture, AI, and Platform Engineering teams to drive technology strategy, architecture, engineering execution, and organizational growth. What You'll Do Lead Through Technical Excellence • Serve as a senior technology leader and role model for engineering teams. • Drive architecture, design, and technology decisions across distributed systems, streaming, AI, and cloud-native platforms. • Engage deeply with engineers, architects, and product leaders to solve complex technical challenges. • Influence engineering standards, software quality, and operational excellence. Build and Scale the Platform • Help shape and deliver a highly scalable, resilient, and secure decisioning platform operating at Mastercard scale. • Balan

KubernetesAIRecruitment
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $295.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Observability team builds the infrastructure that empowers engineers to understand, operate, and improve the Roblox platform and ecosystem. Our team owns the end-to-end observability stack across telemetry, distributed tracing, logging, profiling, storage systems, and developer-facing visualization tools. We are looking for an Engineering Manager to lead the next generation of AI-powered observability platforms. In this role, you will help build intelligent systems that leverage AI to revolutionize CI/CD, testing, and DevOps workflows — enabling engineers to move faster, improve reliability, and operate large-scale distributed systems with greater efficiency and confidence. This is a highly impactful leadership role at the center of Roblox infrastructure. Your work will directly improve developer productivity, platform reliability, and operational excellence across the company. You will partner closely with infrastructure, product engineering, and AI platform teams to shape the future of developer tooling and autonomous operations at scale. You Have 3+ years of engineering management experience with a proven track record of hiring, mentoring, and growing high-performing teams. Strong ex

AWSCI/CDGitAI
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

We are looking for a 100% hands-on Storage Services Software engineer to join the block storage group. You will be a member of a team that builds the next-generation block storage capabilities and architects a proprietary distributed file system solution from its inception. You will work closely with a variety of teams and architects including the networking team, and external customers. You will take part in defining the software architecture and implementation of the most advanced storage services! Services that will need to meet extreme performance and scalability demands! We have crafted a team of extraordinary people stretching around the globe, whose mission is to push the frontiers of what is possible today and define the platform of tomorrow. At NVIDIA, we work, think and learn as a team. We thrive in a deeply strong environment, and we're passionate about a culture that demands innovation and the highest standards. The rewards are sweet and include collaborating with some of the smartest people in the industry, an aggressive compensation plan that rewards top performers, and the opportunity to work on products that transform the way people work and play. What you’ll be doing: 100% hands-on coding role in C language, Kernel and Userspace Access advanced AI tools and a token budget for code development provided by NVIDIA, the world's AI factory leader. Research, design, implement and test, new and existing, distributed storage services and features of NVIDIA’s block and file storage solution, in both Host and DPU environments. Acquire understanding of the algorithms, the technicalities and the interaction with other components across NVIDIA’s block and file storage ecosystem. Analyze and solve challenging bugs and customer cases in la

L
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At Linear, we're building the product development system for teams and agents. AI is fundamentally changing how software gets built, and we’re shaping the tools this new era requires. Founded in 2019, Linear has become the platform of choice for more than 40,000 companies (including OpenAI, Coinbase, and Ramp) to plan, build, and ship their products. Today, our team is distributed across North America, Europe, and Australia, and we’re continuing to grow internationally. What unites us is relentless focus, fast execution, and a deep care for software craftsmanship. As a small team, we’re all generalists that work across the full stack (built in Typescript end-to-end). We’re looking for experienced engineers that thrive in an environment of autonomy and individual responsibility to help us build the future of product development. Location & work mode Linear is a remote-first company, with optional co-working offices in San Francisco, New York, and London. This role is open to candidates based in the US and Europe. You can work from anywhere within these regions. We value deep focus and async collaboration, with intentional moments to connect in person through team off-sites, optional co-working, and occasional travel. What you'll do Work closely with founders and design to implement new concepts and ideas Build AI-powered functionality into the core of Linear Update our realtime collaborative content editor used across all internal surfaces Build new user-facing features with beautiful and scalable UI components Obsessively improve application performance Refine our software development processes to keep the team operating at high velocity What we're looking for 5+ years of experience building customer-facing products at a high-quality software company Strong React and TypeScript fundamentals, with experience across the full stack (Browser technologies, Node, GraphQL, PostgreSQL) Track record of driving complex, end-to-end features (not just incremental improvemen

TypeScriptReactSQLPostgreSQL
L
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At Linear, we're building the product development system for teams and agents. AI is fundamentally changing how software gets built, and we’re shaping the tools this new era requires. Founded in 2019, Linear has become the platform of choice for more than 40,000 companies (including OpenAI, Coinbase, and Ramp) to plan, build, and ship their products. Today, our team is distributed across North America, Europe, and Australia, and we’re continuing to grow internationally. What unites us is relentless focus, fast execution, and a deep care for software craftsmanship. As a small team, we're all generalists that work across the full stack (built in TypeScript end-to-end). We’re looking for engineers that thrive in an environment of autonomy and individual responsibility to help us build the future of product development. Location & work mode Linear is a remote-first company, with optional co-working offices in San Francisco, New York, and London. This role is open to candidates based in North America. You can work from anywhere within this region. We value deep focus and async collaboration, with intentional moments to connect in person through team off-sites, optional co-working, and occasional travel. What you'll do Work closely with founders, product, and design to implement new concepts and ideas Build AI-powered functionality into the core of Linear Update our collaborative content editor used across all internal surfaces Build new user-facing features with beautiful and scalable UI components Obsessively improve application performance Refine our software development processes to keep the team operating at high velocity What we're looking for 2-5 years of experience building customer-facing products at a software company with a high engineering bar Strong React and TypeScript fundamentals, with experience across the full stack (Browser technologies, Node, GraphQL, PostgreSQL) Track record of driving complex, end-to-end features (not just incremental improvement

TypeScriptReactSQLPostgreSQL
🔔

Get new distributed systems engineer data platform delivery database retrieval jobs in United States by email

Daily job updates · Unsubscribe anytime