Jobs in United States

Distributed Systems Engineer Data Platform Delivery Database Retrieval in United States

432 active opportunities · Updated October 2026

Explore current distributed systems engineer data platform delivery database retrieval jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

We’re looking for a Staff Software Engineer to help shape AI governance for developer tooling at Coder. This role sits on our AI Governance team, which builds and maintains two enterprise-grade components of Coder's AI governance stack. AI Gateway is a centralized LLM gateway that sits between coding agents and providers such as OpenAI or Anthropic, providing organizations with audit trails, token tracking, cost control, and centralized authentication. Agent Firewall wraps those agents with default-deny network policies, controlling which domains and methods they can reach inside workspaces. This team works across the full stack - from Go backend and React frontend to integrating with LLM provider APIs. Day to day, you'll be shipping features, hardening security boundaries, collaborating with enterprise customers on real-world policy needs, and contributing to Coder's open-source ecosystem. What you'll do here Design and build product features that push the standard for remote development in self-hosted environments Create and improve upon popular open source projects that integrate with VS Code, JetBrains, and other developer tools Champion best practices to both internal team members and external contributors Collaborate with Product and Design teams at Coder, as well as with partners like JetBrains, to execute key product integrations Document the design, implementation, and operations of systems for knowledge sharing within the team Work alongside Customer Success teams to support Coder’s enterprise user base Work with cutting-edge AI technologies to create seamless, painless developer experiences Rapidly iterate from prototype to implementation in a highly adaptive, reactive team environment What we're looking for 8+ years of full-stack experience writing code in a professional setting, with 1+ year(s) writing Go (ideally in current or most recent position) Proficiency in building distributed systems in Go Excellent verbal and written communication skills Excep

TypeScriptReactAWSGCP
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $196.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a senior software engineer on the Cell Platform team at Roblox, you will build systems that Roblox engineers use to create and deploy resources onto Kubernetes. Our engineers deploy their services in a complex, hybrid, multiple-cluster, K8s environment. The Cell Platform manages this complexity for our users, with tools, APIs, K8s controllers, and UX, simplifying infrastructure for our internal customers. You Have: A desire to work on critical, large-scale distributed systems An appreciation of observability and instrumentation and tooling to make your life easier 3+ years of experience as software engineer Bachelor's degree in Computer Science or an equivalent field You will: Build our Roblox-wide control plane using Kubernetes primitives (and plenty of custom resources) Work on the interface of the few hundred person infrastructure organization to the thousands of Roblox engineers Write and review high quality code and tests (largely Golang) Work on a team that cares about inclusivity and shipping For roles that are based at our headquarters in San Mateo, CA: The starting base pay for this position is as shown below. The actual base pay is dependent upon a variety of job-related

AWSKubernetesGitAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Communications team, you'll own the backend systems that power chat and messaging for millions of players on Roblox and support billions of messages per day. Your day-to-day will focus on building and scaling the infrastructure that makes in-game communication fast, safe, and reliable. That means driving architectural decisions, creating communication modes that don't exist yet, and defining developer-facing APIs so creators can build richer experiences. You'll work closely with frontend engineers, platform, product, and design, and you'll have real influence over the technical and product direction of the team. If you're an experienced engineer who gets excited about large-scale distributed systems and wants to shape the way millions of people connect inside virtual worlds, we'd love to talk. You Will Build and scale backend infrastructure for millions of concurrent users, architecting high-throughput distributed services, integrating ML models, and enabling rapid product experimentation Equip creators with the APIs and tools to build deeply integrated social experiences in their games Own projects end-to-end: from design and architecture through produc

SQLPostgreSQLMySQLAWS
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role We’re building the observability product for OpenAI—from scalable infrastructure to a rich, AI-powered UI. Our systems ingest over petabytes of logs and billions of time series metrics across our fleet. We're now layering intelligence on top—think agents that summarize SEVs, auto-generate dashboards, or help engineers debug through notebook-like UIs. We’re hiring software engineers across the stack—infra, backend, and product. You’ll join a small, gritty team building both foundational infra and novel internal tools to make OpenAI's production systems reliable, performant, and observable. What You’ll Do Own core observability infrastructure, including distributed logging, time series, and trace storage Build AI-native tools that help engineers detect, understand, and resolve issues autonomously. Contribute to UI experiences like dashboards, notebooking, or interactive debugging Collaborate closely with engineers, researchers, user ops, and other teams across the company to build the next generation observability product You Might Be a Fit If You: Have operated large-scale distributed systems in production. ( especially logging systems or some other time series databases) Thrive in ambiguous environments and roll up your sleeves to solve unscoped problems. Have full-stack chops or product sensibilities—you're excited to build real tools people use. Have strong fundamentals in systems, networking, and cloud infra (Kubernetes, AWS, etc). Bonus : built or contributed to observability systems (e.g. Prometheus, OpenTelemetry, etc). Why This Team We’re b

AWSKubernetesRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Cloud Agents team builds product infrastructure for long-running agents in the cloud: orchestration, sandboxing and isolation, secure environment connectivity, secrets and identity, observability, reliability, and cost controls. These agents securely connect to diverse developer and customer environments and use tools to accomplish goals. We partner closely with product, research, and infrastructure teams to turn agentic capabilities into dependable platforms for OpenAI products and developers building on OpenAI. About the Role We are looking for an experienced software engineer to help build and scale our cloud agent platform. You will design and operate systems for orchestrating agents at scale. You will work closely with product engineers on ChatGPT, API, and Codex to define the right abstractions and enable them to ship products quickly. Strong backend or infrastructure experience is important; experience with Python, Rust, distributed systems, cloud infrastructure, or product platforms is especially helpful. In this role, you will: Design and scale the orchestration, sandboxing and storage systems that run agentic workloads for Codex, ChatGPT, and the OpenAI API. Partner with product engineers to build a platform that enables them to ship quickly and turn feedback into robust abstractions. Improve reliability, security, performance, and cost efficiency for long-running agents. Deploy services that can operate across different environments and clouds. Your background might look something like: 9+ years of professional engineering experience, excluding internships, in relevant roles at technology and product-driven companies. Experience leading large-scale backend, platform, or infrastructure projects from ambiguous problem statements to production systems. Proficiency in one or more backend languages such as Python, Go, Rust, TypeScript, or similar, and the ability to move across service, platform, and product boundaries. Strong understanding

TypeScriptPythonAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun

AWSAzureGCPKubernetes
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -93.3%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Product Security team ensures that Snowflake products are built and shipped with the highest level of security. Our team drives the security posture of Snowflake products and is responsible for embedding security into every stage of the product lifecycle, from design through deployment and beyond. We design and build frameworks, systems and services that keep Snowflake secure. As a Principal Software Engineer II on the Product Security team, you will be the senior technical authority for Product Security and play a critical leadership role in shaping and advancing Snowflake’s security. This is a unique opportunity to define and influence our long-term security strategy and have a direct impact on the security of the Snowflake platform and the trust of our customers. You will operate across organizational boundaries, guiding major security initiatives, influencing architectural decisions at the highest levels, setting the technical direction for the organization, and ensuring consistent security excellence across all product teams while working closely with business leaders to advance Snowflake’s business. The role requires deep expertise in security, software engineering, distributed systems, software infrastructure, AI/ML, applied cryptography, threat modeling and clou

PythonJavaAIC++
C
📍 Dallas 8000 Frankford Road, United States
✓ Quality checkedCompany trend +340.2%

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary We are seeking a highly experienced Principal / Director-level Full-Stack Software Development Engineer to join our Digital Caremark organization and lead the architecture, design, and delivery of next-generation digital solutions. This role spans AI-enabled applications, scalable digital platforms, and enterprise integrations that power critical healthcare and pharmacy experiences. This is a senior technical leadership role for a hands-on engineer who can operate across the full stack—from intuitive front-end applications to resilient backend services—while setting architectural direction, influencing engineering standards, and mentoring teams. The ideal candidate combines deep technical expertise, platform thinking, and strong collaboration skills to build secure, scalable, API-first solutions in a highly regulated environment. Key Responsibilities: Architecture & System Design • Define and drive architecture for large-scale distributed systems and digital platforms • Lead design reviews and set architecture standards and best practices • Champion API-first, microservices, and event-driven architecture patterns • Ensure systems meet scalability, reliability, security, and compliance requirements • Balance performance, cost, and speed in technical decision-making Fu

TypeScriptPythonJavaReact
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role We are looking for an engineer who wants to take the world's largest and most capable AI models and optimize them for use in a high-volume, low-latency, and high-availability production and research environment. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Work alongside researchers to enable advanced research through awesome engineering. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 5 years of professional software engineering experience. Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc. Have experience architecting, building, observing, and debugging production distributed systems. Bonus point if worked on performance-critical distributed systems. Have need

AWSAzureRestMachine Learning
P
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity The Forward Deployed Engineering (FDE) team tackles some of Postman's most strategic technical challenges. We partner directly with enterprise customers to solve problems that don't fit neatly into existing product boundaries, building production-grade solutions that often become core product capabilities. As a Sr. Forward Deployed Engineer, you'll operate at the intersection of engineering, product, and customer success. You'll deploy Postman's critical infrastructure into customer environments, solve complex distributed systems challenges, and build the enterprise foundations that enable customers to safely adopt AI at scale. This isn't consulting or professional services. You're an engineer building production systems alongside customers, turning real-world deployments into durable platform capabilities used by thousands of organizations. This team will build 0-to-1 products from scratch - innovative, strategic initiatives designed to unlock hundreds of millions of dollars in new revenue. If you enjoy difficult engineering problems, working directly with customers, and building products from the field

PythonJavaLinuxRest
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $1.9M/yr

Quick readStrong listing-quality and freshness signals

At Datadog, we're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. APM at Datadog is on its way to redefine how users interact with their telemetry. We are integrating intelligence directly into troubleshooting workflows to help engineers find root causes faster, navigate complex distributed systems seamlessly, and optimize application performance with minimal cognitive load. APM provides deep visibility from end-user interactions to backend services and we are now expanding this foundation with new AI-driven insights, guidance, and automation. As a Product Manager II for APM, you will work with world-class engineers, designers, and partner product teams to shape the future of Distributed Tracing, Performance Analysis, and Intelligent Troubleshooting. You will help build advanced capabilities that scale to thousands of customers and make sophisticated observability workflows accessible to every engineer, from experts to beginners. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What you will do: Develop a deep understanding of APM customers, their performance challenges, telemetry workflows, and competitors Lead conversations with design partners and strategic customers to uncover real-world performance issues, validate product assumptions, and guide solutions from early prototypes through General Availability Define and deliver the next generation of APM features with engineering and design, especially agentic on

MicroservicesAIGoRust
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b

TypeScriptPythonAWSDocker
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b

TypeScriptPythonAWSDocker
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $196.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. WHY SAFETY? At Roblox, we strive to connect a billion people with optimism and civility, and the Safety organization's mission is to become the leader in civil immersive online communities. We systematically and proactively detect, remove, and prevent problematic content and behavior, and we make Roblox accounts secure and free from compromise. We also keep the platform compliant for changing regulations and growth markets. We cover a broad area of the tech spectrum, including machine learning, experimentation, automation, highly scalable distributed backend systems, detection workflows, and AI-powered text filters. Aligned and partnering with product teams, we use this tool belt to discover new opportunities, influence and shape the product roadmap and prioritization, build safety products, and measure the impact on our community of users and developers. In doing so, we keep Roblox safe, civil, and inclusive, and we foster positive relationships between people around the world. WHY CONTENT SUITABILITY? Join the Content Suitability team and play a pivotal role in shaping the future of content on Roblox. Our team is at the forefront of building tools and systems that enable creators to launc

JavaAWSGitRest
S
📍 Bellevue, Washington, United States· Full-time
✓ High-confidence listingCompany trend -93.3%
Quick readStrong listing-quality and freshness signals

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for talented systems developers and researchers to join the Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization . Our mission is to build the next generation of high-performance and intelligent inference systems . We optimize not only how fast and efficiently models run, but also how quickly inference systems can adapt to new models, architectures, hardware, and workloads. Our work spans the full inference stack—from distributed serving and runtime systems to GPU kernels and model-system co-design. We explore techniques such as adaptive parallelism, speculative and parallel decoding, disaggregated inference, scheduling and batching, KV-cache optimization, model swapping, quantization, and GPU kernel optimization to push the frontier of latency, throughput, scalability, and cost. Beyond optimizing individual models, we are building intelligent and adaptive inference systems that can automate performance optimization—rapidly profiling new models and workloads, identifying bottlenecks, selecting effective execution strategies, and adapting system configurations with minimal manual tuning. We embrace AI-native engineering , using AI not only as the workload we optimize, but also as a tool to accelerate system deve

Machine LearningAISwiftGo
🔔

Get new distributed systems engineer data platform delivery database retrieval jobs in United States by email

Daily job updates · Unsubscribe anytime