Jobs in United States

Distributed Systems Engineer in United States

426 active opportunities · Updated October 2026

Explore current distributed systems engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Billing & Payments Platform team builds Snowflake's central data repository and infrastructure for customer resource consumption, revenue processing, invoicing, and reporting. Our systems power Snowflake's business and enable every other engineering team — and the architectures we ship double as reference patterns for our customers building on Snowflake. Computing Snowflake's bills, at its core, is a challenging distributed systems problem: real-time usage metering across every cloud and region, and supporting an ever-evolving catalog of pricing models — including the new commercial constructs we are inventing for Cortex AI, Snowflake Intelligence, and the broader agentic AI portfolio . Our applications must meet strict requirements for accuracy, auditability, and low-latency processing. This is a deeply cross-functional role. You will partner daily with Product, Finance, Legal, Growth, Go-to-Market Systems, Snowsight UI, Cortex AI, and product engineering teams across Snowflake to deliver experiences that customers and internal stakeholders depend on every day. What You'll Do As a Senior Software Engineer on Billing Platform, you will: Own medium-sized projects end-to-end — from design through launch and operation — and contribute as a key engineer on large, multi-

PythonJavaSQLCI/CD
SF
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -100%

From $136K/yr

Quick readStrong listing-quality and freshness signals

About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As an ML Platform Engineer at Stitch Fix, you will play a key role in building and maintaining the critical infrastructure that powers machine learning and AI across our organization. You will design, develop, and support scalable, resilient services and frameworks for ML model training and deployment, feature engineering and serving, candidate generation, AI agent deployment and observability, and other core platform capabilities. In this role, you'll contribute to the day-to-day operations of the ML Platform team, ensuring the smooth functioning of existing systems while driving improvements. You’ll collaborate closely with full-stack data scientists, offering consultation and support to help them unlock the full potential of our platform. With significant autonomy, you’ll have the opportunity to shape the future of ML and AI at Stitch Fix. Your ideas and expertise will drive improvements, codify best practices, and influence how we approach machine learning and AI systems at scale. Responsibilities: Collaborate with cross-functional teams, including data scientists, engineers, and business partners, to solve complex distributed systems and business challenges at scale. Be part of a team with high visibility across the organization, driving impactful solutions that make a difference. Share your ideas and help guide the team’s investments toward high-value opportunities. Foster a culture of technical collaboration and contribute to the development of scalable, resilient systems. About You You bring

PythonRedisAWSRest
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $345K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Creator Content Platform Team provides a secure, scalable, and extensible foundation for ingestion, processing, storing, managing, and serving user content. As a Principal Software Engineer (Backend, Distributed Systems), you will design and build backend services to help game creators reach the broadest audience possible. You will help us build & scale large distributed systems, data processing and analysis pipelines, an access control system, Open Cloud APIs all of which are central for all content created in Roblox. \ You Will: Solve on a variety of unique technical challenges Have the independence, opportunity and the end-to-end responsibility to design, build, test and deploy services within the Roblox ecosystem Guide the future technical direction of the team and have impact on engineering Be a technical bar-raiser for high code quality, architectural designs, and long-term approaches Mentor and develop fellow engineers on the team Design systems and services that are scalable and resilient Collaborate with passionate, Engineers, Product Managers and other Roblox team members, cross-functionally You have: Experience: You have 13+ years of experience working on backend, d

PythonJavaAWSGit
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Orchestration pod within Engine Productivity, you'll design and run the platform that executes large-scale end-to-end and integration tests, running the real, shipping client on real hardware, across Roblox's data centers, cloud, and our own device labs, so our engineering teams can ship the engine, clients, Studio, and more with speed and confidence. Every Roblox engine, client, and Studio change, along with the experiences built on top of them, should ship with confidence, and the Orchestration team is the layer that makes that possible. We build large-scale distributed services that turn thousands of test suites into a reliable, push-button pipeline: fanning work out across fleets of machines and real devices, moving artifacts to where they're needed, managing single- and multi-client test state, and giving test owners and maintainers a system to validate their own runs. It looks a lot like building a specialized cloud platform, with capacity-aware scheduling, isolation and sandboxing, and smart retry and backoff, plus the classic distributed systems problems (fairness, efficiency, failure handling, and reliability) at Roblox scale. You Will: Design a

PythonAWSGitAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $196.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As an Economy Fraud engineer, you defend Roblox from all types of fraud, including theft, scams, money laundering, and payment fraud. Roblox is a high-growth, unique product environment. You will be developing anti-fraud and abuse solutions for web, mobile, and 3D environments. This high impact work and your innovation is critical for the well-being of our community and to the future of our company. We aim for our users to have peace of mind that their communities and transactions are protected. Our defenses also protect our company’s rapid expansion and safeguard billions in revenue. Roblox’s virtual marketplace handles over 4 million transactions a day, and enables our top developers to make millions of dollars a year. Our team’s challenges are not just regular day-to-day technical challenges. Fraud and abuse approaches need to shift over time, depending on the current behaviors of fraudsters. As an Economy Fraud engineer, you will be in a data-driven environment developing both classical and novel approaches to detect and prevent this bad behavior. You Have: 4+ years of professional experience working with scalable, distributed systems Strong experience in large-scale, data-driven

SQLAWSGitMachine Learning
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Communications team, you'll own the backend systems that power chat and messaging for millions of players on Roblox and support billions of messages per day. Your day-to-day will focus on building and scaling the infrastructure that makes in-game communication fast, safe, and reliable. That means driving architectural decisions, creating communication modes that don't exist yet, and defining developer-facing APIs so creators can build richer experiences. You'll work closely with frontend engineers, platform, product, and design, and you'll have real influence over the technical and product direction of the team. If you're an experienced engineer who gets excited about large-scale distributed systems and wants to shape the way millions of people connect inside virtual worlds, we'd love to talk. You Will Build and scale backend infrastructure for millions of concurrent users, architecting high-throughput distributed services, integrating ML models, and enabling rapid product experimentation Equip creators with the APIs and tools to build deeply integrated social experiences in their games Own projects end-to-end: from design and architecture through produc

SQLPostgreSQLMySQLAWS
NR
📍 Atlanta, Georgia, United States
✓ Quality checkedCompany trend -73.9%

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Do you get excited about complex distributed systems? Do you love data? We’re looking for an experienced technology leader to lead and grow our global engineering team and help our customers build better software by building the next generation of Application Performance Monitoring (APM) by leading our Agent teams. New Relic is committed to giving our customers valuable insights into their systems, and New Relic’s APM agents are used by tens of thousands of companies to evaluate and improve the performance of their most important business applications. Opportunity to work from a remote office may be available depending on the applicant's location. What you'll do Hands on with data analysis and technical problem solving Leverage open source tools like OpenTelemetry to acquire data Help design and build our APM agents that run in our customer environments and give engineers deep insight into application performance, and business line owners actionable intelligence on their business performance Build processes that ensure reliability, scalability, team growth and execution Work across teams to build engineering plans and execute, getting things done in a distributed, large-scale, and fast-growing environment This role requires 4+ years of experience leading people and teams 10+ years of experience in Software Engineering Proven track record of leading and scaling strong technical teams Experience handling high performing self-directed remote teams Capable of div

JavaAIRecruitment
P
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -85.6%
Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is seeking a Staff Software Engineer, Capacity Engineering. The team is responsible for efficiently managing one of the largest-scale cloud-native infrastructures in the world. This role is highly impactful, as efficiency is an ongoing strategic priority for Pinterest. The role has direct visibility across Pinterest Engineering and with Engineering and company leadership. The team is looking for a candidate with a strong background in implementing performance and efficiency projects on large scale distributed systems. In this individual-contributor role you will own and drive performance and efficiency for a core area of Capacity Engineering, partnering with the company-wide efficiency lead and collaborating with performance and efficiency leaders across the organization. What you’ll do: Drive efficiency in large-scale shared environme

PythonJavaAWSKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Data Platform at OpenAI owns the foundational data stack powering critical product, research, and analytics workflows. We operate some of the largest Spark compute fleets in production; design, and build data lakes and metadata systems on Iceberg and Delta with a vision toward exabyte-scale architecture; run high throughput streaming platforms on Kafka and Flink; provide orchestration with Airflow; and support ML feature engineering tooling such as Chronon. Our mission is to deliver reliable, secure, and efficient data access at scale and accelerate intelligent, AI assisted data workflows. Join us to build and operate these core platforms that underpin OpenAI products, research, and analytics. We’re not just scaling infrastructure – we’re redefining how people interact with data. Our vision includes intelligent interfaces and AI-assisted workflows that make working with data faster, more reliable, and more intuitive. About the Role This role focuses on building and operating data infrastructure that supports massive compute fleets and storage systems, designed for high performance and scalability. You’ll help design, build, and operate the next generation of data infrastructure at OpenAI. You will scale and harden big data compute and storage platforms, build and support high-throughput streaming systems, build and operate low latency data ingestions, enable secure and governed data access for ML and analytics, and design for reliability and performance at extreme scale. You will take full lifecycle ownership: architecture, implementation, production operations, and on-call participation. You’ve supported Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms. You’re well-versed in infrastructure tooling like Terraform, experienced in debugging large-scale distributed systems, and excited about solving data infrastructure problems in the AI space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per wee

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Cloud Agents team builds product infrastructure for long-running agents in the cloud: orchestration, sandboxing and isolation, secure environment connectivity, secrets and identity, observability, reliability, and cost controls. These agents securely connect to diverse developer and customer environments and use tools to accomplish goals. We partner closely with product, research, and infrastructure teams to turn agentic capabilities into dependable platforms for OpenAI products and developers building on OpenAI. About the Role We are looking for an experienced software engineer to help build and scale our cloud agent platform. You will design and operate systems for orchestrating agents at scale. You will work closely with product engineers on ChatGPT, API, and Codex to define the right abstractions and enable them to ship products quickly. Strong backend or infrastructure experience is important; experience with Python, Rust, distributed systems, cloud infrastructure, or product platforms is especially helpful. In this role, you will: Design and scale the orchestration, sandboxing and storage systems that run agentic workloads for Codex, ChatGPT, and the OpenAI API. Partner with product engineers to build a platform that enables them to ship quickly and turn feedback into robust abstractions. Improve reliability, security, performance, and cost efficiency for long-running agents. Deploy services that can operate across different environments and clouds. Your background might look something like: 9+ years of professional engineering experience, excluding internships, in relevant roles at technology and product-driven companies. Experience leading large-scale backend, platform, or infrastructure projects from ambiguous problem statements to production systems. Proficiency in one or more backend languages such as Python, Go, Rust, TypeScript, or similar, and the ability to move across service, platform, and product boundaries. Strong understanding

TypeScriptPythonAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op

PythonAWSRestAI
O
📍 Seattle, Washington, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team Online Data builds and operates Habitat, the single product surface of Online Data and the system of record for OpenAI’s online user data. As OpenAI’s scale and product requirements evolve, Habitat is becoming a full-stack, one-size-fits-most database platform with end-to-end ownership of: Provisioning and developer experience APIs and guardrails Scaling, performance, and reliability Data movement, caching, routing, and placement Privacy enforcement and access control Change Data Capture (CDC) as a first-class primitive The foundation for future storage backends You’ll work on the core online database platform behind OpenAI’s products, building and operating Habitat services that handle high-QPS, latency-sensitive workloads across regions. You’ll partner closely with internal platform and product teams to ship safe, reliable systems, then push them to be faster and more cost-efficient through better caching, routing, observability, and operational tooling. This is a critical role for engineers who like owning hard distributed-systems problems end to end and sweating the details from p99 latency to production operations at massive scale. In this role, you will Design and build core abstractions spanning storage, caching, routing, CDC, and privacy enforcement Own a major surface area end to end, from product and API design to operational excellence Improve latency, correctness, and cost efficiency for real production workloads at massive scale Build strong instrumentation, debugging workflows, and developer-first tooling Collaborate closely with internal product and infrastructure teams to understand requirements and ship pragmatic solutions Participate in an on-call rotation and raise the bar on reliability while aggressively improving performance and usability You might thrive in this role if you have A strong track record building and operating high-scale backend or data-intensive distributed systems in production Excellent systems judgment and the a

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About OpenAI OpenAI is dedicated to ensuring that artificial general intelligence (AGI) benefits all of humanity. Our mission requires building not only world-class AI models, but also the infrastructure that enables those models to be deployed reliably, efficiently, and at global scale. As demand for AI continues to grow, we are expanding the ways OpenAI can bring high-performance inference capacity online across a diverse hardware ecosystem. About the Team The GPT Infrastructure team builds software that turns advanced inference and optimization research into production products. One focus is enabling strategic infrastructure partners and accelerator vendors to qualify and onboard new compute without a bespoke porting and optimization effort for every hardware platform. We build the control planes, APIs, secure partner-side execution environments, evaluation systems, artifact pipelines, and operational tooling that make these workflows repeatable and trustworthy. The work sits at the intersection of distributed systems, AI inference, compilers and runtimes, performance engineering, security, and external partnerships. About the Role We are seeking an experienced systems generalist who can work comfortably across the stack to help build an automated inference optimization platform. Given a workload, target hardware profile, compiler and runtime context, and a trusted verifier, the system runs durable optimization campaigns that generate, compile, execute, grade, and improve candidate kernels, runtime configurations, and serving-stack changes. You will design both the OpenAI-hosted control plane and the partner-side software that evaluates candidates on real accelerator hardware. The product must keep long-running workflows reliable, make performance results reproducible, and maintain clear trust boundaries around sensitive model and hardware information. This is a deeply cross-stack role, combining strong software engineering fundamentals with systems thinking and

PythonAWSLinuxRest
E(
📍 San Francisco Bay Area, California, United States· Full-time
✓ High-confidence listingCompany trend -100%

$135K – $225K/yr

Quick readStrong listing-quality and freshness signals

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are We are seeking an experienced DevOps Engineer to join our growing team and play a pivotal role in designing and building our platform and infrastructure as we continue to scale our product and user base. As a part of our team, you will be working in a dynamic, fast-paced environment to ensure the reliability, scalability, and performance of our systems, while focusing on service architecture and deployment, query optimization, distributed systems, data and machine learning infrastructure, and security and authentication. Most importantly, you are excited to be part of a mission-oriented, fast-paced, high-growth startup that can create a lasting impact. You will: Partner with product teams to architect, design, and build the foundational infrastructure for our products. Design, develop, and deploy highly available and scalable Multi-tenant SaaS solutions on any one of the public cloud networks like AWS, Azure and GCP. Leverage technologies such as Kubernetes, Helm, Terraform, and Istio to achieve infrastructure resilience. Drive the automation of infrastructure tasks, from provisioning to configuration management and deployment, utilizing tools like Terraform, Ansible, a

AWSAzureGCPKubernetes
S
📍 Mclean, Virginia, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer, Snowflake Natsec Running Snowflake in public sectors in different countries and regions, even in different industry verticals, requires us to build a compliant, secure, and auditable infrastructure. Many key design decisions are deeply rooted in the Snowflake product architecture. As a Senior Software Engineer, you will be responsible for leading several key areas and collaborating with various engineering groups in addition to the Public Sector team. To be successful in the area, you will need to have (and continue to build) a broad and in-depth knowledge base on cloud infrastructure, privacy, and governance, compliance controls, data security and data residency in various aspects of Snowflake. AS A SENIOR SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Solve real business needs at large scale by applying your software engineering and analytical problem solving skills. Design, implement and maintain scalable distributed systems for our cloud automation platform that include cloud control plane, Kubernetes container platform and traffic and networking. Work directly with customers to quickly understand their critical problems and design and implement solutions Deploy and maintain availability of cloud compute servers and Kubernetes cluster that power the

PythonJavaAWSAzure
🔔

Get new distributed systems engineer jobs in United States by email

Daily job updates · Unsubscribe anytime