Jobs in United States

Distributed Systems Engineer in United States

426 active opportunities · Updated October 2026

Explore current distributed systems engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $192K/yr

Quick readStrong listing-quality and freshness signals

Datadog's Application Performance Monitoring (APM) provides deep visibility into the health, performance, and lifecycle of modern distributed applications, tracing requests from end-user devices (web and mobile) through to backend services. Our goal is to help customers detect root causes faster, optimize application performance, and improve resource efficiency at scale. As the Engineering Manager for APM Serverless, you will help define and deliver the end-to-end serverless APM experience, from auto-instrumentation through troubleshooting, and ensure that OpenTelemetry and Datadog-native customers alike have a frictionless and performant journey. You will also lead efforts to expand coverage of cloud-managed services across providers, ensuring customers can seamlessly trace and monitor critical services in all major and emerging cloud environments. We’re looking for an experienced engineering leader who thrives at the intersection of infrastructure and developer experience. You should care about well-designed APIs, observability-first thinking, and building systems that empower other developers. This is a high-leverage role that will influence how developers across the industry understand and instrument their serverless workloads. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Lead a polyglot team of 8-9 engineers and partner closely with Product and Engineering teams across Datadog to deliver industry-leading serverless capabilities that power consistent, scalable, and intuitive instrumentation across languages. Drive a domain that is technically rich: Lambda, Azure Functions, GCP, OTel billing, Rust, durable functions, distributed tracing across managed services. Engineers on this team work

AWSAzureGCPAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The ChatGPT Model Flywheel team unified goal is to transform model advancements into great ChatGPT user experiences through reliable serving, rapid experimentation, safe deployment, and continuous improvement. Team Focus Areas Model Experimentation: Enable rapid, safe model validation for ChatGPT and Codex products through experiment automation and lifecycle management. Model Deployment: Ensure safe, scalable deployment of model capabilities with robust rollout and operational tooling. Automate capacity management and incorporate platform-wide health monitors. Model Measurement: Build comprehensive evaluation and measurement systems for model quality, from user signals to launch scorecards. Improve end-to-end feedback loops for continual model improvement. Key Partnerships Collaborate cross-functionally with teams including Model Measurement DS, Research, Codex, Fleet, Inference, and API. In this role, you will: Elevate and consolidate ChatGPT’s harness, context management, and system prompt frameworks. Drive expansion and improvement of multi-tier model experiences. Support and scale self-serve experiment capabilities and automated guardrails. Lead model rollout automation, capacity management, and health monitoring. Shape end-to-end measurement systems (evals, grader signals, user feedback, etc.). You might thrive in this role if you have: Proven experience leading engineering teams in complex, cross-functional environments. Demonstrated success shipping production systems at scale (ideally for AI or large backend services). Deep understanding of model-driven product development, deployment lifecycle, and measurement tooling. Excellent communication and collaboration skills—experience interfacing directly with engineering, research, and product stakeholders. Prior involvement with large language models, distributed infrastructure, or experimentation platforms is a plus. Why Work With Us Tackle highly impactful technical challenges at the cutting edg

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Post-Training Frontiers team is responsible for training the frontier agents OpenAI ships to the world (GPT-Next). We train the flagship agentic models behind Codex, ChatGPT, and the API through large-scale reinforcement learning. The team’s work spans four areas. First, execution and science: working with teams across OpenAI to decide what can go into the final model and how, using scientific experiments and evals that are representative of the final pipeline so issues can be recognized early. Second, RL scaling: executing the final large-scale reinforcement learning run, making sure GPUs are used efficiently and training stays healthy. Third, research: improving horizontal capabilities like instruction following, factuality, memory, and multi-agent behavior, where the team’s broad visibility helps identify cross-cutting improvements across teams and domains. Fourth, engineering: maintaining the infrastructure stack and internal tools to ensure that both the final run and all integrations go as smoothly as possible and that the systems are easy to work with. About the Role This role focuses on keeping our frontier RL training runs fast, reliable, and unblocked. You will work across engineering and infrastructure problems as they emerge, from scaling and orchestration issues to inference bottlenecks, numerical problems, and hardware failures, as well as supporting large horizontal integrations in the big run, like multi-agent capabilities or memory. This is a role for a strong generalist who quickly learns anything needed for the task, has high attention to detail, debugs deeply, and is motivated by fixing the highest-impact problem in front of the team. In this role, you will: Keep large-scale async RL training runs moving by jumping into the most urgent engineering and infrastructure problems. Debug issues across training systems, inference, orchestration, scaling, and distributed infrastructure. Improve the reliability and efficiency of RL trai

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Workload Networking team is responsible for the collective communication stack used in our largest training jobs. Using a combination of C++ and CUDA we work on novel collective communication techniques that enable efficient training of our flagship models on our largest custom built supercomputers. The models we train are key ingredients to the AI research progress at OpenAI and the field as a whole, and we continually incorporate learnings from our entire research org into our training platform. About the Role As a Software Engineer, Networking you will design and implement custom networking collectives that are tightly integrated into our training stack. We’re looking for people who have a background in low level performance critical software. Experience with collective communication is a bonus. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Collaborate closely with ML researchers to design and implement efficient collective operations in C++ and CUDA. Ensure that our largest training jobs take full advantage of the different network transports used in our supercomputers. Work on simulations to inform our future supercomputer network designs. You might thrive in this role if you: Have written distributed algorithms using RDMA in the past. Are comfortable writing low level performance sensitive CPU and/or GPU code. Are familiar with network simulation techniques. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voic

AWSRestAIC++
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $399.4K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Content Platform team at Roblox powers the infrastructure behind every asset used across the Roblox ecosystem—enabling creators and developers to bring their visions to life at global scale. From 3D models and images to videos and audio, our platform manages the complete lifecycle of all assets essential for immersive experiences, supporting one of the largest services in the world at over 100+ million requests per second. Our mission is to deliver a seamless, reliable, and innovative content system that empowers creators, supports record-breaking games, and ensures the highest standards of performance and safety for our community. As the Technical Director for Content Platform, you will lead multidisciplinary engineering teams responsible for the technical and product vision of Roblox’s asset infrastructure. You will own the lifecycle of every asset—from creation and upload to storage, indexing, delivery, and rendering in the game client. Your leadership will be critical in scaling our systems, optimizing distributed infrastructure, and enabling new possibilities for creators and players alike. You Will: Define and drive the long-term strategy, architecture, and priorities for the Cont

AWSGitAIGo
M
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -67.9%
Quick readStrong listing-quality and freshness signals

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on the distributed object storage system that underpins every container image, volume, and checkpoint on Modal: hundreds of petabytes of data, replicated across multiple cloud object stores and a CDN, cached on local NVMe across a large fleet of workers in many datacenters, and shared peer-to-peer within each datacenter. You'll make cold starts feel local when the data is hundreds of milliseconds away, designing the caching, preloading, and peer-to-peer layers that hide object-store latency and keep public ingress off saturated uplinks. You'll own durability and cost at petabyte scale, from streaming and batch replication between origins, to garbage collecti

M
📍 New York, new york, United States· Full-time
✓ High-confidence listingCompany trend -67.9%
Quick readStrong listing-quality and freshness signals

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for a strong technical lead to guide the engineers designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. You'll lead the team responsible for Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll own the full lifecycle of a machine, from accepting and benchmarking new hardware from a growing set of providers, to network bring-up, kernel and image management, GPU and disk health tracking, and automated remediation of unhealthy hosts. You'll manage a team of 3–8 engineers while staying hands-on across the stack which involves BMCs, firmware, PXE, bootloaders, Linux networking, drivers, and distributed control-plane services, and you'll shape our long-

M
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -67.9%
Quick readStrong listing-quality and freshness signals

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for a strong technical lead to guide the engineers designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. You'll lead the team responsible for the distributed object storage system that underpins every container image, volume, and checkpoint on Modal: hundreds of petabytes of data, replicated across multiple cloud object stores and a CDN, cached on local NVMe across a large fleet of workers in many datacenters, and shared peer-to-peer within each datacenter. You'll set technical direction for the primitives that other teams (filesystems, training, sandboxes) build on, balancing durability, latency, throughput, and cost. You'll own the roadmap from today's hardest problems (garbage collection at petabyte scale, active-active replication, rate limiting that protects the upstream without wasting ut

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $280K/yr

Quick readStrong listing-quality and freshness signals

Datadog is seeking a Director of Product Management to lead our AI Observability portfolio and shape how organizations build, monitor, and scale AI systems in production. This role leads LLM Observability and helps define the next wave of innovation across GPU Monitoring, Distributed AI Monitoring, and emerging research-oriented tooling such as Model Lab. You will set the vision and strategy for this rapidly growing area, expanding established products while incubating new capabilities that deliver deep visibility into AI infrastructure, model performance, and distributed AI environments. As AI becomes core to modern applications, this team plays a critical role in ensuring customers can deploy and scale AI with confidence. We’re looking for a builder-minded product leader with strong technical depth and hands-on curiosity - someone who has built or worked closely with AI-powered products and understands the realities of production AI. You will lead a team of product managers and partner closely with engineering and design to advance Datadog’s leadership in AI observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the vision and strategy for AI-driven products, ensuring alignment with overall company goals and customer needs. This will include managing our embed program to enhance the capabilities of existing products as well as developing dedicated and independent AI products. Lead and mentor a team of product managers, helping them grow and advance their careers while ensuring the delivery of high-quality, AI-powered features. Collaborate with cross-functional teams including engineering, data science, marketing, and sales to deliver AI product solutions that meet customer needs and business objectives. Identify new opportunities for

Machine LearningAIGoRust
E(
📍 San Francisco Bay Area, California, United States· Full-time
✓ Quality checkedCompany trend -100%

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. The Role Many enterprise deals Ema closes starts with a BDR earning five minutes with the right person. Our Business Development Representatives run a signal-led, account-based hunting motion against a defined enterprise ICP to book qualified meetings (SQLs) that convert into demos and paid POCs. The motion works: the team has a real playbook, real signal scoring in our tools, and real enterprise logos in the pipeline. What it needs now is a dedicated leader to run it with daily rigor. As Enterprise BDR Leader, you own the BDR function end to end. You will manage a distributed team of Business Development Representatives across the US and India, instill the metrics discipline and coaching cadence that turns activity into qualified pipeline, and rebuild onboarding and enablement so new reps ramp fast and hit quota on schedule. This is a newly created role — the BDR function has been running without a dedicated, full-time manager, and you are the person who changes that. You report into our GTM organization, and work closely with our Account Executives, Regional VPs, and Marketing/demand gen to keep the pipeline both full and qualified. What You Will Own 1. The team. Hire, ramp, and

P
📍 Colorado Springs, Colorado, United States
✓ High-confidence listingCompany trend +26.9%

$151K – $241K/yr

Quick readStrong listing-quality and freshness signals

Job Title Service Operations Manager Job Description Your role: Lead service operations for a global interventional cardiology and vascular device portfolio, owning field performance, delivery execution, and operational governance across a geographically distributed team of field service engineers, product support specialists, and service coordinators. You will set the operating rhythm for corrective, preventive, and installation service across North America while partnering with international service leaders to align standards globally. Own the service training and technical education function end to end, including training program management, instructor-led and digital course delivery, field certification programs, and training center operations across three continental sites. You will ensure engineers maintain verified competency on legacy, newly acquired, and next-generation product lines in a regulated medical device environment where patient safety depends on technician proficiency. Drive measurable improvement in service performance using indicators such as time to system restoration, first-visit resolution, installation quality, contract capture, and preventive maintenance completion. You will lead root-cause analysis on systemic delivery issues, design operational improvements with engineering and supply chain partners, and translate field data into changes that improve customer uptime and reduce repeat service events. Manage service quality and compliance activities including complaint escalation, field execution of corrective actions and compliance tracking, product lifecycle support for systems approaching end of service, and regulatory documentation across all served markets. You will work directly with quality, regulatory, and product engineering teams to ensure service processes satisfy medical device requirements while maintaining the speed and respo

Supply Chain
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -93.7%

From $92K/yr

Quick readStrong listing-quality and freshness signals

We are hiring a Senior Technical Product Marketing Manager to lead positioning and messaging and to grow adoption of MongoDB Search and Vector Search as foundational components of our platform – the retrieval layer powering the next generation of grounded AI applications and agents. This is a high-impact role for a marketer who thinks like a builder. As developers architect increasingly sophisticated systems – RAG pipelines, agentic workflows, multi-modal search experiences – retrieval has moved from an implementation detail to a core design decision. You’ll join a high-performing, globally distributed team and partner closely with Marketing, Builder Relations, Product Management, Engineering, Partners, and Sales to develop, measure, and achieve cross-functional goals. The role requires technical depth in information retrieval — lexical and vector search, hybrid approaches, embeddings, re-ranking, agentic retrieval loops, and the tradeoffs that matter in production systems — paired with the product marketing instincts to turn that depth into crisp, differentiated messaging for distinct user and buyer personas. Hands-on experience building or shipping AI-enabled products is a strong advantage. Individuals with prior experience in technical sales, developer relations, or technical marketing are encouraged to apply. This person is a voracious consumer of AI research and pays close attention to shifting patterns in application architectures and development, including agentic systems. This individual is confident in communicating with technical practitioners and non-technical decision makers in one-to-few and one-to-many engagements for internal and external audiences. We are looking to speak to candidates who are based in the US for our hybrid working model. What You’ll Do Drive Strategy & Execution: Act as a strategic partner for high-impact initiatives that align with MongoDB’s long-term business goals in collaboration with Marketing, Developer Relations, Product

MongoDBAWSAzureAI
H
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Field Service Specialist provides installation support, commissioning, maintenance, troubleshooting, and on-site customer support for deployed KARNO Power Modules. This is the first dedicated field service role supporting KARNO and is a foundational position within Hyliion's field service organization, which is being built to support installations across the United States. Early deployments focus on defense and data center applications. The position begins with an intensive training period of approximately three to four months in Milford, OH, working alongside the research, development, and engineering teams as KARNO units are built and serviced during final testing, including training on heat engine fundamentals, PLCs, HMIs, and advanced control systems. The position then transitions to field installation, commissioning, and long-term on-site support at customer locations, initially across the West Coast and expanding to other regions as the installed base grows. As the service organization scales, this position helps define its processes and standards. AI at Hyliion At Hyliion, AI is core to how we work. We equip every team member with leading AI tools and count on you to use them — to move faster, solve harder problems, and help us realize the full potential of KARNO technology for the world. Duties and Responsibilities Provide on-site maintenance, troubleshooting, and break-fix support for deployed KARNO units, with remote support from the engineering team. Diagnose issues using remote monitoring systems and troubleshooting logs. Troubleshoot controls, instrumentation, and high-voltage electrical system issues. Super

AIHuman ResourcesHR
C
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why This Role? We're hiring a People Project Manager to join our HR PMO and help deliver an ambitious portfolio of HR programs, products, and initiatives as we scale globally. This is not a business-as-usual project management role . We're a hypergrowth company building at speed, and our HR team is building alongside it. We're looking for someone who gets genuine energy from bringing order to chaos, creating structure where there isn't any yet, and getting important things shipped. Hypergrowth, for real. Things move fast, priorities evolve, and there's always something meaningful to jump into. If you love pace, variety, and figuring things out as you go, you'll have a huge canvas here. A chance to build. Our HR PMO is taking shape within a hugely ambitious HR team. You'll help create the workflows, rhythms, and tools we need to build a world-class function at scale. Real ownership. Take projects from idea to delivery — creating structure, driving actions, managing dependencies, and keeping teams moving. Global from day one. Work across our distributed HR team and with stakeholders around the world. You'll thrive if you: Create o

GitAIGoExcel
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for a Customer Marketing Manager who can own the full customer evidence motion at Baseten: building the systems that capture customer stories, running the co-marketing programs that amplify them, and developing the channels and assets that get those stories in front of the right people. Our customers are ML engineers and AI teams deploying serious workloads — and the stories they tell about what they've built matter. We've earned trust with some of the most demanding technical teams in the industry, and this role exists to turn that trust into evidence. RESPONSIBILITIES Co-Marketing Execution Serve as the DRI for every customer co-marketing launch end to end — managing timelines, coordinating internal and external stakeholders, and driving the process from first outreach to final publication Own the single source of truth for what's in flight across all customer co-marketing activity Coordinate with design, social, and sales to ensure every asset is built, approved, and distributed correctly Customer Evidence & Asset Library Own the customer evidence library: written case studies, video stories, customer quote repository, logo library, and sales snippets ensuring all assets stay current and are tagged and accessible for sales and marketing use Run the monthly operating rhythm: new logo additions from closed-won opportunities, asset updates, and customer health checks Programs & Channels I

Machine LearningAIGoRust
🔔

Get new distributed systems engineer jobs in United States by email

Daily job updates · Unsubscribe anytime