About the Team The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world. Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure. About the Role We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology. You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more. There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future. In this role, you will: Build, lead, and grow high-performing infrastructure engineering teams. Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search. Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach. Champion pragmatic use of agent technology to amplify execution velocity. Reduce operational toil and incident frequency through better abstractions, gua
Jobs in United States
Distributed Systems Engineer in United States
426 active opportunities · Updated October 2026
Showing
15 jobs
Explore current distributed systems engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
We’re looking for a Software Engineer to architect and build backend systems that enforce data privacy and automate compliance at scale. You’ll work closely with product, infrastructure, security, and legal teams to embed privacy-by-design into our data and access layers. This is a hands-on, high-impact role for an experienced engineer who is passionate about protecting user data while enabling innovation. What You’ll Do Design, build, and operate backend services that enforce policy-driven data access, lifecycle controls, and privacy protections. Develop distributed authorization and identity-aware enforcement mechanisms integrated directly into data services and control planes. Implement auditability, policy hooks, and enforcement observability to ensure compliance is continuously verifiable. Partner with Security, Legal, and Compliance to convert privacy requirements into scalable technical designs and developer-friendly APIs. Harden data platforms and backend services through schema-level controls and data handling constraints by default. Collaborate with infrastructure teams to ensure consistent enforcement across systems while minimizing duplicated implementations. Contribute patterns, libraries, and education that elevate trustworthy data access patterns across the organization. You Might Thrive in This Role If You Have 5+ years of industry experience building and operating backend or infrastructure systems in production. Strong software engineering fundamentals , with fluency in at least one major programming language (e.g., Python, Go, Rust, C++, Java). Experience with distributed authorization, RBAC/ACL systems, encryption-based access, or policy engines. Familiarity with global privacy regulations and their architectural implications. Ability to influence and collaborate with teams across legal, compliance, product, and engineering. A bias toward practical, impactful solutions that balance privacy protections with product needs. Nice to Have Experience wi
About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training Performance Engineer, you’ll drive efficiency improvements across our distributed training stack. You’ll analyze large-scale training runs, identify utilization gaps, and design optimizations that push the boundaries of throughput and uptime. This role blends deep systems understanding with practical performance engineering — analyzing GPU kernel performance, collective communication throughput, investigating I/O bottlenecks, and sharding our models so we can train them at massive scale. You’ll help ensure that our clusters are running at peak performance, enabling OpenAI to train larger, more capable models with the same compute budget. This role is based in San Francisco, CA. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you will: Profil
OpenAI’s charter calls on us to ensure the benefits of AI are distributed broadly and safely. Our Health AI team focuses on expanding access to high-quality medical expertise and aims to set a high standard for deploying AI responsibly in high-stakes domains. Improving health will be one of the defining impacts of AGI. Today, millions of people lack access to reliable medical information, and clinicians around the world face increasing time and resource constraints. We are building AI systems that support patients, clinicians, and health workers, while meeting the highest standards for safety, reliability, and privacy. We are seeking full stack software engineers to help build and scale products used by consumers and care providers globally. You will work closely with product, design, and research teams to ship real systems in a fast-moving, high-impact environment. In this role, you will: Design and build scalable fullstack systems for consumer and enterprise health. Own end-to-end feature development—from early design and implementation through deployment, monitoring, and iteration. Build and maintain data pipelines and services that meet strict privacy, security, and compliance requirements (e.g., HIPAA). Collaborate closely with researchers and safety teams to integrate reliability, evaluation, and guardrails into production systems. Debug, optimize, and harden systems to support high availability, performance, and global scale. Take ownership of ambiguous problems and drive them to practical, high-quality solutions. You might thrive in this role if you: Are deeply motivated by improving health outcomes and expanding access to medical expertise. Are a strong engineer who enjoys building durable, well-designed systems. Have 5+ years of experience writing maintainable, production-quality code. Can operate with high agency—owning problems end-to-end with minimal supervision. Enjoy working in fast-moving, cross-functional teams with engineers, product managers, desi
Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Senior Manufacturing Engineer serves as the senior technical authority for assembly process and equipment design within the Industrialization team. This position designs the assembly lines, fixtures, and material flow systems that convert KARNO prototype and development builds into repeatable, scalable production operations — and then systematically removes waste, labor content, and variation from those operations. Working closely with industrialization leadership, design engineering, production, quality, and supply chain, the Senior Manufacturing Engineer owns the most complex assembly value streams end to end: line and station design, fixture and tooling design, material presentation and handling, process qualification, and continuous waste reduction. This position sets the technical standard for how assembly processes are designed and documented at Hyliion, provides mentorship and design review for other manufacturing engineers, and is accountable for measurable improvement in cycle time, first-pass yield, labor content, and ergonomics across the assembly areas. AI at Hyliion At Hyliion, AI is core to how we work. We equip every team member with leading AI tools and count on you to use them — to move faster, solve harder problems, and help us realize the full potential of KARNO technology for the world. Duties and Responsibilities Assembly Line and Cell Design : Design assembly lines, cells, and workstations for KARNO core, module, and subassembly operations. Establish work sequencing, balance work content to takt, define station layouts and footprints, and design lines that accommodate planned rate increases rather t
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We’re looking for an Infrastructure Security Engineer to design and secure the core systems that power our platform. This role focuses on building security directly into our infrastructure—from container isolation and orchestration to identity and secrets management in a multi-tenant, cloud-native environment. You’ll work closely with engineering teams to define secure primitives and ensure our platform is resilient, scalable, and trustworthy by design. This is a hands-on, deeply technical role focused on real systems, not compliance or policy. What You'll Do: Platform & Runtime Security Design and improve isolation mechanisms for multi-tenant workloads (containers, sandboxing, execution environments) Strengthen boundaries between customers, workloads, and internal systems Identify and mitigate risks in distributed, dynamic compute environments Container &
What you’ll do Act as the technical lead for large parts of the scanner platform: system architecture, codebase structure, and long-term maintainability. Own core runtime foundations: distributed control, state management, fault handling, and reliability. Drive engineering rigor: testability, code quality, review standards, performance regression prevention, and release processes. Build robust observability: logs, metrics, traces, and replayable diagnostics (with privacy constraints). Collaborate with hardware and recon/ML teams to define interfaces, data contracts, timing/synchronization, and failure modes. Lead complex refactors (e.g., message passing / RPC boundaries, modularization, concurrency model) without halting forward progress. What we’re looking for Deep software architecture experience for real-world systems: robotics, instrumentation, medical devices, or other complex distributed products. Strong Python and concurrency background (asyncio, multiprocessing, profiling, performance engineering). Track record of shipping systems that are observable, debuggable, and resilient. Strong technical leadership: clarity, pragmatic trade-offs, and mentoring. Useful experience Building but rock-solid systems: clear interfaces (gRPC/protobuf or equivalent), strong state modeling, and failure handling. High-leverage engineering habits on a lean team: good tests, CI, reproducible dev environments, and fast code review. Practical performance + concurrency work in Python (asyncio, profiling, multiprocessing) and comfort debugging distributed behavior. Security-minded device software: safe defaults, encrypted data paths, and disciplined handling of PII/PHI. Operational thinking: remote updates/management, excellent logging, and diagnostics that make real hardware debuggable.
About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a software engineer on the Scaling team, you’ll help build and optimize the low-level stack that orchestrates computation and data movement across OpenAI’s supercomputing clusters. Your work will involve designing high-performance runtimes, building custom kernels, contributing to compiler infrastructure, and developing scalable simulation systems to validate and optimize distributed training workloads. You will work at the intersection of systems programming, ML infrastructure, and high-performance computing, helping to create both ergonomic developer APIs and highly efficient runtime systems. This means balancing ease of use and introspection with the need for stability and performance on our evolving hardware fleet. This role is based in San Francisco, CA, with a hybrid work model (3 days/week in-office). Relocation assistance is available. In this role, you will: Design and build APIs and runtime components to orchestrate computation and data movement across heterogeneous ML workloads. Contribute to compiler infrastructure, including the development of optimizations and compiler passes to support evolving hardware. Engineer and optimize compute and data kernels, ensuring correctness, high performance, and portability across simulation and production environments. Profile and optimize system bottlenecks, especially around I/O, memory hierarchy, and interconnects, at both local and distributed scales. Develop simulation infrastructure to validate runtime b
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali
Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Mechanical Engineer is responsible for end-to-end hardware ownership of components and sub-systems for the KARNO generator, taking designs from CAD through prototype, test, and validation. Working across mechanical, electrical, software, and performance teams, this role designs and troubleshoots complex thermal and mechanical systems that must perform reliably across extreme operating environments and a wide range of fuels. The position exists to advance the development of Hyliion's fuel-agnostic power generation technology through hands-on, test-driven engineering and disciplined design execution. Duties and Responsibilities Own hardware components and sub-systems end-to-end—from concept through durability, manufacturability, serviceability, cost, weight, and validation—taking designs from CAD to hardware running on a test stand. Design components and sub-systems that must survive extreme thermal environments, perform across a wide range of fuels (20+), and push the boundaries of metal additive manufacturing. Create 3D models in NX and generate 2D prints with full GD&T per ASME Y14.5. Perform design checking and print review to ensure tolerances, processes, and material specifications align with Hyliion's GD&T standards (ASME Y14.5). Conduct fluid and thermal systems design and optimization. Perform structural and thermal FEA (ANSYS or equivalent). Install, calibrate, and read instrumentation for pressure, temperature, flow, strain, and acceleration in lab environments. Execute prototype build, test, and validation cycles early and often to identify and resolve issues in the lab rather than the field. Collaborate cross-functionally
Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Manager, Electrical Engineering is responsible for the electrical systems of the KARNO generator, including high-voltage power electronics, battery systems, low- and high-voltage architecture, wiring harnesses, and the hardware that converts linear motion into electrical output. This is a working manager role: the position leads and develops a team of electrical engineers while remaining directly involved in technical execution, including circuit architecture, schematic review, and hardware bring-up in the lab. The Manager is accountable for the technical excellence, safety, and reliability of the electrical engineering function, and for establishing the design standards and review practices the team works to. The position plans team capacity, owns hiring and development for the electrical engineering staff, and partners with mechanical, controls, supply chain, and program management on system integration. KARNO systems are deployed in data center, military, and industrial applications. AI at Hyliion At Hyliion, AI is core to how we work. We equip every team member with leading AI tools and count on you to use them — to move faster, solve harder problems, and help us realize the full potential of KARNO technology for the world. Duties and Responsibilities Own critical electrical designs personally, including regular time at the bench and in the test cell, while leading the team as a practicing engineer. Lead the electrical engineering team in the design and development of KARNO generator electrical systems, including high-voltage power electronics, battery systems, linear generator power stages, and low-voltage controls hardwar
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 The Mission The Foundry is ClickUp's internal AI innovation lab — embedded inside GTM Systems and accountable for turning AI capabilities into production-grade, internally deployed products that make every GTM function faster and smarter. We build the infrastructure that powers AI-first work across Sales, Marketing, Post-Sales, and Revenue Operations. As the Senior Software Engineer on this team you will own the technical delivery of our MCP server platform, agent orchestration layer, and internal tooling — shipping production systems used daily by hundreds of ClickUp employees, and scaling your own throughput by treating AI tools as first-class engineering collaborators. What You'll Own MCP Server Platform Design, build, and operate Model Context Protocol servers that expose CRM, ticketing, analytics, and communication data to AI agents across the GTM stack Implement Okta PKCE authentication flows and RBAC policy enforcement so agents access only the data they're authorized to touch Maintain deployment infrastructure on AWS (Bedrock, Lambda, ECS, API Gateway) and contribute to GCP workloads where applicable Own observability: structured logging, distributed tracing, latency SLOs, and on-call runbooks for every production server Agent Orchestration & AI-Native Products Build and maintain multi-step autonomous agents that execute end-to-end GTM workflows — lead qualification, deal room assembly, onboarding automation, support triage, and more Architect prompt engineering frameworks, tool-call schemas, and agent evaluation harnesses that make AI behavior predictable and auditable Integrate with LLM p
At Linear, we're building the product development system for teams and agents. AI is fundamentally changing how software gets built, and we’re shaping the tools this new era requires. Founded in 2019, Linear has become the platform of choice for more than 40,000 companies (including OpenAI, Coinbase, and Ramp) to plan, build, and ship their products. Today, our team is distributed across North America, Europe, and Australia, and we’re continuing to grow internationally. What unites us is relentless focus, fast execution, and a deep care for software craftsmanship. We’re looking for experienced engineers who have shipped applied AI systems to production and want to define what the agent-native future looks like. We are building intelligence into the core of Linear, enabling the product to orchestrate coding, proactively move work forward, and power-up every software team. You’ll work closely with product and design to transform foundation models into structured, reliable workflows embedded deeply in the core of Linear. We care deeply about keeping Linear fast, intuitive, and opinionated—AI is no exception. Location & work mode Linear is a remote-first company, with optional co-working offices in San Francisco, New York, and London. This role is open to candidates based in the North America. You can work from anywhere within this region. We value deep focus and async collaboration, with intentional moments to connect in person through team off-sites, optional co-working, and occasional travel. What you'll do Build AI-powered product features that feel native, fast, and delightful to use Work with product and design to prototype and iterate on intelligent workflows and user interactions Design backend services to power natural language interfaces, smart suggestions, agentic workloads, and more Optimize prompts, fine-tune model behavior, and evaluate performance Help to guide our agent platform, allowing third parties to bring agents into the core Linear experience
From $295.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox Ads & Discovery business growing at a rapid rate, we are building large scale ads machine learning infrastructure to deliver more value to our users and our advertisers. As a Machine Learning Infrastructure Engineer, you’ll build scalable, reliable, and high-performance infrastructure that powers ML systems across our organization. You’ll operate at the scales of hundreds of billions of engagements, and redefine how we deliver performance ads to hundreds of millions of users. You will: You will co-design models and systems, working at the intersection of model architecture and ML infrastructure, partnering closely with core modelers, data and AI infrastructure engineers, and product teams to push the boundaries of large-scale training and serving. Your work will span recommendation, search, and agentic applications, including large transformer architectures, LLMs, generative rankers, and efficient offline and online content-understanding systems. You will investigate model, data, and systems tradeoffs end to end—from data pipelines and distributed training to low-latency inference and production serving. This includes designing efficient KV-cache strategies, applying p
Other cities to consider
More places hiring for this role
Get new distributed systems engineer jobs in United States by email
Daily job updates · Unsubscribe anytime