Jobiba hiring network

Distributed Systems Engineer Jobs

1,306 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

M
Mongodb
📍 San Francisco• Full-time
1mo ago

Join our MongoDB Search Systems Engineering team at the forefront of building next-generation database infrastructure at scale. You'll work alongside some of our most talented and deeply technical engineering teams who are architecting distributed systems that power highly performant search and AI applications for enterprise customers worldwide. This team operates at the intersection of infrastructure innovation and customer impact—building the foundational platform capabilities that will define how organizations leverage search and AI at scale for years to come. In this role, you'll serve as both strategic partner and technical translator, helping brilliant engineers navigate the path to operational excellence while maintaining the forward vision to anticipate market needs 18-24 months ahead of current development cycles. What Makes This Role Unique: This is about architecting the future of search infrastructure. You'll be working with engineers who live and breathe distributed systems, helping them channel that brilliance toward platforms that don't just work today, but anticipate the architectural needs of tomorrow's AI-native applications. If you light up at the challenge of seeing around corners in infrastructure and can hold your own in technical debates about consensus protocols while keeping the team focused on customer value—this is your role. This role can be based out of our San Francisco office or remotely in USA region. Role Responsibilities Contribute to the Future of Search Infrastructure: Drive the long-term technical vision for MongoDB Search, anticipating emerging patterns in vector, hybrid, and AI-native applications to build conviction around platform investments 12–18 months ahead of market demand Champion the Customer & Innovation: Act as the voice of customers building mission-critical AI applications, using deep technical discovery to uncover latent needs around scalability and consistency, ensuring our architecture leads the industry in

reactmongodbaws
View job →
L
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft Infrastructure builds the systems engineers depend on to ship stable, scalable, and efficient services. We're hiring a Senior Technical Program Manager to run cross-functional programs across our infrastructure and data platform teams. This role blends program delivery with product sense: you'll own the roadmap for your area, set priorities, and act as the voice of the customer back into how we build. Responsibilities Run infrastructure programs end to end, from kickoff through delivery Own the roadmap for your platform area: shape the strategy, sequence the work, and make the prioritization calls Drive data platform migration and modernization work, coordinating across engineering, data, and platform teams to keep dependencies and timelines under control Be the voice of the customer: partner with engineering teams across Lyft, surface their pain points, and feed that back into priorities and roadmaps Define success metrics and adoption goals, gather and document customer requirements, and make sure what ships actually solves the problem Build feedback loops with customer teams and turn what you hear into concrete improvements Partner with engineering and infrastructure leads to build plans, call out risks early, and keep stakeholders aligned Own program health: track milestones, surface blockers before they slip, and keep decision-makers in the loop Use your technical background in distributed systems and data infrastructure to ask sharp questions and build plans the team believes in Share in the team's release oncall rotation Experience 5+ years in Technical Program Management or a TPM/PM hybrid role A background in software, data, or systems engineering, enough to go deep with engineers Experience owning a roadmap: setting strategy, prioritizing across competing demands, and defining what suc

E
Enigma
📍 New York• Full-time• $180K – $295K/yr
14 days ago

The Opportunity This is a critical and exciting time at Enigma. Our customers consistently tell us that our data products create tremendous value and are deeply aligned with their most important workflows. As demand grows, we have an urgent opportunity to improve both the intelligence of our data and the systems through which customers access it. We are looking for an experienced Senior/Staff Machine Learning Engineer to join our Match Team and help shape the next generation of Enigma’s customer-facing data products. In this role, you will combine advanced statistical and machine learning research with the engineering systems required to power fast, relevant, and reliable search experiences at scale. This is a uniquely high-impact role sitting at the intersection of information retrieval, ranking systems, semantic search, distributed systems, and customer data delivery. The Role At the core of Enigma’s product is our data, which makes both data science and delivery systems central to what we build. As a Senior/Staff ML Engineer on the Match Team, you will lead efforts that improve the relevance, latency, and scalability of our customer-facing data products. You’ll work across the full lifecycle: framing retrieval and ranking problems, developing models and experimentation strategies, evaluating results using real-world signals, and implementing high-throughput search and retrieval systems. This role is ideal for someone who is excited by both hard ranking/search problems and the systems challenges of turning those solutions into low-latency, production-grade retrieval systems. What You'll Do Develop innovative solutions to complex problems in information retrieval, ranking, semantic search, query understanding, and recommendation systems Build and optimize low-latency, high-throughput search APIs, indexing pipelines, and retrieval systems using Python, Typesense, and AWS Evaluate and evolve our search technology stack, driving technical design decisions across index

pythonawsmachine learning
View job →
R
Roblox
📍 San Mateo• Full-time• From $153K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As an Early Career Software Engineer at Roblox, your story begins with supportive mentorship and immediate, global impact. You’ll work alongside seasoned engineers and curious experts, designing, coding, and deploying real features that reach millions of people globally. By the end of your first year, you’ll have directly influenced the future of our platform and the experience of our global community with millions of daily active users. You Will: Join a community of curious, supportive engineers, actively engaging in architectural discussions and system design. Investigate and experiment with cutting-edge technologies, like machine learning frameworks and large language models (LLMs), to solve complex technical challenges and improve our engineering systems. Design, code, and test innovative features, navigating the full development lifecycle from initial design to production deployment. Partner closely with cross-functional teams, including Design, Product, Data, QA, and DevOps, to deliver cohesive products and features. Support the continuous evolution of our distributed systems, operating at our massive scale of 2 trillion analytics events a day. Engage in our team mat

pythonjavanode.js
View job →
O
1mo ago

About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun

awsazuregcp
View job →
R
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Software Engineer Intern at Roblox, your 12-week journey is designed around accelerated learning and tangible global impact. With the support of a dedicated mentor, you’ll apply your knowledge to production-scale reality as you take end-to-end ownership of a project tackling some of the hardest technical challenges, including distributed systems, real-time communication, 3D co-experience, data processing, rendering, and more. You Will: Join our supportive community of engineers, receiving dedicated mentorship as you deliver live production projects. Investigate and experiment with cutting-edge technologies, such as machine learning frameworks, agentic coding tools, and large language models (LLMs), to solve complex technical challenges and improve our engineering systems. Own a project from beginning to end, from coding and testing to deploying it to production, and presenting your work to peers and leaders across the company. Partner closely with cross-functional teams, including Design, Product, Data, QA, and DevOps, to deliver cohesive products and features. Build a close community with fellow interns and full-time builders alike, while gaining a broad perspective on the compa

pythonjavanode.js
View job →
C
Coinbase
📍 Brazil• Full-time• Remote
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Join the EAA Compliance CXAE team within the Platform group as a Software Engineer building AI-first platforms that transform how CX Compliance agents work. This team owns the tools, services, and applications that streamline compliance and KYC processes, helping resolve customer issues faster with greater accuracy. You'll build full-stack applications using Golang and React that enable compliance agents to increase productivity, drive automation, and deliver impact at scale through custom development and third-party integrations. What you'll do: Build end-to-end user-facing features using Golang, React, and cloud technologies for CX compliance agent platforms Lead assessment, selection, and implementation of third-party tools and vendor integrations that extend platform capabilities Drive cross-functional outcomes on complex problems in collaboration with product, design, security, data, and peer engineering teams Own technical architecture decisions and implement scalable, high-traffic services across event-driven and distributed systems Mentor engineers on design techniques, coding standards, testing practices, and deployment processes Required Skills and Experience: 4+ years of production software engineering experience building large-scale systems with Golang, React, and cloud technologies (AWS, Kubernetes, Terraform) Demonstrated experience integrating third-pa

REMOTEreactsqlaws
View job →
O
1mo ago

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

awsrestai
View job →
O
OpenAI
📍 United States• Full-time
1mo ago

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

awsrestai
View job →
C
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Washington DC, Austin, NYC, San Francisco About the Dept Cloudflare’s engineers build and operate the software that helps power 25+ million Internet properties and millions of businesses around the world. Across our engineering organizations, we have opportunities for high caliber, curious and empathetic people to take on big challenges and build some of the best skills in the industry. We’re looking for talented team me

awskubernetesai
View job →
D
Datadog
📍 New York• Full-time• From $192K/yr
28 days ago

Coordination Systems provides foundational distributed systems building blocks for internal Datadog platforms. Our services cover sharding, consensus, resource protection, configuration distribution, and much more. We are looking for a manager to lead the Coordination Systems - Storage team. This team provides essential configuration storage and distribution systems that are depended upon by almost every service and pod at Datadog. We power critical runtime configuration (e.g. feature flags), complex control planes (e.g. dynamic sharding configuration), and much more. Storage is one of four subteams within Coordination Systems. If successful, the candidate will have opportunities to lead other growing and impactful areas such as Resource Protection. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: (Describe role responsibilities here/max 6 bullets) Lead a core team of 5 engineers (distributed, with majority in NYC) Lead ceremonies, prioritize and delegate project Stay hands-on with the code, e.g. isolated features, small remediations, investigation follow ups Stay actively involved in operations, incidents, root cause analysis, etc. Constantly promote a culture of operational excellence, organizing gamedays, conducting operational reviews, staying proactive with reliability Who You Are: (Describe role qualifications here/max 6 bullets) Strong distributed systems skills, able to understand and account for a variety of failure modes, well-versed in end-to-end o11y, validation testing, simulation setup, etc. Worked on platform teams before, providing critical infrastructure to internal stakeholders Experienced in handling significant incidents, both as a responder and follow-up ow

airustexcel
View job →

Datadog's Software Engineers with Systems depth leverage their experience with systems and tooling to build software that ensures Datadog remains reliable, performant, and secure. For this track, their Software Engineering experience may resemble the Distributed Systems track, but is typically applied in combination with their systems experience to build and run internal platforms and tools that our products are built on. These people typically have deep experience building and managing large cloud infrastructure deployments, or leading reliability efforts for orgs similar to ours, or building release machinery to allow hundreds or thousands of devs to do their jobs without stepping on each others' toes. The systems and tooling where they may have experience depth may include (but not limited to): bazel, build tooling, cassandra, CDN, chef, configuration management, container orchestration, consul, docker, elasticsearch envoy, haproxy, kafka, kubernetes, load balancing, network architecture, postgres, redis, release management, RPC frameworks, service discovery, spinnaker, terraform, zookeeper. Bonus: You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. #LI-KM5 Datadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan. Th

postgresqlredisdocker
View job →

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Team Creator Services Machine Intelligence Team : The Machine Intelligence team is building an NPC system that can (1) play any Roblox game and (2) perform real-time inference efficiently enough to support deployment to all Roblox players. ML Platform Team : The Foundation AI Group is on a mission to establish Roblox as the standard for 3D foundational models (3DFMs), democratizing creation by making it simple for anyone to generate high-quality, immersive 3D experiences using AI. The AI Platform team is a foundational part of this vision, supporting hundreds of ML use cases and billions of inferences daily across Discovery, Safety, Engine, and more. We are seeking exceptional PhD new graduates to drive innovation across three critical areas: AI Platform, Distributed Inference Systems. What You Will Do As a Senior Machine Learning Engineer, you will be a key contributor to building the cutting-edge systems that power AI at Roblox. Creator Services Machine Intelligence Team Develop Scale Data Pipelines: Design, build and maintain robust data pipelines to collect complex 3D game states and real-time player actions across the platform. Train Novel Architectures: Solve the feature e

awsazuregcp
View job →

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Team Creator Services Machine Intelligence Team : The Machine Intelligence team is building an NPC system that can (1) play any Roblox game and (2) perform real-time inference efficiently enough to support deployment to all Roblox players. ML Platform Team : The Foundation AI Group is on a mission to establish Roblox as the standard for 3D foundational models (3DFMs), democratizing creation by making it simple for anyone to generate high-quality, immersive 3D experiences using AI. The AI Platform team is a foundational part of this vision, supporting hundreds of ML use cases and billions of inferences daily across Discovery, Safety, Engine, and more. We are seeking exceptional PhD new graduates to drive innovation across three critical areas: AI Platform, Distributed Inference Systems. What You Will Do As a Senior Machine Learning Engineer, you will be a key contributor to building the cutting-edge systems that power AI at Roblox. Creator Services Machine Intelligence Team Develop Scale Data Pipelines: Design, build and maintain robust data pipelines to collect complex 3D game states and real-time player actions across the platform. Train Novel Architectures: Solve the feature e

awsazuregcp
View job →
🔔

Get new distributed systems engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More distributed systems engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.