Jobiba hiring network

Distributed Systems Engineer Jobs

1,306 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

R
Roblox
📍 San Mateo• Full-time• From $295.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Engineering Manager, Home Infrastructure The Home Infrastructure team builds the mission-critical backend and data systems that power Roblox’s Homepage and Experience Details Page, two of the highest-traffic surfaces on Roblox. These surfaces reach the vast majority of Roblox’s daily active users and are core drivers of discovery, engagement, retention, and platform growth. We are a full-stack product infrastructure team responsible for content distribution across Roblox. Our systems support multiple modes of user interaction, including exploratory browsing, directed discovery, and personalized content recommendations across the many types of content that make up the Roblox ecosystem. This team sits at the intersection of large-scale distributed systems, machine learning-powered personalization, data infrastructure, and product experimentation. We partner closely with Machine Learning, Data Science, Product, Design, Frontend, Ads, Marketplace, Virtual Economy, and other teams across Roblox to build the platforms that help users find the most relevant and engaging content. As Engineering Manager for Home Infrastructure, you will lead a team of Backend and Data Engineers responsible for the e

awsgitmachine learning
View job →
M
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Requirements: 5+ years of experience writing high-quality production code Experience building high-performance distributed systems at a large scale (the more battle scars, the better) Strong cloud skills Strong knowledge of low-level operating system foundations (Linux kernel, file systems, containers, etc.) Experience with performance engineering (tell us a story of when you shaved off a few milliseconds!) Ability to work in-person in our NYC or SF office. Prior experience with Rust is nice to have, but not required. Ability to participate in on-call rotation and respond to production incidents.

linuxrestai
View job →
M
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Requirements: 5+ years of experience writing high-quality production code Experience building high-performance distributed systems at a large scale (the more battle scars, the better) Strong cloud skills Strong knowledge of low-level operating system foundations (Linux kernel, file systems, containers, etc.) Experience with performance engineering (tell us a story of when you shaved off a few milliseconds!) Ability to work in-person in our Stockholm office. Prior experience with Rust is nice to have, but not required. Ability to participate in on-call rotation and respond to production incidents.

linuxrestai
View job →
M
Mongodb
📍 Toronto• Full-time• From C$137K/yr
1mo ago

We are hiring a Senior Software Engineer to join our Server Security team. The Server Security team is a development-focused group within MongoDB's core engineering organization. Operating "close to the bottom of the stack," the team builds features that enable database users to secure their data globally. You will work on critical components including: Cryptography: Queryable Encryption , at-rest data encryption, and fundamental cryptographic principles. Identity & Access: Authentication and authorization systems, TLS, and X.509 certificate management Network Security: High-performance, low-latency networking protocols (PKI, Hashing, CRLs) System Integrity: Resilience, observability, and compliance assurance within a large-scale distributed database Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies distributed systems fundamentals to deliver core features of a popular database, join us! Let’s change what’s possible for application developers, system architects, and database operators. The role As a Senior Engineer, you will apply distributed systems fundamentals to deliver core security features. You will be a leader in improving MongoDB's security posture by owning features and leading investigations into complex areas of the codebase. What you’ll do: Build and test new security features in a large, feature-rich C++ codebase Work across engineering, cloud services, and support teams to coordinate feature rollouts and changes Stand for code quality and security best practices, assisting fellow engineers in writing well-reasoned, secure code Use strong diagnostic intuition to solve thorny technical issues related to distributed systems, concurrency, and OS internals This role can be remote or hybrid anywhere in the USA or Canada. We will prioritize candidates who are already located in one of these countries. Candidate Profile We are looking for a highly technical engineer w

javamongodbaws
View job →
M
Mongodb
📍 New York City• Full-time• From $126K/yr
1mo ago

We are hiring a Senior Software Engineer to join our Server Security team. The Server Security team is a development-focused group within MongoDB's core engineering organization. Operating "close to the bottom of the stack," the team builds features that enable database users to secure their data globally. You will work on critical components including: Cryptography: Queryable Encryption , at-rest data encryption, and fundamental cryptographic principles. Identity & Access: Authentication and authorization systems, TLS, and X.509 certificate management Network Security: High-performance, low-latency networking protocols (PKI, Hashing, CRLs) System Integrity: Resilience, observability, and compliance assurance within a large-scale distributed database Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies distributed systems fundamentals to deliver core features of a popular database, join us! Let’s change what’s possible for application developers, system architects, and database operators. The role As a Senior Engineer, you will apply distributed systems fundamentals to deliver core security features. You will be a leader in improving MongoDB's security posture by owning features and leading investigations into complex areas of the codebase. What you’ll do: Build and test new security features in a large, feature-rich C++ codebase Work across engineering, cloud services, and support teams to coordinate feature rollouts and changes Stand for code quality and security best practices, assisting fellow engineers in writing well-reasoned, secure code Use strong diagnostic intuition to solve thorny technical issues related to distributed systems, concurrency, and OS internals This role can be remote or hybrid anywhere in the USA or Canada. We will prioritize candidates who are already located in one of these countries. Candidate Profile We are looking for a highly technical engineer w

javamongodbaws
View job →
M
Mongodb
📍 New York City; United States• Full-time• From $106K/yr
1mo ago

Join the MongoDB Networking & Observability team and help build the core of a distributed database! Our team focuses on creating and enhancing components which facilitate communication between distributed processes and make these processes, and their communication, easily observable. Networking Observability’s responsibilities include improving MongoDB networking, improving the efficiency of resource utilization, and building low-overhead observability features. Our team includes engineers located in New York City and fully remote engineers. We operate close to the bottom of the stack, and have a lot of influence over the availability, performance, and robustness of our open source database. Recently, we’ve improved connection handling, explored new networking architectures, and integrated OpenTelemetry to make issues easier to diagnose and connect MongoDB to modern observability tools. We are planning to further improve our networking’s stack performance, availability and scalability as well as further enhance our observability stack using open observability frameworks. Are you excited to help the MongoDB engineering team build a better database? We are! Join us today, and we can build a faster, more reliable, exceptionally observable, database system together. This role can be based out of our New York City office or remotely within the United States and Canada. Candidate Profile 3+ years of experience building distributed systems Passionate about delivering and deploying a product with cross-team stakeholders Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Hands-on experience with building production-level code. Experience in C++ is required Interest in furthering their knowledge of networking, observability and how computer architecture and internals impact the availability of SaaS Solid verbal and written communication skills and highly motivated to collaborate with colleagues Po

mongodbawsazure
View job →
M
Mongodb
📍 Alberta• Full-time• From C$122K/yr
1mo ago

Join the MongoDB Networking & Observability team and help build the core of a distributed database! Our team focuses on creating and enhancing components which facilitate communication between distributed processes and make these processes, and their communication, easily observable. Networking Observability’s responsibilities include improving MongoDB networking, improving the efficiency of resource utilization, and building low-overhead observability features. Our team includes engineers located in New York City and fully remote engineers. We operate close to the bottom of the stack, and have a lot of influence over the availability, performance, and robustness of our open source database. Recently, we’ve improved connection handling, explored new networking architectures, and integrated OpenTelemetry to make issues easier to diagnose and connect MongoDB to modern observability tools. We are planning to further improve our networking’s stack performance, availability and scalability as well as further enhance our observability stack using open observability frameworks. Are you excited to help the MongoDB engineering team build a better database? We are! Join us today, and we can build a faster, more reliable, exceptionally observable, database system together. This role will be based remotely in Canada. Candidate Profile 3+ years of experience building distributed systems Passionate about delivering and deploying a product with cross-team stakeholders Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Hands-on experience with building production-level code. Experience in C++ is required Interest in furthering their knowledge of networking, observability and how computer architecture and internals impact the availability of SaaS Solid verbal and written communication skills and highly motivated to collaborate with colleagues Position Expectations Understand and improve the current funct

mongodbawsazure
View job →
V
Vercel
📍 San Francisco• Hybrid
9 days ago

About Vercel: Vercel is the agentic infrastructure company, freeing people and agents to ship what's next. For more than a decade we've helped builders move from idea to production with speed, security, and exceptional developer experience. Now we're scaling our products for both agents and people to ship and run software, built in the open and trusted by OpenAI, PayPal, Ramp, Supreme, and millions of developers worldwide. About the Role: The Scheduled Tasks team builds the platform primitives that let applications and agents run work now, later, or for a long time. We own Vercel Workflows, Queues, and Cron: the systems developers use to build long-running, event-driven, and scheduled applications. You will join a team working at the intersection of developer experience and distributed systems. Together, we are building and scaling the products that make it straightforward for developers to coordinate background work, move messages between services, and schedule work with confidence. You will collaborate closely with engineers across Vercel to make these powerful capabilities feel simple, composable, and native to the platform. What You Will Do: Design, build, and operate platform capabilities across Workflows, Queues, and Cron. Help developers build applications and agents that coordinate background, event-driven, and scheduled work. Build APIs, SDKs, and tooling that make it easy to define, run, and manage durable work at scale. Work on the distributed-systems foundations behind scheduling, message delivery, execution, retries, and state. Partner with product, developer experience, and infrastructure teams to turn customer needs into clear, useful developer primitives. Raise the engineering bar through thoughtful design reviews, well-tested code, production ownership, and clear technical communication. Engage with developers and the open-source community to understand where background jobs, queues, and scheduling create friction—and use that feedback to improve th

javascripttypescript
View job →
T-
15 days ago

About the Role: Tubi's content platform is the engine behind one of the largest free streaming services in the world. Every play, every deal, every creator, every frame of video flows through systems CPE owns, and the surface area is enormous. Distributed services running on the hottest path of Tubi's traffic. Video pipelines processing one of the largest workloads in streaming. Workflow engines automating the operations that used to consume entire teams. Creator-facing products turning a back-office process into a real platform. And on top of all of it, an AI-native rebuild of the CMS that most companies aren't willing to attempt. This isn't a single-domain role. It's a platform where backend, frontend, video, infrastructure, and applied AI all collide at the scale where decisions actually matter, where an architectural choice ripples across millions of titles and billions of requests, and where the difference between "good enough" and "great" shows up in revenue. We're looking for builders who want to range across domains — backend one quarter, frontend the next, applied AI the one after that — and who want their work to be felt: by viewers when a title plays instantly, by creators when they go live the same day, by Content Ops when a workflow runs itself, and by the business when the platform stops being a cost center and starts being a force multiplier. The infrastructure is already there. The mandate is already there. What's missing is the people who want to build the thing, not talk about it. Come build it. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: You'll work on systems that sit at the heart of Tubi's business, where the content pipeline meets the viewer, the creator, and increasingly, the AI agent. The work spans the full stack of a modern content platform: distributed services, video infrastructure, workflow automation, and applied AI, all running at

typescriptpythonreact
View job →
M
9 days ago

Enterprise Advanced is a distributed team across Europe and India that builds the software running MongoDB on any infrastructure, at global scale — from on-prem data centers to private cloud. You'll work primarily on Ops Manager and Automation, the systems that let customers deploy fault-tolerant, globally distributed MongoDB clusters in minutes. Our software manages some of the largest self-managed MongoDB deployments in the world, with production clusters running hundreds of shards and nodes under a single deployment. The main focus of this team is to adapt our software to manage MongoDB clusters which are deployed in data centers or private cloud platforms. You will work on the core functionality for all of our products, mainly on the Ops Manager , and Automation products. Our team's end users are some of the largest businesses in the world, deploying massive clusters and processing huge amounts of data. This role is based in our Gurgaon office, and can work in a hybrid fashion. This role will report to the Senior Engineering Manager also based in our Gurgaon office. What you’ll do Design, implement, test, and release features for Ops Manager Own end-to-end delivery of complex projects, from design through incremental shipping Troubleshoot and resolve issues surfaced in customer deployments running at scale Apply engineering judgment and MongoDB's core values across planning, design, and code review A great fit for this role will be You enjoy distributed-systems problems; consistency, fault tolerance, and scale are the daily reality, not edge cases People who like ambiguity and are comfortable defining their own approach with guidance, not step-by-step instruction You're flexible! You're willing to take on a wide variety of responsibilities, learning as you go You're a self-starter! You're comfortable organizing your own time, acting on feedback and prioritizing with guidance from senior members of your team Requirements 4+ years experience with a language

javascriptpythonjava
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

REMOTEkubernetesaigo
View job →

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale is seeking a Staff Software Engineer to lead the technical vision for our Infrastructure team. As a Staff Engineer, you will be responsible for the architectural evolution of our control plane and data plane, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. You will act as a force multiplier, setting the standards for Kubernetes-based cloud-native infrastructure while mentoring engineers and driving cross-functional alignment across the Ray open-source community and our proprietary product teams. Key Responsibilities Architectural Leadership: Define and drive the multi-year technical roadmap for services that orchestrate Ray clusters across diverse cloud and on-premises environments. Systemic Optimization: Lead the design and optimization of high-performance control plane components specifically tailored for large-scale, heterogeneous AI/ML workloads. Platform Reliability: Establish the organization-wide standards for the reliability, scalability, and observability of Anyscale-managed infrastructure. Strategic Integration: Direct the long-term strategy for accelerator integration (GPUs, TPUs) and container management to ens

pythonawsazure
View job →
M
Mongodb
📍 United States• Full-time• From $127K/yr
1mo ago

The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim

mongodbawsazure
View job →
M
1mo ago

The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our Toronto or Vancouver offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design primitives of at least one of AWS, Azur

mongodbawsazure
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

kubernetesaigo
View job →
🔔

Get new distributed systems engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More distributed systems engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.