A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Apollo is Palantir’s autonomous software management and deployment platform. It enables seamless, continuous delivery of mission-critical software (Foundry, Gotham, AIP) across a vast range of environments: on-prem, public cloud, disconnected (air-gapped) networks, and highly regulated settings (including IL-5 and FedRAMP). As a Software Engineer on the Apollo team, you’ll build and operate a large-scale distributed system to allow the remote operation and maintenance of Kubernetes clusters. Our mission is to extract the entire state of a cluster into a portable, high-performance artifact within minutes, enabling full and almost instant cluster reconstruction from the ground up—all while pushing the limits of speed, reliability, and scale. You’ll design and implement backup and restore solutions for Kubernetes, leveraging proprietary compression infrastructure tailored to Palantir’s unique deployment models. You’ll also build and optimize our container artifact store, which is based on the OCI (Open Container Initiative) distribution spec—the industry standard for storing and distributing container images and artifacts. You’ll own the backbone of every environment Apollo supports, from hyperscalers to Army trucks. If you’re excited by challenges at the intersection of container technologies like OCI and docker, storage, and distributed systems, you’ll find opportunities here to dive deep into storage formats and low-level optimizations, where milliseconds matter. As we increasingly automate cluster creation and management on diverse hardware, you’ll play a key role in scaling Palantir’s presence at the edge and solving tough distributed systems proble
Jobiba hiring network
Distributed Systems Engineer Jobs
1,306 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Apollo is Palantir’s autonomous software management and deployment platform. It enables seamless, continuous delivery of mission-critical software (Foundry, Gotham, AIP) across a vast range of environments: on-prem, public cloud, disconnected (air-gapped) networks, and highly regulated settings (including IL-5 and FedRAMP). As a Software Engineer on the Apollo team, you’ll build and operate a large-scale distributed system to allow the remote operation and maintenance of Kubernetes clusters. Our mission is to extract the entire state of a cluster into a portable, high-performance artifact within minutes, enabling full and almost instant cluster reconstruction from the ground up—all while pushing the limits of speed, reliability, and scale. You’ll design and implement backup and restore solutions for Kubernetes, leveraging proprietary compression infrastructure tailored to Palantir’s unique deployment models. You’ll also build and optimize our container artifact store, which is based on the OCI (Open Container Initiative) distribution spec—the industry standard for storing and distributing container images and artifacts. You’ll own the backbone of every environment Apollo supports, from hyperscalers to Army trucks. If you’re excited by challenges at the intersection of container technologies like OCI and docker, storage, and distributed systems, you’ll find opportunities here to dive deep into storage formats and low-level optimizations, where milliseconds matter. As we increasingly automate cluster creation and management on diverse hardware, you’ll play a key role in scaling Palantir’s presence at the edge and solving tough distributed systems proble
About the Team The Storage teams build and operate online stateful systems and abstractions that are reliable, efficient, secure and easy to use for DoorDash Engineering. The teams are responsible for understanding Product Engineering’s evolving needs and developing platform and infrastructure capabilities to serve them. The team currently supports CockroachDB, Cassandra, Kafka and Redis as well as data abstraction services to reduce the complexity of interacting with storage systems for Product Engineers. About the Role The Storage team is building and operating a high-performance, scalable, and reliable data abstraction layer that optimizes both efficiency and reliability. Our goal is to create a platform that manages itself and fades into the background—empowering engineers to focus on delivering product experiences our customers love. This role is available across two teams within Storage, each solving unique and high-impact challenges: One team is building the orchestration layer for DoorDash’s storage platform—unifying lifecycle management, operations, and self-serve APIs for databases and streaming systems, turning complex, stateful infrastructure into reliable, developer-friendly services used across the company. One team builds and operates the distributed data platform powering DoorDash's largest stateful workloads -- including Cassandra, which backs critical product surfaces across DoorDash, Wolt, and Roo. You'll design high-throughput data abstractions, smart clients, and platform services that make distributed data reliable and easy to work with at multi-petabyte, multi-million-QPS scale, with opportunities to go deep on distributed systems internals and contribute to the open-source Cassandra ecosystem. If you're passionate about distributed systems, developer experience, and building foundational infrastructure at scale, we'd love to hear from you. You must be located in San Francisco, Sunnyvale, Seattle, or the New York Metro Area for this hybrid pos
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend. About the Ray Core Team The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray. A snapshot of projects you can work on: Optimizing performance of large-scale workloads on Ray Stability and stress testing infrastructure Improving fault tolerance (HA) As part of this role, you will: Leading cross-team projects while mentoring junior team members Develop high quality open source software to simplify distributed programming (Ray) Identify, implement, and evaluate architectural improvements
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend. About the Ray Core Team The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray. A snapshot of projects you can work on: - Optimizing performance of large-scale workloads on Ray - Stability and stress testing infrastructure - Improving fault tolerance (HA) As part of this role, you will: Develop high quality open source software to simplify distributed programming (Ray) Identify, implement, and evaluate architectural improvements to Ray core Improve the testing process for Ray to make re
About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems. The Team An exciting opportunity to join a new team within the Software Operations group. The Build Engineering team is a new function within Software Infrastructure, which focuses on the overall process of building and integration of the Machine Learn ing S oftware S tack. You will work closely with the QA and development teams to get an understanding of how our ML SW stack is built, helping to ensure good build practices, and proving that the stack works together and is reproducible in secure, sandboxed environments. Responsibilities and Duties Developing our internal t
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Issue Workflow is Sentry's primary product surface. Our issue platform processes billions events daily and turns them into actionable insights that help millions of developers fix bugs faster. As a Staff Software Engineer on the Issue Workflow team, you'll architect the systems that power this experience. You'll work at the intersection of high-scale distributed systems and product engineering, building real-time data pipelines, search backends, and analysis systems that surface signal from noise. This is product engineering at massive scale—where every architectural decision impacts millions of debugging sessions. You'll be the technical leader who shapes how Sentry groups issues, how we make search lightning-fast, how we enable sophisticated agentic workflows, and how we ensure that the product is performant even at billions-of-events scale. Your work will define what's possible for the most trafficked part of Sentry's platform. In this role you will Drive technical strategy and roadmap. Partner with engineering leadership, product, and design to shape the multi-quarter technical vision for Issue Workflow platform. Make strategic calls about architectural direction, technology choices, and technical debt. Ensure the team is building a strong foundation to scale with Sentry's growth. Solve complex performance and scalability challenges. Champion product quality and user experience. Build features that don't just work—they delight. You understand that milliseconds matter in the developer experience. You sweat the details of interfaces, error messages, loading states, and edge cases. You instrument everything s
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Issue Workflow is Sentry's primary product surface. Our issue platform processes billions events daily and turns them into actionable insights that help millions of developers fix bugs faster. As a Staff Software Engineer on the Issue Workflow team, you'll architect the systems that power this experience. You'll work at the intersection of high-scale distributed systems and product engineering, building real-time data pipelines, search backends, and analysis systems that surface signal from noise. This is product engineering at massive scale—where every architectural decision impacts millions of debugging sessions. You'll be the technical leader who shapes how Sentry groups issues, how we make search lightning-fast, how we enable sophisticated agentic workflows, and how we ensure that the product is performant even at billions-of-events scale. Your work will define what's possible for the most trafficked part of Sentry's platform. In this role you will Drive technical strategy and roadmap. Partner with engineering leadership, product, and design to shape the multi-quarter technical vision for Issue Workflow platform. Make strategic calls about architectural direction, technology choices, and technical debt. Ensure the team is building a strong foundation to scale with Sentry's growth. Solve complex performance and scalability challenges. Champion product quality and user experience. Build features that don't just work—they delight. You understand that milliseconds matter in the developer experience. You sweat the details of interfaces, error messages, loading states, and edge cases. You instrument everything s
About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems. The Team An exciting opportunity to join a new team within the Software Operations group. The Build Engineering team is a new function within Software Infrastructure, which focuses on the overall process of building and integration of the Machine Learn ing S oftware S tack. You will work closely with the QA and development teams to get an understanding of how our ML SW stack is built, helping to ensure good build practices, and proving that the stack works together and is reproducible in secure, sandboxed environments. Responsibilities and Duties Developing our internal t
About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems. The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible. Responsibilities and Duties Develop, own, and maintain tools and services to support the software build and release process Deploy and maintain services with Kub
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry's issue platform processes billions of events every day to help millions of developers find and fix bugs. The hard part is deciding which events point to a problem, which belong together, and what a developer needs to know to investigate. As a Senior Software Engineer on the Issue Detection team, you'll design, build, and operate the systems that make those decisions. You'll work on real-time processing pipelines, monitors, and analysis systems that detect problems and turn them into issues. The work combines distributed systems with product engineering. Choices about detection accuracy and processing latency affect which problems developers see and how soon they can act. You'll help shape how developers monitor their applications and how Sentry groups related events into issues. You'll also build the context developers and AI agents need to investigate what went wrong. Keeping these systems reliable and fast as Sentry grows is part of the job, alongside making the issues they produce more useful. In this role you will Build and scale features on a product surface handling billions of events daily, where both query latency and correctness are immediately visible to users. Own the design and delivery of substantial projects end to end, scoping alongside product and design, making the technical calls within your scope, shipping, and instrumenting what you ship so the team can measure it. You will contribute to meaningful technical product decisions : grouping quality, search performance, migrations and backfills against enormous datasets, and making the surface work well for both humans and agents. Champi
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Infrastructure Engineer Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a senior software engineer on the Cell Platform team at Roblox, you will build systems that Roblox engineers use to create and deploy resources onto Kubernetes. Our engineers deploy their services in a complex, hybrid, multiple-cluster, K8s environment. The Cell Platform manages this complexity for our users, with tools, APIs, K8s controllers, and UX, simplifying infrastructure for our internal customers. You Have: A desire to work on critical, large-scale distributed systems An appreciation of observability and instrumentation and tooling to make your life easier 3+ years of experience as software engineer Bachelor's degree in Computer Science or an equivalent field You will: Build our Roblox-wide control plane using Kubernetes primitives (and plenty of custom resources) Work on the interface of the few hundred person infrastructure organization to the thousands of Roblox engineers Write and review high quality code and tests (largely Golang) Work on a team that cares about inclusivity and shipping For roles that are based at our headquarters in San Mateo, CA: The starting base pay for this position is as shown below. The actual base pay is dependent upon a variety of job-related
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . We're seeking an exceptional Staff Software Engineer to join our Observability team at Pinterest. This role combines deep technical expertise in distributed systems and data engineering with a product-oriented mindset to build world-class observability solutions that empower our engineering organization. As a Staff Engineer on the Observability team, you'll be responsible for designing and building the infrastructure and tools that provide visibility into Pinterest's large-scale distributed systems, helping thousands of engineers understand, debug, and optimize their services. What you'll do: Define and execute the observability roadmap, treating it as a product. Understand engineering team needs and translate them into technical solutions with measurable impact Architect, build, and scale distributed observability infrastructure (me
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Platform Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning/sharding strat
Get new distributed systems engineer jobs by email
Daily job updates · Unsubscribe anytime