Jobiba hiring network

Distributed Systems Engineer Jobs

1,306 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

R
Roblox
📍 San Mateo• Full-time• From $295.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's data infrastructure processes petabytes of data daily, powering analytics, ML, and product decisions for a platform serving 200M+ daily active users. As a Principal Software Engineer in our Data Infra org, you will be the primary technical leader driving the strategic vision, long-term architecture, and massive scalability of our distributed data platforms that power Roblox. You will own and drive the next-generation architecture of our core platforms, which span Kafka, Flink, Spark, Trino, Druid, Airflow and Data Catalog. This role operates under high ambiguity, demanding unparalleled ownership to redefine the limits of infrastructure handling exabyte-scale workloads, and providing a unique opportunity to lead the future evolution of our global data ecosystem. You Will: Define Multi-Year Technical Strategy: Own and drive the end-to-end architectural vision for Roblox's core data platforms spanning Kafka, Flink, Spark, Trino, Druid, Airflow, and Data Catalog systems. Turn multi-year company strategies into concrete, production-grade infrastructure blueprints. Lead Cross-Functional Alignment: Partner closely with executive leadership, platform governance, data science, and product e

javaawsgcp
View job →
R
Roblox
📍 San Mateo• Full-time• From $243.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's Cache team is building a next-generation caching solution designed to deliver sub-millisecond average latency, horizontal scalability, and high efficiency—all at a drastically lower cost. Our ultimate vision is to shape a caching infrastructure capable of supporting 1 billion Daily Active Users while reducing costs by 90%. We are turning hours of onboarding and capacity expansion into seconds, freeing service owners entirely from managing cluster lifecycles. As a Senior Engineer on the Cache team (part of the Infra Storage org), you will innovate and operate large-scale, in-house distributed systems to solve Roblox's ever-growing caching challenges. You will report directly to the Engineering Manager for the Cache team. (Check out our recent engineering blog post here to learn more about the team's latest work!) You will: Lead the architectural transition to a next-generation, multitenant caching service built on ValKey, ensuring strict data, resource, and failure isolation for all tenants. Drive systemic optimizations to mitigate head-of-line blocking, manage hot keys, and maximize CPU and memory utilization across physical machine clusters. Design and build robust frameworks to a

redisawskubernetes
View job →
D
Datadog
📍 Massachusetts• Full-time• From $100K/yr
1mo ago

We’re looking for Software Engineering Interns to help build and scale the systems that power Datadog’s observability and security platform. Interns contribute directly to real-world engineering challenges across backend, frontend, infrastructure, data engineering, and developer tooling while working alongside experienced engineers and mentors. You’ll help design, build, and improve systems that process and analyze massive volumes of metrics, logs, and application data in real time. Whether you’re interested in distributed systems, Kubernetes, AI-powered products like Bits AI, or developer platform tooling, you’ll work on meaningful projects that deliver impact to customers at global scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Contribute to production systems that process and analyze large-scale observability and application data in real time Build and improve distributed systems across backend infrastructure, developer platforms, and cloud-native services Help identify and solve performance, reliability, and scalability challenges in critical services supporting Datadog’s growing customer base Own and deliver technical projects from design through deployment with support from experienced engineers and mentors Develop technical expertise through hands-on experience with technologies such as Kubernetes, distributed systems, and cloud-native infrastructure Collaborate with fellow interns, mentors, and engineers while building software that delivers impact at global scale Who You Are: Pursuing a degree in Computer Science, Software Engineering, or a related technical field, or have equivalent practical experience Targeting a 2028 full-time start date Demonstrate strong computer science fundamentals, including data struc

kubernetesgitrest
View job →
M
Mongodb
📍 Gurugram• Full-time
1mo ago

The Infrastructure Engineering team is responsible for building and maintaining a self-service internal development platform that enables MongoDB engineering teams to reliably deploy and operate their own production services and products. We work with numerous engineering teams across the company to understand their infrastructure requirements and development workflows, develop broadly applicable self-service platform services and tooling, continuously monitor how platform services are being utilized, and look for ways to improve developer productivity through automation and education. We are big open source enthusiasts and use a number of open source tools in our stack (contributing upstream whenever possible). Some of the tools we use regularly include Go, AWS, Kubernetes, Crossplane, Terraform, Helm, Drone, Prometheus, and Grafana. However, technology is nothing without a stellar team of engineers that are focused on doing high quality work and working as a team to solve complex distributed computing and platform engineering problems. This is where you come in! We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Our ideal candidate 2+ years of experience managing and mentoring a team of 3+ engineers Has 5+ years of experience owning the design and implementation of large software/infrastructure projects Has built and operated large-scale distributed systems in cloud providers (AWS strongly preferred) Has a strong backend programming background. Fluency in Go is strongly preferred; deep experience with another compiled or strongly-typed backend language is acceptable Pragmatic, detail-oriented, self-motivated, and understands the benefits of collaboration Strong experience operating production Kubernetes clusters, not just deployed to it Has practical experience defining and operating against SLI/SLOs for services they owned Strong experience with observability tooling: metrics, logging, traces, Prometheus, Grafana, OpenTe

mongodbawsazure
View job →
D
1mo ago

Role Description As a Software Engineer on the Metadata team, you’ll build and operate the large-scale distributed databases that every Dropbox service depends on. Metadata systems are mission-critical, in the live path for all user operations and must meet stringent requirements for latency, durability, and transactional consistency. You’ll design and evolve the core infrastructure that manages Dropbox’s databases at scale, enabling fast, reliable access to data for millions of users and hundreds of internal services. This work spans distributed systems, replication, caching, and transactional database systems. You’ll collaborate closely with engineers across Infrastructure and Product teams to ensure the metadata layer meets business needs and continues to scale with Dropbox’s growth. This is an opportunity to leverage your expertise in distributed systems and grow into broader technical leadership. Our Engineering Career Framework is viewable by anyone outside the company and describes what’s expected for our engineers at each of our career levels. Check out our blog post on this topic and more here . Responsibilities Design and maintain distributed database systems providing low-latency, strongly consistent data access Implement and optimize replication, consensus, and caching mechanisms to meet availability and performance goals Operate production systems, including participating in the on-call rotation, ensuring high availability and data durability Collaborate with infrastructure and product teams to assess current and future use cases and requirements, supporting the development of a mid- to long-term roadmap that reflects these needs Contribute to system design reviews, postmortems, and reliability improvements Write high-quality, efficient code in Go and Rust for performance-critical systems On-call work may be necessary occasionally to help address bugs, outages, or other operational issues, with the goal of maintaining a stable and high-quality experienc

REMOTEsqlmysqlredis
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training: ML Framework Engineer, you will work on improving the training throughput for our internal training framework, while enabling researchers to experiment with new ideas. This requires good engineering (for example designing, implementing, and optimizing state-of-the-art AI models), writing bug-free machine learning code (surprisingly difficult!), and acquiring deep knowledge of the performance of supercomputers. In all the projects this role pursues, the ultimate goal is to push the field forward. We’re looking for people who love optimizing performance, understanding distributed systems, and who cannot stand having bugs in their code. Since our training framework is used for large runs with massive numbers of GPUs, performance improvements here will have a large impact. This role is based in San Francisco, CA. We use a

pythonawsrest
View job →
O
OpenAI
📍 San Francisco• Full-time• From $230K/yr
1mo ago

About the Role The Engineering Acceleration Delivery / Continuous Deployment team builds and operates the systems that safely ship OpenAI’s infrastructure and product code to production. We own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across OpenAI to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale infrastructure orchestration. In This Role, You Will Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces, and alerts to detect regressions and trigger safe rollbacks. Build systems that support agent-assisted or autonomous deployment workflows using modern AI tooling. Technologies commonly used in this environment include: Kubernetes for large-scale container orchestration and runtime infrastructure Python and FastAPI for internal services Terraform for infrastructure as code GitOps-based deployment workflows (e.g., ArgoCD, Flux, or similar systems) Buildkite for CI orchestration You may be a strong fit if you: Have worked with Kubernetes-based deployment systems at scale Have experience building or operating continuous deployment platforms Are familiar with GitOps tooling such as

pythonawskubernetes
View job →
S
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We're hiring Software Engineers for our Data Platform team to build and evolve Snowflake's real-time stream processing and data transformation platform. If you've built or researched high-throughput streaming systems or scalable transformation engines - in industry or graduate work - we want to talk. AS A SOFTWARE ENGINEER, DATA PLATFORM AT SNOWFLAKE, YOU WILL: Design and implement low-latency stream processing and in-flight data transformation systems at global scale Own correctness and performance of streaming execution (watermarks, exactly-once delivery, out-of-order handling) Build transformation primitives and operators that run reliably at cloud-scale throughput Contribute to architectural decisions for our next-generation streaming and transformation platform Write production-quality systems code and collaborate across product and engineering teams OUR IDEAL CANDIDATE WILL HAVE: BS/MS/PhD in Computer Science or related field — graduate research in streaming, distributed systems, or query/transformation engines is a strong differentiator Deep distributed systems fundamentals: fault tolerance, consistency, state management Hands-on or experience with stream processing or transformation systems (Flink, Kafka Streams, Spark Structured Streaming, or similar) Proficiency i

pythonjavasql
View job →

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary Health100 is America's trusted front door to health and care. The Health100 platform integrates any participating health plan, PBM, pharmacy (retail and specialty), provider, digital health point solution provider, and employer, and addresses the top health care challenges for the consumer. The Senior Manager, Software Engineering will lead engineering teams building a best-in-class consumer health experience on the Health100 platform, focused on identifying, prioritizing, shaping, and executing complex platform initiatives. As a key member of our engineering organization, you will drive innovation, manage cross-functional teams, and deliver scalable cloud-native and AI-enabled solutions that improve consumer experiences and business outcomes. This role works within a top-notch organization of software developers who identify, design, and deliver technology solutions using Java, distributed systems, cloud platforms, and AI-powered capabilities to achieve defined business value. You will guide the integration and validation of complex technology solutions, drive engineering excellence, and leverage emerging technologies including Generative AI and intelligent automation to accelerate innovation and efficiency *Remote eligible within the United States. Preference for candidates in close proximity to our Woonsocket, RI headquarters. Resp

javagcpdocker
View job →
I
Illumina
📍 California• $141.6K – $212.4K/yr
10 days ago

What if the work you did every day could impact the lives of people you know? Or all of humanity? At Illumina, we are expanding access to genomic technology to realize health equity for billions of people around the world. Our efforts enable life-changing discoveries that are transforming human health through the early detection and diagnosis of diseases and new treatment options for patients. Working at Illumina means being part of something bigger than yourself. Every person, in every role, has the opportunity to make a difference. Surrounded by extraordinary people, inspiring leaders, and world changing projects, you will do more and become more than you ever thought possible. Summary The Staff Data Engineer is a seasoned, hands-on engineer who designs, builds, and scales data products on our cloud lakehouse, powering analytics, reporting, and AI/ML across Illumina. We are looking for someone with strong proficiency in Python, SQL, and data modeling, a solid understanding of distributed systems and system design who has built and scaled data products on modern cloud platforms such as Databricks and Snowflake. This is a hands-on, senior individual-contributor role with end-to-end ownership and leadership spanning multiple domains such as Supply Chain, Manufacturing and Quality, including mentoring engineers on our global (India-based) team. Responsibilities Partner across business, AI, and platform teams translating domain needs (e.g., SAP, Manufacturing, Quality) into well-modeled, governed and scalable data products. Design, build, and scale end-to-end data products on Databricks (and interoperating with Snowflake) — from ingestion through curated, analytics-ready datasets following a medallion (Bronze/Silver/Gold) architecture. Develop reusable frameworks, libraries, and

pythonsqlaws
View job →
NR
10 days ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity We are a global team of innovators shaping the future of observability. Our intelligent platform gives customers real-time insight into complex systems so they can innovate faster and operate reliably in an AI-first world. If you’re excited by high-throughput distributed systems and want to contribute to one of the largest and fastest-growing observability platforms, we’d love to hear from you. Join a backend engineering team focused on building and operating JVM-based services that ingest, process, and serve massive volumes of telemetry data. You’ll work on high-scale, low-latency systems that power mission-critical observability features used by engineers worldwide. What you'll do Design, build, and operate JVM-based microservices (primarily Java and Kotlin) with a focus on performance, scalability, and reliability. Own services end-to-end: architecture, implementation, deployment, monitoring, on-call participation, and continuous improvement. Apply strong concurrency and performance practices: asynchronous programming, backpressure, efficient I/O, memory management, and GC tuning. Build and evolve event-driven systems; work with Kafka for streaming, partitioning, consumer groups, and schema evolution.Instrument services for deep observability (metrics, logs, traces), define SLIs/SLOs, and use e Experience with Kafka or similar streaming technologies (topic/partition strategy, consumer lag, idempotency, schema compatibility) strongly preferred. Proficiency w

javarediskubernetes
View job →
T-
Tubi - Canada
📍 Toronto• Full-time• From C$1.4M/yr
15 days ago

About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission is to engineer resilience from the ground up, enabling our product teams to innovate rapidly while ensuring our users have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through a culture of data-driven decision-making, blameless learning, and relentless automation. As a Senior Site Reliability Engineer, you are a hands-on engineer who blends deep software development expertise with a passion for operational excellence. You will be responsible for designing, building, and running the resilient, scalable, and increasingly self-healing systems that power our products. You will apply sound engineering principles to solve our most complex reliability challenges, with a mandate to automate everything, eliminate toil, and write robust, maintainable code. You will be a force multiplier, mentoring other engineers and elevating the site reliability bar for the entire organization. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: System Architecture & Design: Design, build, and maintain scalable, highly available, and fault-tolerant distributed systems. Partner with development teams as a reliability consultant, reviewing designs and influencing architectural decisions to ensure new services are built with reliability, observability, and performance as core principles, not afterthoughts. Automation & Software Development: Write robust, performant, and maintainable code to automate operational tasks, and CI/CD pipelines. Build the internal tools, libraries, and frameworks that enable engineering teams to self-service their

typescriptpythonaws
View job →
H
15 days ago

Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Advanced Engineer, Power System Controls plays a key role in the design, development, and implementation of control software for Hyliion’s Karno Power Module. This position focuses on systems involving combustion, thermal management, pressure regulation, high-voltage, and power management. The engineer will be responsible for developing and validating control algorithms, tuning system parameters, and analyzing data to ensure performance meets engineering specifications. Additional responsibilities include preparing technical documentation, supporting root cause analysis, and ensuring timely, high-quality software delivery. The role requires cross-functional collaboration and occasional travel to support system testing and troubleshooting. Duties and Responsibilities Design, develop, and implement high-quality control software for Hyliion’s Karno Power Module, which includes combustion, thermal, pressure, high-voltage and power management systems. Define and conduct tests to verify software and tune control parameters to meet key performance indicators. Prepare reports and technical documentation related to system performance, control strategies, and compliance. Process and analyze data to verify software against engineering specifications, support root cause analysis and for optimizing performance. Ensure on time delivery with quality. Assist product team in defining customer requirements and generate corresponding engineering specifications. Qualifications Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions. Qualifications include: Education, Experience and Cert

pythonaic++
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $216K/yr
15 days ago

At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. Scale Frontier Data is the organization behind the training and evaluation data that frontier labs depend on. We build the systems, tooling, and expert workflows that turn hard human expertise into signals that models can learn from, across reasoning, coding, agentic tool use, and domain expertise. About our Customer Platform team: Our Customer Platform Team plays a pivotal role in integrating our platform with external systems and ensuring seamless, reliable connectivity for both internal users and customers. As the leader of this team, you’ll drive the strategy, architecture, and development of our connectivity solutions, focusing on API integration, distributed systems, and a robust data platform. Your role will be crucial in maintaining and enhancing our platform’s ability to meet the needs of both our internal and external stakeholders. Responsibilities: Own large areas within our product Comfortable working cross functionally, whether that be internal or external customers Build features end-to-end: front-end, back-end, system design, debugging and testing Deliver experiments at a high velocity and level of quality to engage our customers Work across the entire product lifecycle from conceptualization through production Influence the culture, values, and processes of a growing engineering team Inspire and mentor less experienced engineers Collaborating with cross-functional teams to define, design, and ship new product features and experiences. Requirements: At least 7-10 years of relevant experience is preferred Track record of shipping high-quality products and features at scale Desire to work in a very fast-paced environment Abil

awsrestai
View job →
CH
15 days ago

Opportunity Overview: We are seeking a Lead Software Engineer to join our Integrations team. In this role, you will be designing, developing, and scaling highly available healthcare integration systems supporting prior authorization workflows across providers, payers, and delegated entities. You'll direct a fast-paced, autonomous,agile team of software engineers in the design, development, and operational support of a growing enterprise integration platform. This is an opportunity to drive technical excellence at the intersection of healthcare interoperability and modern distributed systems. What you’ll do: Technical Leadership: Provide technical leadership across architecture, system design, platform scalability, reliability, and operational excellence. Platform Engineering: Design and build scalable, resilient, and high-performing systems that support critical business workflows and enterprise integrations. Integration Solutions: Lead the development and maintenance of secure integrations with internal and external platforms, partners, and third-party systems. Cloud & Automation: Drive cloud infrastructure, deployment automation, and software delivery practices that enable reliable and efficient releases. Distributed Systems: Design and support event-driven and distributed architectures that enable scalable and fault-tolerant processing. Operational Excellence: Establish monitoring, observability, and incident response practices to ensure system reliability, performance, and availability. Quality Engineering: Champion automated testing, quality assurance, and engineering best practices throughout the software development lifecycle. Production Support: Lead the resolution of complex production issues and drive continuous improvement in platform stability and operational efficiency. Cross-Functional Collaboration: Partner with product, operations, data, security, and business stakeholders to deliver solutions aligned with organizational goals. Agile Delive

javaawsdocker
View job →
🔔

Get new distributed systems engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More distributed systems engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.