Jobiba hiring network

Distributed Systems Engineer Jobs

1,301 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale with trillions of data points per day, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams for tens of thousands of companies globally. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team The Datadog Security Libraries team owns the customer-side integrations behind our run-time security products App & API Protection , Workload Protection , and Code Security . Our libraries let customers automatically manage application security risk with continuous, real-time monitoring of vulnerabilities and threats against their web applications, serverless applications, and APIs, in production. Automatically integrated with Application Performance Monitoring (APM) distributed tracing and code-level context, our software empowers development, operations, and security teams to build and run secure applications. As a polyglot team we ship and maintain the security capabilities of Datadog's tracing libraries across .NET , Java , Go , Node.js , Python , Ruby , and PHP , on top of a shared C++ core and a set of HTTP proxy integrations (primarily Envoy, NGINX, and HAProxy). Our code runs inside thousands of production applications around the world. Recent work spans exploit prevention (RASP) and WAF detections, API Security, code security (IAST and SCA), and AI-assisted ("agentic") onboarding, always measured by real product outcomes and operational telemetry. The Opportunity We're looking for a senior, polyglot engineer to contribute across several of our security libraries, with .NET or Java expertise. You'll design and build security integrations and detection features, take them from prototype to production-hardened, and own them operationally as they instrument thousands of applications. As a se

pythonjavanode.js
View job →

About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale with trillions of data points per day, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams for tens of thousands of companies globally. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team The Datadog Security Libraries team owns the customer-side integrations behind our run-time security products App & API Protection , Workload Protection , and Code Security . Our libraries let customers automatically manage application security risk with continuous, real-time monitoring of vulnerabilities and threats against their web applications, serverless applications, and APIs, in production. Automatically integrated with Application Performance Monitoring (APM) distributed tracing and code-level context, our software empowers development, operations, and security teams to build and run secure applications. As a polyglot team we ship and maintain the security capabilities of Datadog's tracing libraries across .NET , Java , Go , Node.js , Python , Ruby , and PHP , on top of a shared C++ core and a set of HTTP proxy integrations (primarily Envoy, NGINX, and HAProxy). Our code runs inside thousands of production applications around the world. Recent work spans exploit prevention (RASP) and WAF detections, API Security, code security (IAST and SCA), and AI-assisted ("agentic") onboarding, always measured by real product outcomes and operational telemetry. The Opportunity We're looking for a senior, polyglot engineer to contribute across several of our security libraries, with .NET or Java expertise. You'll design and build security integrations and detection features, take them from prototype to production-hardened, and own them operationally as they instrument thousands of applications. As a se

pythonjavanode.js
View job →
M
1mo ago

The MongoDB Customer Observability Team is a diverse group of contributors working together to help our users manage MongoDB at global scale. The team is responsible for MongoDB Atlas: our database-as-a-service offering and fastest-growing product, which allows users to deploy globally distributed MongoDB clusters in just minutes. We're seeking a Senior Software Engineer to join our team to tackle exciting challenges within the Observability space. You'll contribute to developing tools and platforms that help our users understand the health and performance of their MongoDB deployments. This includes collecting metrics, monitoring slow queries, and offering actionable insights such as index and schema suggestions that improve the speed, efficiency, and overall reliability of their databases. This role provides a unique opportunity to drive engineering excellence across both dimensions of observability, leveraging technologies being developed within the Customer Observability group and contributing directly to the success of our customers and our product teams. If you're passionate about large-scale systems, digging deep into telemetry data, and building tools that make a real impact both internally and externally, we’d love to have you on board! We are looking to speak to candidates who are based in Dublin for our hybrid working model. We're looking for someone who Has at least 5 years of experience as a backend or full stack engineer Enjoys collaboration and being part of a team Is approachable, curious, and intellectually honest Is a backend engineer with a willingness to take on frontend tasks or a full-stack developer with a bias towards backend Has written backend systems in a compiled language (Java, C#, Go, etc.) Has experience with the design and architecture of a modern, scalable web application Enjoys chasing down difficult problems in a distributed environment and on an database diagnostic/operation level Always strives to expand their knowledg

typescriptjavareact
View job →
M
1mo ago

The MongoDB Query Execution Team is hiring software engineers who want to join us in developing a high performing, reliable and modular distributed query system. Our engineers work on implementing and maintaining execution algorithms, building new query language features, tuning database performance, and more to power our customers' critical workloads. This role can be based out of our Dublin office or remotely in Ireland. Relocation can be supported. Position Expectations Understand and improve current functionality of the MongoDB query engine Contribute high quality C++ code and give and solicit feedback in code reviews Identify, design, implement, test, and support new features related to query performance and robustness, query language enhancements, diagnostics for query performance problems, and integration with other products and tools Work constructively with peers to deliver excellent technical solutions Candidate Profile 5+ years of experience in systems programming Experience in databases and/or data management systems is a huge plus, but not a requirement Hands-on experience building industrial-strength software Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Experience with large code bases, preferably in C++, C, Rust or a similar compiled language B.Sc in Computer Science or similar field, or equivalent practical experience Interest in the theory and practice of database query engines. Hands-on experience or M.Sc./Ph.D in the domain is a plus Success Measures In three months you’ll have contributed to the development of a project slated for the next major version, as well as fixed a few bugs in a minor version of our latest stable release series In six months, you’ll have taken on code review responsibilities and are independently delivering complex functionality and squashing bugs independently In twelve months, you’re leading the development of a new major feature and are h

mongodbawsazure
View job →
EI
17 days ago

100% Remote | Senior Frontend Engineer | Fintech SaaS Firm About the Role We’re looking for a Senior Frontend Engineer to build and maintain scalable, high-performance user interfaces for our communication platform. You’ll work closely with backend engineers, designers, and product managers to deliver exceptional user experiences while keeping performance, maintainability, and scalability at the core. What You’ll Do Develop and maintain responsive UIs using React JS, TypeScript, JavaScript, HTML5, and CSS. Collaborate with cross-functional teams to design and deliver high-quality features. Write clean, maintainable, and well-documented code. Optimize performance with caching and other best practices. Review code, mentor peers, and uphold coding standards. Debug and troubleshoot production issues promptly. Stay current with frontend trends and bring innovative ideas to the team. Job qualifications: 3–8 years’ experience in web development with a focus on scalability. Expert in React JS, JavaScript, TypeScript, HTML5, and CSS. Strong grasp of responsive design, performance optimization, and client-side session management. Familiarity with Git, CI/CD, and distributed development. Excellent problem-solving and collaboration skills. Preferred/Bonus Skills Experience with React Native or other mobile development frameworks. Familiarity with state management libraries like Redux or Zustand. Experience with modern build tools such as Webpack or Vite. A strong portfolio or active GitHub profile showcasing previous work. Why Join Eltropy? Join a high-impact team building mission-critical backend systems for financial institutions. Work on modern technology stacks in a fast-growing SaaS company. 100% remote work with a collaborative, engineering-led culture. Opportunity to own and influence core backend architecture. About Eltropy Eltropy is a rocket ship FinTech on a mission to disrupt the way people acc

javascripttypescriptjava
View job →
DC
17 days ago

Role Overview You’ll be the Principal Software Engineer driving the next generation of a large-scale enterprise SaaS platform. In this role, you combine deep hands-on engineering with high-impact technical leadership, shaping how cloud-native and AI-enabled products are designed and built. You’ll design and deliver secure, scalable, serverless systems on AWS using TypeScript and Node.js, modernize critical platform components, and set the technical direction for multiple teams. You’ll also lead how AI capabilities are integrated across the product ecosystem, ensuring they are transparent, observable, and compliant. If you enjoy system-level thinking, complex distributed architectures, and mentoring senior engineers while still staying close to the code, this role gives you company-wide impact and the opportunity to define the long-term technical vision. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture and delivery of secure, scalable, serverless applications on AWS using TypeScript/Node.js. Define and evolve the platform architecture, driving modernization, performance, resilience, and maintainability. Design and operate distributed, event-driven systems using services like Lambda, DynamoDB, Aurora, S3, and EventBridge. Shape and implement AI-enabled solutions, embedding governance, observability, and responsible AI practices into the platform. Own Infrastructure as Code (e.g., Terraform, AWS CDK, CloudFormation) to reliably provision and manage cloud infrastructure. Mentor senior engineers, influence technical decisions across teams, and clearly communicate complex concepts to diverse stakeholders. These are the essentials you’ll need to get an interview Extensive experience (typically 12+ years) building secure, production-grade software systems. Proven track record architecting and delivering cloud-native, serverless applications on AWS. Strong expertise in Node.js, TypeScript, REST API design, and at leas

typescriptreactnode.js
View job →

About the Role Adobe is seeking a Machine Learning Engineer to join the Adobe Genuine Engineering team. This group protects Adobe's ecosystem from fraud, abuse, and misuse using intelligent systems worldwide. In this position, you will build and develop machine learning models from scratch, including custom transformer-based frameworks, to identify fraudulent actions, stop account sharing, and protect the experience of hundreds of millions of users. You will manage the entire model lifecycle: raw behavioral data and feature engineering, architecture development, large-scale GPU training, deployment, and monitoring. The team is actively building in-house behavioral foundation models that learn identity-preserving representations from long sequences of user activity. This is a role for an engineer who wants to own deep learning systems end-to-end — not consume pre-built ones. Key Responsibilities Build and train deep learning models from scratch, including custom transformer and attention-based architectures for long behavioral event sequences. Own the full training stack: event tokenization, temporal and positional embeddings, self-supervised pretraining (e.g., masked modeling, contrastive learning), and downstream fine-tuning. Train large models efficiently on GPU infrastructure using mixed-precision training, gradient accumulation/checkpointing, efficient attention, and distributed strategies (DDP, FSDP, or equivalent). Build and optimize feature pipelines on Databricks and Spark, transforming raw behavioral events into high-quality model inputs. Translate prototypes into production ML systems — scalable, reliable, and observable — and drive inference performance through architectural and serving-side optimization. Contribute to MLOps practices: experiment tracking, model versioning, CI/CD, automated retraining, and

pythonmachine learningai
View job →
A
1mo ago

About the Role Adobe is seeking a Machine Learning Engineer to join the Adobe Genuine Engineering team. This group protects Adobe's ecosystem from fraud, abuse, and misuse using intelligent systems worldwide. In this position, you will build and develop machine learning models from scratch, including custom transformer-based frameworks, to identify fraudulent actions, stop account sharing, and protect the experience of hundreds of millions of users. You will manage the entire model lifecycle: raw behavioral data and feature engineering, architecture development, large-scale GPU training, deployment, and monitoring. The team is actively building in-house behavioral foundation models that learn identity-preserving representations from long sequences of user activity. This is a role for an engineer who wants to own deep learning systems end-to-end — not consume pre-built ones. Key Responsibilities Build and train deep learning models from scratch, including custom transformer and attention-based architectures for long behavioral event sequences. Own the full training stack: event tokenization, temporal and positional embeddings, self-supervised pretraining (e.g., masked modeling, contrastive learning), and downstream fine-tuning. Train large models efficiently on GPU infrastructure using mixed-precision training, gradient accumulation/checkpointing, efficient attention, and distributed strategies (DDP, FSDP, or equivalent). Build and optimize feature pipelines on Databricks and Spark, transforming raw behavioral events into high-quality model inputs. Translate prototypes into production ML systems — scalable, reliable, and observable — and drive inference performance through architectural and serving-side optimization. Contribute to MLOps practices: experiment tracking, model versioning, CI/CD, automated retraining, and

pythonmachine learningai
View job →

About the Role Adobe is seeking a Machine Learning Engineer to join the Adobe Genuine Engineering team. This group protects Adobe's ecosystem from fraud, abuse, and misuse using intelligent systems worldwide. In this position, you will build and develop machine learning models from scratch, including custom transformer-based frameworks, to identify fraudulent actions, stop account sharing, and protect the experience of hundreds of millions of users. You will manage the entire model lifecycle: raw behavioral data and feature engineering, architecture development, large-scale GPU training, deployment, and monitoring. The team is actively building in-house behavioral foundation models that learn identity-preserving representations from long sequences of user activity. This is a role for an engineer who wants to own deep learning systems end-to-end — not consume pre-built ones. Key Responsibilities Build and train deep learning models from scratch, including custom transformer and attention-based architectures for long behavioral event sequences. Own the full training stack: event tokenization, temporal and positional embeddings, self-supervised pretraining (e.g., masked modeling, contrastive learning), and downstream fine-tuning. Train large models efficiently on GPU infrastructure using mixed-precision training, gradient accumulation/checkpointing, efficient attention, and distributed strategies (DDP, FSDP, or equivalent). Build and optimize feature pipelines on Databricks and Spark, transforming raw behavioral events into high-quality model inputs. Translate prototypes into production ML systems — scalable, reliable, and observable — and drive inference performance through architectural and serving-side optimization. Contribute to MLOps practices: experiment tracking, model versioning, CI/CD, automated retraining, and

pythonmachine learningai
View job →
M
Modal
📍 New York• Full-time
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We’re looking for an Infrastructure Security Engineer to design and secure the core systems that power our platform. This role focuses on building security directly into our infrastructure—from container isolation and orchestration to identity and secrets management in a multi-tenant, cloud-native environment. You’ll work closely with engineering teams to define secure primitives and ensure our platform is resilient, scalable, and trustworthy by design. This is a hands-on, deeply technical role focused on real systems, not compliance or policy. What You'll Do: Platform & Runtime Security Design and improve isolation mechanisms for multi-tenant workloads (containers, sandboxing, execution environments) Strengthen boundaries between customers, workloads, and internal systems Identify and mitigate risks in distributed, dynamic compute environments Container &

awsgcpkubernetes
View job →
M
Midjourney
📍 San Francisco• Full-time
1mo ago

What you’ll do Act as the technical lead for large parts of the scanner platform: system architecture, codebase structure, and long-term maintainability. Own core runtime foundations: distributed control, state management, fault handling, and reliability. Drive engineering rigor: testability, code quality, review standards, performance regression prevention, and release processes. Build robust observability: logs, metrics, traces, and replayable diagnostics (with privacy constraints). Collaborate with hardware and recon/ML teams to define interfaces, data contracts, timing/synchronization, and failure modes. Lead complex refactors (e.g., message passing / RPC boundaries, modularization, concurrency model) without halting forward progress. What we’re looking for Deep software architecture experience for real-world systems: robotics, instrumentation, medical devices, or other complex distributed products. Strong Python and concurrency background (asyncio, multiprocessing, profiling, performance engineering). Track record of shipping systems that are observable, debuggable, and resilient. Strong technical leadership: clarity, pragmatic trade-offs, and mentoring. Useful experience Building but rock-solid systems: clear interfaces (gRPC/protobuf or equivalent), strong state modeling, and failure handling. High-leverage engineering habits on a lean team: good tests, CI, reproducible dev environments, and fast code review. Practical performance + concurrency work in Python (asyncio, profiling, multiprocessing) and comfort debugging distributed behavior. Security-minded device software: safe defaults, encrypted data paths, and disciplined handling of PII/PHI. Operational thinking: remote updates/management, excellent logging, and diagnostics that make real hardware debuggable.

pythonaigo
View job →
M
1mo ago

Join the MongoDB Server Query team, and help us build a world-class distributed open source query engine. Our team plays a crucial role in the experience and performance of data processing. We are responsible for the MongoDB Query Language and the lifecycle of each query, from parsing to optimization to plan selection and finally execution. This also includes our geospatial search and update subsystems. Our global team is growing fast. In North America, we have a presence across the US and Canada including New York, West Coast, Toronto. In Europe, we have a presence in Dublin, Germany, France, Netherlands, UK, Bulgaria, Spain and Italy currently. We support office-based and remote work and align projects with convenient work hours for each time-zone. We have tons of interesting problems to solve with direct impact on users for transactional, time-series and analytical workloads. We need your help to design and build the heart of a distributed, flexible schema, document database. We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. Candidate Profile 3+ years of experience in data intensive environments Hands-on experience building industrial-strength software Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Experience with large code bases B.Sc in Computer Science or similar field, or equivalent practical experience Experience in C++ and in developing database systems is a plus Interest in the theory and practice of database query engines. Hands-on experience or M.Sc./Ph.D in the domain is a plus Position Expectations Understand and improve current functionality of the MongoDB query engine Identify, design, implement, test, and support new features related to query performance and robustness, query language enhancements, diagnostics for query performance problems, and integration with other products and tools Work with other engineers to

mongodbawsazure
View job →
C
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Machine Learning Engineer on the CX Intelligence team within Enterprise Applications and Architecture, you'll build the AI-powered conversational systems that connect Coinbase's Help Center, chatbots, and agent workflows. The team owns the multi-agent platform powering Coinbase Chat and agent tooling, partnering with Conversation Design, CX, and Engineering to deliver secure, scalable automated support. You'll lead the design and implementation of a unified orchestration layer that coordinates interactions between vendor AI, internal multi-agent systems, and human participants, directly improving how millions of customers get help. What you'll do: Architect and deploy the orchestration layer that manages state transitions, context sharing, and intent routing across vendor and internal LLM frameworks in a distributed conversational environment. Build production-grade Python services that bridge advanced ML/AI research with reliable, measurable customer-facing products. Lead end-to-end project execution for complex ML initiatives, managing priorities, technical trade-offs, and cross-functional dependencies from design through delivery. Establish best practices for system design, coding standards, and AI/ML development workflows across the team. Mentor engineers on architectural integrity and modern AI/ML patterns, raising the technical bar for the broader team. Co

REMOTEpythonawsmachine learning
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i

pythonreactnode.js
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a software engineer on the Scaling team, you’ll help build and optimize the low-level stack that orchestrates computation and data movement across OpenAI’s supercomputing clusters. Your work will involve designing high-performance runtimes, building custom kernels, contributing to compiler infrastructure, and developing scalable simulation systems to validate and optimize distributed training workloads. You will work at the intersection of systems programming, ML infrastructure, and high-performance computing, helping to create both ergonomic developer APIs and highly efficient runtime systems. This means balancing ease of use and introspection with the need for stability and performance on our evolving hardware fleet. This role is based in San Francisco, CA, with a hybrid work model (3 days/week in-office). Relocation assistance is available. In this role, you will: Design and build APIs and runtime components to orchestrate computation and data movement across heterogeneous ML workloads. Contribute to compiler infrastructure, including the development of optimizations and compiler passes to support evolving hardware. Engineer and optimize compute and data kernels, ensuring correctness, high performance, and portability across simulation and production environments. Profile and optimize system bottlenecks, especially around I/O, memory hierarchy, and interconnects, at both local and distributed scales. Develop simulation infrastructure to validate runtime b

pythonawsrest
View job →
🔔

Get new distributed systems engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More distributed systems engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.