Jobiba hiring network

Distributed Systems Engineer Jobs

1,301 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

L
Lyft
📍 New York, NY• Full-time• ₹1.5L – ₹1.9L/yr
1mo ago

At Lyft, our mission is to improve people’s lives with the world’s best transportation. To do this, we start with our own community by creating an open, inclusive, and diverse organization. We have an ambitious goal to strengthen our international presence by growing a life-changing product, and your efforts will play an essential role in our collective success. Growth teams are at the heart of our products and decision-making to drive business growth. We’re looking for passionate, driven engineers to build systems that empower our users (both Drivers and Riders) to make the most effective use of Lyft’s products and experiences by making them more predictive, personalized, and adaptive. We’re looking for someone who can build & own innovative applications, passionate about solving problems with distributed computing in building reliable systems, and is excited about working in a fast-paced, innovative, and collegial environment. As a Senior Software Engineer, with your technical expertise you will build and own innovative applications with a desire to achieve exceptional design fidelity, usability, accessibility, and performance. You will be responsible for sound technical execution of projects through hands-on development, quality-assurance, and prototyping. These projects will require close collaboration with our product managers, user interface designers, brand producers, data scientists, and engineering teams. You’ll be managing multiple exciting projects at once, so being able to juggle these initiatives is key! We’re also counting on you to be a champion for best engineering practices, guiding engineers, designers, and management along the way. We are looking for candidates who are self starters and have a proven track record of delivering software solutions that can solve critical business needs. We’re looking for a candidate who thrives in tackling complex, ambiguous challenges and can

pythonsqlaws
View job →
L
Lyft
📍 San Francisco, CA• Full-time• ₹1.5L – ₹1.9L/yr
1mo ago

At Lyft, our mission is to improve people’s lives with the world’s best transportation. To do this, we start with our own community by creating an open, inclusive, and diverse organization. We have an ambitious goal to strengthen our international presence by growing a life-changing product, and your efforts will play an essential role in our collective success. Growth teams are at the heart of our products and decision-making to drive business growth. We’re looking for passionate, driven engineers to build systems that empower our users (both Drivers and Riders) to make the most effective use of Lyft’s products and experiences by making them more predictive, personalized, and adaptive. We’re looking for someone who can build & own innovative applications, passionate about solving problems with distributed computing in building reliable systems, and is excited about working in a fast-paced, innovative, and collegial environment. As a Senior Software Engineer, with your technical expertise you will build and own innovative applications with a desire to achieve exceptional design fidelity, usability, accessibility, and performance. You will be responsible for sound technical execution of projects through hands-on development, quality-assurance, and prototyping. These projects will require close collaboration with our product managers, user interface designers, brand producers, data scientists, and engineering teams. You’ll be managing multiple exciting projects at once, so being able to juggle these initiatives is key! We’re also counting on you to be a champion for best engineering practices, guiding engineers, designers, and management along the way. We are looking for candidates who are self starters and have a proven track record of delivering software solutions that can solve critical business needs. We’re looking for a candidate who thrives in tackling complex, ambiguous challenges and can

pythonsqlaws
View job →
S
Stripe
📍 New York• Full-time
1mo ago

Who we are About the team Link is a digital wallet designed for effortless and secure online payments and digital transactions. With Link, consumers enjoy convenience and peace of mind—it works on any device or browser, is backed by the highest security mechanisms, offers purchase protections on eligible items, and ensures seamless and quick payments. Across the Link Engineering org, we focus on building delightful payment experiences and allowing our global consumer base to pay with their preferred payment methods. Our team's work spans the entire stack from front-end experiences, to infrastructure that supports low-latency transactions, to intelligent systems that help protect consumers and merchants from bad actors. Team matching for one of the subteams will begin during final stages. Please note we may also consider you for different orgs based on your experience, location, etc. More information on our team matching process can be found here . What you'll do We're looking for engineers who want to make an impact on payments at a global scale. You'll play a key role in expanding our product suite and the infrastructure that supports it. Our team collaborates with many cross-functional teams and many other teams across Stripe to deliver innovative solutions that address evolving user needs. Responsibilities Build and design the next generation of Stripe products to meet the high-growth needs of our company and customers for years to come Debug and solve critical production issues across services and multiple levels of the stack Mentor engineers to help them grow Collaborate with stakeholders across the company to build new features at large-scale, while improving internal engineering standards, tooling, and processes Collaborate effectively in a distributed and hybrid team, maintaining open communication and strong connections with colleagues Who you are We're looking for someone who meets the minimum requirements to be considered for the role. If you meet these r

awsdockerkubernetes
View job →
S
Stripe
📍 Chicago• Full-time• Remote
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Stripe Infrastructure is responsible for the reliability, scale, performance, and cost of Stripe's systems and the productivity and sentiment of Stripe's people. You may work on a wide variety of critical business areas including: Core Infrastructure—We're the home for Stripe's critical tier0 infrastructure systems (Compute, Networking, DocumentDB, Distributed Caching and High assurance engineering). We build the foundational platform for Stripe products and services to allow them to operate at scale. We drive reliability, availability, efficiency, and scalability of these systems. Developer Infrastructure—We're responsible for the productivity of all developers at Stripe. Ensure Stripe's engineers have a reliable, fast, and easy-to-use inner dev loop to maximize productivity while building everything from low-latency microservices to large-scale data pipelines and machine learning models. Data Infrastructure—We're responsible for offering data serving infrastructure spanning across data warehouse analytics, streaming analytics, and search capabilities. The stack is supported by a collection of internally developed large-scale distributed services and several popular open-source technologies like Trino/Presto, Apache Pinot, Hive Metastore, ElasticSearch etc. The systems we own support all of the data serving needs of high-scale services and thousands of individual Stripes across the company. Admin Platform—We empower Stripes to quickly

REMOTEjavarestmicroservices
View job →
O
1mo ago

About the Team The Monetization team is a new cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products—including next-generation ads experiences—that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to help build the machine learning infrastructure that powers OpenAI’s monetization and ads systems. In this foundational role, you’ll design and develop the platform layer that enables teams to build, train, deploy, serve, monitor, and continuously improve machine learning models used across advertising and monetization products. You’ll work across the full ML lifecycle, from large-scale data pipelines and feature infrastructure to training systems, model serving, experimentation platforms, and monitoring frameworks. The systems you build will support high-throughput, low-latency advertising workloads while maintaining strict standards for reliability, privacy, security, and performance. This role sits at the intersection of machine learning systems, distributed infrastructure, and monetization, offering the opportunity to shape the core platforms that help translate model innovation into m

awsrestmachine learning
View job →
O
1mo ago

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

pythonawslinux
View job →
O
1mo ago

About the Team Training Runtime designs the core distributed runtime that powers everything from early research experiments to frontier-scale model runs. We work on building robust, scalable, high performance components to support our distributed training workloads. Our priorities are to maximize the productivity of our researchers and our hardware, with the goal of accelerating progress towards AGI. Within Training Runtime, the Process Management team develops the distributed OS responsible for launching, coordinating, and supervising the large numbers of processes that make up modern training workloads. Our runtime sits beneath training frameworks and on top of research infrastructure, ensuring jobs run reliably across massive clusters while maintaining performance, stability, and observability. Success for us is measured by both system reliability and researcher velocity - enabling ideas to scale from experiments to production training runs. About the Role As a Training Runtime: Process Management Engineer , you will work on the software that ties thousands of computers together and exposes them as a unified system. This system has to serve individual researchers running multiple parallel experiments, as well as our largest training runs spanning 100’s of thousands and even millions of machines and accelerators. This requires easy to use, introspectable systems that can promote a fast debugging and development cycle, as well as relentless optimization for scale while maintaining stability and performance throughout. You will work primarily in Rust , building high-performance asynchronous systems with a strong emphasis on performance, correctness, and scalability. Working at this scale and at the frontier of AI development poses novel challenges. Out-of-the-box approaches often don’t work. The problems you will be working on are highly ambiguous and require strong design judgment as well as proficient execution to advance the state of our infrastructure. We’re loo

pythonawslinux
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing

pythonawskubernetes
View job →
PE
Private Employer
📍 Bangalore, Karnataka• Full-time• Hybrid
1mo ago

About the team DevOps - because even the best developers need heroes. Whether you agree or not, a Senior DevOps Engineer is definitely nothing less than a hero to us - and with good reason! When 5% of Indian households shop with us, it’s important to build resilient systems to manage millions of orders every day. We’ve done this – with zero downtime! 😎 Sounds impossible? Well, that’s the kind of Engineering muscle that has helped Meesho become the e-commerce giant that it is today. We value speed over perfection, and see failures as opportunities to become better. We’ve taken steps to inculcate a strong ‘Founder’s Mindset’ across our engineering teams, making us grow and move fast. We place special emphasis on the continuous growth of each team member - and we do this with regular 1-1s and open communication. As Senior DevOps Engineer, you will be part of self-starters who thrive on teamwork and constructive feedback. We know how to party as hard as we work! If we aren’t building unparalleled tech solutions, you can find us debating the plot points of our favourite books and games – or even gossipping over chai. So, if a day filled with building impactful solutions with a fun team sounds appealing to you, join us. About the role As DevOps Engineer with us, your primary responsibility will be to create systems that serve as the brains of complex distributed products. You will closely mentor younger engineers on the team. You will work on code modularity, scalability, reusability and contribute to team building in the Engineering org.

aigodevops
View job →
N
9 days ago

We are now looking for a Senior Deep Learning Software Engineer, PyTorch. NVIDIA is hiring software engineers to design and build tools used by AI engineers across the world to design, develop, and deploy AI applications scalable across thousands of GPUs. This position will embed you in an ambitious and diverse team that influences all areas of NVIDIA's AI platform as well as directly contributes to PyTorch, a premiere deep learning framework. In this role you will work with multiple teams at NVIDIA across fields, as well as collaborate internationally with the PyTorch community to develop the best AI platform in the world. What you will be doing: Design and build PyTorch components that run efficiently on supercomputers with 1000s-100ks of GPUs. Collaborate with NVIDIA’s hardware and software teams to improve the overall GPU performance in PyTorch. Design, build and support production AI solutions used by enterprise customers and partners. Work with internal applied researchers to improve their AI tools. What we need to see: BS in Computer Science or Engineering (or equivalent experience). 3+ years professional experience in deep learning. Proficient with C++ programming. Strong understanding of systems software and interfaces. Demonstrated experience with Thread and Distributed Parallel Programming Demonstrated background developing large software projects. Strong verbal and written communication skills Ways to stand out from the crowd: Contributions and participation in the open source community. Familiarity with deep learning compilers. Familiarity with deep learning modeling trends. Background with CUDA Programming as well as Python.

pythonartificial intelligenceai
View job →

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler’s high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You’ll Do (Role Expectations) Maintain h

pythonkuberneteslinux
View job →

About Us Are you ready to build the future of the supply chain? At Gather AI, we’re not just creating software; we’re pioneering a new era of warehouse intelligence. We’ve developed a groundbreaking, vision-powered platform that uses autonomous drones and existing equipment to capture real-time data, completely digitizing workflows that have historically been manual and error-prone. This means facilities operate smarter, safer, and more efficiently, ultimately redefining “on-time, in full” delivery. If you’re looking for an opportunity to contribute to truly transformative technology and make a significant impact in a vital industry, Gather AI is the place for you. We’re leading the charge in the rapidly evolving robotics industry, and we invite you to join us in reshaping the global supply chain, one intelligent warehouse at a time. About the Team You’ll be part of our Engineering organization, working closely with our Full Stack and WMS integration teams. The group builds and operates the production systems behind Gather AI’s warehouse intelligence platform, including APIs, integrations, data pipelines, databases, and application services used by our customers and internal teams every day. This is a small, distributed engineering team where engineers work across a broad technical surface area and have the opportunity to see their work through the full software development lifecycle, from design and development through deployment, production monitoring, troubleshooting, and improvement. About the Role We are hiring an SDE II, Full Stack, for our India-based team. You'll build and run the web platforms behind Gather AI's products - Drone Vision, MHE Vision and SAGE, from the customer-facing dashboards down to the APIs, Integration layers and data pipelines that serve them. This is a fully remote, hands-on engineering role for someone early in their career who already has experience building production software and wants broader ownership. You’ll work alo

REMOTEpythonnode.jssql
View job →
AC
17 days ago

Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. The Mission As a Lead Software Engineer working on the Appian platform, your mission is to be the technical anchor for our most ambitious engineering initiatives. You will ensure Appian is always fast, scalable, and resilient, regardless of the complexity our customers demand. While a Senior Engineer focuses on feature health, you are focused on platform evolution . You will lead the design and orchestration of high-stakes systems, integrating cutting-edge AI into the core fabric of our product. This role demands a balance of elite coding mastery, visionary system architecture, and the ability to align technical roadmaps with Appian's overarching business strategy. Key Responsibilities The Lead Software Engineer isn't just "the best coder"—they are the architect of the team’s success and the guardian of our technical future. Technical Strategy & Roadmap: Work alongside Product Leadership to transform long-term business goals into scalable technical milestones. You don't just build features; you define the technical "North Star." Advanced Architecture: Lead the design of distributed, event-driven, and cloud-native systems. You are responsible for making high-level decisions on system boundaries, data consistency, and integration patterns. AI Orchestration & Innovation: Drive the adoption of AI-native engineering and an AI-native culture within the team. This includes using AI tools to improve developer productivity and implementing AI features into the platform to make the Appian platform more intelli

javasqlaws
View job →
V
Verse
📍 San Francisco• Full-time• $150K – $240K/yr
17 days ago

Location: San Francisco, CA (Remote/Hybrid Available) What is Verse? The race to AI has become the race to power. Every breakthrough in artificial intelligence depends on one thing: access to electricity. But across the country, aging grid infrastructure and years-long interconnection queues are slowing the deployment of the data centers that will power the next generation of innovation. Solving this challenge isn't just about energy—it's about unlocking the future of AI. At Verse, we're building the energy intelligence platform for the AI economy. Our software helps the world's largest energy consumers achieve faster, cheaper, and cleaner power by combining real-time control of energy assets with complete visibility into their energy portfolio. Backed by Bessemer Venture Partners, GV, Coatue, and NVIDIA, and built by pioneers in grid-scale batteries, energy markets, and enterprise software, we're redefining how the world's most ambitious organizations access and manage energy. The Role As a Software Engineer focused on Fleet Telemetry & Control at Verse, you will be working closely with our energy solutions partners to design, implement, and test distributed energy resource controls and telemetry software on customer hardware at sites around the world. You will be part of a dynamic, high-performance team building applications directly on bare-metal or on hardware-level virtualization platforms. As an advanced technical leader in network programming and state management development, engineering teams will look to you for best standards and practices for interfacing with on-premises grid assets using solutions you will build and maintain. Key Responsibilities Foster a culture and mindset of well-designed systems, test-driven software, and transparent communication with a high caliber of mutual respect and consideration for stakeholders Mentor and support career and junior level engineers in their fleet telemetry and control software career development

pythonsqlai
View job →
Z
Zscaler
📍 Netherlands• Full-time• Remote
17 days ago

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer to join our Cloud Infrastructure & Operations team. This is a remote role based in the Netherlands, reporting to the Senior Director, Software Engineering. As a Staff SRE, you will leverage your expertise in Linux/UNIX System Administration to build scalable infrastructure and manage platforms like Kubernetes using automation and high security standards. You will troubleshoot complex Linux networking and security issues, manage firewall technologies, and ensure secure access across our global platforms and applications. What you’ll do (Role Expectations) Create and maintain highly scalable solutions based on KVM LINUX, Kubernetes, and Public Cloud Providers Analyze and troubleshoot systems performance and issues across the OS and Applications Maintain platform security and observability using nftables and robust monitoring tools Manage and deploy systems and s

REMOTEpythonawsdocker
View job →
🔔

Get new distributed systems engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More distributed systems engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.