Jobiba hiring network

Distributed Systems Engineer Jobs

1,306 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

SA
Scale AI
📍 San Francisco• Full-time• From $189.6K/yr
15 days ago

Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering MLEs, researchers, data scientists and operators for fast and automatic training and evaluation of LLM's, as well as evaluation of data quality. Scale is uniquely positioned at the heart of the field of AI as an indispensable provider of training and evaluation data and end-to-end solutions for the ML lifecycle. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development. You will be building and optimizing the platform to enable our next generation of LLM training, inference and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you! You will: Build, profile and optimize our training and inference framework Collaborate with ML teams to accelerate their research and development and enable them to develop the next generation of models and data curation Research and integrate state-of-the-art technologies to optimize our ML system Ideally you’d have: Strong excitement about system optimization Experience with multi-node LLM training and inference Experience with developing large-scale distributed ML systems Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc. Strong written and verbal communication skills and the ability to operate in a cross functional team environment Nice to haves: Demonstrated expertise in post-training methods &/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the positi

awsrestai
View job →
E
10 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Everpure is seeking a highly motivated and strategic Pre-Sales Systems Engineer (SE) to join our team in Austria. This role is not simply about selling storage in the cloud or on premise; it's about partnering with critical government entities and Enterprise clients to design, architect, and deploy the secure, simplified, and sustainable data foundation required to support Everpure's continued rapid growth in the region. There's no better way to do that than with Everpure's Enterprise Data Cloud. WHAT YOU'LL DO Develop an exhaustive understanding of what drives a customer’s business and what motivates their decision making Connect the dots from technology solutions, inclusive of the Everpure portfolio and others from the ecosystem, to measurable customer business outcomes Partner closely with account managers, specialists and channel partners to create a seamless and holistic customer experience and strategy to drive revenue growth and net new business Passionately bring to light the advantages of an Everpure solution Refine sales strategy and tactics, taking command of technical responsibilities Delight customers and teammates with your technical leadership and domain expertise on storage products, distributed storage architectures, file systems, and competitive storage offerings in the DAS, NAS and SAN product spaces Take control of evaluations, benchmarks and system configurations Build and deliver techni

REMOTEsqlawsazure
View job →
C
Cloudflare
📍 Hybrid• Full-time• Hybrid• $168K – $231K/yr
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Senior Systems Engineer Available Locations Austin Atlanta Denver Seattle About the Role This is a hands-on software engineering role on the Cloudflare One Appliance team. You will shape and build the Linux-based edge runtime and distributed control plane that connect customer sites to Cloudflare. As a Senior Systems Engineer, you will work across low-level systems, distributed services, and API design while providing technical leadership for majo

typescriptawslinux
View job →

NVIDIA is seeking a Senior Software Engineer to help us develop distributed storage services for AI/ML. In this role you will work closely with the broader NVIDIA team to design and build a reliable, scalable, and efficient storage-as-a-service tailored to AI applications that can be deployed anywhere and scale without limitations. This service supports the whole NVIDIA critical business from graphics drivers to autonomous vehicles to deep learning frameworks. To achieve this goal, we are looking for an engineer with a deep understanding of distributed systems, outstanding design skills, and a track record in building and delivering large-scale distributed services. What you will be doing: Leading the overall architecture and design of our distributed storage service optimized for AI/ML Develop and maintain distributed, robust and scalable Go programs deployed to state of the art open-source ecosystems, including Kubernetes. Develop and maintain user-space applications, containers, Go-bindings, and CLI tools. Building features for a distributed storage service to enhance availability and reliability for large-scale deployments Engaging and collaborating with NVIDIA Research, Computing, Product teams, cross-functional teams, and external customers to deliver Cloud services. Automating distributed storage service end-to-end, including deployment, management, and monitoring What we need to see: Bachelor’s of Science in Computer Science, or related field (or equivalent experience) with 8+ years of industry experience Strong background in developing distributed systems involving Golang, Kubernetes, and Cloud Service Provider integrations Strong track record of delivering distributed services in a variety of distributed computing environments Experience in i

kubernetesartificial intelligenceai
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
15 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As our TT-Distributed Software Engineer, you will develop and optimize distributed software systems that power the most efficient and highest-performing AI and HPC clusters. In this role, you'll work on distributed programming across multiple nodes, utilizing systems programming, inter-node communication, and Tenstorrent’s scalable architectures to advance the state-of-the-art distributed inference and training infrastructure. This role is hybrid, based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Strong C or C++ engineer with solid foundations in systems programming, operating systems, and distributed systems principles. Enthusiastic about distributed computing, including IPC, socket programming, and cluster resource coordination. Comfortable reasoning about scalability, fault tolerance, and performance across multi-node environments. Curious and first-principles thinker who challenges conventional approaches to distributed system design. Motivated to grow into a deep technical expert in large-scale distributed AI infrastructure. What We Need Architect, implement, and optim

awsaic++
View job →
M
Modal
📍 New York• Full-time
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role We're looking for an Engineering Manager to lead a team of highly experienced engineers building the infrastructure that powers Modal's serverless GPU platform. This is a hands-on leadership role — expect to split your time between technical contribution and people management depending on what the team needs. You'll set direction, remove blockers, and build a strong engineering culture as your team tackles hard problems in distributed computing, large-scale data handling, and performance optimization. Who You Are You're an experienced engineering leader who stays close to the work and builds alongside your team when it counts. You earn trust through technical depth, not title. You communicate clearly, help strong engineers move fast without cutting corners, and stay calm and pragmatic under pressure. You care as much about how your team gets to an answer as the answ

javalinuxai
View job →
M
Modal
📍 Sweden• Full-time
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We’re looking for an Engineering Manager to lead a group of highly experienced engineers. This is a hands-on leadership role where you’ll spend roughly half your time on technical contribution and half on people management, depending on the need. You’ll work closely with the team to set direction, remove blockers, and foster a strong engineering culture as they tackle complex systems challenges in distributed computing, large-scale data handling, and performance optimization. Who You Are: We think you are an experienced engineering leader who thrives close to the work and enjoys building alongside their team when needed. You earn trust through technical depth, communicate with clarity, and help great engineers move fast and make sound decisions. You thrive in a fast paced environment, you are pragmatic, calm under pressure, and focused on impact. Requirements: At l

javalinuxai
View job →
PE
24 days ago

About the Role We are a small team of AI builders in Paytm Labs. As a Staff AI Platform Engineer, you will work across inference and agentic systems. You will contribute to Paytm's AI inference platform (Pi), serving internal teams and enterprise customers - running our own coding and domain-specific models (voice, vision, risk, fintech workflows) as well as third-party models. You will also architect and build the platform that enables autonomous AI agents to operate safely and reliably in production - the runtime, orchestration, and developer tooling for agents to reason, plan, use tools, and execute complex multi-step workflows, automating both software development and business processes. You will work at the intersection of LLMs, distributed systems, and production fintech infrastructure, helping define how inference and agentic AI are built and deployed across payments, risk, fraud, collections, support, and developer experience.

E
Everpure
📍 Ontario• Full-time• Remote
15 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE: Are you passionate about driving innovation in the data storage industry? Are you skilled in crafting technical solutions that exceed customer expectations? If so, we have an exciting opportunity for you! Everpure (formerly Pure Storage), a leader in the data storage and flash technology space, is seeking a talented and motivated Senior Pre-Sales Systems Engineer to join our dynamic team. As a Senior Pre-Sales Systems Engineer, you will play a crucial role in understanding our customers' unique challenges and tailoring Everpure solutions to meet their specific needs. Collaborating closely with the sales team, you will act as a technical expert during the sales process, helping to showcase the value of our products and services. WHAT YOU’LL DO: Develop an exhaustive understanding of what drives a customer’s business and what motivates their decision making Connect the dots from technology solutions, inclusive of the Everpure portfolio and others from the ecosystem, to measurable customer business outcomes Partner closely with account managers, specialists and channel partners to create a seamless and holistic customer experience and strategy to drive revenue growth and net new business Delight customers and teammates with your technical leadership and domain expertise on storage products, distributed storage architectures, file systems, and competitive storage offerings in the DAS, NAS and SAN product spaces

REMOTEsqlawsazure
View job →
M
Mongodb
📍 San Francisco• Full-time• From $122K/yr
1mo ago

MongoDB’s mission is to empower innovators to create, transform, and disrupt industries by unleashing the power of software and data. We enable organizations of all sizes to easily build, scale, and run modern applications by helping them modernize legacy workloads, embrace innovation, and unleash AI. Our industry-leading developer data platform, MongoDB Atlas, is the only globally distributed, multi-cloud database and is available in more than 115 regions across AWS, Google Cloud, and Microsoft Azure. Atlas allows customers to build and run applications anywhere—on premises, or across cloud providers. With offices worldwide and over 175,000 new developers signing up to use MongoDB every month, it’s no wonder that leading organizations, like Samsung and Toyota, trust MongoDB to build next-generation, AI-powered applications. Atlas Search is a multi-cloud service that allows users to execute complex full text and vector search queries using the MongoDB Query Language . Our users are free to focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team builds and maintains the instances and supporting infrastructure powering Atlas Search. This platform deploys and monitors search deployments, providing a highly scalable yet observable system for customers and engineers. The Atlas Search product is quickly gaining traction with customers and we are shipping core infrastructure components that enable this growth. This role is based in San Francisco, CA with an in-office or hybrid work model. Successful candidates will have the following qualities: 2+ years of hands-on experience designing, building, testing, and maintaining industrial-strength backend software and automation in complex codebases Experience developing distributed systems and multithreaded applications Familiarity with public cloud platforms, distributed infrastructure, and metric-based development Experience with at least one modern statically typed program

javamongodbaws
View job →
M
Mongodb
📍 Toronto• Full-time• From C$108K/yr
1mo ago

Atlas Search is a multi-cloud service that allows users to execute complex full text and vector search queries using the MongoDB Query Language . Our users are free to focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team is building the cloud-based distributed systems software responsible for the lifecycle of search indexes including: data ingestion, index building, partitioning, performance, availability, and backup management. Our product is quickly gaining traction with customers and we are making core architectural improvements that you will contribute to. This role is based in Toronto, ON hybrid. Successful candidates will have the following qualities: 2+ years of hands-on experience designing, building, testing, and maintaining industrial-strength backend software in a complex codebase Experience developing distributed systems and cloud services Experience with at least one modern statically typed programming language, and interest in working with Java Excellent verbal and written technical communication skills and enthusiasm for collaborating closely with colleagues A growth mindset and the desire to learn quickly through taking on challenges, reflecting on outcomes, and incorporating feedback A strong sense of ownership over their work, from initial design all the way through maintaining code in production You will: Contribute to the design, implementation, and support of projects that improve the scalability of Atlas Search to make using it a seamless experience for even the largest workloads Work with a collaborative team that prioritizes sound technical decision-making and building systems that our customers love and that we are proud of as engineers Have the opportunity to lead projects and own subsystems Provide input on the team’s roadmap and help determine the architecture of our system Success measures: In 3 months you’ll have a solid high-level understanding of what our team does and how we operate.

javamongodbaws
View job →
M
Mongodb
📍 San Francisco• Full-time• From $106K/yr
1mo ago

Atlas Search is a multi-cloud service that allows users to execute complex full text and vector search queries using the MongoDB Query Language . Our users are free to focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team is building the cloud-based distributed systems software responsible for the lifecycle of search indexes including: data ingestion, index building, partitioning, performance, availability, and backup management. Our product is quickly gaining traction with customers and we are making core architectural improvements that you will contribute to. We are looking to speak to candidates who are based in San Francisco, CA for our hybrid working model. Successful candidates will have the following qualities: 2+ years of hands-on experience designing, building, testing, and maintaining industrial-strength backend software in a complex codebase Experience developing distributed systems and multithreaded applications Experience with at least one modern statically typed programming language, and interest in working with Java Excellent verbal and written technical communication skills and enthusiasm for collaborating closely with colleagues A growth mindset and the desire to learn quickly through taking on challenges, reflecting on outcomes, and incorporating feedback A strong sense of ownership over their work, from initial design all the way through maintaining code in production You will: Contribute to the design, implementation, and support of projects that improve the scalability of Atlas Search to make using it a seamless experience for even the largest workloads Work with a collaborative team that prioritizes sound technical decision-making and building systems that our customers love and that we are proud of as engineers Have the opportunity to lead projects and own subsystems Provide input on the team’s roadmap and help determine the architecture of our system Success measures: In 3 months you’ll have a

javamongodbaws
View job →
R
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Engineering Manager, Home Infrastructure The Home Infrastructure team builds the mission-critical backend and data systems that power Roblox’s Homepage and Experience Details Page, two of the highest-traffic surfaces on Roblox. These surfaces reach the vast majority of Roblox’s daily active users and are core drivers of discovery, engagement, retention, and platform growth. We are a full-stack product infrastructure team responsible for content distribution across Roblox. Our systems support multiple modes of user interaction, including exploratory browsing, directed discovery, and personalized content recommendations across the many types of content that make up the Roblox ecosystem. This team sits at the intersection of large-scale distributed systems, machine learning-powered personalization, data infrastructure, and product experimentation. We partner closely with Machine Learning, Data Science, Product, Design, Frontend, Ads, Marketplace, Virtual Economy, and other teams across Roblox to build the platforms that help users find the most relevant and engaging content. As Engineering Manager for Home Infrastructure, you will lead a team of Backend and Data Engineers responsible for the e

awsgitmachine learning
View job →
BE
15 days ago

About Backblaze Backblaze is the object storage leader in the open cloud movement, fueling customer success with cloud storage built purposefully to unlock budgets, unburden administrators, and unleash innovators. Together with our partners, we’re helping customers break free from the restrictive, overpriced legacy solutions that hold them back, and blaze forward with the full power of the open cloud in their hands. Founded in 2007, we scaled the business with less than $3 million in outside funding until 2021, when we did a traditional IPO on the Nasdaq stock exchange. Today, Backblaze generates over $136M ARR and is the leading specialized storage cloud, managing over three billion gigabytes of data storage for 500K+ customers in 175+ countries, including businesses, developers, IT professionals, and individuals. But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us. We’re seeking a Sr. Reliability Engineer ll (DBA) to join our team! About the Role We are seeking a Site Reliability Engineer (SRE) with a DBA (Database Administration) focus to help ensure the stability, scalability, and reliability of our production database systems - primarily Vitess (distributed MySQL) and Cassandra - alongside the rest of our services and infrastructure. This role operates within procedures and runbooks established by our senior DBA SREs, and focuses on building automation, maintaining observability, and supporting incident response to keep customer-facing systems performing at their best. The SRE will collaborate with engineering, product, and operations teams to embed reliability practices into day-to-day development and operations while contributing to tools and processes that improve efficiency and reduce manual effort Key Responsibilities Database Administration Operating and maintaining high-availability database systems — primarily Vitess (distributed MySQL) and Cassandra — against established architecture and runbooks. Op

REMOTEpythonsqlmysql
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $180K/yr
15 days ago

Scale GP is Scale's enterprise Generative AI platform—APIs and infrastructure for knowledge retrieval, inference, evaluation, and intelligent automation. We power mission-critical workflows for leading enterprises, helping teams turn complex data and models into reliable, production-ready AI systems. We're building a new AI Enablement team to create the next generation of agent-powered tools that ground AI in real operational workflows. Our goal: help internal teams demystify their own workflows, then deploy agentic systems that reason over data, take action, and deliver measurable outcomes. We don't build in a vacuum. You'll use our own platform to solve real business problems internally—then selectively commercialize that same stack for customers. What we run on is what we sell. This is a 0→1 team. We're looking for a sharp, product-minded engineer who thrives in ambiguity, moves fast, and loves building systems from scratch alongside customers and cross-functional partners. You'll work closely with product, forward-deployed engineers, data scientists, and applied AI teams to turn real-world problems into scalable production solutions. If you like shipping fast, owning outcomes, and working across the stack—from polished frontends to distributed backends to LLM integrations—this role is for you. What You’ll Do Own full-stack features and projects end-to-end — from design through production deployment — within a larger product area Sample surfaces - Accounting Agents, Finance Copilots, GTM Agents, Agentic Experimentation Platforms Develop reliable backend services in Typescript/Python, work with distributed systems, data pipelines, and AI/ML infrastructure Integrate LLMs, vector databases, and agentic frameworks to power intelligent workflows Ship quickly through tight experimentation loops while maintaining high quality and reliability Adapt across the stack and learn new tools as needed to solve real problems end-to-end Ideal Experience 3+ years of full-tim

typescriptpythonaws
View job →
🔔

Get new distributed systems engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More distributed systems engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.