At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Lyft Business Product Platform team builds the systems and experiences that power Lyft's B2B products — enabling companies, organizations, and their employees to seamlessly access Lyft's transportation network. We sit at the intersection of product and platform, owning both the customer-facing features and the underlying infrastructure that makes them reliable at scale. Our work directly impacts how businesses integrate with Lyft, how admins manage their programs, and how millions of riders get where they need to go. Responsibilities: Drive architecture and technical design for systems that are highly available, scalable, and built to last — not just for today's requirements but for where the product is heading Own features end-to-end: from shaping the technical spec and design through to production rollout and operational health Think critically about how AI capabilities can be incorporated into Lyft Business products to improve the experience for business admins and riders — and bring that perspective into roadmap and architecture conversations Make well-reasoned trade-off decisions and communicate them clearly to peers, leads, and cross-functional partners Write clean, well-tested, maintainable code and hold a high bar for the same in code reviews Partner across engineering, product, and design to align on direction and get buy-in on technical approaches Proactively engage in incident response, contributing both to resolution and to long-term reliability improvements Grow the team's technical culture through design reviews, tech talks, and mentorship Experience: 5+ years of software engineering experience, with a track record of designing and shipping production systems at scale Strong system design instincts — you can reason through distributed systems trade-offs, identify failure modes, and
Jobiba hiring network
Distributed Systems Engineer Jobs
1,306 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Marketplace teams are at the heart of our products and decision-making, owning everything from rider pricing to driver earnings, incentives, and efficient matching. We’re looking for passionate, driven engineers to build systems that empower our riders and drivers to have the best transportation experience possible through prediction, adaptivity, and personalization. We’re looking for someone who is excited about working in a fast-paced, innovative, and impactful environment to create reliable solutions to distributed computing, ML, and data problems. The Pricing team is a centerpiece of Lyft’s Marketplace org, determining prices for all rideshare products and supporting new initiatives. We work with Product & Science to solve and implement complex pricing requirements, balancing the needs of riders, drivers, and the business goals. As an owner of one of the most critical flows in the company, you will work on a wide array of challenges such as latency-sensitive concurrency problems, large scale distributed systems, and experimentation. If you’re interested in playing a large part in demand / supply management and improving the Lyft customer experience, this could be a great fit for you. Responsibilities: Help define the roadmap and architecture based on technology and business needs Unblock, support, effectively communicate, and obtain buy-in across teams to achieve results Lead projects of multiple people from idea to positive execution Write clear, scalable and clear design documentation Write well-crafted, well-tested, readable, maintainable code Utilize your expertise in Python, Golang, AWS to deliver robust and scalable solutions Participate in code reviews to ensure code quality and distribute knowledge Proactively participate in resolving ongoing incidents Share your kno
About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. Our team builds the software, tooling, and operational systems that help manage this fleet at scale. We work across production engineering, distributed systems, capacity management, and operational automation to improve reliability, reduce manual work, and make better use of available compute. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will develop the systems that help manage the GPU fleet powering ChatGPT, including tooling for fleet health, capacity planning, operational automation, and incident response. You will work closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and compute utilization. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference. Develop systems for capacity planning, fleet health monitoring, and resource utilization. Automate operational workflows, including incident detection, diagnosis, and response. Identify and address bottlenecks affecting fleet reliability, scalability, and performance. Partner with infrastructure, research, and product engineering teams to improve the compute platform. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems. Have a background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering. Have built software that automates operational workflows and reduces manual work. Have worked with distributed infrastructure, cluster orchestration, or large-scale int
By applying to this role, you will be considered for Research Engineer roles across all teams at OpenAI. About the Role As a Research Engineer here, you will be responsible for building AI systems that can perform previously impossible tasks or achieve unprecedented levels of performance. We're looking for people with solid engineering skills (for example designing, implementing, and improving a massive-scale distributed machine learning system), writing bug-free machine learning code, and building the science behind the algorithms employed. The most outstanding deep learning results are increasingly attained at a massive scale, and these results require engineers who are comfortable working in large distributed systems. We expect engineering to play a key role in most major advances in AI of the future. We expect you to: Have strong programming skills Have experience working in large distributed systems Be excited about OpenAI’s approach to research Nice to have: Interested in and thoughtful about the impacts of AI technology Past experience in creating high-performance implementations of deep learning algorithms About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment
About the Team OpenAI’s Applications Engineering organization builds and operates the products (such as ChatGPT & Codex) that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role We’re hiring Backend Software Engineers to design and implement safe services and infrastructure that power our core products. What You’ll Do Architect, build, and improve scalable backend systems and APIs. Drive performance, reliability, and safety across distributed services. Implement data storage, retrieval, compute, and integration solutions. Participate in long-term architectural planning and technical design reviews. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. You Might Thrive Here If You: Have strong experience with distributed systems, APIs, and backend languages (e.g., Go, Python, Rust, C++). Have experience setting up and maintaining production backend services and data pipelines. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Enjoy building resilient services that handle large scale and complexity. Are self-directed and enjoy figuring out the best way to solve a particular problem Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role On the Accelerators team, you will help OpenAI evaluate and bring up new compute platforms that can support large-scale AI training and inference. Your work will range from prototyping system software on new accelerators to enabling performance optimizations across our AI workloads. You’ll work across the stack, collaborating with both hardware and software aspects - working on kernels, sharding strategies, scaling across distributed systems, and performance modeling. You'll help adapt OpenAI's software stack to non-traditional hardware and drive efficiency improvements in core AI workloads. This is not a compiler-focused role, rather bridging ML algorithms with system performance - especially at scale. In this role, you will: Prototype and enable OpenAI's AI software stack on new, exploratory accelerator platforms. Optimize large-scale model performance (LLMs, recommender systems, distributed AI workloads) for diverse hardware environments. Develop kernels, sharding mechanisms, and system scaling strategies tailored to emerging accelerators. Collaborate on optimizations at the model code level (e.g. PyTorch) and below to enhance performance on non-traditional hardware. Perform system-level performance modeling, debug bottlenecks, and drive end-to-end optimization. Work with hardware teams and vendors to evaluate alternatives to existing platforms and adapt the software stack to their architectures. Contribute to runtime improvements, compute/communication over
About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are the Snowflake Metadata team. We own Snowflake’s metadata systems that make it easy for customers to query, modify and manage their petabyte-scale data. We develop distributed systems that store and maintain metadata, transaction frameworks that power Snowflake’s query and DML capabilities, asynchronous systems that provide time travel and lifecycle management capabilities and entity metadata supporting DDL capabilities. We also build foundational capabilities that deliver global features like cross-region replication (Snowgrid), data sharing, and data marketplace. AS A PRINCIPAL SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Solve real business needs at large scale by applying your software engineering and analytical problem solving skills. Design, develop and support fault-tolerant scalable distributed systems for our Snowgrid and Data Sharing teams. Create architecture and design, influence our product roadmap, and take ownership and responsibility over new projects. Analyze fault-tolerance and high availability issues, performance and scale challenges, and solve them. Mentor and grow junior engineers. Understand trade-offs between consistency, performance and costs to build solutions which can meet the demands of rapidly growing services. Ensure operational readiness of
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Forward Deployed Infrastructure Engineers (FDIEs) build, operate, and maintain the infrastructure that powers Palantir’s platforms and production deployments. As an FDIE intern, you’ll work alongside full-time FDIEs to deploy and operate Palantir software across real production environments, automate manual processes, and develop novel solutions to infrastructure challenges using tools like Foundry and Apollo. Every day looks different — you might be debugging a distributed systems issue, building automation to replace a manual runbook, or designing infrastructure improvements that scale across multiple deployments. You’ll be treated as a full member of the team, with real ownership over the work you take on. Core Responsibilities As an FDIE intern, your responsibilities look similar to those at a small startup, with the resources, stability, and mentorship of an established tech company. You’ll work in small teams with minimal supervision and own end-to-end execution of real infrastructure projects. Your day might span discussing systems architecture with fellow engineers, debugging a production issue, building automation to eliminate a manual process, or deploying new Palantir products across production environments. FDIE interns are treated just like full-time engineers, with significant freedom and ownership over their work. Specifically, you can expect to: Deploy and operate Palantir software across production environments, including monitoring, alerting, configuration management, and upgrades Debug, improve, and optimize Palantir’s services and infra
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Forward Deployed Infrastructure Engineers (FDIEs) build, operate, and maintain the infrastructure that powers Palantir’s platforms and production deployments. As an FDIE intern, you’ll work alongside full-time FDIEs to deploy and operate Palantir software across real production environments, automate manual processes, and develop novel solutions to infrastructure challenges using tools like Foundry and Apollo. Every day looks different — you might be debugging a distributed systems issue, building automation to replace a manual runbook, or designing infrastructure improvements that scale across multiple deployments. You’ll be treated as a full member of the team, with real ownership over the work you take on. Core Responsibilities As an FDIE intern, your responsibilities look similar to those at a small startup, with the resources, stability, and mentorship of an established tech company. You’ll work in small teams with minimal supervision and own end-to-end execution of real infrastructure projects. Your day might span discussing systems architecture with fellow engineers, debugging a production issue, building automation to eliminate a manual process, or deploying new Palantir products across production environments. FDIE interns are treated just like full-time engineers, with significant freedom and ownership over their work. Specifically, you can expect to: Deploy and operate Palantir software across production environments, including monitoring, alerting, configuration management, and upgrades Debug, improve, and optimize Palantir’s services and infra
We are seeking a highly skilled and hard-working Senior Test Developer / test engineer to join our multifaceted Enterprise Software QA team. This role offers an outstanding opportunity to leave your mark on the design, construction, optimization and testing of large-scale infrastructure for various foundational NVIDIA unified cloud services and data center offerings. If you are a dedicated engineer with strong expertise in cloud infrastructure and distributed systems and want to apply your skills with AI tools, this role could fit you perfectly. You will thrive in an exciting, innovative environment. What you'll be doing: Work with development teams on test plans for all layers of SW stack for cloud infrastructure, execution, reviews, failure analysis and assessing overall quality and risk. Work with customer PMs on software issues including technical feedback from OEMs and CSPs. Develop key benchmarks to track execution and deploy process improvements to improve efficiency Leverage AI skills to expedite the test scope, test plan, execution and automation workflows. Lead NVIDIA Cloud and Data Center bring up activities which will involve validation, reporting, working with engineering to debug issues, providing design input at times, adding coverage in different areas. Design, develop and maintain CI/CD pipelines for continuous testing in cloud environments when needed. Perform performance, scalability, and reliability testing of cloud services. Implement and maintain test environments in cloud platforms such as AWS, Azure, or Google Cloud. Supervise the infrastructure to alert on significant events, ensuring the highest level of system performance and reliability. Work with various different partner teams to ensure availability of clusters to test on and take the lead in resolve all issues. Working with tea
We're looking for an ML Data & Platform Engineer to own the infrastructure that powers our speech AI models: the pipelines that source and prepare training data, and the platform that trains, evaluates, and serves them in production. Speech AI has a data problem most ML teams don't, and you'll be at the centre of solving it, working as part of our ML team to remove friction across the entire lifecycle and get better models into production faster. This is a broad, cross-functional role suited to someone who enjoys working across the full stack: data infrastructure, distributed systems, and production ML, and who takes ownership of problems end to end rather than waiting to be told what to fix. What you'll do Designing, building, and maintaining scalable data pipelines for ingesting, transforming, validating, and storing large datasets used to train our models Developing and maintaining web scraping and data acquisition solutions to keep training datasets fresh, high-quality, and available at scale Building and operating the infrastructure that lets the ML team deploy and evaluate new models quickly, and that serves models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning data infrastructure, training, and inference Continuously improving our data and MLOps practices, and helping shape the roadmap for how our platform evolves as we scale What you'll need Strong proficiency in Python and SQL, with a solid backend or data engineering foundation Hands-on experience with containerisation and orchestration (Docker, Kubernetes), and working with a major cloud provider Experience building data pipelines and ETL/ELT processe
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Exa team and lead the charge in redefining enterprise storage by unifying block, file, and object protocols across hybrid-cloud environments. You will combine deep technical expertise in distributed systems with hands-on people leadership to guide architectural decisions and mentor high-impact engineers. This is a unique opportunity to build new engineering teams from the ground up and drive industry-leading innovation alongside Product and Architecture partners. Your work will directly impact how customers consume, scale, and operate mission-critical storage infrastructure. WHAT YOU'LL DO Establish and scale a new engineering organization focused on critical Exa platform services, ensuring a foundation of long-term success, technical excellence, and a high-performing culture. Own the successful delivery of complex, high-scale engineering features for the Exa platform, ensuring world-class security, reliability, and availability across multi-array and hybrid-cloud deployments. Define the technical vision and execution roadmap in close partnership with Product Management and Architecture, translating customer needs into a measurable business impact for Pure Storage. Drive a culture of operational rigor, owning the refinement of engineering processes around observability, CI/CD, and incident response, while actively mentoring the next generation of technical leads. WHAT YOU BRING Leadership and Scaling: Pro
Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. Software Engineer (Senior) – Platform Engineering & AI-Powered Process Automation Location: Chennai, India | Work Model: On-site Architect the foundational systems powering the next generation of global software execution. In this role, you will build high-performance, reusable platform services and modern developer tooling that directly elevate developer velocity and scale enterprise software globally. About the Team The Platform Engineering, Foundations & Enablement team builds the bedrock of Appian’s technical ecosystem. By delivering resilient shared services, automated developer tooling, and core frameworks, this team enables engineering organizations to build scalable enterprise solutions and accelerates Appian’s growth in the AI-Powered Process Automation market. The Opportunity Top-tier engineers join Appian to tackle deep distributed systems challenges at massive enterprise scale. This role gives you direct ownership over the core platform architecture and internal developer tools that power critical workflows worldwide. If you want to eliminate engineering friction, set architecture standards, and shape high-concurrency systems using cutting-edge enterprise tech, this opportunity offers high visibility and immediate technical influence. What You’ll Do Architect and construct high-performance, resilient platform features, robust APIs , and shared frameworks to power mission-critical operations. Execute technical spikes to clear architectural runways, ensuring system stability, modularity, and seamless long-term scalability. Elevate engine
About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
Get new distributed systems engineer jobs by email
Daily job updates · Unsubscribe anytime