Jobiba hiring network

Lead Software Engineer Distributed Systems Jobs

6,753 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead software engineer distributed systems jobs. Use filters to narrow by work mode, employment type, experience and date posted.

S
1mo ago

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role At Sentry, Support is an engineering discipline. Our customers are the greatest technical minds in the world—developers at elite enterprises building the future of software—and they deserve answers that go deeper than a knowledge base link. We're looking for an APAC Technical Support Engineer based in San Francisco to join our global Support Engineering team. This role is designed to provide APAC coverage to our users; with the shift being Sunday through Thursday 4PM-12AM PST. We are architecting the Technical Support engine . We’re looking for an experienced engineer to help us redefine the standard of technical support by combining deep human expertise with autonomous agentic systems. You are a debugger of both code and systems. You will treat support volume as a data signal to build automated resolution paths, ensuring our human engineers only touch the most complex, high-impact architectural puzzles. Sentry Support Engineers aren't just clearing queues; they are Orchestrators . You will engage with our users across GitHub, Discord, and our internal systems, while acting as the Technical Lead for our Agentic Ops. You ensure that when a developer asks a complex question, our systems have the right context and a seamless "Human-in-the-Loop" path to you when deep, nuanced expertise is required. In this role you will Master the Sentry Ecosystem & Support Elite Developers Deep-Dive Debugging: Perform root-cause analysis on complex issues and distributed tracing gaps across polyglot environments. Support the Great Minds: Act as a strategic consultant for senior engineers at our largest enterprise customers, s

javascriptpythonjava
View job →
S
1mo ago

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role At Sentry, Support is an engineering discipline. Our customers are the greatest technical minds in the world—developers at elite enterprises building the future of software—and they deserve answers that go deeper than a knowledge base link. We're looking for an APAC Technical Support Engineer based in Australia to join our global Support Engineering team. This role is designed to overlap with our San Francisco headquarters; Monday thru Friday 9AM-5PM AEST. We are architecting the Technical Support engine . We’re looking for a veteran engineer to help us redefine the standard of technical support by combining deep human expertise with autonomous agentic systems. You are a debugger of both code and systems. You will treat support volume as a data signal to build automated resolution paths, ensuring our human engineers only touch the most complex, high-impact architectural puzzles. Sentry Support Engineers aren't just clearing queues; they are Orchestrators . You will engage with our users across GitHub, Discord, and our internal systems, while acting as the Technical Lead for our Agentic Ops. You ensure that when a developer asks a complex question, our systems have the right context and a seamless "Human-in-the-Loop" path to you when deep, nuanced expertise is required. In this role you will Master the Sentry Ecosystem & Support Elite Developers Deep-Dive Debugging: Perform root-cause analysis on complex issues and distributed tracing gaps across polyglot environments. Support the Great Minds: Act as a strategic consultant for senior engineers at our largest enterprise customers, solving high-stakes archite

javascriptpythonjava
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali

awsrestai
View job →
P
Point72
📍 Bengaluru• Full-time
19 days ago

JOB TITLE Site Reliability Engineer A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO You will play a highly critical operational role where you will apply a combination of software and systems engineering skills to develop and maintain a complex set of distributed, real-time systems that serve critical stakeholders in Point72’s Global Macro business. You will focus on optimizing the operations of existing systems and infrastructure in an efficient manner, through a strict adherence to automation and tooling Specifically, you will: Build out foundational technical components of an extensive SRE program across multiple complex systems, both new and existing • Collaborate with our development and quant teams to ensure that ongoing change is consistent with a pre-determined, measurable set of SLOs spanning multiple complex user interactions with our systems • Monitor system capacity and performance, identifying and addressing potential future bottlenecks and sources of instability before they become impactful to our stakeholders • Review and provide feedback on automation code developed by peers to maintain high standards of code quality and efficiency • Troubleshoot and resolve system issues, analyzing their impact on infrastructure and service operations • Participate in or lead design reviews with peers and stakeholders, evaluating and selecting the best technologies and automation strategies for our needs WHAT’S REQUIRED We are looking for highly motivated, proactive engineers

pythonawsdocker
View job →
G
19 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe

pythonaigo
View job →

The AI platform is responsible for all AI infrastructure across Datadog. Our mission is to provide tools and platforms that enable data scientists and engineers to conduct large-scale training and inference with ease. We support products such as Bits AI , LLMObs and all our AI research . As an engineering manager for the Training & Serving team, you’ll join a new and fast growing team and organization. You will support building and scaling the team, define our technical vision and help shape the roadmap. Your team will lead the charge on multiple critical technical challenges: distributed training of foundation models, serving at scale, designing the user experience. You’ll work closely with sister teams in the AI platform organization ensuring a seamless AI development cycle. You’ll also partner with the Applied AI org and with Datadog infrastructure & tooling teams to build out systems from the ground up. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Manage and grow the Training & Serving team, directly managing 10+ engineers Define our technical roadmap in alignment with AI platform goals and the Applied AI team roadmap. Work with our core platform teams to tailor Datadog's storage, infrastructure and data pipelines to our needs Create a strong team culture aligned with our engineering standards and our customer focus Participate in hands-on work: Code reviews, design reviews and some coding Who You Are: Previous experience (1+ years) leading software engineering teams, as a tech lead or people manager Strong technician with a mix of backend, data engineer and infrastructure experience who is interested in remaining a hands-on leader Excellent leader with strong

restaigo
View job →
A
1mo ago

About the Role & Team We’re looking for an Engineering Manager to lead the Data Infrastructure team within Statsig Experiment at Amplitude. You will lead a multidisciplinary team of software engineers, data engineers, and data scientists responsible for the systems that power experimentation at scale. The team owns three critical areas: Data ingestion: Collecting and importing experiment exposures, custom events, OpenTelemetry data, and real user monitoring data across SDKs, streaming systems, cloud storage, and customer data warehouses. Data computation: Building distributed computation systems that transform raw data into accurate, timely experiment results. Stats engine: Developing and productionizing the statistical methods that help customers make trustworthy decisions from their experiments. This is not a traditional data engineering management role. We are looking for a leader with a solid data science and statistical foundation who can connect advances in experimentation methodology with scalable production systems. You will help set our technical and scientific direction, translating new statistical methods and machine learning research into capabilities that customers can use reliably at scale. You’ll partner closely with data scientists, engineers, product managers, and customers to advance the state of experimentation. The ideal candidate is equally comfortable discussing causal inference and statistical power with data scientists, distributed computation architectures with engineers, and experimentation strategy with customers. What You’ll Do Lead and grow the team responsible for Statsig’s data ingestion, experiment computation, and stats engine. Define the technical and scientific strategy for advancing experimentation across both Statsig Cloud and warehouse-native deployments. Partner with data scientists and engineers to turn new statistical and causal inference methods into scalable, reliable product capabilities. Evolve our data and computatio

restmachine learningai
View job →
A
Amplitude
📍 San Francisco• Full-time• From $24K/yr
19 days ago

Amplitude is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Winter 2026 Report, Amplitude is the best-in-class solution for product, data, and marketing teams. Learn more at amplitude.com . As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do. Amplitude’s Commitment to Diversity Equity & Inclusion (DEI): Amplitude believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive. About the Role & Team We’re looking for an Engineering Manager to lead the Data Infrastructure team within Statsig Experiment at Amplitude. You will lead a multidisciplinary team of software engineers, data engineers, and data scientists responsible for the systems that power experimentation at scale. The team owns three critical areas: Data ingestion: Collecting and importing experiment exposures, custom events, OpenTelemetry data, and real user monitoring data across SDKs, streaming systems, cloud storage, and customer data warehouses. Data computation: Building distributed computation systems that transform raw data into accurate, timely experiment results. Stats engine: Developing and productionizing the statistical methods that help customers make trustworthy decisions from their experiments. This is not a traditional data engineering management ro

gitrestmachine learning
View job →
R
Roblox
📍 San Mateo• Full-time• From $295.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Engineering Manager, Home Infrastructure The Home Infrastructure team builds the mission-critical backend and data systems that power Roblox’s Homepage and Experience Details Page, two of the highest-traffic surfaces on Roblox. These surfaces reach the vast majority of Roblox’s daily active users and are core drivers of discovery, engagement, retention, and platform growth. We are a full-stack product infrastructure team responsible for content distribution across Roblox. Our systems support multiple modes of user interaction, including exploratory browsing, directed discovery, and personalized content recommendations across the many types of content that make up the Roblox ecosystem. This team sits at the intersection of large-scale distributed systems, machine learning-powered personalization, data infrastructure, and product experimentation. We partner closely with Machine Learning, Data Science, Product, Design, Frontend, Ads, Marketplace, Virtual Economy, and other teams across Roblox to build the platforms that help users find the most relevant and engaging content. As Engineering Manager for Home Infrastructure, you will lead a team of Backend and Data Engineers responsible for the e

awsgitmachine learning
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Snowhouse Foundation team builds our globally distributed data warehouse. We manage a vast array of petabyte scale data sets that are continuously ingested, processed and replicated from across all Snowflake environments and external data sources. Snowhouse powers all of Snowflake’s core business, engineering and data science needs and provides customers with full visibility into their account activities, usage, and resource consumption from all their global environments. The team is investing in multiple critical areas, including a pipeline authoring platform, high performance/high efficiency data export, ingestion and data layout. Our team is also responsible for a fundamental product for Snowflake’s customers: the Snowflake system database/application that provides customers with all usage insights they need to reason about their global Snowflake footprint as well as 1st party business logic such as ML powered functions and Budgeting applications. AS A PRINCIPAL SOFTWARE ENGINEER IN SNOWHOUSE FOUNDATION, YOU WILL: Design and implement innovative highly available distributed platforms and pipelines and enhance the overall Snowflake data infrastructure Lead and drive projects from idea formulation to design, implementation and successful productionization. Collaborate

awsazuregcp
View job →
L
Lyft
📍 Toronto• Full-time• From C$136K/yr
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft is looking for experienced software engineers from a variety of disciplines. We are growing our team with people who want to build, improve and incorporate technologies that make the lives of our community more enriched. As an engineer at Lyft, you'll collaborate with teams like product, data science, analytics, and operations on code that empower us to iterate quickly, while focusing on delighting our passengers and drivers. As a Senior Software Engineer on the Marketplace team, you will lead work streams to improve business operations in Lyft’s Marketplace. You'll design AI driven data analytics platforms, data pipelines and metric governance systems that our business leaders use on a daily basis to make key strategic decisions. You will partner with business leaders and data scientists across our organizations to enable running the business more efficiently. Responsibilities: Help define the roadmap and architecture based on technology and business needs Drive AI innovation for business analytics and operations Write well-crafted, well-tested, readable, maintainable code Have a good grasp and ability to explain the various tradeoffs made in decisions Participate in code reviews to ensure code quality and distribute knowledge Lead projects from idea to positive execution Incorporate considerations for business context and failure modes in your work Proactively participate in resolving ongoing incidents Unblock, support, effectively communicate, and obtain buy-in across teams to achieve results Share your knowledge by giving brown bags, tech talks, and evangelizing appropriate tech and engineering best practices Experience: 5+ years of software engineering industry experience with a high level programming language (bonus points for experience with Python or Go) AI Experience in Agen

pythonmicroservicesai
View job →
S
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Issuing team owns the Stripe card issuing product end to end from the APIs and backend services that power spend cards, debit cards, and charge cards, to the dashboard and embedded components that businesses use to manage their card programs. We build the infrastructure and interfaces that let platforms and businesses instantly create, distribute, and control payment cards at scale, and we're responsible for the full cardholder lifecycle across commercial and consumer issuing. We're at an inflection point. Issuing is expanding into new geographies, new card categories (stablecoin, healthcare, consumer credit), and deeper partnerships with major platforms. This is a high-impact, high-visibility role at the intersection of financial infrastructure, developer-facing APIs, and end-user product experience, and it requires someone who can lead technically across a complex, fast-moving domain. What you'll do Define and drive the technical strategy for Issuing, including multi-year architecture decisions across backend services, APIs, and user-facing surfaces. Own end-to-end technical solutions for critical systems, authorization flows, spending controls, cardholder management, and card program configuration, ensuring they are reliable, scalable, and operationally sound. Identify and lead infrastructure investments that unlock new issuing models, including geographic expansion, new card program types, and platform-level extensibility

L
Lyft
📍 Toronto• Full-time• From C$108K/yr
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. As a Frontend Software Engineer on the Operator Core Tooling pod, you'll play a vital role in building the robust services that power our critical operations tooling platform. Your work will directly empower our micromobility operations teams by providing them with intuitive and efficient tools, significantly improving their daily workflows as they manage our fleet. You'll collaborate closely with business leaders, front-end developers, and data scientists across Lyft to achieve this impact. Responsibilities: Help define the roadmap and architecture based on technology and business needs Write well-crafted, well-tested, readable, maintainable code Have a good grasp and ability to explain the various tradeoffs made in decisions Participate in code reviews to ensure code quality and distribute knowledge Lead projects from idea to positive execution Incorporate considerations for business context and failure modes in your work Proactively participate in resolving ongoing incidents Unblock, support, effectively communicate and obtain buy-in across teams to achieve results Share your knowledge by giving brown bags, tech talks, and evangelizing appropriate tech and engineering best practices See the direct impact of your work on the efficiency of our operating teams Experience: 3+ years of software engineering industry Advanced knowledge of JavaScript Experience working with modern JavaScript frameworks, like React Experience working with NodeJS and Express applications Experience working with design systems (e.g. Bootstrap, Salesforce Lightning, GitHub Primer) Good understanding of web performance and how browsers and DOM work Experience with unit, integration, and end-to-end testing Experience designing, building and improving a set of team owned components Culture of investigating and solvin

javascriptjavareact
View job →
L
Lyft
📍 Mexico City• Full-time
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft AV brings the autonomous future to life by matching our community of riders with self-driving vehicles to meet their transportation needs today. The program is powered by a technical team innovating to build new features, experiment, and evolve approaches for emerging business needs. At the same time, we prioritize production quality because our customers put their trust in us for safety and reliability. As a Software Engineer for Autonomous Vehicles, you’ll be responsible for executing integrations with partners, making tradeoffs between technical investments and product work, and collaborating with other engineers on system design. You will help shape the product direction by developing a deep understanding of the customer and working closely with cross-functional partners from Product, Design, Science, and Operations. You will work with a group of talented engineers and help the team to deliver significant business impact while being open to change through constant experimentation in an ambiguous emerging product. If you enjoy collaborating with technical and nontechnical partners in a fast-moving space with real-world impact, this is the role for you. Responsibilities: Help establish roadmap and architecture based on technology and understanding of customer needs Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and distribute knowledge Share your knowledge by giving brown bags, tech talks, and promoting appropriate tech and engineering best practices Can help lead large projects from idea to positive execution Unblock, support and communicate with internal partners to achieve results Experience: BS/MS or equivalent in Computer Engineering, Computer Science, or related field or relevant work experience Experience in distributed sy

pythonsqlai
View job →
L
Lyft
📍 Mexico City• Full-time
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft is a global company that connects people to transportation to change the way we live and get around our communities. Our drivers and passengers entrust Lyft with their personal information and travel details to get where they are going and expect us to keep that data safe. Lyft's Privacy team is composed of engineers and analysts dedicated to ensuring appropriate data protections are applied to our drivers' and passengers' data. Our goal is to make privacy the default setting for our internal engineers and our customers. We help design architectures, build a scalable data platform, and define policies to reduce the exposure of user data and ensure appropriate privacy and security protections are applied across the company. We are looking for a highly motivated, collaborative, team-focused and technically strong Software Engineer to join our Privacy team. As a member of this team, you will design and build systems that make privacy the default at Lyft. Every day, you'll partner with engineering, product, legal, and compliance stakeholders on high-impact projects — from building data redaction pipelines to enabling users to exercise their privacy rights to deploying data management systems across a diverse set of Lyft services. You'll bring strong engineering instincts, a passion for privacy, and the ability to lead large projects from idea to execution. Responsibilities: Write well-crafted, well-tested, readable, and maintainable code Promote appropriate tech and engineering best practices Partner with teams across Lyft on organizational privacy initiatives Partner with senior engineers to design, build, and maintain systems that enhance privacy across the organization Independently lead tasks from idea to execution Participate in code reviews to ensure code quality and distribute knowledge Parti

pythonjavasql
View job →
🔔

Get new lead software engineer distributed systems jobs by email

Daily job updates · Unsubscribe anytime