At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Build Core Data Engineering Primitives at Cloud Scale Data pipelines are foundational infrastructure — when they're fast, correct, and maintainable, customers build on them with confidence. If you've spent the bulk of your career building large-scale data infrastructure — designing streaming or transformation primitives, reasoning hard about consistency and fault tolerance, and owning the systems that run under millions of customer workloads — this role might be for you. You'll be working on the streaming and transformation layer at Snowflake: the constructs that define how customers move, shape, and maintain data. AI has a real presence in this work — in how customers use these pipelines and in how we think about building them — but the core job is hard distributed systems engineering, and that's what we're hiring for. About the Team We build the core data engineering primitives that power Snowflake's streaming and transformation capabilities. From the constructs customers use to define real-time pipelines to the execution fabric that makes those pipelines reliable and cost-efficient at cloud scale, our team owns the full stack of declarative data engineering. We're a small, high-ownership team operating close to the product — which means your decisions ship, your architec
Jobiba hiring network
Distributed Systems Engineer Jobs
1,306 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions
At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions
About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability Team Focus Areas Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely. Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads. Databases: Building high performance, distributed database systems that power all of OpenAI's product stack. In this
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: We are seeking talented distributed systems engineers who are passionate about building innovative solutions for application deployment. Your mission will be to enhance the capabilities of Replit Infrastructure, optimize performance across global regions, and drive efficiency while delivering an exceptional user experience. If you have a strong foundation in software development, a deep understanding of cloud technologies, and a track record of delivering high-quality code, we want to hear from you. In this role you will: Expand Replit's cloud infrastructure offerings: Launch new cloud products to be used by Replit Agent to build complex apps. Collaborate with cross-functional teams to design and implement these features, empowering developers with a comprehensive suite of tools to build and deploy their applications efficiently. Enhance reliability and scalability: Identify bottlenecks, optimize critical paths, and implement robust monitoring and alerting systems. Work closely with the SRE team to ensure high availability and minimal downtime. Enable our customers to seamlessly scale their applications to meet the demands of their growing user base. Improve utilization of cloud infrastructure: Analyze our infrastructure costs and identify opportunities for optimization. Implement strategies to reduce cloud expenses without compromising performance or reliability. This could involve techniques such as resource provisioning, auto-scaling, cost-aware scheduling, and data lifecycle management. Your efforts will directly contribute to the financial efficiency of our cloud services. Required skills and experience: Distributed systems: Track record of working with platform-as-a-service, distributed storage, o
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is building the world’s fastest, most efficient AI compute clusters. TT-Fabric is the high-performance nervous system of this platform: the low-level networking layer that lets thousands of RISC-V and AI processors snap together into a single, massively parallel distributed supercomputer. If you love squeezing nanoseconds out of hot paths, designing protocols that move data at absurd scale, and turning messy hardware constraints into elegant distributed systems, this is an opportunity to shape the fabric that future AI models will run on This role is hybrid based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who We Are Strong systems engineer with deep C or C++ experience and comfort working in low-level or bare-metal environments. Passionate about hardware-software interaction, performance tuning, and eliminating inefficiencies at the protocol level. Curious about networking, synchronization, and communication across large clusters. Comfortable reasoning from first principles and challenging industry conventions. Motivated by building infrastructure that directly impacts large-scale
NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel
About the Team API Frontiers turns OpenAI’s frontier models into production APIs that developers can use to build reliable products and agents. We own the core path connecting models to developers through the Responses API, with a focus on safety, reliability, and speed. Working closely with Research, Safety, Codex, and other API teams, we bring new model capabilities into production and improve them through developer feedback. About the Role We are looking for a backend software engineer to build and operate the services behind the Responses API. You will shape API behavior, bring new capabilities from research into production, and make long-running agent workflows dependable and fast. The work combines distributed systems engineering with product judgment: designing useful developer interfaces, managing staged rollouts, and following production issues through to durable fixes. In this role, you will: Design, build, and operate APIs and backend services that bring frontier model capabilities to developers. Partner with Research, Safety, Codex, and API teams to define API behavior and deliver safe, staged launches. Build API capabilities for agent workflows, including task delegation, context sharing, and parallel execution. Strengthen long-running request reliability across timeouts, cancellation, streaming, and background execution. Improve request-processing performance and tail latency through profiling, efficient systems code, and persistent connections. Turn developer feedback and production failures into better observability, diagnostics, and lasting product improvements. Your background might look something like: 5+ years of experience building and operating backend services or developer-facing APIs in production. Strong software engineering fundamentals, with practical knowledge of distributed systems, concurrency, and asynchronous execution. Ability to diagnose production failures and performance bottlenecks using observability data and profiling. Product
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Position Overview: We are seeking a highly technical Staff Observability Site Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to delivering a world class, comprehensive, scalable Observability Platform that enables our SRE teams and business partners. You will treat infrastructure as code —utilizing Terraform and strong coding proficiency in Go, Python, or Ruby —to automate the deployment of agents and collectors across complex distributed systems. Key Responsibilities Automated Infrastructure: Design, build, and maintain scalable observability infrastructure using tools like Terraform. Splunk Engineering: Optimize the collection, processing, and storage of log data to ensure high reliability and low latency of our Splunk services Incident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "observability-driven development." Automation: Eliminate "toil" by automating the deployment and scaling of observability agents and collectors. Required Skills & Experience (The Essentials) Log Management: Minimum 5+ Experience scaling and managing Splunk Cloud at scale (1000+ SVCs), including Workload Management (WLM) and HEC optimization. Visualization: Expertise in creating intuitive, actionable Splunk dashboards that correlate data across multiple sources. SRE Mindset: Minimum 5+ years of experience in an SRE, Dev
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As an Observability team member, you are responsible for the operation and maintenance of our logging and metrics infrastructure. You ensure all teams at Lyft are aware of the operational health of their products by monitoring system availability and take a holistic view of our platform performance. You build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you find opportunities to improve our systems in order to push our platform forward. You provide our partners with the support they need to help them build robust large scale distributed systems. We count on the reliability of our infrastructure to empower Lyft teams to provide our customers rich experiences that are highly available with rock solid performance to ensure our transportation platform continues to connect people and places. As we grow our team, we are seeking experienced Infrastructure Engineer to ensure that as our Infrastructure continues to scale, our platform continues to provide an essential and dependable service that transports millions of people every day. Specifically we are searching for someone who brings fresh perspectives, enjoys collaborating with cross-functional teams in order to continually improve our products and services for our customers. Responsibilities: Maintain, improve, and develop tooling and systems that enhance the reliability, scalability, and efficiency of our platform. Assist engineering teams in defining service-level objectives (SLOs) and provide the necessary toolin
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team The Developer Productivity group is responsible for making Stripe’s developers happy and productive. We work on tools, processes, and code refactoring to accelerate Stripe engineering as Stripe scales. We’re looking for people with an interest in building the tools to improve the day to day experience of Engineers in Stripe. The ideal candidate will have a passion for solving developer experience problems, and a pragmatic ability to ship results iteratively—powered by a mix of technical expertise across some or all of: language processing tools; version control systems; build systems; and distributed systems engineering. You’ll be working on a mix of engineer-facing systems, CI infrastructure platforms, and big-data engineering. What you’ll do We have a ton of important work to do, which is why we’re hiring! Our active projects change all the time, but here are a few examples of recent projects so you can get an idea of the types of work we do: Build and manage systems to handle CI at massive scale—including batching, speculative stacking, merge-race inhibition, and more. Build and manage systems to handle our enormous CI test suite–identifying and managing flaky tests, assessing and reproducing flakiness, and optimizing for test effectiveness. Enhance our CI systems for reliability, including adaptive response to available capacity, resilience and self-recovery from outages, and highly-leveraged observability. Detect and isolate code breaka
Senior Software Engineer - Analytics Compute Platform Team About the Role & Team Every chart, insight, and experiment result a customer sees in Amplitude passes through the compute layer that this team owns. Our team sits between Amplitude's front-end analytics products/ rest APIs/ MCPs and Nova, our proprietary analytics database, and owns the compute API that translates a user's question into a fast, correct answer. Our mandate: enable a highly performant, reliable, and flexible way to compute insights. Those three goals are often in tension the more flexible we make the system, the harder it is to keep it fast and simple, and a lot of the interesting engineering work on this team lives in that tradeoff. Day to day, you'll work closely with engineers on data management, marketing analytics, Session Replay, Experiment, CDP, Query, and various product teams, since they're all consumers of what we build. As a Senior Software Engineer, you will Design and build core parts of the query/compute engine that powers analysis across Amplitude's product suite Evolve the compute API and semantic layer that other engineering teams build features on top of, so that it can support a more generic table Help unify and evolve our core data model, including how non-event data (profile properties, lookup tables, etc.) is represented alongside event data Improve the performance, reliability, and scalability of query planning and execution on the Amplitude query engine Partner with teams across Session Replay, Experiment, CDP, Query, and frontend to understand their needs and shape the compute layer around them Take ownership of projects end-to-end, from design through rollout You'll be a great addition to the team if you have Strong backend or distributed-systems engineering experience, ideally touching OLAP databases, query engines, or data infrastructure Experience with SQL, query planning/execution, or building APIs that many other engineering teams depend on Comfort reasoning
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Data Access team within Infra / Storage builds EaaS (Entities-as-a-Service), Roblox’s large-scale managed OLTP access and management platform powering tens of millions of QPS across thousands of services. EaaS abstracts the complexity of distributed databases and caching systems behind a consistent and productive developer experience, enabling teams across Roblox to safely build and operate stateful systems at massive scale without requiring deep database expertise. The team operates at the intersection of distributed systems engineering, reliability engineering, and platform architecture — solving hard infrastructure problems that directly impact Roblox-wide scale and stability. As a Senior Software Engineer you will work on some of Roblox’s hardest backend infrastructure challenges around scalability, reliability, workload governance, adaptive flow control, and distributed systems. You Will Build and evolve EaaS, Roblox’s managed OLTP access platform powering tens of millions of QPS across hundreds of services. Design infrastructure that abstracts distributed databases and caching systems behind a consistent, safe, and highly productive developer experience. Drive reliability and scal
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. We are building and maintaining a highly scalable asynchronous platform that empowers our organization to handle critical business cases. As a software engineering team, our mission is to create robust and innovative solutions that drive the success of our business and deliver unparalleled value to our customers. We adopt Infrastructure as Code practice to automate the provisioning and configuration of our resources, which helps reduce manual configuration and improve consistency. Our team culture is built on collaboration, open communication, and a supportive environment where each member's ideas are valued and contributions are recognized. We believe in the importance of fostering a positive workplace culture that inspires innovation and creativity. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and production readiness reviews, platform management and capacity planning ceremonies with cross-functional teams Document Infrastructure operations process and insights, identify repeatable actions and ruthlessly automate repetitive tasks Participate in our teams on-call rotations, respond to incidents and support other teams mitigate customer impacting events Experience: 5+ years experience working on teams responsible for software development, automation and systems engineering Experience building large-scale infrastructure, distributed systems or networks. Knowledge with SQS,
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . People use Pinterest to find ideas and brands that they love. We aspire to help our advertisers and partners reach their audiences with inspiring content. The API team is responsible for ensuring our first party clients (Android, iOS, and Web) have a stable API to develop on top of. Our customers are Pinterest developers who want to build product features, and who rely on a highly available system to do so. As the Engineering Manager for API, you will lead a talented and growing engineering team responsible for growing the existing portfolio. The ideal candidate should have experience building backend web applications, be driven to become an expert in their domain, have some knowledge of distributed systems engineering, and have a passion for leadership. What you’ll do: Collaborate with stakeholders across the organization to architect so
Get new distributed systems engineer jobs by email
Daily job updates · Unsubscribe anytime