Jobiba hiring network

Software Reliability Engineer Jobs

6,428 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

R
Roblox
📍 San Mateo• Full-time• From $196.8K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Consumer Apps Consoles team owns the console app foundation that makes Roblox feel fast, fluid, and reliable for millions of players and is dedicated to delivering superior performance, reliability, and user experience on PlayStation and Xbox. As a Senior Software Engineer on the Consoles team, you'll be a driving technical force — leading complex client work end-to-end, raising the quality bar across our codebase, and partnering closely with product, design, and platform teams to ship experiences that are polished and performant at scale. You Will: Shape & improve the PlayStation and Xbox experience that millions of players see every day Translate ambiguous product ideas into well-scoped, high-quality engineering plans in close partnership with product, design, and data science. Raise the engineering bar through rigorous code review, proactive mentorship of junior engineers, and the standards you set in your own code. Drive platform-level investments that benefit the broader engineering organization. Work with our Xbox and PlayStation counterparts to ensure feature parity and platform-specific excellence. You Have: 5+ years of experience building and shipping client software, with

reactawsgit
View job →
R
Roblox
📍 San Mateo• Full-time• From $243.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on our Release Engineering team, you will build the systems that hundreds of engineers across Roblox use every day to safely ship code to tens of millions of concurrent players. Your work will empower teams to ship bold, high-impact changes quickly and confidently on every device Roblox runs on. If you enjoy building developer-facing infrastructure where reliability and blast radius directly impact end-users, you will be right at home on our growing team. You Will: Design and develop backend services and automation that power our release and experimentation process across desktop, mobile, console, VR, and servers Work in C++ engine code to improve telemetry reporting, enhance feature rollout and automatic abort capabilities, and extend release functionality Build progressive client and server rollout, regression detection, and automated rollback systems that keep our weekly multi-platform releases safe at scale Work directly with engineering customers to turn pain points into durable, flexible, and safe-by-default tooling You Have: 5+ years building backend services with C#, Python, TypeScript, or similar Familiar with and comfortable working with C++ Familiar

typescriptpythonaws
View job →
R
Roblox
📍 San Mateo• Full-time• From $243.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's data infrastructure processes petabytes of data daily, powering analytics, ML, and product decisions. As a Senior Software Engineer in our Data Infra org, you will design, build, and scale the distributed data infrastructure platforms that power Roblox. You will own and drive the next-generation architecture of our core platforms, which span Kafka, Flink, Spark, Trino, Druid, Airflow and Data Catalog. This role combines high ambiguity and ownership to push the boundaries of what our infrastructure can handle at massive scale, giving you the unique opportunity to steer the evolution of the data landscape. You Will: Own and Scale Core Platform Components: Take responsibility for the design, architecture, and implementation of 1–2 key data platform frameworks within our stack Collaborate and Align: Partner with infra, data science, and product engineering teams to ensure your target platform's capabilities are directly guided by platform governance and product requirements. Optimize Performance at Scale: Dive deep into engine internals, query planning, state management, memory optimization, serialization efficiency to maximize throughput and reliability under heavy load. Drive Infrast

javaawsgcp
View job →
A
Asana
📍 Reykjavik• Full-time
1mo ago

The Infrastructure team builds the foundation we need to support our web and mobile applications, as well as our robust API. We build and operate the software that enables Asana's security, scalability, and speed. Each day, we combine industry best-practices and innovation to support this product-focused company. We're looking for a Junior Infrastructure Software Engineer (New Grad) with a growing passion for Site Reliability and Operability. You will work with a world-class team of engineers on deploying and operating existing systems, and building new ones for challenges that are unique to our problem space. You will have a unique opportunity to learn how to design, develop, and operate services that power Asana, working on projects that help define the future of our infrastructure and how we architect and operate critical services at scale. This role is based in our Reykjavik office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. What You’ll Achieve Learn and Execute: Experience growth and development by being paired with a mentor who will support and guide you through opportunities to stretch and learn. Engineering Craftsmanship: Participate in Asana’s robust technical boot camps to learn our standards for high-quality, maintainable code. You will eventually own specific technical domains through our Areas of Responsibility (AoR) system. Build Foundations: Contribute to projects that define the future of infrastructure for Asana, learning how we architect and operate critical services at scale. Global Collaboration: Partner with other infrastructure teams in San Francisco, New York, and Warsaw to understand and contribute to our service-oriented architecture while navigating cross-timezone workflows. Tooling & Frameworks: Help develop framework

awsgitrest
View job →
D
Datadog
📍 New York• Full-time• From $192K/yr
1mo ago

Senior Software Engineer - Streaming Platform Client Data streams are mission-critical at Datadog, powering near real-time communication across the vast majority of our services. Our Streaming Platform group builds the core infrastructure and abstractions that ensure Datadog remains a trusted partner for engineers worldwide. See our blog post . The Streaming Platform Client team sits at the heart of this ecosystem. We own the Rust client library (producers and consumers) with language bindings for Java, Go, and Python. We focus on building intuitive APIs and abstractions that make a powerful distributed system easy to adopt and operate for the hundreds of internal users of our library. Our library runs on critical data paths that handle hundreds of millions of messages per second making performance, observability, and reliability paramount. We also develop and operate the service that bridges the clients fleet with the platform's control plane, handling complex balancing, scaling, and static stability challenges. We are seeking a Senior Software Engineer to help us evolve these features. You will collaborate directly with our users, tackle performance-critical code, and solve complex distributed systems challenges across the control plane, client libraries, and data plane. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Work within a distributed, high-impact team spanning Europe and the US, building critical technologies that power data pipelines for dozens of internal teams and hundreds of services. Architect and implement resilient interactions between our client libraries and the control plane. Optimize our high-throughput, low-level streaming library to push the boundaries of performance and efficiency. Champion the developer experience by providing

pythonjavagit
View job →

We are looking for a strong technical leader to join the Private Action Runner team, part of the larger Action Platform group and help shape one of the core execution layers behind Datadog’s action-taking and remediation capabilities. Private Action Runner (PAR) enables Datadog products and AI agents to securely run actions inside customer infrastructure with controls for authentication, permissions, auditing and safe execution. The role will be hands-on, covering architecture, implementation, reliability and collaboration with teams integrating PAR across Datadog. It also offers leadership exposure through leading a team of 3 engineers, with the expectation that the role will quickly transition into a formal Engineering Manager 1 position as the team grows. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: (Describe role responsibilities here) Solve a scaling bottleneck in a critical service Deploy a new feature to production, progressively rolling it out with feature flags Investigate and fix a production issue from a service your team owns Design a way to scale up a service for more traffic With your team, plan the most important projects to work on next Mentor and lead a small team of 3 engineers Who You Are: (Describe role qualifications here) You have significant experience in one or more languages You value code simplicity and performance You can design architecture to solve problems at high scale You have a BS/MS/PhD in a scientific field or equivalent experience You want to work in a fast, high-growth startup environment that respects its engineers and customers You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how You have demonstrated ability to use AI

aigorust
View job →
D
Datadog
📍 New York• Full-time• From $192K/yr
1mo ago

Senior Software Engineer - Streaming Platform Client Data streams are mission-critical at Datadog, powering near real-time communication across the vast majority of our services. Our Streaming Platform group builds the core infrastructure and abstractions that ensure Datadog remains a trusted partner for engineers worldwide. See our blog post . The Streaming Platform Client team sits at the heart of this ecosystem. We own the Rust client library (producers and consumers) with language bindings for Java, Go, and Python. We focus on building intuitive APIs and abstractions that make a powerful distributed system easy to adopt and operate for the hundreds of internal users of our library. Our library runs on critical data paths that handle hundreds of millions of messages per second making performance, observability, and reliability paramount. We also develop and operate the service that bridges the clients fleet with the platform's control plane, handling complex balancing, scaling, and static stability challenges. We are seeking a Senior Software Engineer to help us evolve these features. You will collaborate directly with our users, tackle performance-critical code, and solve complex distributed systems challenges across the control plane, client libraries, and data plane. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Work within a distributed, high-impact team spanning Europe and the US, building critical technologies that power data pipelines for dozens of internal teams and hundreds of services. Architect and implement resilient interactions between our client libraries and the control plane. Optimize our high-throughput, low-level streaming library to push the boundaries of performance and efficiency. Champion the developer experience by pro

pythonjavagit
View job →
M
1mo ago

The MongoDB Customer Observability Team is a diverse group of contributors working together to help our users manage MongoDB at global scale. The team is responsible for MongoDB Atlas: our database-as-a-service offering and fastest-growing product, which allows users to deploy globally distributed MongoDB clusters in just minutes. We're seeking a Senior Software Engineer to join our team to tackle exciting challenges within the Observability space. You'll contribute to developing tools and platforms that help our users understand the health and performance of their MongoDB deployments. This includes collecting metrics, monitoring slow queries, and offering actionable insights such as index and schema suggestions that improve the speed, efficiency, and overall reliability of their databases. This role provides a unique opportunity to drive engineering excellence across both dimensions of observability, leveraging technologies being developed within the Customer Observability group and contributing directly to the success of our customers and our product teams. If you're passionate about large-scale systems, digging deep into telemetry data, and building tools that make a real impact both internally and externally, we’d love to have you on board! We are looking to speak to candidates who are based in Dublin for our hybrid working model. We're looking for someone who Has at least 5 years of experience as a backend or full stack engineer Enjoys collaboration and being part of a team Is approachable, curious, and intellectually honest Is a backend engineer with a willingness to take on frontend tasks or a full-stack developer with a bias towards backend Has written backend systems in a compiled language (Java, C#, Go, etc.) Has experience with the design and architecture of a modern, scalable web application Enjoys chasing down difficult problems in a distributed environment and on an database diagnostic/operation level Always strives to expand their knowledg

typescriptjavareact
View job →
M
Mongodb
📍 Gurugram• Full-time
1mo ago

We’re looking for a Software Engineer 3 to join our Marketing Technology Engineering team within Marketing Operations - Technology & Automation. This is a hands-on engineering role for someone who enjoys building reliable, scalable digital platforms and internal systems that improve how teams ship, measure, and optimize web experiences. You’ll work across application development, integrations, experimentation, data-informed decision making, and AI-enabled workflows that help the team move faster and deliver better outcomes. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. What you’ll do Build and maintain production-quality software that powers martech experiences Design and implement backend services, internal tools, automations, and integrations across the martech ecosystem Improve system reliability, observability, maintainability, and developer experience across the team’s platforms Contribute to experimentation, personalization, and data workflows that support better customer and developer experiences Help evaluate and apply AI-enabled capabilities and tools where they can improve engineering velocity, quality, or user experience Participate in code reviews, technical design discussions, and team planning Take ownership of projects from implementation through rollout, monitoring, and iteration What we’re looking for 3+ years of professional software engineering experience building and supporting production systems Strong coding skills in one or more modern programming languages such as JavaScript or Python Experience building web applications, backend services, APIs, data pipelines, or internal platforms Solid understanding of software engineering fundamentals including testing, debugging, code quality, and maintainability Experience working with cloud services, CI/CD workflows, and modern development practices Ability to work across systems and collaborate effectively with cross-functional partners Strong writte

javascriptpythonjava
View job →
F
Figma
📍 Ca New York• Full-time• From $153K/yr
1mo ago

Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! The Production Engineering team focuses on end-to-end reliability, durability, scalability, and performance of Figma products and services. We’re looking for an experienced Engineer to scale and drive the initiatives and programs that support Figma’s production engineering efforts, allowing our product teams to rapidly deliver new features. This is a critical part of success for both our ability to build new products and features, support the growth of our user base, and enable innovation for all engineering (infrastructure and product teams, alike). We’re looking for a generalist with a strong grasp of Computer Science fundamentals. The ideal candidate will have experience building and running complex large scale services. Additionally, they will have a passion for building tools that simplify operating large-scale infrastructure and for improving the operational maturity of Figma. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: Work closely with the engineering team to define standard methodologies and goals around reliability, durability, scalability, and performance Address common operational challenges through better telemetry and by building tools / services Debug production issues across services and levels of the stack Participate in design reviews and production reviews for new features, products or infrastructure components Plan for the growth of Figma’s infrastructure Operate and maintain AWS Infr

awsazurerest
View job →

Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! The Production Engineering team focuses on end-to-end reliability, durability, scalability, and performance of Figma products and services. We’re looking for an experienced Engineer to scale and drive the initiatives and programs that support Figma’s production engineering efforts, allowing our product teams to rapidly deliver new features. This is a critical part of success for both our ability to build new products and features, support the growth of our user base, and enable innovation for all engineering (infrastructure and product teams, alike). We’re looking for a generalist with a strong grasp of Computer Science fundamentals. The ideal candidate will have experience building and running complex large scale services. Additionally, they will have a passion for building tools that simplify operating large-scale infrastructure and for improving the operational maturity of Figma. What you'll do at Figma: Work closely with the engineering team to define standard methodologies and goals around reliability, durability, scalability, and performance Address common operational challenges through better telemetry and by building tools / services Debug production issues across services and levels of the stack Participate in design reviews and production reviews for new features, products or infrastructure components Plan for the growth of Figma’s infrastructure Operate and maintain AWS Infrastructure We’d love to hear from you if you have: Have 5+ years of experience operating infrastruct

awsazurerest
View job →
F
Figma
📍 Ca New York• Full-time• From $153K/yr
1mo ago

Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! AI Platform teams build the core frameworks, abstractions, and systems that support AI features across Figma. We create new capabilities that AI product teams can build on, while collaborating with teams from around the company to improve our performance, reliability, and technical quality. We’re looking for strong infrastructure and platform-minded engineers to contribute to our agent infrastructure, context retrieval & ranking platform, and core AI services, in order to accelerate our most critical company-wide AI initiatives. Here are just a few areas our platform teams work on: Evals for design : Building evaluation frameworks for design generation quality that are used across every Figma AI feature. Agentic search: Providing relevant context from throughout the Figma ecosystem to agents via search tools, in order to improve agent quality. Figma MCP : Making our MCP server faster, more reliable, and easier for internal & external engineers alike to develop on. Agent infrastructure : Iterating on the sandboxes, harnesses, and tools leveraged by the agents that power Figma Make and the Figma Design Agent. Preview & publishing platforms : Creating the shared platforms for building, previewing, and publishing code written by agents via Figma Make. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you’ll do at Figma: Support end-to-end AI feature development by designing, building, and maintaining systems that are scalable, reliab

typescriptpythonaws
View job →
I
Instacart
📍 United States - Remote• Full-time• Remote• From $265K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview The Commerce Platform organization at Instacart is seeking an experienced Staff Software Engineer to lead foundational platform initiatives on the Orders Platform team. The Orders platform is the orchestration layer for commerce at Instacart, coordinating the end-to-end lifecycle of every order across customers, shoppers, retailers, payments, fulfillment systems, and partner integrations at scale. In this role, you will help define and drive the architectural strategy for systems that support hundreds of millions of transactions each year. You will work on the reliability, correctness, scalability, and auditability of order workflows across diverse use cases and fulfillment models, while unlocking new experiences for customers, shoppers, retailers, and partners. The Orders Platform team sits at the center of Instacart’s commerce stack, which means this rol

REMOTEpythonjavasql
View job →
I
Instacart
📍 United States - Remote• Full-time• Remote• From $265K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Instacarts Data Infrastructure organization builds and operates the systems that power our company’s data ecosystem, including a modern open data lakehouse on Apache Iceberg, a multi-engine compute platform for stream and analytical workloads, and self-serve tooling that helps Product, Data Science, ML, Ads, Finance, and engineering teams move fast with data. We’re looking for a Staff Software Engineer, Data Infrastructure to join our Data Governance and Foundations Team. In this role, you’ll serve as a senior technical leader owning the architecture and delivery of our open lakehouse foundation, governance and access patterns, and multi-engine compute strategy—balancing today’s reliability with the next three to five years of scale, maturity, and cost efficiency. You’ll collaborate closely with engineering leadership and stakeholders across Data Science,

REMOTEpythonsqlaws
View job →
L
Lyft
📍 San Francisco• Full-time
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. We are hiring a Machine Learning Engineer to join our ETA team. Our team builds and maintains Lyft's system responsible for estimating/predicting ETAs for every ride request on our platform. ETAs play a critical role in matching decisions, pricing estimates and overall user experience. Low latency, high reliability and high accuracy are paramount for our success. If you are a critical thinker with experience in machine learning workflows and writing reliable code, passionate about solving business problems using data and working in a dynamic, creative, and collaborative environment, we are searching for you. Our technology stack runs on AWS, Kubernetes, Go, Spark, Python and Apache Airflow. In this role, you will work with incredibly passionate and talented colleagues from software engineering, machine learning and data science on building rideshare experiences that delight millions of riders and drivers. Responsibilities: Perform data analysis and build proof-of-concept to explore and compare ML and non-ML solutions Be able to make effective tradeoffs between model accuracy, its productization complexity and runtime performance Develop statistical, machine learning, or optimization models Write production quality code that can scale well to serve millions of requests per day Participate in code reviews, design reviews, production on-call support and incident triaging process. Write well-crafted, well-tested, readable, maintainable code Experience: B.S., M.S., or Ph.D. in Computer Science or other quantitative fields or related work experience 3+ years of Machine Learning experience Nice-to-have: Experience with big data processing / distributed data pipelines and tools such as Apache Airflow and Spark Ability to work in distributed teams spread across time zones. (North America and

pythonawskubernetes
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime