Jobiba hiring network

Software Reliability Engineer Jobs

6,428 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

L
Lyft
📍 San Francisco• Full-time
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Lyft Business Product Platform team builds the systems and experiences that power Lyft's B2B products — enabling companies, organizations, and their employees to seamlessly access Lyft's transportation network. We sit at the intersection of product and platform, owning both the customer-facing features and the underlying infrastructure that makes them reliable at scale. Our work directly impacts how businesses integrate with Lyft, how admins manage their programs, and how millions of riders get where they need to go. Responsibilities: Drive architecture and technical design for systems that are highly available, scalable, and built to last — not just for today's requirements but for where the product is heading Own features end-to-end: from shaping the technical spec and design through to production rollout and operational health Think critically about how AI capabilities can be incorporated into Lyft Business products to improve the experience for business admins and riders — and bring that perspective into roadmap and architecture conversations Make well-reasoned trade-off decisions and communicate them clearly to peers, leads, and cross-functional partners Write clean, well-tested, maintainable code and hold a high bar for the same in code reviews Partner across engineering, product, and design to align on direction and get buy-in on technical approaches Proactively engage in incident response, contributing both to resolution and to long-term reliability improvements Grow the team's technical culture through design reviews, tech talks, and mentorship Experience: 5+ years of software engineering experience, with a track record of designing and shipping production systems at scale Strong system design instincts — you can reason through distributed systems trade-offs, identify failure modes, and

sqlawsazure
View job →
L
Lyft
📍 Toronto• Full-time• From C$108K/yr
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. We are building and maintaining a highly scalable asynchronous platform that empowers our organization to handle critical business cases. As a software engineering team, our mission is to create robust and innovative solutions that drive the success of our business and deliver unparalleled value to our customers. We adopt Infrastructure as Code practice to automate the provisioning and configuration of our resources, which helps reduce manual configuration and improve consistency. Our team culture is built on collaboration, open communication, and a supportive environment where each member's ideas are valued and contributions are recognized. We believe in the importance of fostering a positive workplace culture that inspires innovation and creativity. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and production readiness reviews, platform management and capacity planning ceremonies with cross-functional teams Document Infrastructure operations process and insights, identify repeatable actions and ruthlessly automate repetitive tasks Participate in our teams on-call rotations, respond to incidents and support other teams mitigate customer impacting events Experience: 5+ years experience working on teams responsible for software development, automation and systems engineering Experience building large-scale infrastructure, distributed systems or networks. Knowledge with SQS,

pythonawsazure
View job →
L
Lyft
📍 Toronto• Full-time• From C$108K/yr
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As an Observability team member, you are responsible for the operation and maintenance of our logging and metrics infrastructure. You ensure all teams at Lyft are aware of the operational health of their products by monitoring system availability and take a holistic view of our platform performance. You build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you find opportunities to improve our systems in order to push our platform forward. You provide our partners with the support they need to help them build robust large scale distributed systems. We count on the reliability of our infrastructure to empower Lyft teams to provide our customers rich experiences that are highly available with rock solid performance to ensure our transportation platform continues to connect people and places. As we grow our team, we are seeking experienced Infrastructure Engineer to ensure that as our Infrastructure continues to scale, our platform continues to provide an essential and dependable service that transports millions of people every day. Specifically we are searching for someone who brings fresh perspectives, enjoys collaborating with cross-functional teams in order to continually improve our products and services for our customers. Responsibilities: Maintain, improve, and develop tooling and systems that enhance the reliability, scalability, and efficiency of our platform. Assist engineering teams in defining service-level objectives (SLOs) and provide the necessary toolin

pythonawskubernetes
View job →
D
Discord
📍 San Francisco Bay Area• Full-time• $196K – $220K/yr
1mo ago

Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. We're looking for a Senior Software Engineer to join the Growth team at Discord. Our team owns how new users discover, understand, and get started with Discord — from the first moment someone encounters us on the web through the onboarding experience that turns them into engaged members. You'll work across the stack to build the systems that acquire and activate users at scale, whether that means improving how Discord shows up across the web or designing the in-product experiences that help new users find their footing. This is a high-impact role where you'll directly influence how Discord grows. This person will report to the Senior Engineering Manager for Growth. What you'll be doing: Build and optimize user-facing experiences that improve how people discover, understand, and get started with Discord Develop scalable backend systems that serve dynamic, high-performance content at scale Work across the full stack — frontend (TypeScript, React) and backend (Python) — wherever the problem leads Design and run A/B experiments to measure the impact of your work on user acquisition and activation Partner with Product, Design, and Data Science to identify high-impact growth opportunities and iterate quickly Maintain performance and reliability standards for user-facing properties Write code that's readable, reviewable, and built with the next engineer in mind What you should have: At least 5 years of professional software engineering experience Strong expertise in TypeScript and React with a track record of delivering high-performance frontend experiences Practical experience with Python and backend framewor

typescriptpythonreact
View job →
D
1mo ago

Role Description As a Software Engineer on the Metadata team, you’ll build and operate the large-scale distributed databases that every Dropbox service depends on. Metadata systems are mission-critical, in the live path for all user operations and must meet stringent requirements for latency, durability, and transactional consistency. You’ll design and evolve the core infrastructure that manages Dropbox’s databases at scale, enabling fast, reliable access to data for millions of users and hundreds of internal services. This work spans distributed systems, replication, caching, and transactional database systems. You’ll collaborate closely with engineers across Infrastructure and Product teams to ensure the metadata layer meets business needs and continues to scale with Dropbox’s growth. This is an opportunity to leverage your expertise in distributed systems and grow into broader technical leadership. Our Engineering Career Framework is viewable by anyone outside the company and describes what’s expected for our engineers at each of our career levels. Check out our blog post on this topic and more here . Responsibilities Design and maintain distributed database systems providing low-latency, strongly consistent data access Implement and optimize replication, consensus, and caching mechanisms to meet availability and performance goals Operate production systems, including participating in the on-call rotation, ensuring high availability and data durability Collaborate with infrastructure and product teams to assess current and future use cases and requirements, supporting the development of a mid- to long-term roadmap that reflects these needs Contribute to system design reviews, postmortems, and reliability improvements Write high-quality, efficient code in Go and Rust for performance-critical systems On-call work may be necessary occasionally to help address bugs, outages, or other operational issues, with the goal of maintaining a stable and high-quality experienc

REMOTEsqlmysqlredis
View job →
D
Dropbox
📍 Poland• Full-time• Remote
1mo ago

Role Description As an Infrastructure Engineer, your role will be crucial in shaping and constructing the robust systems that not only support our current flagship products but also lay the groundwork for the next wave of engineering innovations. From optimizing user experiences across various projects to ensuring seamless scalability and data integrity, you'll be at the forefront of shaping the technological backbone of our platform. Collaborating closely with cross-functional teams, you'll leverage your expertise to tackle audacious challenges and push the boundaries of what's possible. Your contributions will directly impact millions of users, as every line of code you write furthers our mission to revolutionize the way people work and collaborate. Join us in redefining the future, where your passion for building scalable, reliable systems will drive meaningful change on a global scale. Our Engineering Career Framework is viewable by anyone outside the company and describes what’s expected for our engineers at each of our career levels. Check out our blog post on this topic and more here . Responsibilities Build infrastructure capable of managing metadata for hundreds of billions of files, handling hundreds of petabytes of user data, and facilitating millions of concurrent connections. Assist in expanding Dropbox's role as the data-fabric, linking hundreds of millions of applications, devices, and services worldwide, while spearheading efforts to improve interoperability and adaptability across various ecosystems. M easur e and optimiz e Dropbox's analytics platform to maintain its status as one of the most advanced in the industry for extracting meaningful insights from vast data volumes. Collaborat e with cross-functional teams to innovate and implement solutions that enhance the performance, reliability, and security of Dropbox's infrastructure, ensuring a seamless experience for users worldwide. On-call work may be necessary occasionally to help address bugs,

REMOTEpythonjavagit
View job →

Who We Are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies — from the world's largest enterprises to the most ambitious startups — use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the Team The Core Change Management group is responsible for the systems that let every Stripe engineer ship code, configuration, and infrastructure changes safely and at high velocity. You will be embedded primarily on the Service Deployments team — the owners of Stripe's end-to-end code deployment platform — with regular collaboration with the Resource Automation and Feature Deployments teams. What Makes This Role Compelling You own the foundation of how Stripe ships software. The deployment platform sits in the critical path of every engineer's workflow at Stripe. The decisions you make affect thousands of deploys per day across hundreds of services, directly determining how fast and safely Stripe's product evolves. Technically rich, architecturally active. The team is executing several concurrent platform transformations: containerizing host-based services at scale, adding intelligent multi-service deploy pipelines, extending real-time anomaly detection to earlier stages of traffic shifts, and rebuilding deployment event infrastructure on top of a durable message bus. This is not maintenance work — the architecture is in motion. Broad surface area, real ownership. You will span the full stack from container scheduling and deployment orchestration business logic to the developer-facing internal platform UI. The problems are multi-layered: reliability, developer experience, performance, and safety all at once. Your judgment prevents incidents.

awsazurekubernetes
View job →
S
Stripe
📍 Seattle• Full-time• From $156.8K/yr
1mo ago

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Design, develop, maintain, and support application programming interfaces (APIs), backend services, and distributed systems using Ruby, Java, Go and Scala. Build, test, deploy, and maintain software infrastructure to support large-scale production systems. Design and develop software systems using analytical techniques and mathematical models to evaluate system performance, predict outcomes, and assess design tradeoffs. Develop and execute software testing, validation, debugging, and documentation procedures to ensure system reliability and performance. Diagnose and resolve production issues across multiple services and layers of the technology stack. Analyze user needs and software requirements to determine technical feasibility, design approaches, and implementation timelines within cost and resource constraints. Collaborate with cross-functional engineering teams to design and implement new features for large-scale systems. Design and build systems to securely store, manage, and modify production data and services. Improve engineering standards, development tools, and software development processes to enhance system quality and efficiency. Who you are Minimum requirements Must have a Bachelor's degree or foreign equivalent in Computer Science, Software Engineering, or a related field, plus 2 years of software development experience. Must also have 2 years of experi

javaairuby
View job →
S
Stripe
📍 South San Francisco• Full-time• $156.8K – $235.2K/yr
1mo ago

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Architect, build, and maintain robust APIs, services, and systems across engineering teams using Ruby and Java; Act as a domain expert of the payments family of Stripe products, and in particular various payout methods and underlying rails used in US / UK; Build and enhance software infrastructure, encompassing the complete development lifecycle from coding to testing, deployment, debugging, and maintenance; Design software systems leveraging scientific analysis and mathematical models to predict and measure design outcomes; Contribute to core interface design and write code; Serve as a role model for how great software should be written; Contribute code changes in every layer of the stack, from launching a net new payout method API to making a UI change in the Stripe Dashboard for the Stripe Connect product; Scope and lead technical projects, laying the groundwork for early-stage products to iteratively evolve and scale; Engineer payments integration with banking partners to launch new markets, payout methods, and capabilities, while ensuring reliability and scalability; Work with engineers across the company to build new features at large-scale; Design and implement secure systems for sensitive data storage; Work with engineers across the company to understand when existing infrastructure can be leveraged vs. when building a bespoke solution is necessary

typescriptpythonjava
View job →
S
Stripe
📍 Seattle• Full-time• From $156.8K/yr
1mo ago

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Design, build, and maintain APIs, services, and systems across Stripe’s engineering teams using Java, Ruby, Scala, and Go Build software infrastructure, including developing, testing, and deploying it Design and develop software systems, using scientific analysis and mathematical models to predict and measure outcome and consequences of design Engineer payments integration with various financial partners software systems Design APIs and underlying data models to support complex financial abstractions and multi-party integrations, enabling flexible billing and settlement configurations across international markets Develop and direct software system testing and validation procedures, programming, and documentation Debug production issues across services and multiple levels of the stack Analyze user needs and software requirements to determine the feasibility of design within time and cost constraints Work with engineers across the company to build new features at large-scale Build new systems to securely store sensitive data; Improve engineering standards, tooling, and processes Integrate observability tools and alerting mechanisms (Datadog, Prometheus) into high-traffic production systems, defining SLAs and real-time diagnostics to meet the reliability standards required by financial partners Mentor junior engineers, lead technical design reviews, and establish best practices in code quality, system design, re

pythonjavasql
View job →
S
Stripe
📍 San Francisco Or New York• Full-time
1mo ago

Who we are About Bridge We're creating an entirely new payments platform, built with stablecoins, to simplify global money movement. Bridge enables faster, cheaper payments and borderless access to dollars via stablecoins. Through our APIs, businesses can send and receive funds across borders faster and cheaper vs. SWIFT and other fiat-only rails. Our virtual accounts enable international consumers and businesses to easily access, store, and spend US dollars. Our payouts infrastructure enables platforms to disburse USD to anyone globally. We believe many trillions of dollars will move and settle through stablecoin payment rails. Bridge is pulling this future forward. We have a small team of people who have previously built financial infrastructure at some of the world's leading companies (Coinbase, Stripe, Square, Brex, Upstart, DoorDash, Airbnb), and each and every one of them chose Bridge because they fundamentally believe that stablecoins will be a critical piece of financial infrastructure that allows for the improvement of global money movement. About the Role As one of the first blockchain engineers at Bridge, you'll have the autonomy to work on projects that are truly global in scale and aim to give customers access to products they've never had access to. Many of the use cases that we work on today didn't exist several months ago, and that's because our team is dedicated to constant improvement for our partners and end users. Our engineers have outsized ownership over projects, so if you're looking for increased autonomy working on a completely greenfield opportunity, Bridge is the place for you. Responsibilities • Build out multi-chain hot wallet infrastructure to support increased reliability of our payments. • Build smart contracts to increase hot wallet payment throughput. • Build smart contract platform to issue custom stablecoins across multiple blockchains (EVM, Solana, Stellar, etc.). • Build smart contract monitoring systems to detect anomalous even

aiswiftgo
View job →
S
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Tax compliance is a hard problem. When millions of payments flow through a marketplace or platform, the rules for what gets reported, to whom, in which format, and by when are complex, high-stakes, and vary by country. The Connect Tax Reporting team builds products to help businesses collect and verify tax information, prepare and deliver tax forms, and manage reporting obligations across jurisdictions. This includes 1099 reporting in the US and digital platform reporting across the EU, Australia, Canada, the UK, and other regions. This is a domain where correctness, reliability, and user trust matter deeply. A late filing, an incorrect form, or a missed reporting threshold can create regulatory risk and erode trust with platforms and their sellers. Our systems handle high-volume seasonal workflows, support complex country-specific rules, and remain dependable under fixed filing deadlines. What you’ll do You’ll build products and systems that simplify tax reporting for platforms and their users. You’ll take projects from discovery and technical design through implementation, launch, and operation. Your work may include improving tax information collection, building APIs and data pipelines, automating reporting processes, and strengthening the reliability of production systems. You’ll collaborate with product managers, tax operations, support, and engineers across Stripe. AI-assisted development will be part of your day-to-day work as you des

typescriptpythonjava
View job →
S
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Big Data Infrastructure operates the critical infrastructure that powers the batch data processing at Stripe. The team supports a variety of use cases, including Payment, Ledger, ML, Fraud Detection, Product Analytics, Regulatory Reporting, Financial Data Reconciliation, and externally facing products like Radar and Sigma. As an example of the scale, the team's systems serve hundreds of teams, thousands of workflows, 100,000+ task executions, O(billion) transformations, and moving terabytes of data processing over 1 GB/second every day. Our users inside Stripe include other engineering teams, Data Scientists, Sales and Operations, Finance, etc. Data Orchestration builds and operates the time-based and event-based orchestration infrastructure that powers and accelerates batch data pipelines. The team operates on a wide range of tech stacks including Airflow, Spark, SQL, Kafka, Flink, Hive MetaStore, Trino, Pinot, Python, Java, Scala, S3, and Iceberg. What you'll do As a Software Engineer on this team, you'll design and build infrastructure that powers batch data processing at Stripe. Responsibilities Design, build, and maintain next-generation and first-generation versions of key Data Platform products, with an emphasis on usability, reliability, security, and efficiency. Design ergonomic APIs and abstractions that build a great customer experience for internal Stripes, that will in turn enhance the experience of millions of Stripe users.

pythonjavasql
View job →

About the Team Enterprise Verticals builds role-specific ChatGPT Work experiences for high-value enterprise workflows. We combine product engineering, plugins and skills, connectors, data, evaluations, and customer evidence to turn useful demos into reliable daily work. This opening sits within the Technology vertical inside Enterprise Verticals. The group focuses on repeatable workflows for people at technology companies, beginning with functions such as data and analytics, sales, and design, and carries the shared platform needs—tool integration, permissions, quality measurement, and safe rollout—across those experiences. We work closely with Design, Research, GTM, Security, and platform teams, as well as with customers and design partners. Success means that people can reach a trustworthy first result, understand what the system did, and keep using the workflow—not merely that a prototype exists. About the Role We are looking for an exceptionally experienced, hands-on full-stack engineer to define and build the next generation of AI-powered enterprise workflows. You will take on the hardest and most ambiguous problems in the Technology vertical: translating real customer needs into product direction, designing the systems behind the experience, and personally writing and shipping production-quality code across the stack. You will own the technical direction and end-to-end delivery of products spanning ChatGPT Work surfaces, backend services, plugins, connectors, enterprise data, permissions, and evaluations. You will make foundational architecture and product tradeoffs; establish patterns other engineers can build on; and hold these experiences to a high bar for reliability, security, observability, and customer value. This is an individual-contributor role for an engineer who leads through technical judgment, direct execution, and influence—not people management. You should be equally comfortable working directly with customers, setting direction with senior cro

typescriptpythonreact
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team API Agents builds the shared agent harness, tools, and infrastructure that turn OpenAI’s frontier models into systems that can reliably complete real work. We carry the capabilities behind Codex into a much broader set of products and workflows across software engineering, research, finance, healthcare, enterprise operations, and more. Our work spans search and connected context, computer use, memory, delegation and multi-agent coordination, and safe execution. Sitting at the intersection of Research, Codex, infrastructure, and applied product teams, we build reusable agent capabilities that compound across the ecosystem. About the Role We are looking for an experienced backend software engineer to build the core systems behind the next generation of agents. You will design reliable services and abstractions that help agents find the right context, use tools and computers, retain knowledge, coordinate over long-running workflows, and take action safely. The role combines deep backend and infrastructure work with strong product judgment, with opportunities to work across agent runtimes, orchestration, search, execution environments, identity and permissions, observability, and evaluations. This is software and systems engineering rather than model training: success comes from strong backend fundamentals, high agency, and the ability to turn fast-moving research capabilities into dependable production primitives. In this role, you will: Design, build, and operate the shared agent harness and backend infrastructure that power long-running, high-value workflows across OpenAI and third-party products. Build reusable capabilities across search and connected context, computer use, memory, tool execution, delegation, subagents, and multi-agent orchestration. Establish the foundations agents need to operate safely in production, including secure execution environments, identity and permissions, observability, evaluations, reliability, and cost and latency effi

typescriptpythonaws
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime