At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Core Services powers the passenger and driver state machines from request and accept through drop-off and payment completion. The team enables new ride variations like autonomous vehicles, taxis, scheduled, business concierge, health and more. The “Core Services” are a suite of distributed Python and Golang systems central to Lyft’s backend. As a software engineer on the team, you will work on integrating new rider and driver products and features onto the state machine and enhancing the performance and reliability of ride state transitions. If you are excited about solving back-end distributed systems problems and owning a mission-critical part of Lyft’s operations, this team is for you. As a Software Engineer at Lyft, you will collaborate with other engineers and cross-functional teams, such as product, data science, and analytics, to lead and execute large projects—from concept to efficient execution. We are looking for motivated engineers who are passionate about solving challenging technical problems and excited to work in a fast-paced, innovative, and cross-functional environment. In this role, you will tackle some of the most interesting and impactful problems in ridesharing. Key traits for success include being passionate about Lyft’s business and product, a quick learner, a collaborative mindset, and an eagerness to drive initiatives both within and across teams. You'll be joining a small, close-knit team with engaged and collaborative co-workers. Responsibilities: Design, develop, deploy, monitor, operate and maintain existing or new elements of the Fulfillment tech stack Write well-crafted, well-tested, readable, maintainable code Have a good grasp and ability to explain the various tradeoffs made in decisions Participate in code reviews to ensure code quality and distribute knowledg
Jobiba hiring network
Software Reliability Engineer Jobs
6,428 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft AV brings the autonomous future to life by matching our community of riders with self-driving vehicles to meet their transportation needs today. The program is powered by a technical team innovating to build new features, experiment, and evolve approaches for emerging business needs. At the same time, we prioritize production quality because our customers put their trust in us for safety and reliability. As a Software Engineer for Autonomous Vehicles, you’ll be responsible for executing integrations with partners, making tradeoffs between technical investments and product work, and collaborating with other engineers on system design. You will help shape the product direction by developing a deep understanding of the customer and working closely with cross-functional partners from Product, Design, Science, and Operations. You will work with a group of talented engineers and help the team to deliver significant business impact while being open to change through constant experimentation in an ambiguous emerging product. If you enjoy collaborating with technical and nontechnical partners in a fast-moving space with real-world impact, this is the role for you. Responsibilities: Help establish roadmap and architecture based on technology and understanding of customer needs Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and distribute knowledge Share your knowledge by giving brown bags, tech talks, and promoting appropriate tech and engineering best practices Can help lead large projects from idea to positive execution Unblock, support and communicate with internal partners to achieve results Experience: BS/MS or equivalent in Computer Engineering, Computer Science, or related field or relevant work experience Experience in distributed sy
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Passport & Commerce builds the trusted account and commerce foundation that lets anyone, anywhere, become part of the Airbnb community — from their first sign-up, through the profile and connections they build, to every booking and business transaction along the way. We're a newly formed org within Guest & Host that brings together identity, account, and commerce foundations under one roof. Our mission is to move Airbnb beyond the transaction — building a world where accounts create trust, value sticks, and what our community earns travels with them, and their businesses, wherever they go. You'll work closely with Payments and Wallet engineering, Identity & Privacy, Profile & Community, and Guest & Host product, design, and data science partners as we stand up this team's roadmap and technical foundations. The Difference You Will Make: As a Staff Software Engineer on the Passport team, you will be a key architect behind the next generation of our account platform, directly influencing how millions of guests and hosts experience Airbnb. You'll set technical direction for how account state, eligibility, and entitlements are computed, stored, and served at scale across every surface where a guest or host interacts with Airbnb. Success in this role looks like a Passport platform that is reliable and extensible enough to support new programs and partner integrations without re-architecture — measured through service reliability (uptime, latency, correctness of entitlement calculations), the speed at which new offerings can launch, and adoption of your platform by othe
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Airbnb is a mission-driven company dedicated to helping create a world where anyone can belong anywhere. It takes a unified team committed to our core values to achieve this goal. Airbnb's various functions embody the company's innovative spirit and our fast-moving team is committed to leading as a 21st century company. The Difference You Will Make: The Service Tools team’s mission is to enable Airbnb backend developers to develop, test, and maintain their code quickly and reliably. This team is responsible for the standard development lifecycle for service owners—everything from AI Integration for service development, to how integration tests work, to building services for deployment. This represents Airbnb’s largest cohort of developers, you will ultimately be responsible for their productivity. A Typical Day: As an engineer on Service Tools, you will work on technologies that help shape an industry-leading end to end developer experience for backend developers. In this role you will be: Building our next-gen build system using the latest technologies (e.g., Bazel). Working on integrations between the build system and CI/CD tooling (e.g., merge queues, code coverage, integration testing). Improving the editor (e.g., IntelliJ) experience for all backend developers. Helping to shape the technical strategy that directly moves our core metrics (Developer Experience, Developer Velocity, Debuggability, Resilience and Reliability) while reducing cost. Partnering with engineering leaders across all Airbnb teams for adoption of the new capabilities. Your customers will be all enginee
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Payments is Stripe's flagship product. Our payments platform processes trillions of dollars annually. This team has the opportunity to expand the reach of Stripe's global payments network, design and implement novel payment capabilities, and deliver best-in-class reliability and performance. We have software engineers in almost every team across Stripe, and in this role, you'll be making some of the most significant decisions for the company. You'll get to work with other engineers to build features that span various parts of the system, as well as our business, sales, and operations teams to understand and solve our users' pain points. What you'll do You'll work on projects that span technologies, systems, and processes where you'll design, build, test, and ship great code every day. Responsibilities Design, build, and maintain APIs, services, and systems across Stripe's engineering teams Work with a wide range of systems, processes, and technologies to own and solve problems end-to-end Collaborate with other engineering teams across Stripe's global offices Uphold our high engineering standards and bring consistency to the many codebases and processes you encounter Improve engineering standards, tooling, and processes Who you are You're energized by solving real problems for real users. You use the word "users" or "customers" a million times a day. You thrive at the intersection of technology and business—equally comfortable diving deep int
At Bolna, we’re building tools that change how businesses leverage voice AI. We’re looking for a Software Engineer to build reliable, scalable systems that power millions of production conversations across languages, industries, and telephony environments. This is a high-impact, high-ownership role where you’ll work on core platform problems across distributed systems, real-time communication, developer infrastructure, and customer-facing products. Our team includes IIT alumni with experience at Bain, Atlassian, Uber, Zomato, and LinkedIn, and is backed by leading investors. Responsibilities Build systems that operate at scale: Design and build backend services that support high-volume, real-time voice AI conversations with strong reliability, performance, and fault tolerance. Own features end to end: Take problems from product requirements and technical design through implementation, testing, deployment, monitoring, and iteration. Improve platform reliability: Build systems that are observable, resilient, and easy to debug. Identify bottlenecks, reduce failure rates, and improve system availability. Work on real-time infrastructure: Solve problems across telephony, streaming audio, webhooks, queues, scheduling, concurrency, and low-latency communication. Build for developers and customers: Improve APIs, SDKs, integrations, dashboards, and internal tools that make the Bolna platform easier to use and operate. Raise the engineering bar: Contribute to technical design reviews, code quality, testing standards, documentation, incident response, and engineering best practices. Required Skills Strong engineering fundamentals: Solid understanding of data structures, algorithms, databases, networking, operating systems, and distributed systems. Backend development experience: 2+ years of experience building and operating production backend systems using Python, Go, Java, Node.js, or a similar language. Production ownership: Experience shipping software to production and own
About the Team The Privacy Engineering team builds secure, reliable systems that help OpenAI meet its legal obligations while protecting user data. We partner closely with Legal and engineering teams across OpenAI to support lawful data access requests and other critical legal workflows. Our work turns complex, high-stakes processes into auditable and dependable technical systems with clear human oversight and strong privacy and security controls. About the Role We’re looking for a full-stack Software Engineer to build the internal tools and data pipelines that power lawful data access request workflows and Legal Operations. You will work across product and data systems to make authorized retrieval and case handling accurate, efficient, and auditable. This role is well suited to someone who enjoys translating ambiguous operational requirements into durable systems, cares deeply about sensitive-data handling, and wants to improve both technical reliability and the day-to-day experience of the people operating these workflows. In this role, you will: Design, build, and operate backend systems and workflow tooling for the full lifecycle of lawful data access requests, from intake and scoping through authorized retrieval, review, preparation, and audit. Build reliable data pipelines and interfaces across products and data stores so authorized teams can locate and handle the right records accurately and reproducibly. Implement least-privilege access, approval gates, provenance, audit trails, data minimization, and safe failure modes for sensitive workflows. Partner with Legal and Legal Operations to translate legal and operational requirements into clear technical designs and intuitive operator experiences. Identify responsible automation opportunities that reduce repetitive work while preserving human review, judgment, and accountability. Own production systems through testing, observability, incident response, documentation, and continuous reliability improvements. Hel
About the Team The Privacy Engineering team builds secure, reliable systems that help OpenAI meet its legal obligations while protecting user data. We partner closely with Legal and engineering teams across OpenAI to support lawful data access requests and other critical legal workflows. Our work turns complex, high-stakes processes into auditable and dependable technical systems with clear human oversight and strong privacy and security controls. About the Role We’re looking for a full-stack Software Engineer to build the internal tools and data pipelines that power lawful data access request workflows and Legal Operations. You will work across product and data systems to make authorized retrieval and case handling accurate, efficient, and auditable. This role is well suited to someone who enjoys translating ambiguous operational requirements into durable systems, cares deeply about sensitive-data handling, and wants to improve both technical reliability and the day-to-day experience of the people operating these workflows. In this role, you will: Design, build, and operate backend systems and workflow tooling for the full lifecycle of lawful data access requests, from intake and scoping through authorized retrieval, review, preparation, and audit. Build reliable data pipelines and interfaces across products and data stores so authorized teams can locate and handle the right records accurately and reproducibly. Implement least-privilege access, approval gates, provenance, audit trails, data minimization, and safe failure modes for sensitive workflows. Partner with Legal and Legal Operations to translate legal and operational requirements into clear technical designs and intuitive operator experiences. Identify responsible automation opportunities that reduce repetitive work while preserving human review, judgment, and accountability. Own production systems through testing, observability, incident response, documentation, and continuous reliability improvements. Hel
About the Team The Privacy Engineering team builds secure, reliable systems that help OpenAI meet its legal obligations while protecting user data. We partner closely with Legal and Engineering teams across OpenAI to support lawful data access requests and other critical legal workflows. Our work turns complex, high-stakes processes into auditable and dependable technical systems with clear human oversight and strong privacy and security controls. About the Role We’re looking for a full-stack Software Engineer to build the internal tools and data pipelines that power lawful data access request workflows and Legal Operations. You will work across product and data systems to make authorized retrieval and case handling accurate, efficient, and auditable. This role is well suited to someone who enjoys translating ambiguous operational requirements into durable systems, cares deeply about sensitive-data handling, and wants to improve both technical reliability and the day-to-day experience of the people operating these workflows. This role is based in San Francisco, CA, with two additional locations under consideration: London, UK, and Dublin, Ireland. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and operate backend systems and workflow tooling for the full lifecycle of lawful data access requests, from intake and scoping through authorized retrieval, review, preparation, and audit. Build reliable data pipelines and interfaces across products and data stores so authorized teams can locate and handle the right records accurately and reproducibly. Implement least-privilege access, approval gates, provenance, audit trails, data minimization, and safe failure modes for sensitive workflows. Partner with Legal and Legal Operations to translate legal and operational requirements into clear technical designs and intuitive operator experiences. Identify responsible automation o
About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man
About the Team OpenAI’s Applications Engineering organization builds and operates the products (such as ChatGPT & Codex) that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role We’re hiring Backend Software Engineers to design and implement safe services and infrastructure that power our core products. What You’ll Do Architect, build, and improve scalable backend systems and APIs. Drive performance, reliability, and safety across distributed services. Implement data storage, retrieval, compute, and integration solutions. Participate in long-term architectural planning and technical design reviews. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. You Might Thrive Here If You: Have strong experience with distributed systems, APIs, and backend languages (e.g., Go, Python, Rust, C++). Have experience setting up and maintaining production backend services and data pipelines. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Enjoy building resilient services that handle large scale and complexity. Are self-directed and enjoy figuring out the best way to solve a particular problem Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
About the Team Our London-based team builds the backend systems that help ChatGPT scale reliably. We work on infrastructure close to the product, partnering with engineering teams to improve the performance, resilience, and operability of critical user-facing systems. Our work combines backend software engineering with distributed systems and production reliability. We build shared capabilities, improve high-traffic workflows, and make it easier to introduce new product functionality without compromising performance or availability. About the Role This role is for software engineers who want to build and evolve backend systems operating at significant scale. You’ll write production code, design shared infrastructure, and solve technical challenges involving performance, distributed systems, and system reliability. You’ll also own how those systems behave in production: how changes are rolled out, how issues are detected and diagnosed, and how recurring operational problems can be addressed through better software and system design. This is a strong fit for backend engineers who enjoy complex systems problems and want a direct connection between the infrastructure they build and the experience of ChatGPT users. In this role, you will: Design, build, and maintain backend systems supporting high-traffic ChatGPT experiences. Develop shared services, APIs, and infrastructure that help product teams build and launch new capabilities safely. Improve the performance, scalability, and efficiency of production systems as usage and product complexity grow. Build and improve systems for asynchronous processing and other large-scale backend workloads. Lead architectural improvements and infrastructure migrations while maintaining correctness, compatibility, and safe rollout and rollback. Strengthen monitoring, alerting, and diagnostics to detect problems early and reduce customer impact. Participate in on-call, incident response, and root-cause analysis, and turn operational lea
Get new software reliability engineer jobs by email
Daily job updates · Unsubscribe anytime