Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! The Production Engineering team focuses on end-to-end reliability, durability, scalability, and performance of Figma products and services. We’re looking for an experienced Engineer to scale and drive the initiatives and programs that support Figma’s production engineering efforts, allowing our product teams to rapidly deliver new features. This is a critical part of success for both our ability to build new products and features, support the growth of our user base, and enable innovation for all engineering (infrastructure and product teams, alike). We’re looking for a generalist with a strong grasp of Computer Science fundamentals. The ideal candidate will have experience building and running complex large scale services. Additionally, they will have a passion for building tools that simplify operating large-scale infrastructure and for improving the operational maturity of Figma. What you'll do at Figma: Work closely with the engineering team to define standard methodologies and goals around reliability, durability, scalability, and performance Address common operational challenges through better telemetry and by building tools / services Debug production issues across services and levels of the stack Participate in design reviews and production reviews for new features, products or infrastructure components Plan for the growth of Figma’s infrastructure Operate and maintain AWS Infrastructure We’d love to hear from you if you have: Have 5+ years of experience operating infrastruct
Jobiba hiring network
Reliability Engineer Jobs
2,049 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Squarespace provides innovative solutions to empower our customers to focus on building their brand and growing their businesses on our platform. The Databases team manages all of the backend infrastructure that Squarespace runs on – MongoDB, CockroachDB, and Kafka clusters, to name a few examples. We are an accomplished, diverse group of people who develop the services that guarantee reliable and scalable infrastructure for both our cross-functional partners in product engineering, as well as our end users on the Squarespace platform. We believe that infrastructure excellence doesn't stop at just building for today; it needs to have a solid foundation of scalability, reliability, and a robust developer experience for the future. This is a hybrid role working from our Dublin office 3 days per week. You will report to the Databases Senior Engineering Manager. You’ll Get To… Nurture high-performing software engineers by guiding navigation when there is ambiguity. Distill the scope of the team and help hire a balanced group of engineers that will excel as a unit. Grow the career development of direct reports through regular 1:1s with direct, actionable feedback. Celebrate wins that motivate the team’s positive culture and robust dynamic. Evaluate consistently to improve team efficiency and effectiveness when required. Evolve a deep understanding of local systems to identify appropriate architectural decisions. Thread with Product, Design & Engineering to champion, define and execute an optimal roadmap. Bond across Engineering, Product, Design, Marketing, Data Science and Business Operations. Who We’re Looking For 3+ years of recent experience managing a Product Engineering team of four or more engineers. 7+ years of industry experience deploying apps across large codebases with many contributors. Ability to fluently translate, document and present technical concepts to non-technical stakeholders. Strong technical foundations to navigate the inherent tra
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! AI Platform teams build the core frameworks, abstractions, and systems that support AI features across Figma. We create new capabilities that AI product teams can build on, while collaborating with teams from around the company to improve our performance, reliability, and technical quality. We’re looking for strong infrastructure and platform-minded engineers to contribute to our agent infrastructure, context retrieval & ranking platform, and core AI services, in order to accelerate our most critical company-wide AI initiatives. Here are just a few areas our platform teams work on: Evals for design : Building evaluation frameworks for design generation quality that are used across every Figma AI feature. Agentic search: Providing relevant context from throughout the Figma ecosystem to agents via search tools, in order to improve agent quality. Figma MCP : Making our MCP server faster, more reliable, and easier for internal & external engineers alike to develop on. Agent infrastructure : Iterating on the sandboxes, harnesses, and tools leveraged by the agents that power Figma Make and the Figma Design Agent. Preview & publishing platforms : Creating the shared platforms for building, previewing, and publishing code written by agents via Figma Make. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you’ll do at Figma: Support end-to-end AI feature development by designing, building, and maintaining systems that are scalable, reliab
ABOUT THE ROLE Peloton is looking for a talented Data Engineering Manager to join the Data Engineering team. In this role, you will lead the DataOps function, driving operational excellence across our data platforms while managing the successful delivery of data initiatives through a combination of internal and offshore engineering resources. You will work closely with business stakeholders, engineering teams, analytics partners, and platform owners to ensure our data ecosystem remains reliable, scalable, and well-governed. This role combines technical leadership, operational management, and stakeholder engagement, to help support Peloton's growing data and AI needs. This role will be hybrid, not remote. YOUR DAILY IMPACT AT PELOTON Lead the DataOps function within the Data Engineering team, driving operational maturity, platform reliability, and process improvements Manage and mentor offshore engineering resources, providing technical guidance, performance feedback, and delivery oversight Partner with stakeholders across multiple business functions to gather requirements, prioritize work, and ensure successful delivery of data solutions Own intake, prioritization, and execution processes for DataOps requests and operational support activities Drive adoption of data engineering standards, ETL best practices, documentation requirements, and operational procedures Serve as an operational owner for key data platforms, including Airflow, Airbyte, and Looker, owning platform governance activities such as user access reviews, compliance reviews, audit support, and operational controls Coordinate incident management, root cause analysis, monitoring, alerting, and operational readiness efforts across the data platform Help shape the long-term operating model for the Data Engineering team as Peloton continues to scale its data and AI initiatives YOU BRING TO PELOTON 5+ years of experience in data engineering, analytics engineering, software engineering, or related technical
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Safety and Customer Care (SCC) team at Lyft manages over 1.7 million monthly human and AI interactions and serves as Lyft's primary direct touchpoint with riders and drivers. We handle critical infrastructure that powers both human associates and AI agents to make riders and drivers feel safe and comfortable while riding or driving with Lyft, transforming every support interaction into a moment of genuine connection. As a Data Engineer on the SCC team, you will have ownership over the data modeling and pipelines that power SCC’s Associate and AI Agent Platform . Your efforts will be critical to the reliability of our pipelines, execution of third party data integrations, accurate reporting of agents performance, and efficiency improvements that can save millions of dollars / year. You will work cross-functionally to bridge Lyft's business goals with data engineering. Your efforts will allow access to business and user behavior insights, using huge amounts of Lyft data to fuel several teams such as Analytics, Data Science, Engineering, and many others. Responsibilities: Owner of the core data pipeline, responsible for scaling up data processing flow to meet the rapid data growth at Lyft Evolve data model and data schema based on business and engineering needs Implement systems tracking data quality and consistency Develop tools supporting self-service data pipeline management (ETL) SQL and MapReduce job tuning to improve data processing performance Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and distribute knowledge Collaborate cross-functionally with product, engineering, data science, and marketing teams to understand business problems and align on prioritization and solutions Experience: Bachelor's degree in Compute
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Core Services powers the passenger and driver state machines from request and accept through drop-off and payment completion. The team enables new ride variations like autonomous vehicles, taxis, scheduled, business concierge, health and more. The “Core Services” are a suite of distributed Python and Golang systems central to Lyft’s backend. As a software engineer on the team, you will work on integrating new rider and driver products and features onto the state machine and enhancing the performance and reliability of ride state transitions. If you are excited about solving back-end distributed systems problems and owning a mission-critical part of Lyft’s operations, this team is for you. As a Software Engineer at Lyft, you will collaborate with other engineers and cross-functional teams, such as product, data science, and analytics, to lead and execute large projects—from concept to efficient execution. We are looking for motivated engineers who are passionate about solving challenging technical problems and excited to work in a fast-paced, innovative, and cross-functional environment. In this role, you will tackle some of the most interesting and impactful problems in ridesharing. Key traits for success include being passionate about Lyft’s business and product, a quick learner, a collaborative mindset, and an eagerness to drive initiatives both within and across teams. You'll be joining a small, close-knit team with engaged and collaborative co-workers. Responsibilities: Design, develop, deploy, monitor, operate and maintain existing or new elements of the Fulfillment tech stack Write well-crafted, well-tested, readable, maintainable code Have a good grasp and ability to explain the various tradeoffs made in decisions Participate in code reviews to ensure code quality and distribute knowledg
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft AV brings the autonomous future to life by matching our community of riders with self-driving vehicles to meet their transportation needs today. The program is powered by a technical team innovating to build new features, experiment, and evolve approaches for emerging business needs. At the same time, we prioritize production quality because our customers put their trust in us for safety and reliability. As a Software Engineer for Autonomous Vehicles, you’ll be responsible for executing integrations with partners, making tradeoffs between technical investments and product work, and collaborating with other engineers on system design. You will help shape the product direction by developing a deep understanding of the customer and working closely with cross-functional partners from Product, Design, Science, and Operations. You will work with a group of talented engineers and help the team to deliver significant business impact while being open to change through constant experimentation in an ambiguous emerging product. If you enjoy collaborating with technical and nontechnical partners in a fast-moving space with real-world impact, this is the role for you. Responsibilities: Help establish roadmap and architecture based on technology and understanding of customer needs Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and distribute knowledge Share your knowledge by giving brown bags, tech talks, and promoting appropriate tech and engineering best practices Can help lead large projects from idea to positive execution Unblock, support and communicate with internal partners to achieve results Experience: BS/MS or equivalent in Computer Engineering, Computer Science, or related field or relevant work experience Experience in distributed sy
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Unified Data Store (UDS) team is the architect of Airbnb’s global system-of-record. We design, build, and operate the mission-critical storage platform that powers every user profile, listing, reservation, and financial transaction on the platform. Supporting over 150 million users worldwide, our work is the bedrock of Airbnb’s reliability and efficiency. As a Staff Engineer in our Brazil Engineering Hub , you will join a high-impact group of technical leaders who value craft and operational excellence. You won't just be managing data; you will be building a modern distributed infrastructure service that enables hundreds of product teams to ship features with total confidence. The Difference You Will Make: We are looking for a hands-on technical leader who thrives on solving deep architectural challenges and leading through ambiguity. As a Staff Engineer, you will serve as the Technical North Star for the UDS Client Stack, ensuring our data access layer is seamless, high-performance, and future-proof. A Typical Day: Define Technical Strategy: Lead the multi-year roadmap and long-term architecture for the UDS client stack, balancing immediate execution with systemic platform evolution. Architect for Scale: Design and operate a high-performance data access layer that abstracts complexities like indexing, replication, and global consistency models. Drive Engineering Excellence: Lead deep-dive design reviews and establish best practices for building fault-tolerant distributed systems across the organization. Empower Developers: Act as a bridge between infrastructure and
Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. The Realtime Infrastructure team is responsible for building and maintaining some of Discord’s highest scale and most critical services. Those systems are at the core of our text chat infrastructure and facilitate the dispatching of every update to our users sessions. This role will have a significant impact on Discord’s overall reliability and performance. It will also help our product teams build new features on top of our infrastructure. This team is small but critical, and its work has a direct impact on Discord's success and ability to scale. This role reports to the Senior Engineering Manager of Realtime Infrastructure. What You'll Be Doing Build and operate large-scale, reliable and performant distributed systems. Collaborate with product teams to create new features. Ensure Discord “just works”. Write code but also manage our infrastructure. Work with a talented team of engineers who have built one of the largest communication platforms in the world. What you should have 2+ years of experience writing and designing backend systems. Experience solving complex distributed system problems. Experience operating and maintaining critical tier 0 services. Knowledge of monitoring and alerting best practices. Familiar with open source software, and not afraid to dig into the source code of a library to find the answer you’re looking for. Bonus Points Experience with Elixir or Rust. Experience working with systems deployed in a cloud environment (GCP, AWS, etc.) Knowledge of devops tools like Salt,Terraform or k8s. You have built or contributed to open source projects. You are a Discord power user and hav
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Web Presence and Platform is organized into two pillars, each of which is grouped into pods that focus on the central tenets of the Stripe public mission. The Presence pillar creates industry-leading designs for Stripe front door surfaces, educating customers about the power of our platform, sharing ideas and expertise with the public, and driving adoption. The Platform pillar builds the internal machinery that powers these surfaces, and is responsible for making our websites fast, stable, and easy to update. Together we design and build stripe.com and other sites that amount to what is, for many, their first impression of Stripe. As such, WPP offers exciting opportunities to have a major impact on Stripe success. We want to make every pixel count, we want it to be enthralling, and we want to help other Stripes seamlessly benefit from our systematic work. While the team's mission anchors to web surfaces, this role lives deeper in the stack. You'll be the engine behind the experiences rather than the pixel-level presentation—APIs, data pipelines, model orchestration, and service reliability that make LLM-driven user experiences possible at scale. What you'll do The Expansion pod brings creativity and executional rigor to attracting new prospects and driving conversion, focusing on user journeys, interactive and highly polished tools for users, and building targeted experiences for our global users. As a full stack engineer, you'll architect a
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Stripe Infrastructure is responsible for the reliability, scale, performance, and cost of Stripe's systems and the productivity and sentiment of Stripe's people. You may work on a wide variety of critical business areas including: Core Infrastructure—We're the home for Stripe's critical tier0 infrastructure systems (Compute, Networking, DocumentDB, Distributed Caching and High assurance engineering). We build the foundational platform for Stripe products and services to allow them to operate at scale. We drive reliability, availability, efficiency, and scalability of these systems. Developer Infrastructure—We're responsible for the productivity of all developers at Stripe. Ensure Stripe's engineers have a reliable, fast, and easy-to-use inner dev loop to maximize productivity while building everything from low-latency microservices to large-scale data pipelines and machine learning models. Data Infrastructure—We're responsible for offering data serving infrastructure spanning across data warehouse analytics, streaming analytics, and search capabilities. The stack is supported by a collection of internally developed large-scale distributed services and several popular open-source technologies like Trino/Presto, Apache Pinot, Hive Metastore, ElasticSearch etc. The systems we own support all of the data serving needs of high-scale services and thousands of individual Stripes across the company. Admin Platform—We empower Stripes to quickly
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Payments is Stripe's flagship product. Our payments platform processes trillions of dollars annually. This team has the opportunity to expand the reach of Stripe's global payments network, design and implement novel payment capabilities, and deliver best-in-class reliability and performance. We have software engineers in almost every team across Stripe, and in this role, you'll be making some of the most significant decisions for the company. You'll get to work with other engineers to build features that span various parts of the system, as well as our business, sales, and operations teams to understand and solve our users' pain points. What you'll do You'll work on projects that span technologies, systems, and processes where you'll design, build, test, and ship great code every day. Responsibilities Design, build, and maintain APIs, services, and systems across Stripe's engineering teams Work with a wide range of systems, processes, and technologies to own and solve problems end-to-end Collaborate with other engineering teams across Stripe's global offices Uphold our high engineering standards and bring consistency to the many codebases and processes you encounter Improve engineering standards, tooling, and processes Who you are You're energized by solving real problems for real users. You use the word "users" or "customers" a million times a day. You thrive at the intersection of technology and business—equally comfortable diving deep int
At Bolna, we’re building tools that change how businesses leverage voice AI. We’re looking for a Software Engineer to build reliable, scalable systems that power millions of production conversations across languages, industries, and telephony environments. This is a high-impact, high-ownership role where you’ll work on core platform problems across distributed systems, real-time communication, developer infrastructure, and customer-facing products. Our team includes IIT alumni with experience at Bain, Atlassian, Uber, Zomato, and LinkedIn, and is backed by leading investors. Responsibilities Build systems that operate at scale: Design and build backend services that support high-volume, real-time voice AI conversations with strong reliability, performance, and fault tolerance. Own features end to end: Take problems from product requirements and technical design through implementation, testing, deployment, monitoring, and iteration. Improve platform reliability: Build systems that are observable, resilient, and easy to debug. Identify bottlenecks, reduce failure rates, and improve system availability. Work on real-time infrastructure: Solve problems across telephony, streaming audio, webhooks, queues, scheduling, concurrency, and low-latency communication. Build for developers and customers: Improve APIs, SDKs, integrations, dashboards, and internal tools that make the Bolna platform easier to use and operate. Raise the engineering bar: Contribute to technical design reviews, code quality, testing standards, documentation, incident response, and engineering best practices. Required Skills Strong engineering fundamentals: Solid understanding of data structures, algorithms, databases, networking, operating systems, and distributed systems. Backend development experience: 2+ years of experience building and operating production backend systems using Python, Go, Java, Node.js, or a similar language. Production ownership: Experience shipping software to production and own
About the Team OpenAI’s Network Engineering team within IT and Security advances the mission of deploying artificial general intelligence (AGI) for the benefit of all by delivering secure, scalable, and resilient network services. We build and operate the connectivity that supports OpenAI’s offices, labs, campuses, cloud environments, people, and devices. By combining strong network fundamentals with security, reliability, automation, and user-centered design, we enable impactful AI research, corporate operations, and product innovation. About the Role As a Network Engineer at OpenAI, you will design, operate, and continuously improve the global networks that connect our offices, labs, campuses, PoPs, cloud environments, people, and devices. The role spans strategic platform engineering and responsive production operations: you will shape architecture, standards, roadmaps, lifecycle plans, and automation while supporting incidents, escalations, and time-sensitive delivery. Operational signals will inform what we stabilize, simplify, standardize, or automate next. We work backward from user needs, investigate root causes, own outcomes end-to-end, and move quickly without compromising security. We are looking for a versatile engineer who can make pragmatic reliability and security tradeoffs, communicate clearly, and turn recurring operational work into durable platforms, tooling, and standards. You will partner across IT, Security, AppEng, Research, Applied, workplace teams, carriers, and vendors. In this role, you will: Design, implement, and operate secure, scalable enterprise networks across offices, labs, campuses, PoPs, cloud connectivity, and hybrid environments. Set strategic direction for network services through architecture, standards, roadmaps, lifecycle planning, capacity strategy, and measurable reliability outcomes. Own production operations, including on-call, incident response, escalations, and time-sensitive delivery, while protecting user experience,
About the Team The Privacy Engineering team builds secure, reliable systems that help OpenAI meet its legal obligations while protecting user data. We partner closely with Legal and engineering teams across OpenAI to support lawful data access requests and other critical legal workflows. Our work turns complex, high-stakes processes into auditable and dependable technical systems with clear human oversight and strong privacy and security controls. About the Role We’re looking for a full-stack Software Engineer to build the internal tools and data pipelines that power lawful data access request workflows and Legal Operations. You will work across product and data systems to make authorized retrieval and case handling accurate, efficient, and auditable. This role is well suited to someone who enjoys translating ambiguous operational requirements into durable systems, cares deeply about sensitive-data handling, and wants to improve both technical reliability and the day-to-day experience of the people operating these workflows. In this role, you will: Design, build, and operate backend systems and workflow tooling for the full lifecycle of lawful data access requests, from intake and scoping through authorized retrieval, review, preparation, and audit. Build reliable data pipelines and interfaces across products and data stores so authorized teams can locate and handle the right records accurately and reproducibly. Implement least-privilege access, approval gates, provenance, audit trails, data minimization, and safe failure modes for sensitive workflows. Partner with Legal and Legal Operations to translate legal and operational requirements into clear technical designs and intuitive operator experiences. Identify responsible automation opportunities that reduce repetitive work while preserving human review, judgment, and accountability. Own production systems through testing, observability, incident response, documentation, and continuous reliability improvements. Hel
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime