About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. About the role As a Senior Software Engineer, you will architect and scale the core systems that power Sentry's billing infrastructure, ensuring accuracy and reliability at massive scale. You will collaborate on building the next generation of Sentry’s usage tracking pipeline, processing hundreds of billions of events daily with low latency and financial-grade accuracy. You will help design flexible pricing primitives that support everything from per-event usage billing to complex enterprise contracts, enabling product and sales teams to experiment rapidly while maintaining revenue accuracy and reduced time-to-market for new products. You will contribute to technical decisions on data consistency challenges unique to billing—like handling event delays, retroactive pricing changes, and distributed count reconciliation across our infrastructure. You'll love this job if you Want to solve the "easy to explain, hard to build" problems—like ensuring a customer's bill matches their usage perfectly, even when processing hundreds of billions of events daily across distributed
Jobiba hiring network
Reliability Engineer Jobs
2,028 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides tools that help developers find and fix issues in their applications. The Developer Infrastructure team owns the systems that make every engineer at Sentry more effective at shipping high quality software: developer environments, CI/CD for our open source and closed source codebases, the golden path for building new services, our SDK and library publishing tooling, and the metrics and dashboards that keep it all healthy. Our goal is straightforward: developers should spend their time thinking about the software they build for customers, not the software they use to build it. That means everything from local and cloud development environments, to the CI/CD pipeline that gets code out safely and quickly, to the tooling that catches flaky tests before they cost someone a day. As AI coding agents become part of how engineers work, we're also investing in making sure our environments, CI, and deployment systems support that shift. As a Software Engineer on Dev Infra, you'll help build and scale this infrastructure end to end. In this role, you will Build and maintain the tooling that powers local and cloud-based developer environments Improve CI for our codebases and CD for our deployment pipeline, keeping both fast and reliable as the org and its infrastructure grow Define and evolve the golden path for building new services and libraries at Sentry Contribute to our SDK and library publishing tooling and release processes Build the metrics, dashboards, and flaky test detection that give the org visibility into deployment health and CI reliability Build tooling that gives engineers, and the AI codin
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry's Infrastructure Engineering team is what makes operating Sentry simple, safe, and seamless for every other engineering team in the company. They build the internal control platforms, configuration systems, traffic routing, and automation that let product engineers operate services safely at scale without needing deep infrastructure expertise themselves. As the Engineering Manager for Infrastructure Engineering, you'll lead a team of engineers building the tools that power Sentry's growth: internal admin and change management tools, configuration automation, and the routing layer that underlies Sentry's architecture. You'll be responsible for technical vision, team health, system reliability, and partnership with engineering teams across the company who depend on your team's tools every day. You'll work closely with leaders across Infrastructure, Platform, and Production Engineering to shape how Sentry scales its operational model as the company grows. In this role you will Lead a team of engineers building the internal control platforms that every engineering team at Sentry relies on to operate services safely. Drive the evolution of Infrastructure Engineering's platform, including configuration management, traffic routing and environment controls Own the team's technical direction, contributing to key decisions on API architecture, internal tooling design, and automation frameworks. Nurture and grow engineers at different levels, providing support through coaching, mentorship, and career development. Foster an inclusive, high-performing team culture focused on ownership, learning, and delivery. Partne
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: Join a team that builds robust, real-time distributed systems for a cutting-edge database. We care about performance, reliability, scalability, and most of all learning and having fun together. Whether you’re a seasoned coder or just getting started, if you’re passionate about technology and eager to learn, you’ll fit right in. Who we are: We show up to work, ready to collaborate and build technologies that make a difference, with people who genuinely care. We chase improvements such as tail latencies, bytes throughput, cache hit rate, and operational cost efficiency. We believe learning is ongoing and that even the most complex problems can have simple solutions. What You’ll Do: Collaborate with teammates to design and build database features that power AI applications. Learn how to tune performance and support reliability in distributed systems (don’t worry, we’ll guide you). Help Pinecone run smoothly on popular cloud providers. Take ownership of your work and grow your skills every day. Have fun. Who You Are: 5+ years of work experience - programming in Rust, Go, C++, or a comparable language. You’re genuinely curious about distributed systems and eager to dive deep into technical challenges. You approach problems with creativity and persistence, and you’re comfortable asking thoughtful questions or seeking feedback. You’re excited to learn, value constructive feedback, and appreciate mentorship. Bonus Points: You have hands-on experience with cloud platforms (AWS, GCP, Azure) or have demonstrated an ability to pick u
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Help power the development of Replit Agent as an engineer in the Replit Cloud organization. The Replit Cloud team builds Replit’s first party cloud infrastructure so users can build, scale, and succeed entirely on Replit. They manage databases, application storage, app publishing and hosting, development/production environment splitting, custom domains, and more. By having a set of first party services that integrate seamlessly, you will power one of Replit’s key product differentiators. You will: Work closely with designers and product managers, to quickly iterate on Replit Cloud to continually grow and improve the product. Drive full-stack feature development from conception to deployment, taking ownership of key product initiatives. Contribute to architectural decisions that shape the future of our product. Ship product and build infrastructure as a true full stack builder using: TypeScript, React, CSS, Postgres, Go, and Terraform. Examples of what you could do: Leverage our unique cloud infrastructure to build differentiated full product experiences, helping non-technical or semi-technical users remove roadblocks to success. Leverage AI agents to proactively optimize or suggest app improvements on latency, reliability, SEO, and more. Be part of engineering leadership, steering teams towards the highest impact work and supporting initiatives across the company. Required skills and experience: Bachelor’s degree in Computer Science or related field, OR equivalent real-world experience in engineering roles. Comfortable building with our tech stack: TypeScript, React, Go Preferred Qualifications Experience building user facing platform as a service products. Experience with AI/agentic systems. Previous e
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s mission is to unlock financial freedom for everyone by making money movement and access to financial data simple and secure. As a Software Engineer, you will design and build the systems that power how millions of people connect to their finances. You will work across the stack, from reliable backend services and APIs to intuitive applications that bring those systems to life. You will collaborate with engineers, product managers, and designers to ship products that make financial services more accessible and transparent. At Plaid, engineers take ownership early, grow quickly, and see their work reach millions of users. Responsibilities: Design & Development: Build and maintain backend services with a focus on performance, reliability and scalability. Collaboration: Work closely with product managers and other stakeholders to define and implement new features that meet product and customer needs. Code Quality: Write clean, maintainable and efficient code. Testing & Debugging: Develop automated tests to ensure the quality and reliability of the codebase. Troubleshoot and resolve issues. Engage in hands-on coding and architectural design, setting and maintaining high technical standard
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s mission is to unlock financial freedom for everyone by making money movement and access to financial data simple and secure. As a Software Engineer, you will design and build the systems that power how millions of people connect to their finances. You will work across the stack, from reliable backend services and APIs to intuitive applications that bring those systems to life. You will collaborate with engineers, product managers, and designers to ship products that make financial services more accessible and transparent. At Plaid, engineers take ownership early, grow quickly, and see their work reach millions of users. Responsibilities: Design & Development: Build and maintain backend services with a focus on performance, reliability and scalability. Collaboration: Work closely with product managers and other stakeholders to define and implement new features that meet product and customer needs. Code Quality: Write clean, maintainable and efficient code. Testing & Debugging: Develop automated tests to ensure the quality and reliability of the codebase. Troubleshoot and resolve issues. Engage in hands-on coding and architectural design, setting and maintaining high technical standard
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s mission is to unlock financial freedom for everyone by making money movement and access to financial data simple and secure. As a Software Engineer, you will design and build the systems that power how millions of people connect to their finances. You will work across the stack, from reliable backend services and APIs to intuitive applications that bring those systems to life. You will collaborate with engineers, product managers, and designers to ship products that make financial services more accessible and transparent. At Plaid, engineers take ownership early, grow quickly, and see their work reach millions of users. Responsibilities: Design & Development: Build and maintain backend services with a focus on performance, reliability and scalability. Collaboration: Work closely with product managers and other stakeholders to define and implement new features that meet product and customer needs. Code Quality: Write clean, maintainable and efficient code. Testing & Debugging: Develop automated tests to ensure the quality and reliability of the codebase. Troubleshoot and resolve issues. Engage in hands-on coding and architectural design, setting and maintaining high technical standard
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Team Description: Plaid is evolving into an AI-first company, and Intelligent Tooling sits at the center of that transformation. The Intelligent Tooling team is being built from the ground up, and our mission is to establish the technical foundations, operating model, and internal platforms that embed AI deeply into Plaid’s coding tools, internal systems, and the entire software development lifecycle. When we are successful, engineers across Plaid will delegate lower-leverage work to AI agents, move faster with confidence, and spend more of their time designing and inventing for customers. Intelligent Tooling owns the platforms and systems that make this possible - from AI coding integrations and SDLC agents to the internal tools that power Plaid’s operations. Role Description: As a Staff Software Engineer on the Intelligent Tooling team, you will build and operate internal systems that directly impact how engineers across Plaid do their work, and own the technical direction for major parts of that surface. This is a hands-on role with significant ownership, where success is measured by real adoption, reliability, and improvements to developer experience. You will work on AI-powered tooling, interna
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid Europe is building Plaid in Europe—the open banking platform that makes international expansion one click away for fintechs who want to scale globally. We are the only transatlantic open banking provider, uniquely positioned to help fintechs scale globally with a single Plaid integration. Our team localises, scales, and operates the bank connectivity layer that powers payments, underwriting, onboarding, and identity—abstracting fragmented European rails and regulations into a single, reliable platform. The team owns four core pillars: - Account-to-Account Payments - Cash Flow & Income Insights for Credit - Bank Connectivity Across Europe - Consumer Experience This team operates at the intersection of distributed systems, financial infrastructure, and consumer UX. We solve hard problems in payments reliability, data quality, regulatory complexity, and cross-market expansion—enabling fintechs to launch in new countries with minimal additional engineering effort. We are a high-impact, cross-functional engineering team working closely with US platform teams, GTM, compliance, and operations to make Plaid the default open banking layer for global fintechs operating in Europe. You will serve as t
What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid Europe is building Plaid in Europe—the open banking platform that makes international expansion one click away for fintechs who want to scale globally. We are the only transatlantic open banking provider, uniquely positioned to help fintechs scale globally with a single Plaid integration. Our team localises, scales, and operates the bank connectivity layer that powers payments, underwriting, onboarding, and identity—abstracting fragmented European rails and regulations into a single, reliable platform. The team owns four core pillars: - Account-to-Account Payments - Cash Flow & Income Insights for Credit - Bank Connectivity Across Europe - Consumer Experience This team operates at the intersection of distributed systems, financial infrastructure, and consumer UX. We solve hard problems in payments reliability, data quality, regulatory complexity, and cross-market expansion—enabling fintechs to launch in new countries with minimal additional engineering effort. We are a high-impact, cross-functional engineering team working closely with US platform teams, GTM, compliance, and operations to make Plaid the default open banking layer for global fintechs operating in Europe. You will serve as t
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know the team The Developer Platform team at Auth0 (an Okta company) owns the platform that developers build identity on. Increasingly they build alongside AI agents, and that raises the bar on everything underneath: interfaces have to hold up whether a person or an agent is calling them, and the systems behind them have to stay reliable and coherent as usage grows. We move fast, we own problems end to end, and we care deeply about the platform we put in front of the developers and agents who depend on it. The opportunity We're hiring a Principal Engineer (P5) to serve as the technical leader and compass for the Developer Platform team. You'll work across the breadth of the platform, tackling the highly complex, vaguely specified problems that span it and turning them into clear technical direction the team can execute against, without day-to-day oversight. Above all, you'll own how the platform is architected to scale: the distributed systems behind it, the reliability and consistency guarantees developers depend on, and the coherence that keeps it easy to build on as usage grows. You'll champion the team's technical execution, raise the engineering bar, mentor the people around you, and partner with tech leads across teams to keep the wider platform aligned. You'll have real influence over how our platform holds up in a world where developers and agents are both first-class consumers. What you'll be doing Own the platform architecture: Set the tech
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Streaming Foundations team builds services and operates data pipeline infrastructure to support event streaming, messaging, and analytics use cases. We are looking for a Software Engineer who is passionate about distributed systems, platform engineering, and solving data-intensive problems at scale. In this high-impact role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Auth0 to scale for years to come. What you’ll be doing Help set the technical direction for the team and influence the engineering roadmap for the Platform’s streaming capabilities Design and lead the implementation of our most complex and critical systems for data-intensive use cases. Research and champion new technologies and architectural patterns to solve strategic challenges and scale the platform. Lead and influence cross-functional initiatives, ensuring technical alignment and successful execution across multiple teams. Improve the operational posture of our systems by designing for observability, reliability, and scalability, and by mentoring others in operational best practices. Coach and mentor senior engineers and act as a technical leader across the engineering organization. Collaborate with different stakeholders like product teams whenever needed. What you’ll bring to our teams 7+ years of software development experience in a fast-paced, agile environment Experience working with Golang or Java is preferred H
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Load and Fault team sits within Robinhood's Developer Infrastructure organization, with a mission to give every engineering team the tools they need to test their services under real-world conditions before those conditions test them in production. We build the platforms and frameworks that enable load testing, fault injection, and resilience validation at scale — treating reliability as a developer productivity problem, not just an operations one. Our work directly raises the quality bar for every service Robinhood ships, and we partner closely with engineering teams across the organization to make resilience testing a seamless part of the development workflow. As a Senior Software Engineer on the Load and Fault Environments team, you will design and build the infrastructure that lets Robinhood's engineers simulate
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime