Jobiba hiring network

Reliability Engineer Jobs

2,028 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

S
Stripe
📍 Nyc Privy• Full-time
1mo ago

Who we are About Privy Our mission is to make privacy and user ownership the default online. To do so, we build simple, flexible APIs and tools for developers that make it easy to build new products on crypto rails. Privy owns the abstractions and infrastructure layer above wallets, integrating across chains, third-party providers, and Stripe products like Treasury and Link. We get to solve hard technical problems while leveraging Stripe's distribution to reach customers like Ramp, Klarna, Deel, Kraken, Hyperliquid, and Fomo — powering experiences for both mainstream users and crypto natives. Learn more about Privy: Privy and Stripe: Bringing crypto to everyone About the team Engineering at Privy is distinguished by: High urgency: Shipping very small iterations, very fast, to learn very quickly. Product taste: Our customers are developers, and to build effective products for them requires technical knowledge - you will often be "the PM". Security mindset: A great portion of our product is trust. While we have a dedicated security team, every engineer brings security to their designs from the start. In practice, we use boring technology like Node, React, and AWS so we can focus our engineering energy entirely on pushing the boundaries of Privy's core product, e.g. through hardware enclaves, multi-region low latency APIs, and blockchain abstractions that are accessible to mainstream developers. What you'll do Design and build the backend systems that power wallets, identity, and onchain infrastructure at scale Create platform primitives and APIs that enable teams across Privy and Stripe to build faster Lead complex technical initiatives across architecture, data, and distributed systems Improve the scalability, reliability, and performance of our core platform Help shape our technical direction through high-leverage engineering work Who you are Minimum requirements 8+ years of experience building and maintaining a production system at scale An understandin

reactawsai
View job →
S
1mo ago

Note: if you are an intern, new grad, or staff applicant, please do not apply using this link and visit our jobs page for those specific postings. Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the Organization Secure Frameworks: We provide security guarantees for common threats, helping teams secure their services by default. We build and maintain core application-layer security libraries and frameworks for development teams, including web security frameworks, cryptography libraries, and other threat mitigations across the Stripe technology stack. More recently, our scope has expanded to protective measures for AI agents. We are a hands-on engineering team that builds security protections and frameworks, rather than conducting security reviews. Core Infrastructure : We’re the home for Stripe's critical tier0 infrastructure systems (Compute, Networking, DocumentDB, Distributed Caching and High assurance engineering). We build the foundational platform for Stripe products and services to allow them to operate at scale. We drive reliability, availability, efficiency and scalability of these systems. Developer Infrastructure : We’re responsible for the productivity of all developers at Stripe. Ensure Stripe’s engineers have a reliable, fast, and easy-to-use inner dev loop to maximize productivity while building everything from low-latency microservices to large-scale data pipelines and machine learning models. Reliability Insights and Excellence : We build tools and frameworks

restmicroservicesmachine learning
View job →
S
Stripe
📍 Seattle• Full-time• From $156.8K/yr
1mo ago

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Design, develop, maintain, and support application programming interfaces (APIs), backend services, and distributed systems using Ruby, Java, Go and Scala. Build, test, deploy, and maintain software infrastructure to support large-scale production systems. Design and develop software systems using analytical techniques and mathematical models to evaluate system performance, predict outcomes, and assess design tradeoffs. Develop and execute software testing, validation, debugging, and documentation procedures to ensure system reliability and performance. Diagnose and resolve production issues across multiple services and layers of the technology stack. Analyze user needs and software requirements to determine technical feasibility, design approaches, and implementation timelines within cost and resource constraints. Collaborate with cross-functional engineering teams to design and implement new features for large-scale systems. Design and build systems to securely store, manage, and modify production data and services. Improve engineering standards, development tools, and software development processes to enhance system quality and efficiency. Who you are Minimum requirements Must have a Bachelor's degree or foreign equivalent in Computer Science, Software Engineering, or a related field, plus 2 years of software development experience. Must also have 2 years of experi

javaairuby
View job →
S
Stripe
📍 South San Francisco• Full-time• $156.8K – $235.2K/yr
1mo ago

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Architect, build, and maintain robust APIs, services, and systems across engineering teams using Ruby and Java; Act as a domain expert of the payments family of Stripe products, and in particular various payout methods and underlying rails used in US / UK; Build and enhance software infrastructure, encompassing the complete development lifecycle from coding to testing, deployment, debugging, and maintenance; Design software systems leveraging scientific analysis and mathematical models to predict and measure design outcomes; Contribute to core interface design and write code; Serve as a role model for how great software should be written; Contribute code changes in every layer of the stack, from launching a net new payout method API to making a UI change in the Stripe Dashboard for the Stripe Connect product; Scope and lead technical projects, laying the groundwork for early-stage products to iteratively evolve and scale; Engineer payments integration with banking partners to launch new markets, payout methods, and capabilities, while ensuring reliability and scalability; Work with engineers across the company to build new features at large-scale; Design and implement secure systems for sensitive data storage; Work with engineers across the company to understand when existing infrastructure can be leveraged vs. when building a bespoke solution is necessary

typescriptpythonjava
View job →
S
Stripe
📍 Seattle• Full-time• From $156.8K/yr
1mo ago

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Design, build, and maintain APIs, services, and systems across Stripe’s engineering teams using Java, Ruby, Scala, and Go Build software infrastructure, including developing, testing, and deploying it Design and develop software systems, using scientific analysis and mathematical models to predict and measure outcome and consequences of design Engineer payments integration with various financial partners software systems Design APIs and underlying data models to support complex financial abstractions and multi-party integrations, enabling flexible billing and settlement configurations across international markets Develop and direct software system testing and validation procedures, programming, and documentation Debug production issues across services and multiple levels of the stack Analyze user needs and software requirements to determine the feasibility of design within time and cost constraints Work with engineers across the company to build new features at large-scale Build new systems to securely store sensitive data; Improve engineering standards, tooling, and processes Integrate observability tools and alerting mechanisms (Datadog, Prometheus) into high-traffic production systems, defining SLAs and real-time diagnostics to meet the reliability standards required by financial partners Mentor junior engineers, lead technical design reviews, and establish best practices in code quality, system design, re

pythonjavasql
View job →
S
Stripe
📍 San Francisco Or New York• Full-time
1mo ago

Who we are About Bridge We're creating an entirely new payments platform, built with stablecoins, to simplify global money movement. Bridge enables faster, cheaper payments and borderless access to dollars via stablecoins. Through our APIs, businesses can send and receive funds across borders faster and cheaper vs. SWIFT and other fiat-only rails. Our virtual accounts enable international consumers and businesses to easily access, store, and spend US dollars. Our payouts infrastructure enables platforms to disburse USD to anyone globally. We believe many trillions of dollars will move and settle through stablecoin payment rails. Bridge is pulling this future forward. We have a small team of people who have previously built financial infrastructure at some of the world's leading companies (Coinbase, Stripe, Square, Brex, Upstart, DoorDash, Airbnb), and each and every one of them chose Bridge because they fundamentally believe that stablecoins will be a critical piece of financial infrastructure that allows for the improvement of global money movement. About the Role As one of the first blockchain engineers at Bridge, you'll have the autonomy to work on projects that are truly global in scale and aim to give customers access to products they've never had access to. Many of the use cases that we work on today didn't exist several months ago, and that's because our team is dedicated to constant improvement for our partners and end users. Our engineers have outsized ownership over projects, so if you're looking for increased autonomy working on a completely greenfield opportunity, Bridge is the place for you. Responsibilities • Build out multi-chain hot wallet infrastructure to support increased reliability of our payments. • Build smart contracts to increase hot wallet payment throughput. • Build smart contract platform to issue custom stablecoins across multiple blockchains (EVM, Solana, Stellar, etc.). • Build smart contract monitoring systems to detect anomalous even

aiswiftgo
View job →
S
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Tax compliance is a hard problem. When millions of payments flow through a marketplace or platform, the rules for what gets reported, to whom, in which format, and by when are complex, high-stakes, and vary by country. The Connect Tax Reporting team builds products to help businesses collect and verify tax information, prepare and deliver tax forms, and manage reporting obligations across jurisdictions. This includes 1099 reporting in the US and digital platform reporting across the EU, Australia, Canada, the UK, and other regions. This is a domain where correctness, reliability, and user trust matter deeply. A late filing, an incorrect form, or a missed reporting threshold can create regulatory risk and erode trust with platforms and their sellers. Our systems handle high-volume seasonal workflows, support complex country-specific rules, and remain dependable under fixed filing deadlines. What you’ll do You’ll build products and systems that simplify tax reporting for platforms and their users. You’ll take projects from discovery and technical design through implementation, launch, and operation. Your work may include improving tax information collection, building APIs and data pipelines, automating reporting processes, and strengthening the reliability of production systems. You’ll collaborate with product managers, tax operations, support, and engineers across Stripe. AI-assisted development will be part of your day-to-day work as you des

typescriptpythonjava
View job →
S
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Big Data Infrastructure operates the critical infrastructure that powers the batch data processing at Stripe. The team supports a variety of use cases, including Payment, Ledger, ML, Fraud Detection, Product Analytics, Regulatory Reporting, Financial Data Reconciliation, and externally facing products like Radar and Sigma. As an example of the scale, the team's systems serve hundreds of teams, thousands of workflows, 100,000+ task executions, O(billion) transformations, and moving terabytes of data processing over 1 GB/second every day. Our users inside Stripe include other engineering teams, Data Scientists, Sales and Operations, Finance, etc. Data Orchestration builds and operates the time-based and event-based orchestration infrastructure that powers and accelerates batch data pipelines. The team operates on a wide range of tech stacks including Airflow, Spark, SQL, Kafka, Flink, Hive MetaStore, Trino, Pinot, Python, Java, Scala, S3, and Iceberg. What you'll do As a Software Engineer on this team, you'll design and build infrastructure that powers batch data processing at Stripe. Responsibilities Design, build, and maintain next-generation and first-generation versions of key Data Platform products, with an emphasis on usability, reliability, security, and efficiency. Design ergonomic APIs and abstractions that build a great customer experience for internal Stripes, that will in turn enhance the experience of millions of Stripe users.

pythonjavasql
View job →
S
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Stripe Infrastructure is responsible for the reliability, scale, performance, and cost of Stripe's systems and the productivity and sentiment of Stripe's people. You may work on a wide variety of critical business areas within Core Infrastructure. We're the home for Stripe's critical tier-0 infrastructure systems (Compute, Networking, DocumentDB, Distributed Caching, and High Assurance Engineering). We build the foundational platform for Stripe products and services to allow them to operate at scale. We drive reliability, availability, efficiency, and scalability of these systems. What you'll do As an Engineering Manager for the Core Performance team at Stripe, you'll lead a team responsible for driving improvements in efficiency and latency across Stripe. Responsibilities • Lead and grow an engineering team in Sydney—hire, coach, mentor, and develop engineers at all levels • Partner with engineering teams across Stripe to identify and implement infrastructure solutions that enhance scalability, reliability, and efficiency • Create team vision, define goals, prioritize work, and deliver high-quality outcomes • Cultivate a culture of engineering excellence, innovation, and continuous improvement • Collaborate with staff engineers and technical leadership on architecture and strategic technical decisions, translating technical strategy into team-level plans • Work effectively with geographically distributed colleagues and remote workers to ens

javaawslinux
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team API Agents builds the shared agent harness, tools, and infrastructure that turn OpenAI’s frontier models into systems that can reliably complete real work. We carry the capabilities behind Codex into a much broader set of products and workflows across software engineering, research, finance, healthcare, enterprise operations, and more. Our work spans search and connected context, computer use, memory, delegation and multi-agent coordination, and safe execution. Sitting at the intersection of Research, Codex, infrastructure, and applied product teams, we build reusable agent capabilities that compound across the ecosystem. About the Role We are looking for an experienced backend software engineer to build the core systems behind the next generation of agents. You will design reliable services and abstractions that help agents find the right context, use tools and computers, retain knowledge, coordinate over long-running workflows, and take action safely. The role combines deep backend and infrastructure work with strong product judgment, with opportunities to work across agent runtimes, orchestration, search, execution environments, identity and permissions, observability, and evaluations. This is software and systems engineering rather than model training: success comes from strong backend fundamentals, high agency, and the ability to turn fast-moving research capabilities into dependable production primitives. In this role, you will: Design, build, and operate the shared agent harness and backend infrastructure that power long-running, high-value workflows across OpenAI and third-party products. Build reusable capabilities across search and connected context, computer use, memory, tool execution, delegation, subagents, and multi-agent orchestration. Establish the foundations agents need to operate safely in production, including secure execution environments, identity and permissions, observability, evaluations, reliability, and cost and latency effi

typescriptpythonaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Data Platform at OpenAI owns the foundational data stack powering critical product, research, and analytics workflows. We operate some of the largest Spark compute fleets in production; design, and build data lakes and metadata systems on Iceberg and Delta with a vision toward exabyte-scale architecture; run high throughput streaming platforms on Kafka and Flink; provide orchestration with Airflow; and support ML feature engineering tooling such as Chronon. Our mission is to deliver reliable, secure, and efficient data access at scale and accelerate intelligent, AI assisted data workflows. Join us to build and operate these core platforms that underpin OpenAI products, research, and analytics. We’re not just scaling infrastructure – we’re redefining how people interact with data. Our vision includes intelligent interfaces and AI-assisted workflows that make working with data faster, more reliable, and more intuitive. About the Role This role focuses on building and operating data infrastructure that supports massive compute fleets and storage systems, designed for high performance and scalability. You’ll help design, build, and operate the next generation of data infrastructure at OpenAI. You will scale and harden big data compute and storage platforms, build and support high-throughput streaming systems, build and operate low latency data ingestions, enable secure and governed data access for ML and analytics, and design for reliability and performance at extreme scale. You will take full lifecycle ownership: architecture, implementation, production operations, and on-call participation. You’ve supported Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms. You’re well-versed in infrastructure tooling like Terraform, experienced in debugging large-scale distributed systems, and excited about solving data infrastructure problems in the AI space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per wee

awsrestmachine learning
View job →
O
1mo ago

About the Team The Applied AI team safely brings OpenAI's technology to the world. We released ChatGPT, Plugins, DALL·E, and the APIs for GPT-4, GPT-3, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate horizon. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. We serve end-users directly through ChatGPT, and serve developers through our APIs, which power product features that were never before possible. About the Role The Engineering Acceleration team designs, builds and maintains the foundational systems that engineers use to build ChatGPT and the API. This is a fast-growing team and you will get a chance to own and define the strategy, vision, and plan for how to increase developer productivity. In this role, you will: Drive the design, development, and implementation of tools, systems, and processes that accelerate engineering velocity, reduce manual effort, and increase the quality of output. Use our latest AI tools to re-think how we can be the most productive team in the industry. Work closely with various teams within OpenAI to understand their workflows, challenges, and needs, and ensure the tools and systems built by the Engineering Acceleration team address these requirements. Bring new features and research capabilities to the world by partnering with product engineers to lay the necessary technical foundations. Guide and advise product engineering teams on best practices for ensuring observable, scalable systems. Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years of experience in engineering, including 3+ years of experience in infrastructure building tooling for developers. Have experi

pythonawskubernetes
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Post-Training Frontiers team is responsible for training the frontier agents OpenAI ships to the world (GPT-Next). We train the flagship agentic models behind Codex, ChatGPT, and the API through large-scale reinforcement learning. The team’s work spans four areas. First, execution and science: working with teams across OpenAI to decide what can go into the final model and how, using scientific experiments and evals that are representative of the final pipeline so issues can be recognized early. Second, RL scaling: executing the final large-scale reinforcement learning run, making sure GPUs are used efficiently and training stays healthy. Third, research: improving horizontal capabilities like instruction following, factuality, memory, and multi-agent behavior, where the team’s broad visibility helps identify cross-cutting improvements across teams and domains. Fourth, engineering: maintaining the infrastructure stack and internal tools to ensure that both the final run and all integrations go as smoothly as possible and that the systems are easy to work with. About the Role This role focuses on keeping our frontier RL training runs fast, reliable, and unblocked. You will work across engineering and infrastructure problems as they emerge, from scaling and orchestration issues to inference bottlenecks, numerical problems, and hardware failures, as well as supporting large horizontal integrations in the big run, like multi-agent capabilities or memory. This is a role for a strong generalist who quickly learns anything needed for the task, has high attention to detail, debugs deeply, and is motivated by fixing the highest-impact problem in front of the team. In this role, you will: Keep large-scale async RL training runs moving by jumping into the most urgent engineering and infrastructure problems. Debug issues across training systems, inference, orchestration, scaling, and distributed infrastructure. Improve the reliability and efficiency of RL trai

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Private Computing team works across product, engineering, security, and safety to build advanced privacy products and infrastructure at OpenAI. Our mission is to provide world-class security features to users so their private data remains private, even from OpenAI. We use technologies like confidential computing, trusted execution environments, and end-to-end encryption to ship product features across ChatGPT, the API, and our future consumer devices. About the Role We’re looking for software engineers to design, build, and scale novel privacy features and infrastructure across ChatGPT, API, and future consumer devices. In this role, you will: Ship fast while balancing difficult trade-offs in complex domains Build core abstractions for trusted execution environments and end-to-end-encryption Build product features for private inference and storage across ChatGPT, API, and future consumer devices Update build systems to increase trust and verifiability Integrate with safety and integrity infrastructure Operate systems at scale with high reliability, including an on-call rotation Collaborate with a diverse set of cross-functional teams across product, engineering, security, safety, policy, and legal You might thrive in this role if you: Care deeply about user privacy and security Have 5+ years of experience in professional software engineering Have experience building and scaling confidential computing or encryption technologies in production environments Have experience with Kubernetes and cloud orchestration systems Take pride in building and operating scalable, reliable, secure systems Can collaborate well and drive alignment in the face of difficult trade-offs Are comfortable with ambiguity and rapid change Workplace & Location This role is based in San Francisco, CA. We follow a hybrid model with 4 days a week in the office and offer relocation assistance to new employees. About OpenAI OpenAI is an AI research and deployment company dedicat

awskubernetesrest
View job →
O
1mo ago

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role We are seeking a Software Engineer, Security Observability to join our Security team. In this role, you will be responsible for building secure, scalable systems that enhance our security observability infrastructure. Leveraging your strong engineering skills, you will collaborate with cross-functional teams to develop, deploy, and maintain robust software solutions that support our security and detection capabilities. This role is open to remote employees, or relocation assistance is available to one of our OpenAI offices in San Francisco, Seattle, or New York City. Due to requirements associated with work this role may support, applicants for this position must be U.S. citizens. In this role, you will: Design and develop scalable software systems that facilitate security observability across our infrastructure. Build and maintain data pipelines that centralize and store security-relevant data from diverse sources. Proactively improve the resilience and reliability of data systems to ensure high platform availability Collaborate closely with Detection & Response (D&R) and other security teams to reduce the company’s security risk. Contribute to data engineering in support of forensic investigations and compliance efforts. You might thrive in this role if you have: Strong software engineering experience, with proficiency in programming languages such as Python, Golang, or similar. A background in infrastructure as code, with exp

pythonawsazure
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.