Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Load and Fault team sits within Robinhood's Developer Infrastructure organization, with a mission to give every engineering team the tools they need to test their services under real-world conditions before those conditions test them in production. We build the platforms and frameworks that enable load testing, fault injection, and resilience validation at scale — treating reliability as a developer productivity problem, not just an operations one. Our work directly raises the quality bar for every service Robinhood ships, and we partner closely with engineering teams across the organization to make resilience testing a seamless part of the development workflow. As a Senior Software Engineer on the Load and Fault Environments team, you will design and build the infrastructure that lets Robinhood's engineers simulate load, inject faults, and validate system behavior under stress — at the scale of a fast-growing financial platform. You'll own meaningful components of the load testing and fault injection platform, write production-quality code, and collaborate with engineers across infrastructure and product teams to ensure the tooling you build gets adopted and drives real
Jobiba hiring network
Reliability Engineer Jobs
2,028 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Auth0 is an easy-to-implement, authentication and authorization platform designed by developers for developers. We make applications’ login boxes safe, secure, and seamless for anyone logging in. The Auth0 Data Engineering Team Within Auth0, the Data Engineering team builds solutions to support analytics needs for the whole organization and is in charge of the Data Platform. It is divided into 3 groups: Pipeline team , sitting closer to the Platform team and data producers and in charge of the efficient data ingestion and provision of quick access to unmodeled data Warehouse team sitting closer to the business teams and data consumers and in charge of modeling the data in the data warehouse to abstract away the complexity for data consumers, simplifying the organization’s analysis, reporting and decision-making Interface team responsible for creating and managing connections between the data platform and various external systems (internal or external facing), making sure everyone has access to consistent and reliable data The Senior Data Engineer Opportunity Reporting to the manager of the Data Engineering Pipeline team, the senior data engineer will be a key player in ensuring the reliability and efficiency of our core data platform. This role is crucial, providing the stable data foundation that not only 'runs the business' day-to-day but also directly empowers the organization to unlock growth and build innovative new products. We are looking for an auto
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta is The World’s Identity Company. We free everyone to safely use any technology anywhere, on any device or app. Our Workforce and Customer Identity Clouds enable secure yet flexible access, authentication, and automation that transforms how people move through the digital world, putting Identity at the heart of business security and growth. At Okta, we celebratea a variety of perspectives and experiences. We are not looking for someone who checks every single box—we’re looking for lifelong learners and people who can make us better with their unique experiences. We are seeking an independent Staff Software Engineer who effectively balances technical excellence with a disciplined approach to the software development lifecycle. What you'll be doing Design, build, and evolve intelligent customer experience solutions that integrate Salesforce CRM with AWS services and modern generative AI capabilities. Architect and develop AI-powered experiences using Amazon Bedrock, LangChain, and modern LLM integration patterns to create context-aware, scalable, and customer-facing solutions. Build robust backend services and integrations to support customer workflows, including LLM orchestration, prompt workflows, RAG pipelines, and other generative AI features. Lead the design and delivery of complex customer experience initiatives from technical discovery through production deployment, ensuring scalability, reliability, and maintainability. Develop and refine prompt e
About the Team OpenAI’s Applied AI Engineering team helps organizations turn frontier AI capabilities into safe, reliable, and high-impact production systems. We work with customer executives, product and engineering teams, security leaders, and transformation teams to identify valuable opportunities, accelerate technical implementation, and scale what works. Enterprise deployments are defined by complexity rather than any one industry: existing architectures, diverse data environments, security and governance requirements, multiple stakeholder groups, and organization-wide change. We turn lessons from these deployments into better products and reusable patterns for customers everywhere. About the Role As an Applied AI Engineer you will partner directly with leading organizations to design, build, and deploy AI systems that deliver measurable business outcomes. You will combine deep technical judgment, hands-on engineering, and customer leadership to take ambitious ideas from use-case selection and architecture through prototyping, evaluation, production launch, and scale. You will write and debug code, build evaluation systems, resolve complex integrations, and guide decisions involving model behavior, reliability, latency, cost, safety, security, governance, and operational readiness. Success is measured by production systems, sustained adoption, and meaningful customer impact—not simply activity or successful demonstrations. This is a rare opportunity to work on consequential real-world deployments at the frontier of AI while directly influencing how OpenAI’s products evolve. This role is based in Tokyo, Japan. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Partner directly with enterprise customers to identify high-value opportunities and translate them into technical architectures, implementation plans, evaluation strategies, and measurable success criteria. Design, build, an
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse using open formats like Apache Iceberg — at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world's leading data platforms. We are hiring a Senior Software Engineer for Observe by Snowflake on the Data Management team. This team is responsible for the tables, views, and materialized views at the core of Observe's architecture. Observe's data lake approach lets customers correlate heterogeneous telemetry — logs, metrics, traces, events — across a unified data model. This role owns that data model: how customers define, shape, and query the semi-structured data that makes cross-si
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Senior Software Engineer to own key components of our AI native External Observability Platform . In this role, you will contribute to the technical road map for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex technical projects, and skilled at collaborating with the brightest technical minds in the industry. Key Responsibilities Develop and Scale Distributed Infrastructure: Design and implement key components of Snowf
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. We are hiring a Senior Frontend Engineer, AI Products team at Observe by Snowflake. As an AI Product engineer you'll always be thinking first about the user experience and how to create the best product, technical choices, and implementation decisions that stem from that product first thinking. This team builds the AI-powered products and developer tooling at the core of Observe's platform, including our flagship AI SRE product, real-tim
About the Team When 5% of Indian households shop with us, it’s important to build resilient systems to manage millions of orders every day. We’ve done this – with zero downtime!Sounds impossible? Well, that’s the kind of Engineering muscle that has helped Meesho become the e-commerce giant that it is today. We value speed over perfection, and see failures as opportunities to become better. We’ve taken steps to inculcate a strong ‘Founder’s Mindset’ across our engineering teams, making us grow and move fast. We place special emphasis on the continuous growth of each team member - and we do this with regular 1-1s and open communication. As Engineering Manager, you will be part of self-starters who thrive on teamwork and constructive feedback. We know how to party as hard as we work! If we aren’t building unparalleled tech solutions, you can find us debating the plot points of our favourite books and games – or even gossiping over chai. So, if a day filled with building impactful solutions with a fun team sounds appealing to you, join us. About the Role We are looking for an Engineering Manager to lead the Platform Team at Meesho.In this role, you will be responsible for building and growing a high-impact platform engineering team, owning the technical roadmap, and delivering a unified database platform used by hundreds of engineers across Meesho. This is not a pure operations role rather a platform engineering leadership role that blends system design, automation, developer experience, and reliability. You will work closely with Backend, Data, Infra, and SRE teams to ensure Meesho’s database ecosystem scales reliably while remaining easy to use and cost-efficient.
About Bazaarvoice At Bazaarvoice, we create smart shopping experiences. Through our expansive global network, product-passionate community & enterprise technology, we connect thousands of brands and retailers with billions of consumers. Our solutions enable brands to connect with consumers and collect valuable user-generated content, at an unprecedented scale. This content achieves global reach by leveraging our extensive and ever-expanding retail, social & search syndication network. And we make it easy for brands & retailers to gain valuable business insights from real-time consumer feedback with intuitive tools and dashboards. The result is smarter shopping: loyal customers, increased sales, and improved products. The problem we are trying to solve : Brands and retailers struggle to make real connections with consumers. It's a challenge to deliver trustworthy and inspiring content in the moments that matter most during the discovery and purchase cycle. The result? Time and money spent on content that doesn't attract new consumers, convert them, or earn their long-term loyalty. Our brand promise : closing the gap between brands and consumers. Founded in 2005, Bazaarvoice is headquartered in Austin, Texas with offices in North America, Europe, Asia and Australia. It’s official: Bazaarvoice is a Great Place to Work in the US , Australia, India, Lithuania, France, Germany and the UK! What you’ll be doing Lead, hire and grow a high-calibre team of frontend & backend engineers, and their line managers. Define and implement a roadmap based on critical business need, that delivers the valuable features, scale, and reliability our clients need, as well making systems more resilient and scalable. Drive business-significant and complex initiatives by collaborating across geographically distributed teams and partners Coach and mentor engineers globally in support of their growth and adherence to best practices. Who you are – Requirements for success in this rol
About BlockTech BlockTech is a fast-paced algorithmic trading firm facilitating global cryptocurrency derivatives and spot trading while expanding into new markets. As we continue to grow rapidly, we are looking for a Senior Software Engineer to take technical ownership of our Core systems and help shape how they evolve through our scale-up phase. This is a role for an engineer who has already built and operated correctness-critical systems at scale and who wants their technical decisions to set direction, not just follow it. What will you do? You will own the systems that keep our trading operation accurate at every level — positions, PnL, instruments, hedging, and the data flowing between them. Across exchanges and asset classes, your work determines whether the numbers our traders and risk team rely on are correct. As a senior engineer, you will also set the technical standards the rest of the team builds against. You will: Own the books. Maintain and evolve our position, trade, and PnL systems so they remain accurate across every venue. Architect the platform that supports trading. Design services that handle real-time exchange flow with the throughput and reliability our quoting and hedging depend on, making the structural calls that keep them fast and correct under load. Own the data flows. Operate and harden the public and private pipelines bringing market prices, account state, and trade events into our systems in real time. Lead venue integrations end-to-end. Deliver exchange and prime-broker integrations across protocol, account model, settlement, reconciliation, and surveillance and define the patterns we reuse for the next one. Model the instrument universe. Own the systems describing every tradable instrument - lifecycle, expiry, margin, RFQs, fast-path quoting - where both correctness and latency matter. Build self-healing systems. Design the detection and repair flows that identify position discrepancies and correct them automatically. Own hedging inf
As a Staff Software Engineer on Coder’s Agentic Engineering team, you’ll shape the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on while setting the team's technical direction. You’ll lead complex work, make sound architectural decisions, and help other engineers do their best work. What you’ll do here Set technical direction across Coder’s agent harness, integrations, and workflows. Design and build production systems in Go, with work across React and TypeScript where needed. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Lead complex projects from early ambiguity through production. Raise the engineering bar through design reviews, code reviews, and technical mentorship. Partner with Product and Design on clear, useful agent experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Deep experience building and operating production software systems. Strong hands-on experience with Go. Experience with React and TypeScript. Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS. A track record of setting technical direction without formal authority. Strong architectural judgment and comfort working through ambiguity. Someone who makes the engineers around them better. Our tech stack Backend: Go, Postgres Frontend: TypeScript, React Infrastructure: AWS, Kubernetes Observability: Prometheus, Grafana CI/CD: GitHub Actions Bonus tacos if you have (Tacos? If you need an ice-breaker, ask how we say thanks by giving tacos!) Experience building coding agents, developer tools, or cloud development environm
About the Team OpenAI’s Applied AI Engineering team helps organizations turn frontier AI capabilities into safe, reliable, and high-impact production systems. We work with customer executives, product and engineering teams, security leaders, and transformation teams to identify valuable opportunities, accelerate technical implementation, and scale what works. Enterprise deployments are defined by complexity rather than any one industry: existing architectures, diverse data environments, security and governance requirements, multiple stakeholder groups, and organization-wide change. We turn lessons from these deployments into better products and reusable patterns for customers everywhere. About the Role As an Applied AI Engineer you will partner directly with leading organizations to design, build, and deploy AI systems that deliver measurable business outcomes. You will combine deep technical judgment, hands-on engineering, and customer leadership to take ambitious ideas from use-case selection and architecture through prototyping, evaluation, production launch, and scale. You will write and debug code, build evaluation systems, resolve complex integrations, and guide decisions involving model behavior, reliability, latency, cost, safety, security, governance, and operational readiness. Success is measured by production systems, sustained adoption, and meaningful customer impact—not simply activity or successful demonstrations. This is a rare opportunity to work on consequential real-world deployments at the frontier of AI while directly influencing how OpenAI’s products evolve. This role is based in London. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees In this role, you will: Partner directly with enterprise customers to identify high-value opportunities and translate them into technical architectures, implementation plans, evaluation strategies, and measurable success criteria. Design, build, and deplo
About the Team OpenAI’s Applied AI Engineering team helps organizations turn frontier AI capabilities into safe, reliable, and high-impact production systems. We work with customer executives, product and engineering teams, security leaders, and transformation teams to identify valuable opportunities, accelerate technical implementation, and scale what works. Enterprise deployments are defined by complexity rather than any one industry: existing architectures, diverse data environments, security and governance requirements, multiple stakeholder groups, and organization-wide change. We turn lessons from these deployments into better products and reusable patterns for customers everywhere. About the Role As an Applied AI Engineer you will partner directly with leading organizations to design, build, and deploy AI systems that deliver measurable business outcomes. You will combine deep technical judgment, hands-on engineering, and customer leadership to take ambitious ideas from use-case selection and architecture through prototyping, evaluation, production launch, and scale. You will write and debug code, build evaluation systems, resolve complex integrations, and guide decisions involving model behavior, reliability, latency, cost, safety, security, governance, and operational readiness. Success is measured by production systems, sustained adoption, and meaningful customer impact—not simply activity or successful demonstrations. This is a rare opportunity to work on consequential real-world deployments at the frontier of AI while directly influencing how OpenAI’s products evolve. This role is based in Abu Dhabi, UAE. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees In this role, you will: Partner directly with enterprise and digital native customers to identify high-value opportunities and translate them into technical architectures, implementation plans, evaluation strategies, and measurable success criteri
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Senior Software Engineer, Simple Trade Experience We're hiring a Senior Software Engineer to join the Simple Trade Experience team within the Consumer & Business group. This team powers the critical buy, sell, and convert journeys on the Simple interface of the Coinbase app and website, supporting multiple trading types and assets for millions of users worldwide. You'll own highly performant, available, and consistent backend and frontend systems that directly enable customers to trade with confidence at web-scale. What you'll do: Own the design and delivery of highly performant, available, and consistent systems powering trading for millions of users worldwide Architect robust, extensible systems that support multiple trading types and assets (Spot, Limit, Recurring, and more) while maintaining reliability at web-scale Partner closely with product and design to define and ship a best-in-class trading experience across web and mobile Drive production service evolution, scaling and maintaining critical trading infrastructure as traffic and product scope grow Strengthen engineering quality across the team through rigorous code reviews, mentorship, and knowledge sharing Required Skills and Experience: 5+ years of experience in software engineering with demonstrated success building and scaling production services handling web-scale traffic Proven track record design
Biohub is the first large-scale initiative bringing frontier AI models, massive compute, and frontier experimental capabilities under one roof. We're building a general-purpose system to accelerate scientific discovery, integrating frontier AI models, biological foundation models, and lab capabilities, with the ultimate goal of curing disease. Our technology powers scientists around the world, translating AI capabilities into tools that accelerate research everywhere. The Team The AI Cluster Production Engineering team is part of the AI Compute Platform organization at Biohub, a non-profit research lab committed to open science and open-source AI. We own the design, operation, and reliability of large-scale multi-GPU AI clusters that power frontier AI biology research: protein language models, genomic foundation models, and scientific reasoning systems built to be shared, not monetized. Our clusters run Slurm on Kubernetes infrastructure and support everything from day-to-day AI researcher workflows to multi-node hero training runs at thousands of GPUs. The team works at the intersection of AI tooling, distributed systems, HPC, and frontier AI, debugging deep AI infrastructure problems and building AI systems critical to the entire AI organization. The Opportunity CZ Biohub's mission is to cure or prevent all human disease. Achieving that requires training frontier-scale AI biology models, and that demands reliable, high-performance compute infrastructure. This is production engineering work at a frontier AI lab, with the twist that the mission is biology and the science is open. You'll keep GPU clusters running at high utilization, debug the toughest distributed systems failures, and build the operational foundations for scaling to multi-thousand GPU hero runs. The technical problems are genuinely hard (e.g., multi-node distributed training, InfiniBand fabrics, large-scale storage, Slurm at scale) inside an organization where the work is aimed at helping peop
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime