Jobiba hiring network

Reliability Engineer Jobs

2,028 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

V
Vanta
📍 United States• Full-time
1mo ago

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta's Core Platform team provides the foundational infrastructure that powers all engineering at Vanta. We're expanding upmarket to support enterprise customers, which requires strategic investment in platform systems that ensure security, reliability, and developer productivity at scale. As we expand upmarket to support enterprise and regulated customers, we’re investing heavily in platform capabilities that scale securely while reducing cognitive load for product teams. As the Engineering Manager, Core Platform at Vanta, you'll own the foundational infrastructure that every engineer builds on, ensuring it scales with company growth while remaining fast, simple, and reliable. This team’s ownership spans shared services infrastructure, observability and monitoring, datastore management, and async work systems. Our Engineering Managers develop and grow high-performing teams that deliver significant value to our customers and enable our business to scale. This role sits at the intersection of technical architecture and team development, with real authority to set direction and grow a world-class platform team. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll do as an Engineering Manager at Vanta: Lead and grow high-performing platform engineerin

mongodbawsrest
View job →

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. The Vendor Monitoring Data team focuses on gathering external data and conducting risk analysis as part of Vanta's Vendor Risk Management (VRM) product. Our work provides comprehensive insights that help customers mitigate third-party risks effectively. As a Senior Fullstack Engineer, you'll drive complex projects across our technical stack while mentoring our talented engineering team. This role offers a unique opportunity to delve into a hyper-focused subject area: external attack surface scanning. You'll tackle unique technical challenges and contribute directly to Vanta's impact by helping customers continuously and comprehensively monitor risks across their vendor supply chain. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We're growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and contribute to a rapidly-scaling company. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll do as a Senior Fullstack Engineer, Vendor risk management at Vanta: Identify, scope, and lead large technical projects, laying the groundwork for core products to evolve and scale into highly performant, reliable, and customizable systems Make effective tradeoffs that consider business priorities, user experience, and a sustainable technical foundation Engineer sophisticated monitoring and alerting systems to guarantee the reliability, speed, and integrity of our security data pipeline. Collaborate with security researchers to rapidly deploy new scanning techniques and

typescriptrestai
View job →
V
1mo ago

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta’s Developer Experience team builds the tools engineers use every day to bring ideas to production rapidly and reliably. You’ll empower other Vanta engineers to leverage cutting-edge technologies and best practices to make Vanta more performant and scalable on a platform level. Example projects include modernizing our CI/CD pipelines, introducing new test frameworks, launching AI-powered dev tools, and scaling developer environments to support a growing engineering team. This team has a wide breadth of impact across all of product engineering. The work we do compounds in value by making it easier for engineers to diagnose and solve bugs, streamline workflows, and ship value to our customers quickly and safely. Vanta engineers design and develop new product functionality and infrastructure leveraging modern frameworks and tooling, including TypeScript, React, Node.js, MongoDB, Github Actions, and various AWS services such as Fargate and ECS. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. We’d love for you to join us! You will: Set direction for critical dev infrastructure, enabling us to stay ahead of continued rapid growth Design and build CI and build systems that ensure Vanta engineers can develop and ship robust products quickly and confidently Improve the efficiency and reliability of our deployment workflows, including tools for hotfixes, rollbacks, and incident mitigation Lead development of tools that accelerate feedback loops — from typechecking and linting to running tests and deploying changes Build and maintain scalable developmen

typescriptreactnode.js
View job →
S
1mo ago

Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. We're looking for an engineer to own the deployment and operational infrastructure of Multigres, our distributed Postgres platform. You'll be responsible for building and maintaining the Multigres Operator, ensuring reliable cloud deployments, and creating the tooling that powers our Kubernetes-based infrastructure. What You’ll Be Responsible for: Build and maintain the Multigres Operator - Maintain our Go-based Kubernetes operator that orchestrates distributed Postgres deployments Architect cloud deployment infrastructure - Design and implement robust deployment patterns for EKS and other Kubernetes platforms Manage storage and networking layers - Work with CSI drivers, persistent volumes, and cross-cloud networking to ensure data reliability and connectivity Develop deployment tooling - Create internal tools and automation for provisioning, scaling, and managing Multigres clusters Ensure operational excellence - Build monitoring, alerting, and diagnostic capabilities into the deployment layer Collaborate across teams - Work with database engineers, SRE, and product teams to deliver seamless deployment experiences You Might Be a Good Fit If You have: Strong systems programming skills - Proficiency in Go and experience building production-grade operators or controllers Deep Kubernetes expertise - Hands-on experience with Kubernetes internals, custom resources, and cloud-managed Kubernetes services (EKS, GKE, AKS) Database operations knowledge - Understanding of database deployment patterns, backup/restore, replication, and high availability Distributed systems experience - Familiarity with consensus protocols, failure scenarios, and designing for resilience Cloud infrastructure background - Experience with cl

kubernetesrestai
View job →
S
Supabase
📍 Remote• Full-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution, including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the team: Management API , written in TypeScript , Nest.js and for other JavaScript technologies, is a central part of every product in the Supabase stack. It allows all Supabase services and Supabase Studio to communicate with each other, programmatically manage your own Supabase projects and organisations, as well as integrate with 3rd party services like Vercel , Resend , or Lovable . We are seeking someone to help us maintain the existing API, expand OAuth applications capabilities, enhance reliability, and improve the overall public API experience. You will: Design, implement, and maintain both internal and public-facing APIs used across various Supabase products, including Studio, CLI, management APIs, and OAuth applications. Integrate with third-party platforms and partners, either by developing custom integrations or providing clear API points for them to connect with Supabase. Collaborate closely with various teams across Supabase (DevOps, Frontend, etc.) to ensure smooth integration and implementation of API functionality for the rest of the platform. Build and enhance testing, debugging, and monitoring tools to ensure public APIs' stability, reliability, and performance. Work with the dev-workflows team to enhance and improve the Branching experience, making it easier and more valuable for Supabase users. You have: 5+ years of experience in backend API development, with strong expertise in TypeScript and JavaScript (Node.js) and familiarity with modern tools and frameworks (e.g., Nest.js, Express, Vitest, Zod). Expertise in designing robust, scalable, and maintainable APIs, and experience with API versioning, pagination, and error handling best practices. Experience with OAuth

javascripttypescriptjava
View job →

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We are looking for a Software Engineer: IaC Platform Experience to join our Interfaces team and own the Terraform provider as a core part of Supabase's developer platform. This is a hands-on engineering role focused on the Go codebase behind the Supabase Terraform provider. You will partner with product and engineering leadership on roadmap priorities, drive technical execution, and ensure we ship a reliable, predictable, and well-documented Terraform experience for developers at scale. You will focus on resource behavior, lifecycle correctness, schema evolution, upgrade safety, and practical migration paths for existing users. This role is ideal for someone who thrives in async, fast-paced environments and enjoys building practical platform primitives that millions of developers can rely on. What You’ll Own Own the Go Terraform provider codebase, including architecture, implementation quality, test strategy, and release readiness. Improve Terraform provider reliability and ergonomics, including resource behavior, data sources, lifecycle edge cases, and upgrade safety. Drive technical strategy for IaC workflows through design docs, RFCs, and iterative delivery. Build practical migration and interoperability paths for existing Terraform users. Partner with product and engineering leadership in a shared roadmap model to define priorities, scope, and outcomes. Monitor customer feedback, OSS issues, and usage signals to continuously improve the Terraform experience. Create clear documentation and examples that make IaC workflows easier to understand and adopt. What You Bring 5+ years of software engineering experience in developer platforms, infrastructure tooling, or distributed sys

typescriptci/cdgit
View job →
S
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We’re looking for a Postgres Deployment Engineer to join our PostgreSQL team and help elevate our PostgreSQL offerings. You’ll work closely with the PostgreSQL team, playing an instrumental role in technical decision-making and refining internal methodologies. This role is ideal for someone who thrives in async, fast-paced environments and is excited about building developer tools that scale to millions. The focus of this role is owning the stability and deployment of our products. You will act as a bridge between product and infrastructure teams to improve the reliability of deployments and upgrades. You will have co-responsibility for builds and deployments via our public supabase/postgres GitHub repository, which bundles features into Docker images and AWS AMIs for cloud and local use. What You’ll Be Responsible For Package software into our supabase/postgres repo using Nix (with flakes), and help us transition our packaging from traditional to Nix packaging more over time. Manage PostgreSQL lifecycles, ensuring timely major, minor, and extension upgrades. Expand platform release systems to allow developers to increasingly self-service. Optimize CI/CD and tooling, specifically expanding GitHub Actions, team tooling, and testing/release approaches. Resolve production issues by proactively identifying and fixing problems in customer deployments. Maintain best practices and tests to ensure enhanced stability and decreased deployment risks. You Might Be a Good Fit If You Have 3+ years of experience with PostgreSQL and its ecosystem, including extensions and performance optimization. Are an Infrastructure Expert with proven experience in management, tooling, and optimization. Are p

javascriptjavasql
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The internal infrastructure team is responsible for building world-class infrastructure and tools used to train, evaluate and serve Cohere's foundational models. By joining our team, you will work in close collaboration with AI researchers to support their AI workload needs on the cutting edge, with a strong focus on stability, scalability, and observability. You will be responsible for building and operating superclusters across multiple clouds. Your work will directly accelerate the development of industry-leading AI models that power Cohere's platform North. Please Note: All of our infrastructure roles require participating in a 24x7 on-call rotation, where you are compensated for your on-call schedule. As a Staff Software Engineer, you will: Build and scale ML-optimized HPC infrastructure : Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads. Optimize for AI/ML training : Collaborate with cloud providers to fine-tune infrastructure for cost efficiency, reliability, and performance , leveraging technologies like R

pythonkubernetesgit
View job →
C
Cohere
📍 European Union• Full-time• From £180K/yr
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? North is Cohere's cutting-edge AI workspace platform, designed for enterprise users. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools, data and applications. The North for Finance team builds specialized AI solutions for finance teams while accelerating adoption of agentic workflows for their mission-critical operations. You'll work at the intersection of AI innovation and financial services, developing features that integrate domain-specific knowledge with enterprise-grade reliability. As an engineer on this team you will help transform complex financial workflows through intelligent automation. You'll extend core agentic platform capabilities with built-in data isolation, structured-data manipulation and collaborative governance, positioning North as the leading AI workspace for finance. As a Senior Software Engineer, you will: Design, build, ship, and maintain customer facing workflows and agentic automations

C
Clickup
📍 United States• Full-time
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Job Summary We are looking for a GTM DevOps Engineer to join our Business Systems team and own the reliability, automation, and delivery infrastructure behind our Go-To-Market (GTM) technology stack. This role sits at the intersection of platform reliability and CI/CD engineering, ensuring that our critical business systems — including Salesforce, NetSuite, MuleSoft, Workato, and an expanding portfolio of AI-powered workloads — are deployed consistently, operate resiliently, and scale with the business. You will partner closely with Business Systems developers, architects, and business stakeholders to build and maintain the pipelines, monitoring frameworks, and operational standards that keep our GTM systems healthy and our release cycles fast and predictable. As our team builds and deploys AI agents across GCP Cloud Run and AWS Bedrock AgentCore, you will serve as the infrastructure and deployment owner for these workloads — bringing engineering discipline to an environment where AI-generated code is increasingly entering production. This is a hands-on engineering role for someone who thrives in complexity, takes ownership of platform uptime, and brings a software engineering mindset to business application operations — directly supporting GTMSOE's broader mission of operational excellence across the GTM org. Key Responsibilities CI/CD & Release Engineering Design, build, and maintain CI/CD pipelines for Salesforce (SFDX/Salesforce CLI), NetSuite (SuiteScript/SuiteBundler), MuleSoft (Anypoint Platform), and Workato; establish branching strategies, environment promotion standards, and release gatin

pythonnode.jsaws
View job →
S
Sentry
📍 San Francisco• Full-time• $155K – $400K/yr
1mo ago

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role As a Senior Software Engineer on Sentry’s AI/ML team, you’ll be responsible for building the evaluation infrastructure that measures the accuracy, reliability, and real-world performance of our AI systems. This role is critical to ensuring that our debugging agents and AI-powered features behave correctly, safely, and predictably as they scale. You’ll design datasets, benchmarks, and test harnesses that turn ambiguous AI behavior into measurable signals, helping the team ship AI with confidence. In this role you will Design and build robust evaluation frameworks to measure accuracy, reliability, regressions, and edge cases in AI systems Create and curate high-quality datasets, golden test cases, and benchmarks grounded in real production data Build automated test harnesses and metrics pipelines to continuously evaluate models, prompts, and agentic workflows Partner closely with applied AI engineers and product leaders to define what “good” looks like and translate it into measurable criteria Own the evaluation lifecycle for major AI initiatives, from early experimentation through production monitoring You’ll love this job if you Care deeply about correctness, rigor, and measurement in AI systems Enjoy turning fuzzy product goals and model behavior into concrete tests and metrics Like building foundational infrastructure that unlocks faster iteration and higher confidence for the entire AI team Thrive in cross-functional environments and enjoy influencing model design through better evaluation Qualifications Minimum 5+ years of professional experience with a Bachelor’s degree in computer science, machine learni

typescriptpythonmachine learning
View job →
R
Ramp
📍 New York City• Full-time• From $10K/yr
1mo ago

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role Over the past few years at Ramp, we've reimagined how businesses manage spend, and the backbone of that vision is Procure-to-Pay. The P2P team builds the systems that connect vendors, invoices, and payments into one seamless, automated flow. It's where billions of dollars move through our platform, and every line of code matters. We're looking for a backend engineer who loves complex systems, obsesses over correctness, and wants to shape how money moves in modern finance software. You'll help design and scale the infrastructure that powers our payment rails, ensures auditability across products, and unlocks the next wave of AI-driven automation. It's a chance to work on systems that actually move money, with a team that cares deeply about precision, reliability, and scale. The problems are tough, the stakes are high, and the impact is massive. What You'll Do Design and build core P2P systems, from invoice ingestion and approval workflows to payment orchestration and reconciliation logic. Enable AI agents to classify, validate, and pro

pythonsqlrest
View job →
R
Ramp
📍 New York• Full-time• From $10K/yr
1mo ago

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About Production Engineering Production Engineering is Ramp's infrastructure ownership layer. We exist to make Ramp faster, more reliable, and more scalable — and we do that by being embedded in the problems, not adjacent to them. A few things that define how we operate: One team, one company, one objective. There is no "infra team" and "product team" — there is Ramp. We share the company's goals as our own. When a product team struggles with reliability or scalability, that is our struggle. If reliability or scalability is at risk, we own it. We don't wait to be invited, and we don't ask whose code it is. If a system is slow, if it breaks, if it won't scale — that's ours to lead, regardless of where it lives in the stack. We go first, and we go fast. When the path isn't obvious, we don't wait for someone else to find it. We move with urgency, propose the solution, align the stakeholders, and stay in until it's done — not until our ticket is closed. We lead the way. We find the next problem before it finds us. And when we solve it, we don't just fix

awsci/cdrest
View job →
R
Ramp
📍 New York City• Full-time• From $10K/yr
1mo ago

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role The Engineering Platform team helps 300+ Ramp engineers move faster without breaking things. We work across the full software development lifecycle, including local development, CI/CD, testing, service frameworks, reliability, architecture, performance, and maintenance. Platform engineers on this team do not wait for a roadmap to be handed to them. They find high-leverage problems, validate them with engineers and data, ship solutions, and iterate until those solutions are trusted, adopted, and durable. Agentic development has expanded our mission. We are building the systems that let engineers and AI agents work together effectively in Ramp’s large, complex codebases: clear abstractions, orchestration, reliable workflows, fast feedback loops, and strong guardrails are now even more important. These platform investments unlock agentic workflows that can plan, execute, test, review, and ship meaningful changes, helping every engineer and agent at Ramp maximize their impact. What You'll Do Improve Ramp’s development workflows across CI/

pythonci/cdrest
View job →
R
Ramp
📍 New York• Full-time• From $10K/yr
1mo ago

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role As a software engineer on the AI Soltuions team, you will co-lead customer engagements with an AI Solutions Strategist . The Strategist owns business discovery, ROI narrative, stakeholder alignment, and rollout planning. The engineer owns technical discovery, solution design, prototyping, implementation, and production readiness. This is a deeply client-facing role. You will spend significant time with customers and end users, moving projects from bootcamp and workflow discovery through implementation, launch, and steady production usage. What You’ll Do Translate customer goals into clear system requirements and non-functional requirements covering security, privacy, reliability, performance, scalability, and cost. Partner directly with customers to understand current workflows, constraints, systems, data quality, and adoption blockers. Create and maintain solution architecture artifacts: System context and data flow diagrams Integration plan across Ramp and customer systems Security model covering permissions, access patterns, and au

javascripttypescriptpython
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.