Jobiba hiring network

Reliability Engineer Jobs

2,049 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

N
1mo ago

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role Build AI-powered developer experience that makes engineers enabled. You’ll join our Developer Experience team, united by the mission of “Engineers can do their work quickly, easily, comfortably, and safely.” You’ll be part of a team spanning the US and India that turns ambiguous goals into shippable projects that meaningfully improve developer throughput and reliability. Your team owns AI experiences, remote and local dev environments, builds and CI, deploys, testing, and verification. You’ll work on a relatively broad scope within a well-resourced team. This role is based in Hyderabad, India. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve You build an AI harness to auto-resolve bugs reported for the Notion app. This is high-leverage work for the engineering org and requires solid engineering to improve resolution accuracy. You improve CI and build reliability so engineers get early feedback when changes break CI. We’re

typescriptci/cdrest
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $218K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Coinbase's Developer Infrastructure - Test team exists for one reason: every Coinbase engineer should get fast, reliable test signals so the company can test and ship faster. As AI accelerates the pace of code generation, test infrastructure is becoming a critical path for how quickly Coinbase delivers value to customers. As a Staff Software Engineer on the Platform team, you'll set the technical direction for how Coinbase tests and ships software, owning the systems that turn testing into a speed advantage instead of a bottleneck. What you'll do: Define and own the technical strategy for test infrastructure across Coinbase engineering, with feedback speed as a core design constraint. Build and operate core test infrastructure services, including test orchestration, smart test selection, sharding, flaky-test detection, and test result analysis. Drive measurable improvements in test feedback speed and signal reliability so engineers can ship with confidence and without reruns. Own systems end to end, including architecture, observability, SLOs, and on-call operations. Partner with engineering teams across Coinbase to identify bottlenecks and turn them into platform improvements. Mentor engineers, raise technical standards, and shape how the organization approaches test infrastructure. Required Skills and Experience: 10+ years building and operating production so

REMOTEawskubernetesgit
View job →
R
Roblox
📍 San Mateo• Full-time• From $196.8K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Why Reliability? Roblox serves over 100 million people every day across a platform that is constantly evolving — and behind every experience is infrastructure that has to work, every time, at massive scale. The Reliability team at Roblox operates at the depth and breadth of the Roblox stack. Availability of the platform is a key company goal. We are hiring our first Senior Machine Learning engineer within our team. As a Senior Machine Learning Engineer within Reliability, you will help set the direction for how machine learning systems/practices can be leveraged to improve the reliability of the overall Roblox platform. You will own the architectural and execution roadmap of leveraging massive data across - logs, traces, metrics, production changes, to proactively detect issues before they become real problems (MTTD) and/or reduce time to resolve incidents (MTTR). You will have the opportunity to cross functionally collaborate with other similar teams at Roblox to define best practices and software. You will: Help define the roadmap for leveraging Machine Learning Engineering to improve Production Systems Reliability at Roblox. Improve r

awsgitmachine learning
View job →
C
Coinbase
📍 - Canada• Full-time• Remote• From C$191.1K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Software Engineer, Core Reliability on the Infra Reliability team within Platform , you'll help Coinbase scale 50x by improving reliability, security, and deployment safety across our production environment. This team owns the systems that secure service configurations and secrets, reduce customer-facing incidents, and strengthen deployment infrastructure supporting thousands of services and hundreds of daily releases. You'll lead high-impact reliability projects that make our entire service environment more resilient and safer for customers. What you'll do: Own the design and delivery of reliability projects and features that improve resiliency across Coinbase's service environment in partnership with other engineering teams. Partner with critical T0/T1 services to understand architecture, improve scalability, and reduce operational toil. Build and enhance systems that securely manage service configurations and secrets at scale. Improve canary-based release systems and expand deployment capabilities to support thousands of services and hundreds of daily deployments with fewer incidents. Drive reliability best practices and strengthen reliability culture across engineering teams at Coinbase. Required Skills and Experience: 5+ years of software engineering experience designing, building, and maintaining production services in service-oriented architectures

REMOTEawsazuregcp
View job →
A
Affirm
📍 Poland• Full-time• Remote• PLN 2.4M – PLN 3.6M/yr
16 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. The Collections team, as part of the Repayments area, is on a mission to build a robust platform that will maximize the recovery of delinquent users by sending the right message to the right people at the right time, while monitoring and detecting problems and ensuring high reliability of the engineering systems. Collaborating closely with our product managers, backbook risk teams and other engineering teams , you will effectively manage loans throughout the delinquency phase of their lifecycle and develop and implement recovery strategies. We are looking for a highly motivated software engineer to build the next generation platform solutions that will allow us to manage our constantly growing volume while maintaining high availability and reliability of the systems. You will work closely with your team mates in the Collections team and Product to build robust collections solutions which will enable us to keep our loans portfolio healthy and help our customers pay on time and recover from any delays. What You'll Do With the support of your team, you will work on tasks that contribute to the team's projects and goals. You will work collaboratively and proactively with your team and stakeholders, bringing them along for your work and helping to create visibility and dialog regarding the risks and trade-offs related to your work. You will strike the right balance of speed and quality in your work, ensuring that we hit our business goals while protecting our systems from downtime. You will contribute to a sense of community on your team by engaging in growth and development activities What We Look For You have previous work or internship experience designing, developing and launching backend systems at scale and are experienced using one of Python or Kotlin. You are familiar with the building b

REMOTEpythonsqlmysql
View job →
A
Affirm
📍 Poland• Full-time• Remote• PLN 3.1M – PLN 4.5M/yr
16 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. The Collections team, as part of the Repayments area, is on a mission to build a robust platform that will maximize the recovery of delinquent users by sending the right message to the right people at the right time, while monitoring and detecting problems and ensuring high reliability of the engineering systems. Collaborating closely with our product managers, backbook risk teams and other engineering teams , you will effectively manage loans throughout the delinquency phase of their lifecycle and develop and implement recovery strategies. We are looking for a highly motivated software engineer to build the next generation platform solutions that will allow us to manage our constantly growing volume while maintaining high availability and reliability of the systems. You will work closely with your team mates in the Collections team and Product to build robust collections solutions which will enable us to keep our loans portfolio healthy and help our customers pay on time and recover from any delays. What You'll Do · With the support of your team’s tech lead and manager, you will break down larger projects into individual tasks, deliver them in multiple phases, and collaborate with others to ensure timely delivery of your work. · You will support your peers and stakeholders in the product development lifecycle by collaborating with product management, design & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. · You will support the operations and availability of your team’s artifacts by creating and monitoring metrics, escalating when needed, and supporting “keep the lights on” & on-call efforts. · You will contribute to a sense of community on your team by engaging in growth and devel

REMOTEpythonsqlmysql
View job →
A
Affirm
📍 Spain• Full-time• Remote• €684K – €1M/yr
16 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. The Collections team, as part of the Repayments area, is on a mission to build a robust platform that will maximize the recovery of delinquent users by sending the right message to the right people at the right time, while monitoring and detecting problems and ensuring high reliability of the engineering systems. Collaborating closely with our product managers, backbook risk teams and other engineering teams , you will effectively manage loans throughout the delinquency phase of their lifecycle and develop and implement recovery strategies. We are looking for a highly motivated software engineer to build the next generation platform solutions that will allow us to manage our constantly growing volume while maintaining high availability and reliability of the systems. You will work closely with your team mates in the Collections team and Product to build robust collections solutions which will enable us to keep our loans portfolio healthy and help our customers pay on time and recover from any delays. What You'll Do With the support of your team, you will work on tasks that contribute to the team's projects and goals. On-Call Rotation - There would be an on-call rotation for this role as a requirement You will work collaboratively and proactively with your team and stakeholders, bringing them along for your work and helping to create visibility and dialog regarding the risks and trade-offs related to your work. You will strike the right balance of speed and quality in your work, ensuring that we hit our business goals while protecting our systems from downtime. You will contribute to a sense of community on your team by engaging in growth and development activities What We Look For You have previous work or internship experience designing, developing and launching backend systems at scale an

REMOTEpythonsqlmysql
View job →
A
Affirm
📍 Spain• Full-time• Remote
16 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. The Collections team, as part of the Repayments area, is on a mission to build a robust platform that will maximize the recovery of delinquent users by sending the right message to the right people at the right time, while monitoring and detecting problems and ensuring high reliability of the engineering systems. Collaborating closely with our product managers, backbook risk teams and other engineering teams , you will effectively manage loans throughout the delinquency phase of their lifecycle and develop and implement recovery strategies. We are looking for a highly motivated software engineer to build the next generation platform solutions that will allow us to manage our constantly growing volume while maintaining high availability and reliability of the systems. You will work closely with your team mates in the Collections team and Product to build robust collections solutions which will enable us to keep our loans portfolio healthy and help our customers pay on time and recover from any delays. What You'll Do · With the support of your team’s tech lead and manager, you will break down larger projects into individual tasks, deliver them in multiple phases, and collaborate with others to ensure timely delivery of your work. · You will support your peers and stakeholders in the product development lifecycle by collaborating with product management, design & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. · You will support the operations and availability of your team’s artifacts by creating and monitoring metrics, escalating when needed, and supporting “keep the lights on” & on-call efforts. · On-Call Rotation - There would be an on-call rotation for this role as a requirement · Y

REMOTEpythonsqlmysql
View job →

Job Title AI Senior Systems Engineer (AI for RAMS & Systems Engineering) Job Description As an AI Senior Systems Engineer, you will help shape the future of Systems Engineering at Philips by driving the application of Artificial Intelligence within Systems Engineering and Reliability, Availability, Maintainability and Safety Engineering (RAMS) practices. You will identify opportunities where AI can enhance engineering activities, improve engineering productivity, and increase the quality, consistency, and traceability of engineering deliverables throughout the product development lifecycle. Acting as a thought leader and trusted advisor, you will help engineering teams understand, adopt, and effectively apply AI technologies within their daily engineering work. Working closely with systems engineers, RAMS engineers, architects, software teams, and AI specialists, you will bridge the gap between engineering challenges and emerging AI capabilities. You will contribute to the evolution of engineering methodologies, tools, and best practices that enable the next generation of AI-enabled Systems Engineering and RAMS Engineering at Philips. Your role: Drive the adoption of AI within Systems Engineering and RAMS practices across Philips. Identify, develop, and scale AI use cases for requirements engineering, system architecture, modelling, verification, validation, traceability, and engineering knowledge management. Identify, develop, and scale AI use cases for RAMS engineering like FMEA, data analysis, HALT/ALT, Modelling Simulation & Analysis, for Hardware and Software Support engineering teams in evaluating and implementing AI-enabled engineering workflows, methods, and tools. Coach and educate Systems Engineers, RAMS Engineers, Architects, and technical leaders on the opportunities, limitatio

machine learningartificial intelligenceai
View job →
N
1mo ago

We are looking for a highly motivated AI/ML Software Engineer to join the Enterprise Agentic AI Platform team within IT. You will work closely with Business Analysts, and Engineering teams to design, develop, and deploy enterprise AI solutions that improve productivity and automate business workflows across Engineering, Operations, and Manufacturing. What you'll be doing: Design, develop, and deploy Agentic AI applications using Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and AI orchestration frameworks. Build scalable AI services and reusable components integrated with enterprise applications such as PLM, SAP, and other business systems. Collaborate with business and IT teams to translate business requirements into AI-driven solutions. Develop secure, scalable APIs and enterprise integrations to enable intelligent workflows and automation. Improve AI solution quality, performance, and reliability through prompt engineering, evaluation, and continuous optimization. Partner with cross-functional teams throughout the Software Development Lifecycle (SDLC), from solution design through deployment and production support. What we need to see: Bachelor's or Master's degree in Computer Science, Information Technology, AI/ML, or a related field. 6+ years of software engineering experience with strong proficiency in Python and backend application development. Hands-on experience with Generative AI, LLMs, RAG, AI agents, REST APIs, and cloud-native application development. Experience integrating enterprise applications and building scalable, production-ready software solutions. Strong analytical, problem-solving, communicatio

pythonazureai
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $176K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . The Enterprise Applications & Architecture (EAA) Business Unit Delivery team builds and operates the internal platforms that help Coinbase scale. Within EAA, the Compliance CX Agent Experience team owns Agent Cockpit, the primary internal platform used by Coinbase's compliance agents and investigators to review cases, communicate with customers, access account and transaction data, and resolve compliance-related support queries. As a Staff Technical Program Manager, you'll own program delivery for this complex, high-impact platform, driving cross-functional execution to ensure agents and investigators have reliable, performant tools every day. What you'll do: Own the program management framework for Agent Cockpit, including roadmap tracking, intake and prioritization, sprint planning, dependency management, and cross-team delivery execution. Drive technical dependency coordination and platform migrations, translating upstream system changes (API deprecations, service integrations, data model changes) into internal delivery plans and engineering asks. Lead launch readiness and operational risk management, owning security review coordination, production approvals, testing governance, and escalation when critical-path items stall. Partner with Engineering on reliability, observability, and performance programs, including system health metrics, load testing, capacity

REMOTEawsagileai
View job →
SC
Sigma Computing
📍 San Francisco• Full-time• $170K – $240K/yr
16 days ago

Senior Software Engineer - Observability and Reliability About the Role We are growing the engineering team and looking for engineers who have the chops to build and deliver world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible. What You Will Be Doing Build observability tools and platforms, including: metrics, logging, distributed tracing, dashboarding, alerting, application performance management Build with modern tools and languages like Go, Open Telemetry and Kubernetes Participate in on-call rotation and ensure uptime of services Create runtime tools/processes that optimize cloud triaging and limit downtime Define best practices around making our systems and services measurable Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies. We expect successful candidates to be coding a majority of their time Qualifications We Need Strong Computer Science fundamentals 5+ years industry experience building and maintaining high-quality software, especially software other engineers use You apply a product mindset to infrastructure systems and feel accomplished enabling others Desire to be a great teammate and have fun at work Strong sense of craftsmanship, and a healthy academic curiosity Qualifications We Want (also, skills you’ll learn!) Experience building systems for data analytics Distributed systems monitoring and profiling skills Knowledge of cloud application security models Administered cloud service infrastructure (GCP, AWS, Azure) Startup experience Additional Job details Additional Job details The base salary range for this position is $170k - $240k annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience. Base pay is one part of the Total Package that is provided to compensate and recognize e

pythonsqlaws
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.