Jobiba hiring network

Reliability Engineer Jobs

2,049 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore fosters continuous learning and innovation. Job Summary Reporting into the Systems Engineering organisation, the Distinguished Engineer, End-to-End Security Architect will define and lead the security architecture for Graphcore’s inference service platform. This role is responsible for establishing a comprehensive security strategy spanning platform, infrastructure, networking, service operations, customer assurance, and compliance readiness. Working across multiple engineering and operational functions, the successful candidate will provide technical leadership, drive security requirements, and ensure the platform delivers robust protection, resilience, and trust for customers. The Team You will work closely with teams across security architecture, infrastructure engineering, networking, site reliability engineering, platform software, firmware, data centre operations, compliance, legal, customer engineering, and customer security. The team collaborates across the business to deliver secure, reliable, and scalable AI infrastructure and services while supporting customer assurance, regulatory requirements, and operational excellence. Responsibilities and Duties Own the end-to-end security a

airustexcel
View job →
G
15 days ago

Overview: Qsight is a high-growth division of Guidepoint focused on building data intelligence solutions for the healthcare sector. Qsight leverages proprietary datasets and rigorous analysis of alternative data sources to generate actionable insights for top-tier institutional investors, medical device manufacturers, and pharmaceutical companies. The Qsight team develops market intelligence products designed to be highly relevant, accurate, and scalable – delivering superior insights to a diverse, global client base. We are seeking an experienced, motivated Tehnical Operations Engineer to join our growing team. This is a multiple-hats role focused on SaaS/platform operations and tier-2 support for client-facing systems. You will own the administration and reliability of key tools, troubleshoot and resolve escalations with clear documentation, and build lightweight automation and reporting to reduce manual work as we scale. You will partner closely with Customer Success, Product, and Engineering to proactively monitor, support, and improve critical systems. Through practical, creative problem-solving, you will strengthen reliability, accelerate time to resolution, and increase operational visibility. Day to day, you will triage and resolve client technical questions, manage vendor license administration and renewals, and produce reporting that informs operational decisions. This role is a launchpad toward an SRE/Platform Engineering track as you grow into deeper automation, reliability engineering, and systems design work. This is a hybrid position based out of our Toronto office. What You’ll Do: Platform Support Own routine ops and configuration changes for critical SaaS platforms – Including Tableau, Freshdesk, Datadog, and our own client facing and internal portals Configure and maintain Freshdesk portals, routing, SLAs, permissions, integrations, etc. based on business requirements. Automate manual operations with Python, PowerAutomate, and shell scri

pythonsqlrest
View job →
A
Amplitude
📍 Remote• Full-time• $198K – $299K/yr
1mo ago

Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Staff Platform Engineer, you'll set technical direction for the platform across teams, lead our highest-complexity and highest-leverage initiatives, and shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll operate across team boundaries — partnering with product engineering, fellow Staff+ engineers, and engineering leadership to make Kubernetes and cloud infrastructure effortless across the entire engineering org. You'll build the self-service automation, shared standards, and scalable AWS and GCP infrastructure that let dozens of product teams ship faster, safer, and with less cognitive load — and you'll multiply the engineers around you while you do it. Key Responsibilities Set technical direction — shape platform and domain-level technical strategy that improves developer experience, reliability, security, and cost, and lead the high-complexity, cross-cutting initiatives that deliver it with measurable impact for the organization. Drive clarity through ambiguity. Take on the most loosely-defined problems, validate the critical assumptions early, and create alignment with stakeholders across teams so others can move quickly and confidently — driving cross-team decisions to a timely close and escalating when needed. Build the AI-augmented platform. Design org-wide tooling, guardrails, and policy-as-code that help every engineer get more out of AI-assisted development — infra primitives an LLM can safely reason about and PR against, automated review, and standards that hold as AI changes how code gets written. Own Infrastructure-as-Code standards for Kubernetes, AWS, and GCP using Terraform, Helm, Kustomize, and emerging tooling — setti

pythonawsazure
View job →

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Staff Engineer on the GitLab Delivery - Upgrades team, you’ll guide the technical direction for GitLab’s self-managed deployment strategy so customers can deploy, upgrade, and run GitLab reliably in their own infrastructure with minimal disruption. You’ll serve as a technical anchor for the team, working closely with your engineering manager, product manager, and partners across Site Reliability Engineering, Release, Security, and Development to shape cloud-native, operator-driven deployment patterns that reduce operational complexity and upgrade friction. In your first year, you’ll help define the a

sqlpostgresqlkubernetes
View job →
G
Gitlab
📍 United Kingdom• Full-time• Remote
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. Summary: Lead the Switchboard team, which owns the Tenant management product for Dedicated environments. This team is diverse in skills (Frontend, Backend and Reliability engineers) and owns the product end to end. The product enables self-service and workflow automations for customers and internal operators of our large scale Dedicated tenant fleet. Key Responsibilities: Lead the fullstack Switchboard team responsible for GitLab Dedicated’s customer-facing control panel, and set clear direction for how the team delivers and evolves that product. There is a strong coordination with the assigned Product Manager. Hire and develop a

REMOTEgitrestai
View job →
S
Smartsheet
📍 Bellevue• Full-time
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. The Grid Service and Platform Engineering team is looking for a highly motivated and collaborative Software Engineering Manager. This role involves leading the engineering of mission-critical, tier 0 service infrastructure, the foundational data platform that powers Smartsheet at scale. You will oversee services that handle millions requests per day, operate at 99.999% availability, and deliver low-latency, high-throughput performance for millions of customers worldwide. We are an agile team that operates iteratively, focused on building high-quality software and adhering to rigorous operational best practices across complex, cross-functional distributed systems. This full-time position reports to the Director, Engineering and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Manage one or more related teams of 6–10+ software engineers, driving development of tier 0 grid services and platform infrastructure that millions of customers depend on daily. Own and uphold 99.999% service availability targets across critical platform services, embedding reliability engineering, incident management, and on-call rigor into team culture. Help architect and guide technical vision to evolve low-latency, high-throughput service platforms capable of sustaining millions requests per day with predictable, consistent performance under load. Guide and mentor engineers on distributed systems architecture, scalability patterns, and platform best pr

vueawsagile
View job →
R
Roblox
📍 San Mateo• Full-time• From $243.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Data Access team within Infra / Storage builds EaaS (Entities-as-a-Service), Roblox’s large-scale managed OLTP access and management platform powering tens of millions of QPS across thousands of services. EaaS abstracts the complexity of distributed databases and caching systems behind a consistent and productive developer experience, enabling teams across Roblox to safely build and operate stateful systems at massive scale without requiring deep database expertise. The team operates at the intersection of distributed systems engineering, reliability engineering, and platform architecture — solving hard infrastructure problems that directly impact Roblox-wide scale and stability. As a Senior Software Engineer you will work on some of Roblox’s hardest backend infrastructure challenges around scalability, reliability, workload governance, adaptive flow control, and distributed systems. You Will Build and evolve EaaS, Roblox’s managed OLTP access platform powering tens of millions of QPS across hundreds of services. Design infrastructure that abstracts distributed databases and caching systems behind a consistent, safe, and highly productive developer experience. Drive reliability and scal

javaawsgit
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron is seeking a highly motivated and experienced Technical Staff Member to join our Quality Engineering team in Boise, ID. Your role will be crucial in ensuring the reliability and quality of wafers shipped from our Boise manufacturing site! Summary The Product Quality and Reliability team ensures Micron delivers high-quality, reliable semiconductor solutions that meet customer expectations and business objectives. The team partners across product engineering, design, technology development, manufacturing, and reliability organizations to identify risks, solve complex technical challenges, and drive continuous improvements in product performance, yield, and customer satisfaction! Position Overview As a Senior Product Quality and Reliability Engineer, you will be a technical leader. You will drive product excellence in quality, dependability, and yield across advanced semiconductor technologies. You will lead complex technical investigations, influence product and technology decisions, and develop innovative approaches that strengthen quality systems and business outcomes. This role provides an opportunity to have broad interpersonal impact through technical leadership, multi-functional teamwork, and strategic problem-solving. Responsibilities Lead complex investigations involving yield excursions, product deviations, and quality issues, driving root cause identification, containment actions, and sustainable corrective solutions Define product quality and reliability moni

airecruitment
View job →
O
Okta
📍 Bengaluru• Full-time
27 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. With the Okta's Auth0 organization’s increased dedication to ensuring customer availability expectations are exceeded in every way, you will play a key role as we evolve our system architecture to meet the demands of enormous growth and support the hundreds of millions of users who rely on us to provide uninterrupted access to business-critical Reporting to the Manager of Engineering, in this role as a SRE Operations Engineer, you will ensure smooth operations of our Customer Identity Cloud at Okta. Working closely with the SRE team, your primary focus will be on ensuring production systems remain operational at all times, while continually setting and achieving long-term operational success for the platform with potential career growth into Site Reliability Engineering. What you’ll be doing Executes operational work including updating/patching and maintaining the Engineering Service Desk queue Responsible for ensuring team requests are triaged and/or actioned in a timely manner Monitors Platform health and take steps to alleviate issues related to deployment and operations Assist with capacity, performance and scalability testing where required Escalation point for Platform issues from customer support teams Execute runbooks and update processes as required Interface with the SRE team to report core issues, required improvements and new feature requests What you’ll bring to the role General platform infrastructure knowledge, including high availability / l

nodejsmongodbaws
View job →

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<

pythonawsazure
View job →
C
Cloudflare
📍 Hybrid• Full-time• Hybrid• $114K – $157K/yr
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Network Deployment Engineer Available locations: Austin, Atlanta, Denver, New York About the department In this role, you will be focused on the build out and expansion of our global network. You'll work closely with Cloudflare’s SRE (Site Reliability Engineering) team, Network Engineering team, and with various vendors and partners (including hardware vendors, datacenter and network providers, and ISPs) to maintain and improve our global infrastr

pythonawslinux
View job →
C
Cloudflare
📍 Hybrid• Full-time• Hybrid
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Location: Singapore About the Team In this role, you will be focused on the build out and expansion of our global network. You'll work closely with Cloudflare’s Network Infrastructure planning team, Network Engineering team, Infrastructure automation team, Site Reliability Engineering (SRE) team, Project Managers and with various vendors and partners (including hardware vendors, logistics, datacenter and network providers, and ISPs) to p

pythonawslinux
View job →
O
Okta
📍 Washington• Full-time• From $147K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Auth0 Platform Observability team owns the observability tooling that monitors the Auth0 Platform, and we are looking for an Observability Engineer to help ensure that our Product and Platform Engineers can monitor and observe our platform while continuing to rapidly ship software that our customers love. Our engineers maintain and automate observability tooling for our entire platform, including metrics, logs, and traces. We are looking for engineers passionate about monitoring, observing, measuring uptime and availability, and ensuring platform stability. If you have experience within the Site Reliability Engineering (SRE) field or working as a Development Operations (DevOps) engineer, and you have a passion for Observability tooling, this position will allow you to further your learning and development in these areas. As a Senior Engineer on this team, you will act as a core technical leader. You will work cross-functionally to help integrate services with our instrumentation libraries, support product teams, and actively investigate incidents to identify our observability gaps. Responsibilities: Proven ability to champion observability best practices, acting as an educator who can effectively correct anti-patterns and teach other engineering teams how to build robust, standardized instrumentation. Be an expert in running services in production environments Contribute to the process of designing services for high growth and high availability. Provisi

node.jsawsazure
View job →
O
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are seeking an experienced Staff Software Engineer to join Okta's Universal Directory Platform team within the Product Platform Pillar. The team serves as the intelligent core of the enterprise security fabric, maintaining the source of truth for all identity assets and their associated relationships. Opportunity This position will be involved into development, design, and maintenance of our highly performant, reliable, and scalable platform, which is critical for managing user lifecycles, groups, and memberships. The successful candidate will possess experience in building and deploying scalable, reliable software and infrastructure within a cloud environment. What you’ll be doing Understand our identity management group codebase and development process: Jira, Technical Designs, Code Review, Testing, and Deployment. Develop and implement frameworks and toolings for our Universal Directory Service platform. Design and implement high-performance distributed scalable and fault-tolerant software components. Quickly deliver high-quality bug fixes and handle customer-reported issues. Conduct quality code reviews and automated testings. Partner with our Product Development, QA, and Site Reliability Engineering teams for scoping the development and deployment work. What you’ll bring to the role The ideal candidate is someone who is experienced building software systems to manage and deploy reliable and performant infrastructure and prod

javasqlpostgresql
View job →

Location Details: Pune, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a hybrid position. You’ll divide your time between working remotely from your home and an office, so you should live within commuting distance. Hybrid teams may work in-office as much as a few times a week or as little as once a month or quarter, as decided by leadership. The hiring manager can share more about what hybrid work might look like for this team. Join our Team Our team builds and operates the foundational infrastructure platforms that power GoDaddy's engineering organization. We own critical services including secrets management, software distribution, host security controls, and live patching for thousands of Linux systems running on OpenStack. This role sits at the intersection of Linux engineering, platform engineering, reliability engineering, and security. You will help define how core infrastructure services are designed, operated, automated, and scaled across the enterprise! What you'll get to do... Design, build, and operate highly available, scalable, and secure infrastructure platforms supporting large-scale Linux environments, with a focus on reliability, resiliency, and operational efficiency Lead the architecture, implementation, and operation of infrastructure services, including OpenStack, enterprise secrets management, package management, software promotion pipelines, and platform lifecycle management Develop and maintain automation solutions using infrastructure-as-code, Ansible, Python, Go, and self-service capabilities to improve efficiency and reduce operational overhead Build and improve observability and reliability practices through monitoring, logging, alerting, dashboards, managing incidents, analyzing underlying causes, disaster recovery, and service health reporting

pythongitlinux
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.