Jobiba hiring network

Software Reliability Engineer Jobs

6,428 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The SRE Leadership Team The SRE Leadership Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great infrastructure is invisible—it just works. Our team champions a culture of continuous learning, data-driven decision-making, and blameless incident response. We work at the intersection of product engineering, architecture, and operations to ensure Auth0 remains the trusted authentication platform for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with a focus on scalability, resilience, and empowering engineers to grow as technical leaders. What You'll Be Doing Lead the SRE team's technical direction , translating organizational vision into actionable roadmaps while driving complex, cross-functional initiatives across product and platform teams Operate at scale through hands-on participation in 24/7 on-call rotations (follow-the-sun weekdays, shared weekends), directly troubleshooting and remediating incidents on critical systems Build infrastructure resilience , designing and implementing monitoring, alerting, and automation improvements that reduce toil and elevate operational efficiency Champion reliability best practices , establishing policies and cultural standards that embed observability, resilience, and software engineering rigor into all engineering efforts Mentor and develop SRE talent , elevati

pythonawsazure
View job →

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore fosters continuous learning and innovation. Job Summary Reporting into the Systems Engineering organisation, the Distinguished Engineer, End-to-End Security Architect will define and lead the security architecture for Graphcore’s inference service platform. This role is responsible for establishing a comprehensive security strategy spanning platform, infrastructure, networking, service operations, customer assurance, and compliance readiness. Working across multiple engineering and operational functions, the successful candidate will provide technical leadership, drive security requirements, and ensure the platform delivers robust protection, resilience, and trust for customers. The Team You will work closely with teams across security architecture, infrastructure engineering, networking, site reliability engineering, platform software, firmware, data centre operations, compliance, legal, customer engineering, and customer security. The team collaborates across the business to deliver secure, reliable, and scalable AI infrastructure and services while supporting customer assurance, regulatory requirements, and operational excellence. Responsibilities and Duties Own the end-to-end security a

airustexcel
View job →
G
15 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Quality leadership within Manufacturing Operations, the Senior Reliability Scientist is responsible for leading reliability activities across complex, high-performance systems. Working closely with established reliability experts and cross-functional teams, this role uses experimental data and advanced modelling to inform design decisions, validate product reliability and optimise serviceability strategies, including spares provisioning. The Team The Quality team within Manufacturing Operations is responsible for ensuring product robustness, reliability and lifecycle performance across Graphcore’s hardware portfolio. The team includes experienced reliability specialists and works closely with technology research, chip, board, system design, platform and operations teams to translate reliability insights into actionable improvements across the product lifecycle. Responsibilities and Duties: · Define and refine reliability requirements across silicon, board and system levels, working in partnership with research and design teams · Apply ad

aigoexcel
View job →
R
Replit
📍 Foster City• Full-time
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role: Join our Infrastructure Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Infrastructure Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff Infrastructure Engineers who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. You Will: Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios. Optimize Performance and Infrastructure: Collaborate with core infrastructure and product teams to performance tune and optimize our cloud deployments (Kubernetes, Docker, GCP). Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions. Elevate Developer Experience: Design and implement improvements to our build, test, and deployment systems to make software delivery faster, safer,

pythongcpdocker
View job →
KH
K Health
📍 Tel Aviv• Full-time
15 days ago

About the Role: We are looking for a Senior DevOps Engineer to join our DevOps team at K Health. You will own and evolve the infrastructure underpinning a healthcare AI platform serving patients and enterprise health system partners. This is a high-ownership role: you will architect and operate cloud environments across K Health and its enterprise partners, lead complex infrastructure migrations, drive disaster recovery programs, and help build the next generation of AI-powered operations tooling. You will also mentor junior engineers and collaborate closely with product and engineering teams across the company. This is a hybrid role based in New York City (4 days/week in office) and includes participation in a daytime on-call rotation. What you will do: Own the design, implementation, and evolution of our GKE-based Kubernetes infrastructure across K Health and enterprise partner environments. Build and maintain our Terraform modular infrastructure library, including reusable modules with automated testing, across GCP, Cloudflare, and AWS. Architect, build, and maintain GitLab CI/CD shared pipeline templates used by all engineering teams (build, test, security scanning, deployment). Own and maintain self-hosted infrastructure software running in-cluster, including GitLab, ArgoCD, Langfuse, DependencyTrack, NGINX Ingress, and others. Implement and support security and compliance controls across infrastructure and the software supply chain - secrets management, pipeline secret detection, container scanning, SOC2 and HIPAA. Drive disaster recovery readiness: design failover scenarios, author runbooks, and lead periodic DR tests. Lead development of AI-powered operations tooling and agentic infrastructure. Monitor, troubleshoot, and improve production system reliability; respond to incidents during on-call shifts. Mentor junior DevOps engineers and establish team-wide engineering standards. What we are looking for: 5+ years of experience in DevOps, platform engineering,

pythonsqlpostgresql
View job →
O
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About the Team The Professional Services R&D team is a new, dynamic group at the forefront of innovation within Okta. Our mission is to design and build reusable, scalable assets and tools that empower our delivery teams and partners. By making customer engagements more efficient, streamlined, and cost-effective, we directly contribute to our customers' success and accelerate their time-to-value with Okta. This is a unique opportunity to join a strategic team from the ground up and shape the future of Okta's professional services. Position Summary As the DevOps Engineer for the R&D team, you will build and own the infrastructure and automation that enables us to develop and release software with speed and confidence. You will be responsible for creating and managing our CI/CD pipelines, defining our infrastructure as code, and ensuring our deployed assets are scalable, secure, and observable. You will be a key enabler of the team's agility, implementing the tools and processes that allow us to innovate and iterate quickly while maintaining a high bar for quality and reliability. Responsibilities Design, build, and maintain the team's CI/CD pipelines to automate the build, test, and deployment of our software assets. Manage and provision cloud infrastructure using Infrastructure as Code (IaC) principles and tools (e.g., Terraform, CloudFormation). Implement and manage monitoring, logging, and alerting solutions to ensure the health and performance of

pythonawsazure
View job →

Job Title Lead Software Technologist I - UI Job Description Minimum required Education: Bachelor's / Master's Degree in Computer Science, Software Engineering, Information Technology or equivalent. Job title: Lead Software Technologist I Your role: Lead development of intuitive clinical user interface applications used by clinicians in high-acuity environments. Collaborate with UX teams and clinicians to ensure applications meet usability and safety requirements. Provide technical leadership across modern UI technologies and supporting C++ layers for embedded devices. Lead and mentor Platform development teams on software architecture, design patterns, and embedded systems best practices through daily hands-on collaboration. Lead and mentor application development teams on application services for physiological data visualization, and clinical decision support features. Establish and enforce quality standards, development methodologies, and coding practices that drive continuous improvement in software reliability and performance. Conduct rigorous code reviews and provide constructive technical feedback to ensure adherence to medical device software standards (IEC 62304, FDA regulations) Optimize application performance by identifying and resolving bottlenecks in resource-constrained embedded Linux environments. Drive the adoption of AI-enabled development tools and demonstrate measurable productivity improvements across the team. Collaborating with cross-functional teams including Product Management, QA, Regulatory, and other engineering leads to define and deliver features. Support software lifecycle management activities including sustaining engineering, defect resolution, and platform evolution.

javascripttypescriptlinux
View job →

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Staff Engineer on the GitLab Delivery - Upgrades team, you’ll guide the technical direction for GitLab’s self-managed deployment strategy so customers can deploy, upgrade, and run GitLab reliably in their own infrastructure with minimal disruption. You’ll serve as a technical anchor for the team, working closely with your engineering manager, product manager, and partners across Site Reliability Engineering, Release, Security, and Development to shape cloud-native, operator-driven deployment patterns that reduce operational complexity and upgrade friction. In your first year, you’ll help define the a

sqlpostgresqlkubernetes
View job →
C
Cloudflare
📍 Hybrid• Full-time• Hybrid• $114K – $157K/yr
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Network Deployment Engineer Available locations: Austin, Atlanta, Denver, New York About the department In this role, you will be focused on the build out and expansion of our global network. You'll work closely with Cloudflare’s SRE (Site Reliability Engineering) team, Network Engineering team, and with various vendors and partners (including hardware vendors, datacenter and network providers, and ISPs) to maintain and improve our global infrastr

pythonawslinux
View job →
C
Cloudflare
📍 Hybrid• Full-time• Hybrid
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Location: Singapore About the Team In this role, you will be focused on the build out and expansion of our global network. You'll work closely with Cloudflare’s Network Infrastructure planning team, Network Engineering team, Infrastructure automation team, Site Reliability Engineering (SRE) team, Project Managers and with various vendors and partners (including hardware vendors, logistics, datacenter and network providers, and ISPs) to p

pythonawslinux
View job →
O
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are seeking an experienced Staff Software Engineer to join Okta's Universal Directory Platform team within the Product Platform Pillar. The team serves as the intelligent core of the enterprise security fabric, maintaining the source of truth for all identity assets and their associated relationships. Opportunity This position will be involved into development, design, and maintenance of our highly performant, reliable, and scalable platform, which is critical for managing user lifecycles, groups, and memberships. The successful candidate will possess experience in building and deploying scalable, reliable software and infrastructure within a cloud environment. What you’ll be doing Understand our identity management group codebase and development process: Jira, Technical Designs, Code Review, Testing, and Deployment. Develop and implement frameworks and toolings for our Universal Directory Service platform. Design and implement high-performance distributed scalable and fault-tolerant software components. Quickly deliver high-quality bug fixes and handle customer-reported issues. Conduct quality code reviews and automated testings. Partner with our Product Development, QA, and Site Reliability Engineering teams for scoping the development and deployment work. What you’ll bring to the role The ideal candidate is someone who is experienced building software systems to manage and deploy reliable and performant infrastructure and prod

javasqlpostgresql
View job →

MDM TechnologyMDM Technology, a company affiliated with Dun & Bradstreet, provides enterprise-grade master data management (MDM) solutions that help organizations create a single, trusted view of business entities. By cleansing, matching, linking, and enriching data, anchored by the global standard D‑U‑N‑S® Number business identifier, MDM Technology enables accurate identity resolution across systems to support analytics, compliance, and AI workflows across industries. The Senior Quality Assurance Engineer ensures our software meets user and product needs. The role creates and maintains project test plans, defines and tracks quality assurance metrics, collects and analyzes data for software process evaluation and improvements, and integrates them into business processes combining deep technical expertise in API and Integrations testing and Automation. This role must show a passion for AI enabled Software Quality practices Leveraging modern AI tools to improve testing efficiency, accelerate release cycles, and enhance software reliability across complex enterprise integration.

Location Details: Pune, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a hybrid position. You’ll divide your time between working remotely from your home and an office, so you should live within commuting distance. Hybrid teams may work in-office as much as a few times a week or as little as once a month or quarter, as decided by leadership. The hiring manager can share more about what hybrid work might look like for this team. Join our Team Our team builds and operates the foundational infrastructure platforms that power GoDaddy's engineering organization. We own critical services including secrets management, software distribution, host security controls, and live patching for thousands of Linux systems running on OpenStack. This role sits at the intersection of Linux engineering, platform engineering, reliability engineering, and security. You will help define how core infrastructure services are designed, operated, automated, and scaled across the enterprise! What you'll get to do... Design, build, and operate highly available, scalable, and secure infrastructure platforms supporting large-scale Linux environments, with a focus on reliability, resiliency, and operational efficiency Lead the architecture, implementation, and operation of infrastructure services, including OpenStack, enterprise secrets management, package management, software promotion pipelines, and platform lifecycle management Develop and maintain automation solutions using infrastructure-as-code, Ansible, Python, Go, and self-service capabilities to improve efficiency and reduce operational overhead Build and improve observability and reliability practices through monitoring, logging, alerting, dashboards, managing incidents, analyzing underlying causes, disaster recovery, and service health reporting

pythongitlinux
View job →

The Engineering Lead Analyst – SonarQube & Code Quality Engineering is a senior-level engineering role responsible for leading static code analysis, automated code quality governance, security vulnerability remediation, and AI-augmented developer enablement across enterprise software delivery pipelines. In this role, you will champion software reliability, maintainability, clean-coding standards, and automated quality gates. You will partner with development teams, system architects, and platform engineering to integrate and manage enterprise-scale code quality platforms (such as SonarQube) both on-premises and in cloud/SaaS environments. Additionally, you will drive modern engineering practices by embedding Behavior-Driven Development (BDD) within your own software delivery and leveraging Agentic AI workers and Model Context Protocol (MCP) architectures to optimize developer experience, streamline code governance, and boost engineering velocity. Key Responsibilities 1. Code Quality & Static Analysis Platform Ownership Lead the architecture, deployment, administration, and continuous enhancement of enterprise Static Application Security Testing (SAST) and Code Quality platforms (e.g., SonarQube , DeepSource, Codacy, Semgrep). Configure, calibrate, and enforce automated Quality Gates, code rulesets, technical debt calculation models, and code-coverage baselines across multi-language enterprise repositories. Oversee version upgrades, patching, high availability, and operational maintenance for on-premises and SaaS/cloud-hosted code quality infrastructure. 2. CI/CD & Pipeline Integration <li style=

javascripttypescriptpython
View job →
G
Gitlab
📍 United Kingdom• Full-time• Remote
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. Summary: Lead the Switchboard team, which owns the Tenant management product for Dedicated environments. This team is diverse in skills (Frontend, Backend and Reliability engineers) and owns the product end to end. The product enables self-service and workflow automations for customers and internal operators of our large scale Dedicated tenant fleet. Key Responsibilities: Lead the fullstack Switchboard team responsible for GitLab Dedicated’s customer-facing control panel, and set clear direction for how the team delivers and evolves that product. There is a strong coordination with the assigned Product Manager. Hire and develop a

REMOTEgitrestai
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime