Jobiba hiring network

Software Reliability Engineer Jobs

6,428 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

N
Nuro
📍 Mountain View• Full-time• From $145.8K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Team The Robotics Reliability Engineering (RRE) team at Nuro focuses on fleet reliability as our AV capabilities and operating footprint grow. We work across software, hardware, infrastructure, and operations to understand fleet behavior and improve operational readiness. When high-severity events impact missions or fleet-wide performance, RRE coordinates investigations and ensures that lessons learned result in durable platform improvements. About the Role As a Software Reliability Engineer at Nuro, you will help design, build, and operate systems that support our autonomous vehicle platform and fleet operations. You will work across the full development lifecycle, from early design and deployment through operations and continuous improvement. You’ll join an on-call rotation to stay close to real-world operations. At its core, the role is building resilient systems through automation, observability, and operational feedback. T

pythonrestai
View job →
N
Nuro
📍 Mountain View, California (HQ)• Full-time
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Team The Robotics Reliability Engineering (RRE) team at Nuro focuses on fleet reliability as our AV capabilities and operating footprint grow. We work across software, hardware, infrastructure, and operations to understand fleet behavior and improve operational readiness. When high-severity events impact missions or fleet-wide performance, RRE coordinates investigations and ensures that lessons learned result in durable platform improvements. About the Role This is a 12 month temporary full-time position with full benefits and potential for extension based on performance and business needs. As a Software Reliability Engineer at Nuro, you will help design,

pythonlinuxrest
View job →
O
Okta
📍 Toronto• Full-time• From C$160K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Staff Software Reliability Engineer - Data Platform About the Team The Data Platform team is responsible for the foundational data services, systems, and data products for Okta that benefit our users. Today, the Data Platform team solves challenges and enables: Streaming analytics Interactive end-user reporting Data and ML platform for Okta to scale Telemetry of our products and data Our elite team is fast, creative and flexible. We encourage ownership. We expect great things from our engineers and reward them with stimulating new projects, new technologies and the chance to have significant equity in a company. Okta is about to change the cloud computing landscape forever. About the Position This is an opportunity for experienced Software Reliability Engineers to join our fast growing Data Platform organization that is passionate about scaling high volume, low-latency, distributed data-platform services & data products. In this role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Okta to scale for years to come. As a member of the Data Platform team, you will be responsible for designing, building, and deploying the systems that power our data analytics and ML. Our analytics infrastructure stack sits on top of many modern technologies, including Kinesis, Flink, ElasticSearch, and Snowflake. We are looking for experienced Software Engineers who can help desi

javaawskubernetes
View job →
A
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: As a Site Reliability Engineer, you will play a crucial role in ensuring the smooth operation of all user-facing services and other Anyscale production systems. Anyscale values diversity and inclusion, and we encourage applications from individuals of all backgrounds. This includes processes for provisioning, negotiating prices, managing costs, seeing opportunities for teams to reduce wastage by finding applications across the company. You will apply sound engineering principles, operational discipline, and mature automation to our environments and the Anyscale codebase as we scale. As part of this role, you will: Develop a unified perspective on how cloud components are utilized across the company, taking into account diverse needs and requirements. Ensure that deployment methodologies align with the company's reliability goals. Build systems that promote understanding of production environments, facilitating quick identification of issues through robust observability infrastructure for metrics, logging, and tracing. Create monitoring and alerting systems at different levels, enabling teams to easily contribute and enhance the overall monitoring capabilities. Establish testing infrastructure to s

machine learningai
View job →

SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . As a Software Dev Senior Engineer , you will own the reliability, scalability, and operational excellence of our Cloud-based services. You will define and enforce reliability standards, drive the adoption of SRE practices across engineering teams, and build the systems and tooling that keep our production infrastructure healthy. We follow a DevOps model: Development and Operations teams are integrated, and the SRE function acts as the reliability layer — setting Service Level Objectives, managing error budgets, and continuously reducing toil through engineering. Key Responsibilities: Define, publish, and continuously refine Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs ) for all critical services, partnering with product and engineering leadership. Own the error budget framework: track consumption, enforce error budget policies, and drive reliability investments when budgets are at risk. Lead the design and implementation of comprehensive observability platforms — metrics, structured logging, and distributed tracing — to ensure full visibility into pro

pythonsqlpostgresql
View job →
A
Asana
📍 Warsaw• Full-time• $372K – $432K/yr
1mo ago

We're looking for a Senior Platform Reliability Engineer who brings strong software engineering skills and a deep understanding of system behavior under load and stress. This role is a good fit for someone who wants to own reliability as a first-class concern – building the foundational systems that protect Asana's platform, not just responding when things go wrong. You'll build core platform systems like load shedding, rate limiting, circuit breakers, and traffic controls that protect Asana under real-world load. This is deep, cross-cutting work that shapes stability and performance of our entire infrastructure – and you'll partner closely with other platform teams to make reliability something that's built in, not bolted on. Our tech stack includes: AWS, Kubernetes (EKS), CloudFront, Istio, Cilium, MySQL (RDS), OpenSearch, DynamoDB, Redis, Terraform, Datadog, TypeScript, Scala, Go, and Python. (Yeah, we know this sounds like buzzword bingo – but we want this post to actually show up in your searches.) Why this role? Reliability as a first-class feature : You won't be patching things up after the fact. You'll build the systems that make Asana resilient by design. Foundational work : Load shedding, traffic management, ingress/egress – these are the building blocks that protect everything else. You'll own them. Strong collaboration, reasonable hours : You'll work closely with infrastructure teams in Warsaw and Reykjavik, making deep collaboration practical without constant timezone gymnastics. Room to grow : This is a new team, and you'll help shape what Platform Reliability Engineering looks like at Asana – whether that means leading projects, mentoring others, or defining our technical direction. In this role, success means shipping systems that other teams rely on by default – because they make the platform safer, not because they're mandatory. We're especially interested in people who think like backend engineers but obsess over failure modes, capacity plan

typescriptpythonsql
View job →

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: We are looking for a Senior Software Engineer to join our Site Reliability Engineering team. As a Senior Software Engineer in Production SRE, you will be responsible for developing and maintaining the tools and systems that enable our engineering teams to operate our services reliably and at scale. You will work closely with our SREs and other engineering teams to ensure our services are properly instrumented and able to scale with our growing business. The Difference You Will Make: In this role, your expertise in developing and maintaining tools and systems will be instrumental in bolstering our services' reliability and improving how the company manages incidents broadly. By collaborating closely with other engineering teams you will help establish a culture of reliability throughout the organization by providing a comprehensive incident management platform that is being used for instrumentation, operability, and around incidents. Your ability to identify opportunities for improvement and drive their implementation will contribute significantly to our overall operational efficiency and growth, ensuring that our services remain resilient as our business continues to expand. Additionally, as an essential part of this role, you will serve as an active member of the Production SRE team, responding to and managing high severity incidents. Your vast technical experience and leadership skills will be invaluable as you step into the role of Incident Commander during these critical events. You will guide cross-functional teams during crisis situations and ensure timely resolution, minimizi

pythonjavaaws
View job →
P
Pagerduty
📍 Atlanta• Full-time• From $98K/yr
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office, you'll help build and operate the foundational infrastructure that powers PagerDuty's real-time digital operations platform. Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You'll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases. Key Responsibilities ● Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems. ● Contribute to the reliability and scalability of PagerDuty's core platform by hardening existing systems and supporting the rollout of new infrastructure

pythonawsazure
View job →
P
Pagerduty
📍 Atlanta• Full-time• $113K – $171.6K/yr
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. As a Site Reliability Engineer II on the Core Infrastructure team in our Atlanta office, you'll help build and operate the foundational infrastructure that powers PagerDuty's real-time digital operations platform. Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You'll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases. Key Responsibilities ● Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems. ● Contribute to the reliability and scalability of PagerDuty's core platform by hardening existing systems and supporting the rollout of new infrastructur

pythonawsazure
View job →
I
Instacart
📍 United States - Remote• Full-time• Remote• From $160K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Are you passionate about technology and ready to dive into the world of Site Reliability Engineering? We are looking for enthusiastic individuals to join our team as Site Reliability Engineer II, where you'll play a crucial role in ensuring the reliability and performance of our platform. This is a top-notch opportunity to learn from skilled engineers and contribute to solving complex technical challenges. You will be involved in monitoring systems, responding to incidents, and developing automation to streamline operations. We are looking for someone eager to grow their skills, learn new technologies, and contribute to a culture of reliability. The Site Reliability Engineering (SRE) team integrates software and systems engineering to design and manage large-scale, distributed, and fault-tolerant systems. The team is responsible for ensuring high reliabili

REMOTEawsazuregcp
View job →
M
Mongodb
📍 Austin; New York City; San Francisco; Seattle; United States• Full-time• From $127K/yr
1mo ago

We are hiring an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build scalable security controls and services within MongoDB Atlas multi-cloud infrastructure. The team sits within the Site Reliability Engineering organization and works with other engineering teams to ensure that our infrastructure adheres to the highest security standards. This role can be based out of our New York City, Austin, Seattle or San Francisco offices, or work fully remotely on standard East Coast business hours. Responsibilities: Design and build core security primitives and services that protect MongoDB Atlas compute, networking, and identity across AWS, Azure, and GCP Build secure-by-default infrastructure using Linux security mechanisms (AppArmor, SELinux, seccomp, cgroups), Kubernetes, and eBPF to enforce runtime policies and gain deep visibility into systems behaviour Develop APIs, automation, and tooling that manage security posture at scale (CSPM, vulnerability management, workload identity) and provide monitoring, logging, and alerting pipelines that integrate with our tooling (Grafana, Splunk, Victoria Metrics.) Integrate security into our CI/CD and infrastructure-as-code workflows (Terraform) so that security controls are versioned, reviewed, and deployed just like any other code Lead complex projects end‑to‑end, from problem discovery and design docs to implementation, rollout, and long‑term ownership Collaborate with SRE, platform and product engineering teams to define secure architectures for new infrastructure and services Qualifications: You might be a great fit if you match some of the following: 5+ years of experience in Software Engineering, Site Reliability Engineering, or similar roles, preferably with relevant security work Proficiency with at least one programming language (Java, Golang, Rust, Python, or C/C++) and experience with infrastructure-as-code tools (Terraform) to automate security configurations

pythonjavamongodb
View job →

We are hiring an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build scalable security controls and services within MongoDB Atlas multi-cloud infrastructure. The team sits within the Site Reliability Engineering organization and works with other engineering teams to ensure that our infrastructure adheres to the highest security standards. This role can be based out of our Dublin office, or work fully remotely in Ireland. Responsibilities: Design and build core security primitives and services that protect MongoDB Atlas compute, networking, and identity across AWS, Azure, and GCP Build secure-by-default infrastructure using Linux security mechanisms (AppArmor, SELinux, seccomp, cgroups), Kubernetes, and eBPF to enforce runtime policies and gain deep visibility into systems behaviour Develop APIs, automation, and tooling that manage security posture at scale (CSPM, vulnerability management, workload identity) and provide monitoring, logging, and alerting pipelines that integrate with our tooling (Grafana, Splunk, Victoria Metrics.) Integrate security into our CI/CD and infrastructure-as-code workflows (Terraform) so that security controls are versioned, reviewed, and deployed just like any other code Lead complex projects end‑to‑end, from problem discovery and design docs to implementation, rollout, and long‑term ownership Collaborate with SRE, platform and product engineering teams to define secure architectures for new infrastructure and services Qualifications: You might be a great fit if you match some of the following: 5+ years of experience in Software Engineering, Site Reliability Engineering, or similar roles, preferably with relevant security work Proficiency with at least one programming language (Java, Golang, Rust, Python, or C/C++) and experience with infrastructure-as-code tools (Terraform) to automate security configurations and processes A deep understanding of Linux and networking concepts,

pythonjavamongodb
View job →
SL
15 days ago

Title: Staff Site Reliability Engineer, Product Area Focus Location: Noida/ Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering team. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedite

pythonjavareact
View job →
SL
15 days ago

Title: Senior Site Reliability Engineer - I, Product Area Focus Location: Noida (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering teams. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedit

pythonjavareact
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime