Jobiba hiring network

Senior Software Reliability Engineer Jobs

7,292 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

What We Do At GoGuardian, we’re helping build a future where all learners are ready and inspired to solve the world’s greatest challenges. Our award-winning system of learning solutions is purpose-built for K-12 and trusted by school leaders to promote effective teaching and equitable engagement while helping empower educators to keep students safe. What It’s Like to Work at GoGuardian We are an outcomes-focused learning company with a steadfast focus on improving learning environments, one classroom at a time. Working with us means joining a remote team of diverse, committed, mission-driven employees who are inspired by our vision, dedicated to our customers, and ready to roll up their sleeves. Guardians put their heads together to solve problems, learn together from experiments that fail, and stand together by their work with full accountability. We balance our diligence with an inclusive culture that invites everyone to bring their whole self to work. Join us and learn why “I love the people here” is one of the most frequent comments we hear from Guardians. What We Do At GoGuardian, we’re helping build a future where all learners are ready and inspired to solve the world’s greatest challenges. Our award-winning system of learning solutions is purpose-built for K-12 and trusted by school leaders to promote effective teaching and equitable engagement while helping empower educators to keep students safe. The Role We’re looking for a Senior Site Reliability Engineer (SRE) to help design, scale, and maintain the infrastructure that powers our core products and services. In this role, you’ll collaborate with engineering teams to drive operational excellence, optimize system performance, and ensure high availability across production environments. This position sits on Tech Foundation, a team that manages core cloud infrastructure, shared data services, and developer tooling to empower our product teams to deliver software efficiently and secu

javascripttypescriptpython
View job →
N
Nvidia
📍 Remote, Poland• Remote
10 days ago

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, passionate, and self-motivated, we want to hear from you! We are looking for an experienced networking software engineer. An awesome candidate is highly technical who is also comfortable with dealing with enterprise customers. You will join a team of Solution Engineers focused on the Mellanox Networking, DGX Platforms, Container Orchestrators, Deep Learning containers, and other Enterprise related system software. SW Solution Engineers spend approximately 50% of their time helping customers with their most complex problems and 50% of their time doing R&D related work. This individual should have proven grasp of datacenter and networking technologies, to provide comprehensive solutions for complex installations, maintenance, or operations for a broad scope of leading-edge networking products. What you'll be doing: Take ownership and drive customer issues with Ethernet or InfiniBand network adapter/DPU deployments from inception to resolution. Develop features and tools as part of solution engineering efforts to support all Enterprise Service offerings including but not limited to Networking products. Work with NVIDIA Enterprise customers and internal users to improve the availability, reliability, and overall experience of working with NVIDIA Networking products. Bring independent analysis, communication, and problem-solving to customer experience. Collaborate with engineering to document, recreate and solve issues. What we need to see: BSc in Computer Science, Electrical Engineering, Computer Engineering, or related field (or equivalent experience). 8+ years system software developm

REMOTEkuberneteslinux
View job →
N
10 days ago

We are seeking a highly skilled and hard-working Senior Test Developer / test engineer to join our multifaceted Enterprise Software QA team. This role offers an outstanding opportunity to leave your mark on the design, construction, optimization and testing of large-scale infrastructure for various foundational NVIDIA unified cloud services and data center offerings. If you are a dedicated engineer with strong expertise in cloud infrastructure and distributed systems and want to apply your skills with AI tools, this role could fit you perfectly. You will thrive in an exciting, innovative environment. What you'll be doing: Work with development teams on test plans for all layers of SW stack for cloud infrastructure, execution, reviews, failure analysis and assessing overall quality and risk. Work with customer PMs on software issues including technical feedback from OEMs and CSPs. Develop key benchmarks to track execution and deploy process improvements to improve efficiency Leverage AI skills to expedite the test scope, test plan, execution and automation workflows. Lead NVIDIA Cloud and Data Center bring up activities which will involve validation, reporting, working with engineering to debug issues, providing design input at times, adding coverage in different areas. Design, develop and maintain CI/CD pipelines for continuous testing in cloud environments when needed. Perform performance, scalability, and reliability testing of cloud services. Implement and maintain test environments in cloud platforms such as AWS, Azure, or Google Cloud. Supervise the infrastructure to alert on significant events, ensuring the highest level of system performance and reliability. Work with various different partner teams to ensure availability of clusters to test on and take the lead in resolve all issues. Working with tea

awsazuredocker
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Everpure Cloud Azure Native is a generally available service that brings enterprise-grade block storage natively to Azure. With the first version already live, we are expanding the service into new use cases and markets while improving its reliability, operability, and customer experience. Our Foundation team owns a Go-based control-plane service that coordinates how customers provision and use storage in Azure. Around it, we work with a modern cloud stack including Temporal and other platform services for workflows, automation, and observability. This is a production cloud service: the code you write directly shapes how customers deploy, scale, and operate storage in their Azure environments. You’ll work on a core storage service in a major public cloud , as part of a joint effort between Everpure and Microsoft. You’ll design and evolve APIs and service behavior in the critical path of real customer workloads, collaborating closely with engineers across both companies. Clear API contracts, long-lived interfaces, test automation, and CI/CD are fundamental to how we build. You’ll have the opportunity to own services end to end and solve complex distributed-systems problems in the public cloud. WHAT YOU'LL DO Own and evolve a production cloud service that powers Everpure Cloud Azure Native, taking features from idea and design through deployment and operation in Azure for real customers. Build new capabilities and i

pythonjavaaws
View job →

DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. About the job DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides an in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. In this role, you will Build core capabilities for our SaaS Platform across multiple clouds Drive development of functional enhancements for Data Discovery, Observability & Governance for both OSS and SaaS offering Lead efforts around non functional aspects like performance, scalability, reliability Lead and mentor junior engineers Work closely with PM, Customers and OSS community Requirements Over 8+ years of experience building and scaling backend systems, preferably in cloud-first or SaaS environments. Solve complex tech

javaci/cdrest
View job →
A
Affirm
📍 Poland• Full-time• Remote• $384K – $576K/yr
15 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions

REMOTEpythonsqlmysql
View job →
A
Affirm
📍 Spain• Full-time• Remote• From €1M/yr
15 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions

REMOTEpythonsqlmysql
View job →

DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You’ll Do: We are looking for a Senior Software Engineer – Platform Operations based in Pune, India, who will play a key role in ensuring the reliability, performance, and operational excellence of DeepIntent's platform and data ecosystem. This role requires a strong engineering mindset with the ability to troubleshoot complex technical issues, understand distributed data architectures, and collaborate across Engineering, Product, Analytics, and Customer-facing teams to deliver timely and effective solutions. As part of the Operations organization, you will work closely with Engineering to support production systems, improve operational processes, and drive platform stability. The ideal candidate is a self-motivated problem solver who is passionate about learning new technologies, improving system reliability, and delivering exceptional customer outcomes through engineering excellence. Serve as the engineering interface between Customer-facing teams, Analytics, Product, and Engineering organizations. Partner with Platform Support, Client Success, and other customer-facing teams to investigate and resolve complex platform-related issues. Analyze application, API, and data pipeline issues to identify root causes and drive timely resolution. Develop and standardize operational tools, and interfaces to support analytical and operational use cases. Monitor data pipeline executions, investigate failures, and implement corrective and pre

pythonjavasql
View job →
DC
15 days ago

Role Overview Build reliable software services that power products, platforms, and business decisions. As a Senior Software Developer, you’ll design and deliver scalable applications, backend services, and integrations that perform well in production and evolve with changing business needs. You’ll apply strong software engineering practices across APIs, data-intensive applications, cloud services, AI-enabled solutions, and deployment pipelines. You’ll help shape technical solutions, improve system reliability, and contribute to a high-quality engineering culture. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design and develop scalable backend services and applications using Python or TypeScript. Lead the development of APIs, integrations, reusable software components, and AI-enabled features. Build reliable solutions for data ingestion, manipulation, service-to-service communication, and intelligent automation. Apply AI technologies and modern software engineering practices to improve product capabilities, developer productivity, and operational efficiency. Make sound technical decisions around architecture, performance, security, scalability, and maintainability. Deploy and operate applications using AWS services and CI/CD practices while improving testing, monitoring, documentation, and delivery standards. These are the essentials you’ll need to get an interview 5+ years of professional experience developing and delivering production software. Strong hands-on experience with Python; TypeScript or similar languages is also valuable. Proven experience building backend services, APIs, integrations, and service-oriented applications. Experience applying AI technologies, such as generative AI, machine learning services, intelligent automation, or AI-enabled application features. Strong understanding of software design principles, testing, debugging, performance optimization, and secure development. Experience working with cloud platfor

typescriptpythonaws
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $180K/yr
15 days ago

Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform providing APIs for knowledge retrieval, inference, evaluation, and more. We are seeking a strong Senior Full-Stack Engineer to help us build, scale, and refine our rapidly growing product. The ideal candidate is deeply grounded in software engineering best practices and experienced in developing and scaling modern web applications end-to-end. You will work across the stack—from React/TypeScript frontends to Python-based backends—while integrating with LLMs and machine learning systems. You will solve complex challenges in scalability, reliability, and product experience while owning significant product areas in a fast-paced environment. What You’ll Do Own major full-stack product areas , driving features from design through production deployment. Build modern frontend experiences using React and TypeScript, ensuring performance, usability, and responsiveness. Develop reliable backend services in Python, working with distributed systems, data pipelines, and ML/LLM components. Integrate with LLMs, vector databases, and AI infrastructure to power intelligent product experiences. Deliver experiments and new features quickly , maintaining high quality and tight feedback loops with customers. Collaborate across product, ML, and infrastructure teams to shape the direction of Scale GP. Adapt quickly —learning new technologies, frameworks, and tools as needed across the stack. Ideal Experience 5+ years of full-time engineering experience , post-graduation. Strong experience developing full-stack applications using React, TypeScript, and Python . Experience scaling or shipping products at high-growth startups . Familiarity with LLMs, vector databases, embeddings, or other modern AI tooling (tinkering or production experience welcome). Proficiency with SQL and modern API development. Experience with Kubernetes , containerization, and microservice architectures. Experience working with at leas

typescriptpythonreact
View job →
C-
CLEAR - Corporate
📍 New York• Full-time• $225K – $300K/yr
15 days ago

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Software Engineer, Data, you will design, build, and operate the next generation of our data platform and products – going beyond ID to power a networked digital identity – while keeping member privacy, security, and reliability at the core. What you’ll do: Build and operate scalable, reliable data systems and pipelines – from ingestion to modeling to visualization – so Analysts and Engineers can self-service changes in an automated, tested, secure, and high-quality manner. Develop and maintain end-to-end data products and pipelines (batch and/or streaming) that collect, clean, transform, and model data, and own the infrastructure that powers them to unlock new business use cases and reporting. Implement and maintain infrastructure-as-code, CI/CD, and shared developer tooling for data products (e.g., Pulumi/Terraform, GitHub, orchestration tools like Dagster/Airflow) to make it easy and safe for teams to build, test, and ship changes across environments. Improve the security, compliance, and cost posture of the data stack through robust dependency management, IAM and secrets hardening, observability, and performance/cost optimizations. Partner with product and other stakeholders to uncover requirements, make architectural decisions, and continuously improve our data platform and processes. How you’ll measure success: Data reliability & SLAs: % successful pipeline runs, adherence to freshness SLAs for core datasets, and reduction in data-related incidents impacting stakeholders. Platform quality & efficiency: Reductio

pythonsqlaws
View job →
C-
CLEAR - Corporate
📍 New York• Full-time• $175K – $215K/yr
15 days ago

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We’re looking for a Senior Software Engineer to join our Network Engineering team to accelerate building and scaling our innovative systems that support our growing identity platform. In this role, you will build the next-generation infrastructure that underpins all systems at CLEAR. The ideal candidate for this role will approach challenges with an eye toward reliability, simplicity, and scalability. What You'll Do: Develop and maintain a streamlined process for engineers to effortlessly build and deploy scalable and reliable software-defined networking solutions on AWS. Enhance our compute platform (Kubernetes) by integrating new functionalities and features, focusing on AWS networking services and concepts such as VPCs, Route Tables, Security Groups (SGs), ALBs/ELBs, and Route53, as well as implementing Kubernetes networking solutions like service mesh (Istio) to optimize service communication and management. Collaborate across engineering teams to advocate for and implement best practices in observability, utilizing tools like Splunk or Datadog to ensure robust network monitoring. Act as a product owner for our infrastructure, collecting feedback and requirements from engineering teams to address pain points and develop solutions, particularly in the realm of AWS networking and cloud-native design principles. What you're great at: 6+ years of extensive experience in infrastructure and platform development, particularly in software-defined networking and AWS cloud services. Proficient in writing production-grade softwar

pythonawskubernetes
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng

pythonsqlpostgresql
View job →
S
Stripe
📍 New York• Full-time• $190.4K – $285.6K/yr
29 days ago

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Lead the technical design and architecture of major platform initiatives, author design documents and build consensus across engineering teams. Define technical roadmaps for complex, multi-quarter projects that span multiple teams. Make critical architectural decisions for company documentation infrastructure, balancing scalability, reliability, and developer experience. Evaluate and set direction for integrating emerging technologies, including AI/LLM capabilities, into company documentation platforms and authoring tools. Establish and evolve engineering standards, best practices and technical guidelines for the team and broader organization. Partner with engineering teams across the company to understand documentation needs and design integrated solutions. Design, build and maintain scalable, reliable and performant services and systems. Contribute high-quality code across the full stack and navigate codebases with different languages and tools. Debug and resolve complex production issues and improve system reliability. Take ownership of system health and incident response. Who you are Minimum requirements Must have a Bachelor's degree or foreign equivalent in Computer Science, Software Engineering, Engineering, or a related field, plus four (4) years of experience in Software Engineering. Must have four (4) years of experience in each of the following: - Working in a full stack environment with a foc

typescriptjavamongodb
View job →
O
Okta
📍 San Francisco• Full-time• From $165K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in a global on-call rotation supporting highly available customer-facing systems. Participate in incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with en

pythonsqlpostgresql
View job →
🔔

Get new senior software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime