Jobiba hiring network

Software Reliability Engineer Jobs

6,326 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold builders and sharp problem-solvers who are wired to deliver great outcomes. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. The DevX team’s mission is to build and operate the core developer infrastructure at Robinhood. Our team owns and scales the systems that thousands of engineers rely on daily, partnering with software developers across the company to make development fast, reliable, and cost-efficient! As a Staff Software Developer, you will act as a technical leader for our build and developer infrastructure, driving the strategy and execution of the systems thousands engineers depend on every day. Your work will span our build systems, CI pipelines, and remote development environments, ensuring engineers can code, test, and build with speed, safety, and reliability at scale. In this role, you will collaborate with teams across Robinhood to eliminate developer friction and raise the bar for engineering productivity. This is a high-visibility leadership opportunity to shape our developer ecosystem and set new standards of engineering efficiency! This role is based in our Toronto, ON office(s), with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. What you’ll do Architect the long-te

pythonawsci/cd
View job →
S
Stripe
📍 Taipei• Full-time
1mo ago

About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe Terminal helps Stripe users extend their online presence into the physical world. The Terminal team’s mission is to make it as easy for businesses to accept in-person payments as the Stripe API has done for online payments. With Terminal, businesses can unlock in-person payments use cases that are right for their business model—whether it’s creating a flagship retail experience, extending their website to a pop-up store, or enabling a mobile point-of-sale at their next event. What we are looking for: As an Android BSP Engineer, you will be responsible for the kernel and driver level system development, which includes building, troubleshooting, and writing automated tests for the Android system on our embedded payments platforms. This team works closely with partner teams throughout the hardware and software product lifecycle, from hardware manufacturing to Android app teams. We also work with external vendors on part selection and initial hardware bring-up. What you’ll do: Bring up new devices and lead debugging and performance tuning exercises that span multiple hardware/firmware/software teams. Design, implement, and maintain drivers and Android services that operate efficiently in a constrained environment and meet the reliability and security requirements of the industry. Own the definition of one or more work streams focused on hardware bring-up, peripheral drivers and communication, and power and performance management and opti

javaaikotlin
View job →
S
Snowflake
📍 Menlo Park• Full-time
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake is expanding the boundaries of the Data Cloud to support mission-critical transactional workloads. Our goal is to deliver OLTP capabilities with the performance, reliability, simplicity, and scale customers expect from Snowflake, while creating a seamless experience across transactional and analytical data. We are looking for a Senior Engineering Manager – OLTP to lead engineering teams and leaders building core transactional database technology and the cloud infrastructure required to operate it at scale. You will help define the architecture and roadmap, grow the organization, and drive technology from design through production. AS A SENIOR ENGINEERING MANAGER – OLTP AT SNOWFLAKE, YOU WILL: Set technical and execution strategy for key areas of Snowflake's OLTP platform, translating product goals into architecture, roadmaps, and team plans. Lead and grow multiple engineering teams, developing managers and senior technical leaders while fostering a culture of ownership, technical excellence, and execution. Drive adoption of AI and agentic development practices to improve engineering velocity, quality, and productivity across the software development lifecycle. Drive architecture and technical decisions in areas such as transactions, concurrency control, low-latenc

awsazuregcp
View job →

Job Details: Job Description: Join an enthusiastic team of engineers in Intel's Networking Solutions Group (NSG) focused on enabling next generation of programmable Infrastructure Processing Units (IPUs) with our lead customers as part of the Customer Experience Support (CES) organization. Intel brings decades of leadership in networking, virtualization, packet processing, storage, and security to a new class of IPU products that accelerate host networking functions and support emerging use cases such as security, virtualization, storage, load balancing, and data path optimization. Working closely with major cloud service providers and Intel development teams, you will help deliver customized IPU based solutions that enhance isolation, security, performance, storage and system management for our customers. A big part of the day-to-day job is to help customers manage feature request processes, enable solutions, and debug issues. Projects and responsibilities include but are not limited to: • Gain our customers' trust, understand their needs, and build POCs to meet them. Work closely with internal and external partners to understand use cases and requirements. • Be the go-to technical resource for customers building complex Datacenters, AI infrastructure as well as helping them understand performance characteristics for solutions. • Prepare and deliver technical content to customers including presentations, workshops, etc. • Contribute across the full IPU lifecycle, including board and platform bring up, low-level device initialization, OS driver and kernel configuration, system management, feature enablement, use case testing, debugging, and verification. • Defines systems implementation and integration solutions and plans to ensure optimum performance and reliability across hardware, firmware and software w

dockergitlinux
View job →
N
Nvidia
📍 Santa Clara, United States
12 days ago

Are you ready to contribute to world-class innovation and push the boundaries of what's possible? At NVIDIA, you'll have the opportunity to be part of a team that is driving groundbreaking impacts across various markets. As a Thermal Solutions Development Engineer, you will play a pivotal role in our Silicon Codesign Group, transforming thermal solution concepts into lab-ready builds and beyond. What you will be doing: Build thermal solutions for engineering characterization and validation of next-gen GPU/SOC products, ensuring flawless delivery from concept to lab. Drive end-to-end development and deployment of thermal solutions, collaborating with internal teams and external vendors on build requirements, prototype evaluation, test system integration, and software automation. Improve thermal design processes by incorporating feedback and findings, developing workflow and maintaining our world-class standards. Work closely with system architects, chip and board designers, and software/firmware engineers in a dynamic and high-energy environment to bring industry-defining products to market. Apply AI-enabled approaches and AI tools to accelerate design iteration, test planning, and characterization/validation triage (e.g., requirements/spec summarization, experiment prioritization, log/telemetry summarization, anomaly/outlier detection), improving cycle time, coverage, and traceability while validating outputs against physics, specs, and lab measurements. Partner with AI/tooling teams as the thermal domain SME to define use-cases, success criteria, and evaluation methods; provide feedback to improve tool reliability and usability. What we need to see:

N
14 days ago

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated

pythonkuberneteslinux
View job →
E
Everpure
📍 Prague• Full-time
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE FlashArray File extends FlashArray with file services, supporting SMB and NFS through core components including the file gateway, object store, and middleware. We are looking for a C++ developer who is interested in building high-performance, highly reliable systems and solving complex technical problems across file protocols, storage, concurrency, quality, and performance. This team owns core file capabilities end to end, with a strong focus on feature delivery, maintenance, and efficient execution in a highly parallel, customer-critical environment. WHAT YOU’LL DO Design, implement, and review new features in FlashArray File, primarily in C++, across core file-services areas such as protocols, reliability, and customer-facing capabilities. Build system software that serves SMB and NFS workloads and helps evolve the file gateway and adjacent services behind FlashArray File. Improve quality, scalability, performance, and resiliency for a platform that handles highly concurrent access patterns and business-critical storage workloads. Take ownership of work from design through delivery, debugging, and long-term maintenance, helping move projects forward from start to finish. Collaborate closely with engineers and cross-functional partners in Prague and beyond to deliver high-quality outcomes and continuously improve how the team works. We are primarily an in-office environment and therefore, you will be expec

awslinuxrest
View job →
E
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As the Engineering Manager for Self-Services Upgrades (SSU) team, you will lead the engineering team responsible for Everpure’s™ core upgrade platform across FlashArray and FlashBlade ecosystems. You will own the strategy and execution that enables safe, zero-downtime software upgrades and mitigations in both connected cloud and air-gapped darksite environments. By bridging distributed backend services, on-box workflows, and cloud infrastructure, your team directly safeguards customer trust and fleet reliability. Through high-visibility cross-geo, cross-business unit, and cross-functional partnerships with Product Management, Release Engineering, Technical Services, and Fleet Insights, you will drive modern upgrade experiences for mission-critical enterprise environments globally. WHAT YOU'LL DO Drive Strategic Upgrade Delivery: Lead end-to-end engineering across backend services, REST APIs, cloud infrastructure, on-box workflows, and release tools to deliver safe, non-disruptive software upgrades for FlashArray and FlashBlade platforms. Lead Cross-Functional Alignment: Collaborate with Product Management, Release Engineering, Technical Services, and Forensics to evaluate technical trade-offs, align roadmaps, and streamline deployment workflows across shared platform stakeholders. Build High-Performing Teams: Hire, onboard, and mentor a high-performing engineering team, cultivating an inclusive culture of t

awsrestagile
View job →
E
Everpure
📍 Bengaluru• Full-time
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... In this role as an Engineering Manager, you will lead a team of engineers located in Bengaluru, India. You will focus on driving and shaping the direction of our observability software and enabling product engineers to deliver high-quality, reliable software to our customers. You will provide technical leadership and direction, mentor engineers on your team, and collaborate with product, engineering, and cross-functional stakeholders to deliver successful outcomes. The successful candidate must understand the dynamics of global R&D, possess deep knowledge of local culture, and have the ability to champion Pure values and leadership attributes. This role requires the ability to lead and influence multiple stakeholders across cross-functional teams and drive alignment across complex, distributed engineering environments. The team will help build and evolve observability capabilities that provide actionable insights into the health, performance, capacity, and reliability of Pure's products and infrastructure. WHAT YOU'LL NEED TO BRING TO THIS ROLE... 12+ years of combined experience as a software developer and manager 3+ years of technical management experience while staying hands-on 7+ years of hands-on software development experience Strong exposure to one or more of the following areas: distributed systems, systems programming, observability/telemetry, data platforms, or solving prob

awsrestai
View job →
F
17 days ago

About Forma.ai: Forma.ai is a Series B startup that's revolutionizing how sales compensation is designed, managed and optimized. We handle billions in annual managed commissions for market leaders like Edmentum, Stryker, and Autodesk. Our growth has been fuelled by our passion for fundamentally changing and shaping how companies use sales intelligence to drive business strategy. We’re welcoming equally driven individuals who are excited about creating something big! Senior Staff Backend Engineer About the Team We build enterprise software that helps organizations optimize sales performance, enabling go-to-market agility. Our engineering organization includes multiple product application teams responsible for delivering core customer-facing capabilities. We are seeking a Senior Staff Backend Engineer to join our application teams and help set technical direction across multiple domains within engineering. You'll work alongside staff, senior, and early-career engineers, and partner closely with engineering leadership to define, evolve, and scale the systems that power enterprise-grade product workflows. This is an opportunity to own complex, multi-domain technical problems and shape product direction beyond a single team. We are low on meetings, high on accountability. Most of the teams are in the EST time zone, but we have a few located in AST, PST, and Central as well. What you'll be doing You will play a pivotal role in shaping the technical direction of our application stack across multiple domains. You will lead development efforts for our most complex initiatives, the kind that span two or more teams or product areas, and serve as a technical benchmark for system design, code quality, and long-term maintainability. You'll operate at the intersection of data modelling, business logic, and enterprise-scale reliability, and your work will often set standards that neighboring teams adopt. This remains a hands-on

javascripttypescriptpython
View job →
B
Bloomreach
📍 Slovakia• Full-time• From €42K/yr
17 days ago

Bloomreach is building the world’s premier agentic platform for personalization .We’re revolutionizing how businesses connect with their customers, building and deploying AI agents to personalize the entire customer journey. We're taking autonomous search mainstream, making product discovery more intuitive and conversational for customers, and more profitable for businesses. We’re making conversational shopping a reality, connecting every shopper with tailored guidance and product expertise — available on demand, at every touchpoint in their journey. We're designing the future of autonomous marketing , taking the work out of workflows, and reclaiming the creative, strategic, and customer-first work marketers were always meant to do. And we're building all of that on the intelligence of a single AI engine — Loomi — so that personalization isn't only autonomous…it's also consistent.From retail to financial services, hospitality to gaming, businesses use Bloomreach to drive higher growth and lasting loyalty. We power personalization for more than 1,400 global brands, including American Eagle, Sonepar, and Pandora. About the Team We're a small, high-impact engineering team embedded inside the Go-To-Market organization. Our job is to make Bloomreach impossible to ignore in a sales cycle by using AI to build demo environments, automate work, and create self-serve tools for presales teams around the world. We move fast, experiment constantly, and build things that end up in front of some of the world's largest retailers and brands. We care about craft, speed, reliability, and whether the field can actually use what we ship. The Role Bloomreach is seeking an AI Demo Engineer who is AI-native by instinct: a flexible software engineer who reaches for AI tools first, learns quickly, and uses technology to solve whatever problem is blocking the field. This is not a traditional engineering role. You are not maintaining produc

javascripttypescriptpython
View job →
B
Bloomreach
📍 Czech Republic• Full-time
17 days ago

Bloomreach is building the world’s premier agentic platform for personalization .We’re revolutionizing how businesses connect with their customers, building and deploying AI agents to personalize the entire customer journey. We're taking autonomous search mainstream, making product discovery more intuitive and conversational for customers, and more profitable for businesses. We’re making conversational shopping a reality, connecting every shopper with tailored guidance and product expertise — available on demand, at every touchpoint in their journey. We're designing the future of autonomous marketing , taking the work out of workflows, and reclaiming the creative, strategic, and customer-first work marketers were always meant to do. And we're building all of that on the intelligence of a single AI engine — Loomi — so that personalization isn't only autonomous…it's also consistent.From retail to financial services, hospitality to gaming, businesses use Bloomreach to drive higher growth and lasting loyalty. We power personalization for more than 1,400 global brands, including American Eagle, Sonepar, and Pandora. About the Team We're a small, high-impact engineering team embedded inside the Go-To-Market organization. Our job is to make Bloomreach impossible to ignore in a sales cycle by using AI to build demo environments, automate work, and create self-serve tools for presales teams around the world. We move fast, experiment constantly, and build things that end up in front of some of the world's largest retailers and brands. We care about craft, speed, reliability, and whether the field can actually use what we ship. The Role Bloomreach is seeking an AI Demo Engineer who is AI-native by instinct: a flexible software engineer who reaches for AI tools first, learns quickly, and uses technology to solve whatever problem is blocking the field. This is not a traditional engineering role. You are not maintaining produc

javascripttypescriptpython
View job →
V
VTS
📍 New York• Full-time• From $170K/yr
17 days ago

Join VTS as an Engineering Manager. As a technical leader, you'll drive high-impact, customer-facing projects across major features in web and mobile. You’ll play a pivotal role in advancing VTS's evolution into an AI-powered platform, further cementing our industry-leading position. In this role, you will be responsible for fostering a healthy and collaborative culture, guiding the team's execution and delivery, and mentoring engineers to grow their careers and capabilities. ** Please note that this opportunity is located in New York, NY, and requires this hire to work from our office 4 days a week. ** To thrive in this role, you have: Proven experience leading software teams using agile development methodologies. A strong technical background with a proven track record of building, deploying, and maintaining scalable web applications, services, and third-party integrations. Experience working on high-impact, business-critical domains, where reliability, performance, and operational stability are essential. Demonstrated ability to recruit and develop talent, strengthening the team's capabilities through effective coaching and mentorship. The ability to engage in strategic technical discussions, adeptly manage technical tradeoffs and risks, and ensure alignment with product and business goals, especially in environments involving multiple systems and external partners. A sense of empathy for the customer and a relentless focus on shipping high-quality products. Excellent communication skills that empower you to effectively convey ideas, coach, mentor, and educate others. A keen interest in building AI-backed features and leveraging AI as an engineering productivity tool. What you'll do: Lead and Develop a High-Performing Team : Foster a healthy, collaborative, and diverse culture that reflects our company and engineering values. Coach and mentor individual team members, creating a structured environment and feedback loop that supports career development and hi

typescriptreactaws
View job →
G
17 days ago

Job Summary Reporting to the Memory Validation leadership team, the Senior Silicon DDR/HBM Validation Engineer will be responsible for the bring-up, validation, characterization and debug of advanced memory subsystems used in next-generation AI compute platforms. The role will focus on DDR and HBM technologies, working closely with silicon design, firmware, characterization, platform and systems teams to ensure robust memory subsystem functionality, performance and reliability. The successful candidate will take ownership of significant validation activities, contribute to debug and root-cause analysis efforts, and help improve validation methodologies, automation and infrastructure. The Team The Memory Validation team sits within the Validation organisation and is responsible for the bring-up, validation, characterization and debug of memory subsystems across Graphcore silicon and platform products. The team supports DDR and HBM validation activities throughout the product lifecycle, from first silicon through production readiness. Engineers work closely with architecture, RTL, firmware, characterization, systems and platform teams to ensure memory technologies meet functionality, performance, reliability and performance objectives. Responsibilities and Duties Execute validation and bring-up activities for DDR and HBM memory subsystems Verify memory bring-up software, firmware and scripts against defined project requirements Debug firmware, hardware and system-level issues and contribute to root-cause analysis activities Analyse system logs, validation data and characterization results to identify failures and performance issues Perform PHY characterization and analog-level analysis during stress testing and validation activities Develop and execute functional, stress, performance and corner-case validation tests Perform signal integrity, voltage, frequency and timing measurements using laboratory instrumentation Char

pythonaigo
View job →
C-
CLEAR - Corporate
📍 New York• Full-time• $275K – $350K/yr
17 days ago

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We are seeking a strategically-minded, technology-focused, and customer-centric Engineering Manager to lead one of our Infrastructure teams here. You will lead a team responsible for building, operating, and scaling the cloud infrastructure and platform systems that underpin CLEAR’s services, ensuring reliability, performance, and security across our environments. A successful candidate brings strong experience in cloud infrastructure, distributed systems, and operational excellence, along with a solid foundation in software engineering. You are an effective communicator who can lead complex infrastructure initiatives from inception through delivery, and thrive in fast-paced environments. This role requires a focus on building resilient, scalable systems, driving automation, and leading and developing high-performing engineering teams. What you'll do: Hire, develop, and grow engineering talent through coaching, mentorship, performance management, and career development planning Set clear goals and expectations, provide regular feedback, and foster accountability across the team Own and execute the roadmap for cloud infrastructure and platform engineering, and reliability initiatives Design, build, and operate a scalable, secure, and highly available cloud platform infrastructure Drive automation across infrastructure provisioning, deployment, and operations to improve efficiency and reduce manual overhead Establish and enforce best practices for system reliability, observability, incident response, and disaster recovery Partner with eng

pythonjavaaws
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime