Jobiba hiring network

Software Reliability Engineer Jobs

6,326 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

kubernetesaigo
View job →

About Ubiquiti At Ubiquiti Inc., we create technology platforms for Businesses, Smart Homes, and Internet Service Providers, driven by our goal to connect everyone, everywhere. To date, Ubiquiti has shipped over 100 million devices worldwide, from ISP networking products to next generation of IT solutions. Our growth is made possible by the dedicated team of hundreds behind the scenes. From software developers and product managers to designers and strategists, Team UI is driven to achieve our common goal: Rethinking IT. At Ubiquiti, you’ll heighten your potential and broaden your horizons - all while shaping the future of connectivity. Responsibilities (What he/she will do after join UI): Design and build the core synchronization and metadata handling logic for UniFi Drive on macOS, enabling reliable and intuitive access to files across local and cloud environments. Build platform-integrated file experiences by working with macOS system behaviors and APIs, with a focus on correctness, performance, and meeting native user expectations. Tackle non-trivial system problems that arise in real-world usage, such as handling change, coordination, and edge cases across devices and environments. Collaborate closely with backend and infrastructure teams to align on data models, APIs, and end-to-end reliability. Work with product and design partners to translate user workflows into robust technical solutions, balancing system constraints with user experience. Continuously improve performance, stability, and maintainability through measurement, profiling, and iterative refinement. Minimum Qualifications (MUST-haves) : 3+ years of professional software development experience, or equivalent experience, building and shipping production systems. Hands-on development experience on Apple platforms (macOS and/or iOS), with familiarity with the Apple development ecosystem and real-world production workflows. Strong computer science fundamentals, with the ability to reason cl

gitswift
View job →

About Ubiquiti At Ubiquiti Inc., we create technology platforms for Businesses, Smart Homes, and Internet Service Providers, driven by our goal to connect everyone, everywhere. To date, Ubiquiti has shipped over 100 million devices worldwide, from ISP networking products to next generation of IT solutions. Our growth is made possible by the dedicated team of hundreds behind the scenes. From software developers and product managers to designers and strategists, Team UI is driven to achieve our common goal: Rethinking IT. At Ubiquiti, you’ll heighten your potential and broaden your horizons - all while shaping the future of connectivity. Responsibilities (What he/she will do after join UI): Design and build the core synchronization and metadata handling logic for UniFi Drive on macOS, enabling reliable and intuitive access to files across local and cloud environments. Build platform-integrated file experiences by working with macOS system behaviors and APIs, with a focus on correctness, performance, and meeting native user expectations. Tackle non-trivial system problems that arise in real-world usage, such as handling change, coordination, and edge cases across devices and environments. Collaborate closely with backend and infrastructure teams to align on data models, APIs, and end-to-end reliability. Work with product and design partners to translate user workflows into robust technical solutions, balancing system constraints with user experience. Continuously improve performance, stability, and maintainability through measurement, profiling, and iterative refinement. Minimum Qualifications (MUST-haves) : 3+ years of professional software development experience, or equivalent experience, building and shipping production systems. Hands-on development experience on Apple platforms (macOS and/or iOS), with familiarity with the Apple development ecosystem and real-world production workflows. Strong computer science fundamentals, with the ability to reason cl

gitswift
View job →
E
13 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE We are seeking an experienced Platform Software Engineer for our Systems Software Team. You will be working as part of a dynamic team and will be responsible for designing, developing, and testing system software functionality for Pure’s upcoming platforms. The work spans the gamut of Systems software and you will have the opportunity to work in a wide range of areas and features ranging from Platform drivers to networking and storage layers. WHAT YOU'LL DO Plan and influence the lifecycle of new Hardware Platforms. Work on problems ranging from design, bring up, to deployment, upgrades and fleet level reliability. Participate in the full lifecycle of new hardware platforms from early bring up through manufacturing release. Work closely with peer teams to debug complex HW/FW of new server hardware, including CPUs, chipsets, and peripheral components. Debug complex HW/FW issues across x86, PCIe, NVMe, and networking using lab tools (oscilloscope, logic analyzer, JTAG) and kernel/driver traces. Design, implement and improve remote server management capabilities (e.g., using standards like Redfish) and enhance Reliability, Availability, and Serviceability (RAS) features. Design, write and maintain software components in C/C++, Python, Golang and RUST. Collaborate with vendors on requirements specification and follow through to system delivery. Work closely with hardware engineers, system architects, and o

pythonlinuxai
View job →

Are you looking for an opportunity to help solve one of today's biggest business challenges? AI is changing the pace of business, and organizations everywhere are struggling to help their workforce, partners, and customers keep up. At Litmos, we're building the Learning Acceleration Platform that helps organizations build human capability faster—and we're looking for people who are passionate about making a meaningful impact for customers while growing alongside a collaborative, people-first team. Litmos is the Learning Acceleration Platform that helps organizations build capability faster, adapt at the speed business changes, and scale learning to anyone, anywhere. Combining an intuitive platform, AI-powered capabilities, trusted content, expert services, and a broad ecosystem of integrations, Litmos helps organizations accelerate workforce productivity, improve customer adoption and retention, enable high-performing partners, and reduce organizational risk through continuous learning. Organizations such as Hewlett Packard Enterprise, Graco, Sabre, and Russell Mineral Equipment trust Litmos to accelerate learning across their workforce, partners, and customers. Today, more than 11K customers with 30 million learners across 150 countries and 37 languages use Litmos to build the capabilities their organizations need to succeed. Backed by Francisco Partners, one of the world's leading technology investment firms, we're investing in the future of learning—and the people who are building it. Learn more at www.litmos.com . We are looking for a Senior Full Stack Engineer to build and evolve our platform using .NET, React, and SQL Server-based systems. You will work across services, APIs, and front-end systems with a focus on scalability, reliability, and maintainability. You will operate within a globally distributed Agile team (Scrum and ShapeUp), where engineers own delivery, from design through deployment, while collaboratin

reactsqlazure
View job →

The NVIDIA PerfTech team is looking for a talented C++ Software Engineer to help build the next generation of AI-powered developer tools. You will apply strong C++ and software-engineering fundamentals while gaining hands-on experience with agentic workflows, retrieval systems, and AI services. In this role, you will contribute to Genie, NVIDIA’s company-wide AI knowledge and developer-productivity service. You will work across C++ tools and AI services to help engineers find information, understand complex systems, and work more effectively. What You’ll Be Doing: Develop production-quality C++ components, APIs, and integrations for NVIDIA’s AI-powered developer-tools ecosystem. Build capabilities connecting native C++ tools with Genie’s retrieval and agentic features. Contribute to agentic workflows, retrieval systems, ingestion pipelines, MCP tools, APIs, and enterprise integrations. Build benchmarks and improve retrieval quality, reliability, performance, and resource usage. Own features from investigation and design through implementation, testing, and delivery. Collaborate with graphics, software, and hardware teams developing performance-analysis and developer tools. What We Need to See: Bachelor’s or Master’s degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience. 5+ years of modern C++ programming skills gained through professional experience, internships, or substantial technical projects. Good understanding of data structures, algorithms, object-oriented design, multithreading, debugging, and testing. Ability and motivation to work across C++ systems and Python-based AI services. Familiarity with AI-powered applications, agentic workflows, retrieval systems, or related technologies. Abil

pythonaic++
View job →
A
14 days ago

This is Adyen Adyen provides payments, data, and financial products in a single solution for customers like Meta, Uber, H&M, and Microsoft - making us the financial technology platform of choice. At Adyen, everything we do is engineered for ambition. For our teams, we create an environment with opportunities for our people to succeed, backed by the culture and support to ensure they are enabled to truly own their careers. We are motivated individuals who tackle unique technical challenges at scale and solve them as a team. Together, we deliver innovative and ethical solutions that help businesses achieve their ambitions faster. Team Lead - Software Engineer As a Software Engineering Team Lead based in Singapore, you will lead a team of highly skilled software engineers responsible for building and evolving Adyen's Global Cards platform for the APAC region. Your team plays a critical role in developing scalable, reliable, and resilient payment capabilities that power card transactions for merchants across APAC. You'll work closely with Product Managers, Architects, and Engineering teams to deliver new functionality across the Cards domain, while continuously improving the performance, scalability, and reliability of our platform. As a people leader, you'll coach and develop engineers, foster a culture of ownership, and help shape the technical direction of one of Adyen's core payment domains. As part of Adyen's global engineering organisation, you'll collaborate closely with teams across Europe, APAC, and North America to build products that scale globally. What you'll do Lead, coach, and develop a team of Software Engineers, supporting both their technical and professional growth. Foster a high-performing engineering culture built on ownership, collaboration, and continuous learning. Partner closely with Product Managers to translate business priorities into scalable technical solutions. Drive the technical direction of the team, balancing product de

W
Wellhub
📍 Brazil• Full-time• Remote
17 days ago

Your wellbeing, our mission. Join a company shaping a healthier world. GET TO KNOW US At Wellhub we're revolutionizing workplace wellness. Our platform connects employees worldwide to the best partners for fitness, mindfulness, therapy, nutrition, and sleep—all in one simple subscription. Headquartered in NYC with team members in Europe, North America and South America, we’re on a mission to make every company a wellness company. We believe work should be fulfilling, inspiring, and balanced. Here, you’ll find a team that values wellbeing, collaboration, and different perspectives, where passion and creativity push boundaries to create real impact. Your contributions will help shape a healthier, more balanced world for you and millions of people globally. Join us in redefining the future of wellbeing! THE OPPORTUNITY We are hiring a Senior Backend Software Engineer to join our Identity Experience (IDX) team in Brazil ! This is a Remote – Brazil position, meaning you can work from anywhere within the country. Please note that this role is only open to candidates in Brazil. We’re looking for an autonomous and resourceful Senior Software Engineer to join our international Identity Experience (IDX) team. The IDX team owns critical services at the heart of our product, including account creation, account management, account recovery, permissions management, and more. In this role, you’ll build features that sit at the spine of the product and power experiences used by millions of users. You’ll tackle complex engineering challenges where performance, reliability, and scalability are essential, while designing and evolving systems built to operate at massive scale. You’ll also play a key role in driving our AI transformation, leveraging your experience with AI-assisted development to raise engineering standards, accelerate delivery, and build seamless, reliable identity experiences. THE OPPORTUNITY Drive and own critical core identity services, including

REMOTEsqlpostgresqlaws
View job →
L
Litmos
📍 India• Full-time• ₹2.6Cr – ₹3.4Cr/yr
17 days ago

Are you looking for an opportunity to help solve one of today's biggest business challenges? AI is changing the pace of business, and organizations everywhere are struggling to help their workforce, partners, and customers keep up. At Litmos, we're building the Learning Acceleration Platform that helps organizations build human capability faster—and we're looking for people who are passionate about making a meaningful impact for customers while growing alongside a collaborative, people-first team. Litmos is the Learning Acceleration Platform that helps organizations build capability faster, adapt at the speed business changes, and scale learning to anyone, anywhere. Combining an intuitive platform, AI-powered capabilities, trusted content, expert services, and a broad ecosystem of integrations, Litmos helps organizations accelerate workforce productivity, improve customer adoption and retention, enable high-performing partners, and reduce organizational risk through continuous learning. Organizations such as Hewlett Packard Enterprise, Graco, Sabre, and Russell Mineral Equipment trust Litmos to accelerate learning across their workforce, partners, and customers. Today, more than 11K customers with 30 million learners across 150 countries and 37 languages use Litmos to build the capabilities their organizations need to succeed. Backed by Francisco Partners, one of the world's leading technology investment firms, we're investing in the future of learning—and the people who are building it. Learn more at www.litmos.com . We are looking for a Senior Full Stack Engineer to build and evolve our platform using .NET, React, and SQL Server-based systems. You will work across services, APIs, and front-end systems with a focus on scalability, reliability, and maintainability. You will operate within a globally distributed Agile team (Scrum and ShapeUp), where engineers own delivery, from design through deployment, while collaboratin

reactsqlazure
View job →

Scale’s rapidly growing Global Public Sector team is focused on using AI to address critical challenges facing the public sector around the world. Our core work consists of: Creating custom AI applications that will impact millions of citizens Generating high-quality training data for custom LLMs Upskilling and advisory services to spread the impact of AI As a Full Stack Software Engineer (Forward Deployed), you’ll collaborate directly with public sector counterparts to quickly build full-stack, AI applications, to solve their most pressing challenges and achieve meaningful impact for citizens. At Scale, we’re not just building AI solutions—we’re enabling the public sector to transform their operations and better serve citizens through cutting-edge technology. If you’re ready to shape the future of AI in the public sector and be a founding member of our team, we’d love to hear from you. You will: Partner with public sector clients to scope, collect feedback and implement solutions for complex problems, including spending up to two weeks per month in client offices for feedback and delivery. Architect production-grade applications that integrate AI models with full-stack frameworks, managing everything from interactive UIs to backend APIs and systems. Deploy and manage infrastructure within cloud environments, ensuring the highest levels of system integrity, security, scalability, and long-term reliability. Contribute to core platform features designed to be reused across diverse international client use cases. Partner with design, product, and data teams to build robust applications aligned with the broader technical architecture. Ideally you’d have: Bachelor’s degree in Computer Science or a related quantitative field 5+ years of post-graduation, full-stack engineering experience with demonstrated proficiency in React (required), TypeScript, Next.js, Python, Node.js, PostgreSQL or MongoDB plus hands-on experience with Docker, Kubernetes, and Azure

typescriptpythonreact
View job →
C-
CLEAR - Corporate
📍 New York• Full-time• $225K – $300K/yr
17 days ago

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We're looking for Senior Fullstack Software Engineers to help build the next generation of CLEAR's identity platform. Beyond verifying identity, we're creating a secure, networked digital identity that enables seamless experiences across travel, enterprise, healthcare, financial services, and beyond. As a Senior Software Engineer, you'll own complex technical problems from design through deployment, partnering closely with Product, Design, Security, and Operations to deliver reliable, scalable solutions. We're looking for engineers with a strong builder mindset who thrive in ambiguity, take ownership, and enjoy turning ideas into production systems. Level and team matching (open roles across the three pillars that make up Technology at CLEAR: Core Identity, CLEAR1 , and CLEAR Travel ) will occur towards the end of our interview process. Tech stack overview: Python / React / Typescript What you’ll do: Design, build, test, and deploy scalable full-stack applications that power CLEAR's identity platform. Own projects end-to-end from technical discovery and architecture through implementation, rollout, and operational support. Partner closely with Product, Design, Data, Security, and Operations to translate business problems into simple, scalable technical solutions. Drive engineering excellence by improving system reliability, performance, testing, observability, and developer experience. Contribute to architectural decisions and continuously improve the scalability, security, and maintainability of our platform. Mentor teammates through t

typescriptpythonreact
View job →
C-
CLEAR - Corporate
📍 New York• Full-time• $225K – $300K/yr
17 days ago

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Fullstack Software Engineer on CLEAR’s Healthcare team, you will build and scale secure, interoperable identity and data solutions that connect patients, providers, and partners. You’ll operate at the intersection of modern web platforms, healthcare interoperability standards, and high-assurance identity systems powering frictionless, trusted healthcare experiences nationwide. A brief highlight of our tech stack: Python / React / Typescript AWS cloud What you’ll do: Design and deliver secure, scalable fullstack solutions that integrate with enterprise EHR systems and national health information exchange frameworks Build and maintain healthcare data integrations leveraging FHIR (RESTful APIs/JSON) and HL7 v2 messaging to enable compliant, real-time data exchange Develop identity resolution and patient matching capabilities using identifiers such as MRNs and NPIs to ensure integrity across disparate clinical systems Partner with Engineering, Security, Product, and Health Information Management teams to implement compliant, audit-ready workflows for regulated healthcare processes Collaborate with external vendors (e.g., Epic Technical Services) to troubleshoot integration issues, manage deployments across TST/PRD environments, and ensure production reliability How you’ll measure success: Successful delivery and stability of FHIR/HL7 integrations across healthcare partners Reduction in data integrity issues related to patient matching and identity resolution High system uptime and successful production deployments across tiered

typescriptpythonreact
View job →
O
24 days ago

About the Team The Emerging Products team is a lean, high-output product lab group that builds products at the forefront of model capabilities. We collaborate across all teams within the company, from research and infrastructure to consumer products. The team is responsible for identifying new product opportunities, building them quickly, dogfooding them internally, and then launching the successful products to users. We use data, user research, and analytics to inform our ideas, and make decisions on what experiments are worth iterating, stopping, or scaling. About the Role We’re looking for a senior, product-minded software engineer to own ambiguous 0-to-1 work from idea through prototype, validation, and handoff. This is a full-stack role with a strong frontend and product emphasis: you will build the interfaces and supporting backend systems needed to test new experiences quickly, while making sound architectural choices that enable successful concepts to scale. This role is based in our Mission Bay office in San Francisco. In this role, you will: Build and ship high-quality, product experiments across the full stack. Turn ambiguous user needs and emerging technical capabilities into testable product concepts, using research and metrics to guide iteration. Own technical direction for 0-to-1 projects, balancing speed, reliability, and a clear path from prototype to scalable product. Partner closely with design, product, research, and engineering teams to dogfood, evaluate, launch, and transition successful experiments. You might thrive in this role if you: Have a track record of building and shipping end-to-end products in fast-moving, startup, founder-led, growth, or other high-ownership environments. Bring strong frontend engineering skills and enough backend and systems depth to make sound full-stack architectural decisions. Pair product intuition with evidence, using user research and product data to identify opportunities and make pragmatic tradeoffs. Operat

awsrestai
View job →
S
Synthesia
📍 Seattle• Full-time• Remote• From $200K/yr
1mo ago

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. About the role You will work on core enterprise platform systems, focusing on services that ensure Synthesia is secure, reliable, and scalable for our largest customers. You will contribute to our new suite of APIs that will unlock our market leading Avatar technology to be utilised by external creative tools. We are an AI native company and use the most powerful assistant tools on a daily basis to increase our speed of development, automate repetitive tasks and widen the scope of what the team can worked on. This includes Claude and Cursor. You will have ownership of projects that span months and multiple teams, requiring you to break down complex, ambiguous problems into clear steps that can be delivered and validated iteratively. Engineers within Synthesia are empowered to contribute heavily to product discussions and build with both user and business considerations in mind. Impact is key to everything we do here. You will work closely with product, security, legal, and infrastructure partners, and will be expected to translate business and regulatory requirements into scalable technical solutions. You will evaluate your work through system health and reliability metrics, leveraging observability and monitori

REMOTEpythonreactaws
View job →
N
1mo ago

NVIDIA is well positioned as the 'AI Computing Company', our GPUs being the brains that power modern Deep Learning software frameworks, accelerated analytics, modern data centers, and driving autonomous vehicles. We are looking for a Senior Software QA Test Development Engineer to join in the mission of crafting a distributed technology for all NVIDIA teams that remotely manage 10s of 1000s of resources in a simple and controlled fashion, allowing engineers to focus on engineering and automation, rather than being burdened by manual operational tasks. SWQA test developer engineers at NVIDIA are responsible for creating test plans, execution, and reporting, as well as developing scripts for test automation, designing and developing tools for the QA team, and developing integration tests for validation. As a test developer, you must identify weak spots and constantly design better and more creative test plans to break software and identify potential issues. You will have a huge impact on the quality of NVIDIA's products. The ideal candidate must have strong programming skills and hands-on experience using AI development tools to improve quality and productivity across the end-to-end QA workflow. This includes leveraging AI assistants for test automation, code generation, debugging, and enhancing testing efficiency. During the interview process, we will assess your ability to effectively use AI development tools and evaluate your programming capabilities to ensure you can deliver high-quality solutions. What you’ll be doing: Architect, implement, and evolve scalable agentic end to end SWQA workflow, automated test frameworks, infrastructure, and tooling for complex software products. Define test strategy and quality gates across functional, integration, regression, reliability, and release-validation workflows. Build and maintain high-value automated coverage for Linux-based, co

dockerkuberneteslinux
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime