Jobiba hiring network

Senior Software Reliability Engineer Jobs

7,292 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

C-
CLEAR - Corporate
📍 New York• Full-time• $225K – $300K/yr
16 days ago

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Fullstack Software Engineer on CLEAR’s Healthcare team, you will build and scale secure, interoperable identity and data solutions that connect patients, providers, and partners. You’ll operate at the intersection of modern web platforms, healthcare interoperability standards, and high-assurance identity systems powering frictionless, trusted healthcare experiences nationwide. A brief highlight of our tech stack: Python / React / Typescript AWS cloud What you’ll do: Design and deliver secure, scalable fullstack solutions that integrate with enterprise EHR systems and national health information exchange frameworks Build and maintain healthcare data integrations leveraging FHIR (RESTful APIs/JSON) and HL7 v2 messaging to enable compliant, real-time data exchange Develop identity resolution and patient matching capabilities using identifiers such as MRNs and NPIs to ensure integrity across disparate clinical systems Partner with Engineering, Security, Product, and Health Information Management teams to implement compliant, audit-ready workflows for regulated healthcare processes Collaborate with external vendors (e.g., Epic Technical Services) to troubleshoot integration issues, manage deployments across TST/PRD environments, and ensure production reliability How you’ll measure success: Successful delivery and stability of FHIR/HL7 integrations across healthcare partners Reduction in data integrity issues related to patient matching and identity resolution High system uptime and successful production deployments across tiered

typescriptpythonreact
View job →
N
1mo ago

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build cloud-native data and storage services for hybrid and multi-cloud infrastructure, including dataset discovery, ingestion, governance, checkpointing, observability, and low-latency access. Develop scalable cloud-native services and APIs that support exabyte-scale, high-performance GPU training and inference workflows. Work closely with product managers, internal AI teams, platform teams, and partner engineering teams to understand requirements and turn them into reliable production systems. Collaborate with SRE, operations, and support teams to improve service reliability, performance, observability, on-call readiness, and operational scale. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, and verification. What we need to see: BS in Computer Science, Information Systems, Computer Engineering, or equivalent experience, with 5+ years of software engineering experience. Strong foundation in algorithms, data structures, distributed systems, and practi

pythonjavaaws
View job →
N
1mo ago

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys

pythonjavakubernetes
View job →

About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep

pythonjavaaws
View job →

About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep

pythonjavaaws
View job →

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Backend Software Engineers at Palantir build software at scale to transform how organisations use data. Our Software Engineers are involved throughout the product lifecycle, from idea generation, design, prototyping, and production delivery. You will collaborate closely with technical and non-technical teammates to understand our customers' problems and build products that solve them. We encourage movement across teams to share context, skills, and experience, so you'll learn about many different technologies and aspects of each product. Engineers work autonomously and make decisions independently, within a community that will support and challenge you as you grow and develop, becoming a strong technical contributor and engineering leader. Your day-to-day workflow will vary, adapting to the requirements of our users and the technical challenges that arise. One day, you may find yourself collaborating with other engineers to architect a new system that enables a novel workflow, the next you could be fine-tuning performance to enable low-latency operational outcomes. Our Product Development organisation is made up of small teams of Software Engineers. Each team focuses on a specific aspect of a product and work collaboratively to build cross functional capabilities, streamline user workflows and continuously improve our software's efficiency and reliability. We’re hiring engineers who are passionate about solving real-world problems and empowering both developers and end-users to work optimally. If you’re motivated to develop reliable, performant, and scalable systems, and to design robust APIs and primitives, this role offers the opportunity to make a sig

PE
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Backend Software Engineers at Palantir build software at scale to transform how organisations use data. Our Software Engineers are involved throughout the product lifecycle, from idea generation, design, prototyping, and production delivery. You will collaborate closely with technical and non-technical teammates to understand our customers' problems and build products that solve them. We encourage movement across teams to share context, skills, and experience, so you'll learn about many different technologies and aspects of each product. Engineers work autonomously and make decisions independently, within a community that will support and challenge you as you grow and develop, becoming a strong technical contributor and engineering leader. Your day-to-day workflow will vary, adapting to the requirements of our users and the technical challenges that arise. One day, you may find yourself collaborating with other engineers to architect a new system that enables a novel workflow, the next you could be fine-tuning performance to enable low-latency operational outcomes. Our Product Development organisation is made up of small teams of Software Engineers. Each team focuses on a specific aspect of a product and work collaboratively to build cross functional capabilities, streamline user workflows and continuously improve our software's efficiency and reliability. We’re hiring engineers who are passionate about solving real-world problems and empowering both developers and end-users to work optimally. If you’re motivated to develop reliable, performant, and scalable systems, and to design robust APIs and primitives, this role offers the opportunity to make a sig

C
Coinbase
📍 - USA• Full-time• Remote• From $253.9K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Staff Software Engineer on the Platform Payments team, you'll define the engineering vision for how fiat moves into and out of Coinbase across 50+ countries and payment rails. This team builds and operates the fiat on- and off-ramps, routing, orchestration, and funds-flow services that power every customer-facing payment experience. You'll own the multi-quarter technical strategy, architect for global scale and reliability, and drive platform-level improvements that directly impact payment success rates, latency, and cost efficiency. What you'll do: Define and drive the multi-quarter technical strategy for Payments spanning rails, orchestration, and transfers to support new products, geos, and significant growth in payment volume. Architect and evolve the core Payments platform (rails integrations, routing, and funds-flow services) for high availability, low latency, and cost efficiency at global scale. Lead end-to-end design and rollout of large, cross-team initiatives (e.g., new global rails, decomp/migrations, resiliency programs), breaking ambiguity into clear milestones and measurable outcomes. Set and enforce technical standards for financial correctness across Payments, including idempotency, reconciliation, failure-mode handling, and auditability for all money-movement paths. Partner with FinHub, Payments Risk, Regulatory Platform, and Produc

REMOTEpythonjavaaws
View job →
C
Coinbase
📍 - Canada• Full-time• Remote• From $253.9K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Payments Platform is at the front door of Coinbase’s financial ecosystem. Our mission is to move old money into crypto and power financial access and payments for the crypto economy globally. As the bridge between the traditional financial world and the crypto economy, we build and operate the fiat on- and off-ramps and payment services that let Coinbase’s customer-facing products deliver reliable, seamless payment processing and connect customer fiat funds into crypto. We’re seeking a deep technical leader to define the engineering vision for Payments, shape architectural cohesion across our global payment rails and funds-flow systems, drive platform-level reliability and cost efficiency, and lead the next evolution of how fiat moves into and out of Coinbase. What you’ll be doing (ie. job duties): Define and drive the multi‑quarter technical strategy for Payments, spanning rails, orchestration, and transfers, to support new products, geos, and significant growth in payment volume. Architect and evolve the core Payments platform (rails integrations, routing, and funds-flow services) for high availability, low latency, and cost efficiency at global scale. Lead end‑to‑end design and rollout of large, cross‑team initiatives (e.g., new global rails, decomp/migrations, resiliency programs), breaking high ambiguity into clear milestones and measurable outcomes. Set and enf

REMOTEpythonjavaaws
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $253.9K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Staff Software Engineer on the Core Automation team within Platform , you'll architect and build the Agentic AI systems that are reimagining how Coinbase delivers customer support and manages compliance at scale. This team is focused on automating high-volume operational processes with AI, then building the primitives and orchestration platform that let every domain at Coinbase leverage those capabilities. You'll own the technical vision for production AI systems that directly improve customer experience, strengthen compliance posture, and drive outsized efficiency gains. What you'll do: Own the design and implementation of Agentic AI systems powering Coinbase's customer support and compliance automation Architect a reusable platform of primitives, orchestration tools, and interfaces that enable other teams to adopt Agentic AI for their domains Drive cross-functional alignment with product, compliance, and operations stakeholders to translate business needs into scalable technical solutions Lead technical strategy for getting LLM-based applications from prototype to production, including grounding, hallucination reduction, and reliability at scale Mentor engineers across the team on system design, coding standards, and AI/ML best practices Required Skills and Experience: 12+ years of backend software engineering experience, with deep proficiency in Golang

REMOTEawsaigo
View job →
D
Dropbox
📍 Poland• Full-time• Remote
1mo ago

Role Description As an Infrastructure Engineer, your role will be crucial in shaping and constructing the robust systems that not only support our current flagship products but also lay the groundwork for the next wave of engineering innovations. From optimizing user experiences across various projects to ensuring seamless scalability and data integrity, you'll be at the forefront of shaping the technological backbone of our platform. Collaborating closely with cross-functional teams, you'll leverage your expertise to tackle audacious challenges and push the boundaries of what's possible. Your contributions will directly impact millions of users, as every line of code you write furthers our mission to revolutionize the way people work and collaborate. Join us in redefining the future, where your passion for building scalable, reliable systems will drive meaningful change on a global scale. Our Engineering Career Framework is viewable by anyone outside the company and describes what’s expected for our engineers at each of our career levels. Check out our blog post on this topic and more here . Responsibilities Build infrastructure capable of managing metadata for hundreds of billions of files, handling hundreds of petabytes of user data, and facilitating millions of concurrent connections. Lead the expansion of Dropbox's function as the data-fabric, connecting hundreds of millions of applications, devices, and services globally, while also driving initiatives to enhance interoperability and adaptability across diverse ecosystems. M easur e and optimiz e Dropbox's analytics platform to maintain its status as one of the most advanced in the industry for extracting meaningful insights from vast data volumes. Collaborat e with cross-functional teams to innovate and implement solutions that enhance the performance, reliability, and security of Dropbox's infrastructure, ensuring a seamless experience for users worldwide. Proactively identify new opportunities and drive imp

REMOTEpythonjavagit
View job →

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op

pythonawsrest
View job →
C
11 days ago

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary CVS Health's Adjudication & Client Experience Engineering organization is seeking a motivated and highly skilled Senior Analyst - Software Development Engineering to join our Application Production Support team. This role will support critical business applications by providing production support, troubleshooting technical issues, and contributing to ongoing application enhancements and stability improvements. As a Sr. Analyst, you will work closely with Lead Engineers, Software Development Engineers, Product Owners, QA teams, and business stakeholders to investigate and resolve production incidents, perform root cause analysis, and implement code fixes. You will be responsible for supporting enterprise applications built on Java, Angular, APIs, and Cloud platforms while ensuring the reliability and performance of systems that serve our PBM (Pharmacy Benefit Management) business. This role is ideal for a hands-on engineer who enjoys solving production challenges, developing software solutions, and collaborating within a fast-paced environment. The successful candidate will contribute to application support activities, system enhancements, and continuous improvement initiatives while growing their technical and business domain expertise. Required Qualifications 5-8 years of experience in software development, application s

REMOTEtypescriptjavaangular
View job →
N
1mo ago

We are looking for a highly motivated AI/ML Software Engineer to join the Enterprise Agentic AI Platform team within IT. You will work closely with Business Analysts, and Engineering teams to design, develop, and deploy enterprise AI solutions that improve productivity and automate business workflows across Engineering, Operations, and Manufacturing. What you'll be doing: Design, develop, and deploy Agentic AI applications using Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and AI orchestration frameworks. Build scalable AI services and reusable components integrated with enterprise applications such as PLM, SAP, and other business systems. Collaborate with business and IT teams to translate business requirements into AI-driven solutions. Develop secure, scalable APIs and enterprise integrations to enable intelligent workflows and automation. Improve AI solution quality, performance, and reliability through prompt engineering, evaluation, and continuous optimization. Partner with cross-functional teams throughout the Software Development Lifecycle (SDLC), from solution design through deployment and production support. What we need to see: Bachelor's or Master's degree in Computer Science, Information Technology, AI/ML, or a related field. 6+ years of software engineering experience with strong proficiency in Python and backend application development. Hands-on experience with Generative AI, LLMs, RAG, AI agents, REST APIs, and cloud-native application development. Experience integrating enterprise applications and building scalable, production-ready software solutions. Strong analytical, problem-solving, communicatio

pythonazureai
View job →

Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team... Here at GoDaddy, the ML Engineering (MLE) team exists as the backbone of our machine learning infrastructure, enabling ML scientists and product teams across Domains to ship models to production reliably, efficiently, and at scale. This team owns the full lifecycle of ML systems — from CI/CD pipelines and model serving infrastructure to GPU workload orchestration and observability. Through disciplined engineering practices, thoughtful system design, and close collaboration with ML scientists, data engineers, and product teams, we deliver the platform that powers domain search, pricing, recommendations, and emerging AI experiences for millions of customers worldwide. We are currently looking for an experienced, highly motivated Senior Engineering Manager to lead our ML Engineering team based in India. This is an established team with existing engineers — we expect the candidate to ramp up quickly on our ML infrastructure stack, build strong relationships with the team, and partner with both India-based teams and US-based teams to drive execution and grow the team further. This individual will join us on our journey to build and scale ML infrastructure that serves real-time predictions at low latency, automates model deployment and promotion, and provides the observability and reliability guarantees that production ML systems demand. Become part of a team that bridges the gap between ML research and production engineering — shipping systems that directly impact GoDaddy's core revenue. What you'll get to do... Lead a team o

typescriptpythonaws
View job →
🔔

Get new senior software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime