Jobiba hiring network

Software Reliability Engineer Jobs

6,428 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Datadog's Software Engineers with Systems depth leverage their experience with systems and tooling to build software that ensures Datadog remains reliable, performant, and secure. For this track, their Software Engineering experience may resemble the Distributed Systems track, but is typically applied in combination with their systems experience to build and run internal platforms and tools that our products are built on. These people typically have deep experience building and managing large cloud infrastructure deployments, or leading reliability efforts for orgs similar to ours, or building release machinery to allow hundreds or thousands of devs to do their jobs without stepping on each others' toes. The systems and tooling where they may have experience depth may include (but not limited to): bazel, build tooling, cassandra, CDN, chef, configuration management, container orchestration, consul, docker, elasticsearch envoy, haproxy, kafka, kubernetes, load balancing, network architecture, postgres, redis, release management, RPC frameworks, service discovery, spinnaker, terraform, zookeeper. Bonus: You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. #LI-KM5 Datadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan. Th

postgresqlredisdocker
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Okta Privileged Access Management (PAM) is an identity-centric approach to a common and critical privileged access use case. Our elegant Zero Trust architecture is purpose-built for the modern cloud and helps customers solve challenging security and operations pain points at scale. We are looking for a software engineer to join our fast-growing team with a focus on scalability, reliability, and enhancing the core building blocks of the product. In this role you will: Be deeply involved in evolving the core architecture of PAM. Work in our product development teams to build scalable, composable components of our platform. Be responsible for designing and implementing scalable architecture patterns. Delight our customers by providing world class UX using our React-based design system Design and build APIs that customers rely on for access to production infrastructure. Work on backend components written in Go and frontend components written in React. You might be a good fit if you: Have 3-5 years of software development experience with a background in Golang or similar programming languages. Proficient in React or similar front-end UI stacks. Experienced working with relational databases like PostgreSQL or similar RDBMS technologies. have the ability to complete a feature end to end from designing database models to backend APIs and frontend UI components. Experienced working with any cloud provider such as AWS, GCP or Azure. Thrive in a collaborativ

reactpostgresqlaws
View job →
O
10 days ago

About the Team Consumer Monetization builds the experiences and systems that power how customers purchase and pay for OpenAI products. Our scope spans purchasing flows, payments, subscriptions, and billing, along with the shared capabilities that support new products, offers, and distribution channels. We own both customer-facing experiences and the underlying platforms that power them. We partner closely with Product, Design, Growth, Data Science, and engineering teams across OpenAI to make purchasing effective and reliable, and new offerings easier to launch and monetize. About the Role We’re looking for experienced Staff+ engineers to evolve the payments, billing, and subscription capabilities that support OpenAI’s growing product portfolio. You’ll tackle problems where correctness, reliability, and flexibility are essential: supporting new billing requirements, managing the billing lifecycle, synchronizing state across internal systems and external providers, and enabling new products and commercial models. Your work may span several areas based on your expertise and team priorities: Billing and monetization capabilities: Extend billing capabilities and improve integrations and state consistency across systems to support new products and business models. Subscriptions: Orchestrate purchases, renewals, plan changes, cancellations, and recovery, ensuring customers are charged correctly and receive the right benefits. Payments: Expand payment capabilities through processor integrations, routing, and broader payment-method coverage. Risk and integrity: Partner with risk and integrity teams to integrate controls into purchasing flows, reducing abuse while protecting legitimate customer experiences. You’ll help set technical direction while remaining hands-on in implementation and delivery. This is an opportunity to solve complex engineering problems at scale, connect architecture decisions to customer and business outcomes, and help other engineers take on broader ow

REMOTEartificial intelligenceaifinance
View job →

About the Team DoorDash Labs is an independent team within DoorDash. We're hiring a backend software engineer to work at the intersection of software engineering and robotics to solve key business problems with elegant technical solutions. If you have a passion for applying robotics solutions to a service loved by millions of people, then we want to talk to you! About the Role We’re looking for Backend Engineers to work on both Product and Product Platform based teams in DoorDash Labs. Product focused Engineers work at the intersection of product and infrastructure to solve key business problems with elegant technical solutions. You'll operate our backend services and architecture that support all product functionality and will be challenged to consider the big picture -- collaborating cross-functionally, as well as evaluating and executing on trade-offs to maximize business impact for the company. You're excited about this opportunity because you will... Design and implement backend services for IoT that integrates with core DoorDash data, focused on reliability, and future extensibility Create a well documented APIs for other departments to integrate with Improve performance, reliability, scalability and security for our backend systems Introduce tools and best practices to accelerate our development process Design and implement backend services for autonomous delivery system that integrate with core DoorDash data. We're excited about you because you have... B.S., M.S., or PhD. in Computer Science or equivalent 6+ years of industry experience as a software engineer Experience with backend for frontend architecture Ability to improve efficiency, scalability, and stability of multiple system resources Experience with service oriented architecture, writing REST API’s, unit testing, and architectural design Understanding of modern web stacks and architecture (HTTP, REST) Experience with SQL Experience with either Java or Kotlin Nice to Have Experience with

javasqlpostgresql
View job →
M
12 days ago

We are seeking a Staff Engineer to join our growing team to provide technical direction and implement core parts of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Staff Engineer on this new team, you will be responsible for providing technical leadership to teams developing cutting edge technologies related to enabling deployment at scale of AI applications. You will take on challenging, high-visibility projects that improve and enhance the performance, scalability, and reliability of the distributed systems infrastructure for this new product. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day. We value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We're looking to speak with candidates based in the New York City area for our hybrid or in-office working models. Position Expectations Work closely with product management, product engineering, product design peers as well as other teams within the company to define the first version and future evolution of the service Design, build and deliver well-tested core pieces of the platform in collaboration with other vested parties Contribute to shaping architecture, code reviews and development practices, developer experience as the teams and product grow Mentor fellow engineers and assume ownership and accountability of projects Qualifications Strong background in building core components for high scale compute and data distributed systems 8+ years experience of building distributed systems, and/or foundational cloud services at scale and an interest in working with Python, Go and Java Proven success in designing, writing, testing, debugging, performance tuning, possessing a strong grip on the foundational materials of computer science and maintaining distributed and/or highly concurrent software s

pythonjavamongodb
View job →
HI
HP IQ
📍 San Francisco• Full-time• $162K – $225K/yr
16 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Senior Software Engineer, Cloud Services, you will create scalable, reliable backend systems that support HP IQ's mission to transform the way people work. We value how we work as much as what we deliver. Our journey of continuous learning and evolution requires a flexible and resilient services platform that fosters innovation, experimentation, and adaptation. To enable this, we focus on designs and tools rooted in strong engineering principles like abstraction, composition, virtualization, automation, and iterative development cycles. What You Might Do Design, develop, and maintain backend services and RESTful APIs using Java and Spring Boot. Write clean, efficient, and well-tested code following established best practices. Collaborate with frontend developers, product managers, and other engineers to deliver end-to-end features. Integrate with databases and external services, ensuring performance, security, and reliability. Participate in code reviews, debugging, and performance optimization efforts. Contribute to CI/CD pipelines and support application deployment in cloud and

javasqlpostgresql
View job →
HI
HP IQ
📍 San Francisco• Full-time• $140K – $225K/yr
16 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As the Senior Software Engineer, Tooling and Development Infrastructure, you will play a critical role in shaping the developer productivity tools and automated testing strategy. You’ll collaborate closely with design, development, and quality teams to plan, design, and implement robust automated tools and services that ensure the quality and reliability of our AI software stack. You will be highly hands-on in your work and collaborate closely with stakeholders. This position offers a unique opportunity to influence the development of cutting-edge automation frameworks, foster a culture of quality, and contribute to the long-term success of the organization. What You Might Do Develop and implement automation frameworks and testing strategies that cover the entire software stack, from backend systems to user-facing features. Identify, evaluate, and integrate new tools that streamline development. This includes everything from code quality tools and to Infrastructure-as-Code (IaC) solutions. Lead continuous improvement efforts for our build, release, and test systems, ensuring a robust

pythonredisci/cd
View job →
HI
HP IQ
📍 San Francisco• Full-time• $179K – $252K/yr
16 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Lead Software Engineer, Cloud Services, you will create scalable, reliable backend systems that support HP IQ's mission to transform the way people work. We value how we work as much as what we deliver. Our journey of continuous learning and evolution requires a flexible and resilient services platform that fosters innovation, experimentation, and adaptation. To enable this, we focus on designs and tools rooted in strong engineering principles like abstraction, composition, virtualization, automation, and iterative development cycles. What You Might Do Lead technical strategy and execution across development and infrastructure, designing and implementing scalable, secure cloud-native systems while establishing architecture standards and engineering best practices. Own end-to-end delivery of platform and cloud services, including infrastructure buildout, application architecture, APIs, deployment pipelines, observability, reliability, and performance, while mentoring engineers and driving cross-functional alignment. Work with modern containerization and cloud technologies, includi

javaredisdocker
View job →
E
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Everpure Cloud Azure Native is already live and generally available—we’ve shipped the first version of the service, and we’re now focused on making it better every quarter: expanding use cases, entering new markets, and raising the bar on reliability and developer experience. Our Foundation team owns a Go‑based control‑plane service that orchestrates how customers consume enterprise‑grade block storage natively in Azure. Around it, we work with a modern cloud stack including Temporal and other platform services for workflows, automation, and observability. This is a production cloud service, not an internal tool: the code you write directly shapes how customers deploy, scale, and operate storage in their Azure environments. This is a rare opportunity to work on a core storage service in a major public cloud , in a joint effort between Everpure and Microsoft. You’ll design and evolve APIs and service behavior that sit in the critical path of real workloads, collaborating closely with engineers across both companies. We move fast without breaking customers: clear API contracts, long‑lived interfaces, solid test automation, and CI/CD are part of the job, not side projects. If you want to own services end‑to‑end, solve real distributed systems problems in the public cloud, and see your work reflected directly in how Azure customers use storage, this is the place to do it. WHAT YOU'LL DO Own and evolve a production clo

pythonjavaaws
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Everpure Cloud Azure Native is a generally available service that brings enterprise-grade block storage natively to Azure. With the first version already live, we are expanding the service into new use cases and markets while improving its reliability, operability, and customer experience. Our Foundation team owns a Go-based control-plane service that coordinates how customers provision and use storage in Azure. Around it, we work with a modern cloud stack including Temporal and other platform services for workflows, automation, and observability. This is a production cloud service: the code you write directly shapes how customers deploy, scale, and operate storage in their Azure environments. You’ll work on a core storage service in a major public cloud , as part of a joint effort between Everpure and Microsoft. You’ll design and evolve APIs and service behavior in the critical path of real customer workloads, collaborating closely with engineers across both companies. Clear API contracts, long-lived interfaces, test automation, and CI/CD are fundamental to how we build. You’ll have the opportunity to own services end to end and solve complex distributed-systems problems in the public cloud. WHAT YOU'LL DO Own and evolve a production cloud service that powers Everpure Cloud Azure Native, taking features from idea and design through deployment and operation in Azure for real customers. Build new capabilities and i

pythonjavaaws
View job →

DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. About the job DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides an in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. In this role, you will Build core capabilities for our SaaS Platform across multiple clouds Drive development of functional enhancements for Data Discovery, Observability & Governance for both OSS and SaaS offering Lead efforts around non functional aspects like performance, scalability, reliability Lead and mentor junior engineers Work closely with PM, Customers and OSS community Requirements Over 8+ years of experience building and scaling backend systems, preferably in cloud-first or SaaS environments. Solve complex tech

javaci/cdrest
View job →

Staff Software Engineer - Testing & Automation Exceptional software engineering is challenging. Amplifying it to ensure that multiple teams can concurrently create and manage a vast, intricate product escalates the complexity. As a Staff Engineer within the Verification Platform team at Sumo Logic, you will drive the implementation and optimization for our verification platform as well as the modernization of our CI/CD pipelines. Your mission is to develop and sustain automated tooling for all testing, verification, and functional requirements, leveraging AI reasoning and machine learning models to predict and prevent delivery issues, while integrating advanced security validation and non-functional requirements into our delivery lifecycle. You will contribute significantly to establishing automated delivery pipelines, empowering autonomous teams to create independently deployable services, and progressing Sumo Logic’s internal Platform-as-a-Service. This role sits at the intersection of Platform Engineering, Quality Engineering, DevSecOps, and Developer Productivity, helping teams deliver secure, reliable, and independently deployable services at scale. Responsibilities Strategy & Leadership: Drive technical direction and design for a modern Quality Engineering platform, driving the adoption of AI reasoning for enhanced automation of all testing, verification, and functional requirements. Pipeline Modernization: Lead the modernization of CI/CD pipelines to include automated security validation, compliance checks, and other critical non-functional requirements, with a focus on integrating AI/ML for intelligent pipeline optimization and risk prediction. Framework Ownership: Own the delivery pipeline and release automation framework for all Sumo services, ensuring improvements in developer productivity, deployment frequency, and release reliability. Cross-Team Collaboration: Educate and collaborate with teams during design and development phases to ensur

pythonjavareact
View job →
SC
Sigma Computing
📍 San Francisco• Full-time• $170K – $235K/yr
16 days ago

About the Role Sigma Computing is redefining business intelligence by making complex data analysis accessible through a high-performance platform built for the modern data stack. The Compiler Team plays a foundational role in this mission by transforming user-driven spreadsheet interactions into highly optimized SQL queries, enabling seamless exploratory analytics on cloud data warehouses. As a member of the Compiler Team, you will join a group of engineers dedicated to building the core systems and abstractions that power Sigma’s intuitive spreadsheet interface, ensuring speed, reliability, and scalability for all users. What You Will Be Doing Tackle core challenges at the intersection of data modeling, query compilation, and large-scale interactive analytics—making it possible for end-users to query data warehouses efficiently without deep technical knowledge Design, build, and maintain sophisticated compiler infrastructure and intermediate representations that translate spreadsheet operations into optimized query plans Apply advanced optimization strategies to improve performance and accuracy across a wide range of query workloads and data architectures Contribute to both backend (Rust) and key frontend foundations (TypeScript), evolving critical abstractions that enable end-to-end workflow optimizations and new features Debug, analyze, and resolve complex issues, ensuring robustness and maintainability in a rapidly evolving product Collaborate with engineers and product stakeholders to review designs and code, driving technical best practices and architectural decisions throughout the team and company Qualifications We Need 5+ years experience engineering high-quality software systems Demonstrated success building and maintaining complex infrastructure or core platform services Deep understanding of Computer Science fundamentals, particularly in compilers, algorithms, SQL Optimization Passion for teamwork, technical ownership, and continually

typescriptpythonsql
View job →

DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You’ll Do: We are looking for a Senior Software Engineer – Platform Operations based in Pune, India, who will play a key role in ensuring the reliability, performance, and operational excellence of DeepIntent's platform and data ecosystem. This role requires a strong engineering mindset with the ability to troubleshoot complex technical issues, understand distributed data architectures, and collaborate across Engineering, Product, Analytics, and Customer-facing teams to deliver timely and effective solutions. As part of the Operations organization, you will work closely with Engineering to support production systems, improve operational processes, and drive platform stability. The ideal candidate is a self-motivated problem solver who is passionate about learning new technologies, improving system reliability, and delivering exceptional customer outcomes through engineering excellence. Serve as the engineering interface between Customer-facing teams, Analytics, Product, and Engineering organizations. Partner with Platform Support, Client Success, and other customer-facing teams to investigate and resolve complex platform-related issues. Analyze application, API, and data pipeline issues to identify root causes and drive timely resolution. Develop and standardize operational tools, and interfaces to support analytical and operational use cases. Monitor data pipeline executions, investigate failures, and implement corrective and pre

pythonjavasql
View job →
D
16 days ago

DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You’ll Do: Play a key role in ensuring product quality, reliability, and performance. Define test strategies, lead automation initiatives. Design and execute manual and automated tests for UI, API, and database layers. Develop and maintain test automation frameworks using Selenium, Playwright, or similar tools. Write, execute, and optimize test scripts integrated with CI/CD pipelines (e.g. Jenkins). Perform API and database validations, analyze logs, debug failures, and report defects clearly. Collaborate closely with developers, product managers, and DevOps for end-to-end quality assurance. Apply strong coding, testing, and problem-solving skills to identify risks and improve test coverage. Who You Are: Bachelor’s degree in engineering (CS / IT) or equivalent degree from a well-known Institute / University. 1+ years of hands-on experience with QA Automation Framework development & Design (Preferred language: Python). Strong understanding of testing methodologies. Experience with Python, Perl, Shell Scripting, Selenium, Test Automation (QA), and Software Testing (QA) (Preferred). Experience in Software Development, SDET (must have). Strong problem analysis, troubleshooting and debugging skills. Experience in databases, preferably MySQL. REST/API testing experience is a plus. Ability to integrate end-to-end tests with CI/CD pipelines and monitor and improve metrics around test coverage. Ability to work in a dynamic and agile de

pythonsqlmysql
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime