Jobiba hiring network

Software Reliability Engineer Jobs

6,326 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role OpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency of our systems. In this role, you'll apply deep technical expertise to optimize infrastructure and application-level performance across mission-critical products like ChatGPT and our developer API. You’ll work cross-functionally with teams building core services, training models, and developing real-time user experiences to push our latency, throughput, and cost-efficiency to the next level. We are looking for engineers who thrive in ambiguous environments, value deep systems understanding, and are motivated by delivering measurable impact. This is a highly technical, individual contributor role focused on root-cause analysis, profiling, instrumentation, and architecture-level performance improvements across our stack. In this role, you will: Analyze and optimize performance across application, middleware, runtime, and infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond. Develop tooling and metrics that provide deep observability into system performance. Collaborate closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements. Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale. Lead investigations into high-impact performance regressions or scalability issues in production. Drive performance testing strategies and help define SLAs/SLOs around latency and throughput for critical systems. You might thrive in this role if you: Have 7+ years of experience in software engineering with a strong tr

pythonawsrest
View job →
S
Sentry
📍 San Francisco• Full-time• Remote
13 days ago

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry's Infrastructure Engineering team is what makes operating Sentry simple, safe, and seamless for every other engineering team in the company. They build the internal control platforms, configuration systems, traffic routing, and automation that let product engineers operate services safely at scale without needing deep infrastructure expertise themselves. As the Engineering Manager for Infrastructure Engineering, you'll lead a team of engineers building the tools that power Sentry's growth: internal admin and change management tools, configuration automation, and the routing layer that underlies Sentry's architecture. You'll be responsible for technical vision, team health, system reliability, and partnership with engineering teams across the company who depend on your team's tools every day. You'll work closely with leaders across Infrastructure, Platform, and Production Engineering to shape how Sentry scales its operational model as the company grows. In this role you will Lead a team of engineers building the internal control platforms that every engineering team at Sentry relies on to operate services safely. Drive the evolution of Infrastructure Engineering's platform, including configuration management, traffic routing and environment controls Own the team's technical direction, contributing to key decisions on API architecture, internal tooling design, and automation frameworks. Nurture and grow engineers at different levels, providing support through coaching, mentorship, and career development. Foster an inclusive, high-performing team culture focused on ownership, learning, and delivery. Partne

REMOTEpythonkubernetesai
View job →
S
StarRez
📍 Australia• Full-time
17 days ago

About StarRez StarRez is the global leader in student housing software, providing innovative solutions for on and off-campus housing management, resident wellness and experience, and revenue generation. Trusted by 1,400+ clients across 25+ countries, StarRez supports more than 4 million beds annually with its user-friendly, all-in-one platform, delivering seamless experiences for students and administrators. With offices in the United States, Australia, the UK, and India, StarRez blends the robust capabilities of a global organization with the personalized care and service of a trusted partner. The Role: We are looking for a Senior Quality Engineer - AI to help shape how StarRez evaluates, validates, and improves AI-powered product experiences. You will play a pivotal role in elevating our quality practices for AI-powered product experiences and tooling. You will be a product expert within our engineering organization, deeply understanding workflows and customer outcomes. This role will balance hands-on individual contributor responsibilities with leading, influencing, and coaching your peers. You’re someone who is passionate about seeing systems holistically and is eager to drive the highest standards of accuracy, safety, usefulness, and reliability. You will help define what "good" AI output means for StarRez, build repeatable evaluation systems, coach reviewers and subject matter experts, and turn subjective feedback into measurable product improvement. You will work at the intersection of AI engineering, product, and domain knowledge — partnering with engineers, product managers, and SMEs to raise the bar on how StarRez evaluates, monitors, and improves AI tools. Key Responsibilities: Lead or contributed to end-to-end AI quality and evaluation strategy for product experiences involving LLMs, RAG, prompts, tool-use, or agent workflows. Provided quality-focused input during ticket grooming, discovery, and feature discussions, with clear guidance on testability, ob

aigorust
View job →
S
17 days ago

SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . We are seeking an experienced Principal Software Engineer to lead the architecture, development, and evolution of a large-scale malware analysis and cybersecurity platform. This role will own the end-to-end technical architecture across Python-based analysis pipelines, message-driven job orchestration systems, Windows VM sandbox environments, and web-based interfaces. The ideal candidate brings deep expertise in software engineering, malware analysis, and distributed systems, with a proven ability to take ownership of complex legacy platforms and drive modernization initiatives. You will lead the design and implementation of advanced detection capabilities, including behavioral analysis, malware sandboxing, machine learning-based classification, and support for new file types, while ensuring platform scalability, reliability, and performance. This position requires strong hands-on development skills in Python with Linux-based production environments, messaging architectures such as RabbitMQ, virtualization technologies, and cloud-native deployment practices. As the technical leader for the platform,

pythonsqlmysql
View job →
R
24 days ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The International engineering team's mission is to expand Robinhood's products globally and scale our platform to support millions of new customers in diverse markets. We build and maintain experiences across onboarding, funding, account setup, localized product experiences, growth incentives, and platform capabilities to operate in new jurisdictions. Our engineering team partners closely with product, design, compliance, marketing teams, along with an array of platform engineering teams to launch new products and optimize user acquisition globally. We prioritize technical rigor, system reliability, and quick execution to deliver reliable products to our growing customer base. Our culture is centered on clear communication, strong team partnership, and a commitment to helping people manage their financial lives! As an Engineering Manager , you will lead a team of software developers to drive some of Robinhood’s most important international launch initiatives. You will be responsible for driving the architecture and execution required to launch crypto and other financial products in new regions, while shaping the reusable systems that make future launches faster, safer, and m

O
1mo ago

About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. About the role We are looking for an Operating Systems Engineer to build and harden the OS foundations for OpenAI products. We are especially interested in experienced, passionate, and innovative operating systems developers who thrive on building foundational platform software and solving hard problems in security, privacy, performance, power, and reliability. You will work across the OS kernel, core OS services, security and privacy primitives, performance and power, and the frameworks that connect applications and UI to the system. This role emphasizes deep debugging and systems ownership from development through production. You will collaborate closely with embedded, firmware, hardware, application, and product engineering teams. Experience with hardware bring-up is a plus, but not required. What you will do Work on end-to-end OS capabilities spanning the OS kernel, userspace services, application frameworks, UI toolkits, and application-facing APIs. Develop, integrate, and maintain OS components, both kernel-bound and in userspace, including scheduling, memory management, filesystems, drivers, IPC/RPC mechanisms, and security-relevant subsystems. Build and maintain core OS services and daemons (init, service management, device discovery, networking primitives, time, logging, update hooks, crash handling, and so on). Design and implement security and privacy mechanisms: Secure boot and measured boot integration points (where applicable). Mandatory access control and sandboxing. Secrets management, secure storage, key handling, and least-privilege service design. Privacy-preserving telemetry, data minimization, and user-consent oriented system behaviors. Establish a perfo

awslinuxrest
View job →

Are you ready to do your life’s work at the heart of the autonomous revolution? NVIDIA’s SWQA organization is seeking a world-class Software QA Test and Tool Developer to join our Automotive Platform team, where the code you validate ensures the safety of millions on the road. In this role, you won't just be testing software; you will be architecting the security and reliability of the next generation of intelligent vehicles. We are looking for engineers who are as comfortable navigating low-level product architecture as they are deep-diving into complex product use cases with passion for quality. This is a high-impact, hands-on position focused on our industry-leading automotive products, offering a rare opportunity to influence the core of our tech stack. You will also build the tools and frameworks that define performance standards for systems running on Linux and QNX. What you’ll be doing: Design, execute, and automate comprehensive test cases and test scenarios to validate our automotive platforms using various test methodologies to identify and track actionable defects and track them to closure. Participate in deep-dive reviews of product requirements and technical designs, providing critical feedback to ensure features are built for testability and security from day one. Partner closely with project management, hardware teams, and software developers to provide rigorous technical analysis of bugs and publish data-driven statistical reports for global team members. Architect and maintain a distributed test automation framework capable of managing high-concurrency workloads across an extensive automation farm of hundreds of concurrent systems. Develop sophisticated test libraries and automation solutions to accelerate development cycles and expand automated test coverage for re

pythonlinuxai
View job →
S
Sentry
📍 San Francisco• Full-time• $220K – $450K/yr
1mo ago

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry's Infrastructure Engineering team is what makes operating Sentry simple, safe, and seamless for every other engineering team in the company. They build the internal control platforms, configuration systems, traffic routing, and automation that let product engineers operate services safely at scale without needing deep infrastructure expertise themselves. As the Engineering Manager for Infrastructure Engineering, you'll lead a team of engineers building the tools that power Sentry's growth: internal admin and change management tools, configuration automation, and the routing layer that underlies Sentry's architecture. You'll be responsible for technical vision, team health, system reliability, and partnership with engineering teams across the company who depend on your team's tools every day. You'll work closely with leaders across Infrastructure, Platform, and Production Engineering to shape how Sentry scales its operational model as the company grows. In this role you will Lead a team of engineers building the internal control platforms that every engineering team at Sentry relies on to operate services safely. Drive the evolution of Infrastructure Engineering's platform, including configuration management, traffic routing and environment controls Own the team's technical direction, contributing to key decisions on API architecture, internal tooling design, and automation frameworks. Nurture and grow engineers at different levels, providing support through coaching, mentorship, and career development. Foster an inclusive, high-performing team culture focused on ownership, learning, and delivery. Partne

pythonkubernetesai
View job →
PE
Private Employer
📍 United Kingdom• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role The Palantir platform is deployed in numerous critical mission environments including combat zones and classified networks—from the back of a Humvee to a command post to the cloud. This means operating in multiple cloud environments, on-prem air-gapped networks, and at the edge—at scale. We are looking for Edge Infrastructure Engineers to build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure. This role demands a deep focus on low-level systems, including the deployment and management of physical bare metal servers in both traditional data centers and edge environments. You will be responsible for physical network engineering and the development of robust infrastructure that ensures performance of the Palantir platform. In addition to ensuring performance and reliability, you will play a critical role in building and scaling new environments in a forward-deployed capacity, including onsite. Edge Infrastructure Engineers combine hardware-level engineering experience with the drive to improve existing systems and the creativity to develop novel solutions for evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate for and participate in sensible, scalable systems design, sharing responsibility for diagnosing, resolving, and preventing production issues across our most demanding deployments.

PE
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role The Palantir platform is deployed in numerous critical mission environments including combat zones and classified networks—from the back of a Humvee to a command post to the cloud. This means operating in multiple cloud environments, on-prem air-gapped networks, and at the edge—at scale. We are looking for Edge Infrastructure Engineers to build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure. This role demands a deep focus on low-level systems, including the deployment and management of physical bare metal servers in both traditional data centers and edge environments. You will be responsible for physical network engineering and the development of robust infrastructure that ensures performance of the Palantir platform. In addition to ensuring performance and reliability, you will play a critical role in building and scaling new environments in a forward-deployed capacity, including onsite. Edge Infrastructure Engineers combine hardware-level engineering experience with the drive to improve existing systems and the creativity to develop novel solutions for evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate for and participate in sensible, scalable systems design, sharing responsibility for diagnosing, resolving, and preventing production issues across our most demanding deployments. Core Responsibilities Maintaining availability of cloud & physical Ku

sqlpostgresqlkubernetes
View job →
H
Hp
📍 Colorado• $83K – $127.8K/yr
10 days ago

Software Developer in Test Description - This role is responsible for ensuring quality, reliability and performance of software applications throughout the development lifecycle primary through software test automation. The role designs, codes, and implements software test automation using appropriate programming languages, frameworks, and tools. The role works closely with cross-functional teams to gather requirements, provide technical insights, and ensure the successful execution of test automation with the main goal of improve quality of the solution. The role also creates and executes comprehensive test plans, test cases, and test scripts based on project specifications. The role sets and provides design guidance to other developers and SQA engineers regarding test automation and the test framework. *Onsite in Ft Collins 4-days a week Responsibilities • Designs quality assurance and test processes for portions of end-user video conferencing/collaboration application, systems software running on android hardware, local, networked, and Internet-based platforms. • Analyzes design and determines test scripts, coding, automation, and integration activities required based on specific objectives and established project guidelines. • Designs and maintains Test automation framework • Executes and writes portions of testing plans, protocols, and documentation for assigned portion of application; identifies and debugs issues with code and suggests changes or improvements. • Identifies opportunities for performance improvements and optimizes code and application performance. • Utilizes latest AI tools and technologies in speeding up test automation • Executes test cases depending on the needs of the project • Provides valuable input into the development of user stories and acceptance criteria, shaping a quality-oriented d

pythonjavadocker
View job →
I
Instawork
📍 Bengaluru• Full-time
17 days ago

Instawork is on a mission to create meaningful economic opportunities for skilled hourly professionals in communities around the globe. Our AI-powered labor marketplace helps local businesses scale, and enables global technology companies to push the frontiers of robotics and AI. Backed by world-class investors like Benchmark, Spark Capital, Craft Ventures, Greylock, Y Combinator, and others, we’re looking for exceptional talent to reimagine the way the world works. About IRL (Instawork Robotics Labs) Researchers at UC Berkeley have identified a “100,000-year data gap”—the gulf between what trained AI language models and what physical robots actually have to learn from. Closing that gap is the defining infrastructure challenge of the physical AI era. IRL is Instawork’s answer to it. We deploy skilled workers into real commercial and residential environments—kitchens, warehouses, hotel floors, and homes—to capture the high-fidelity task data that the world’s leading robotics labs use to train their foundation models. About the Role Instawork Robotics creates the highest-quality, highest-diversity datasets for robotics and physical AI. We work with leading robotics builders and research labs to solve one of the most important challenges in robotics: closing the data gap. As a Platform Engineer, you will build the products, services, tools, and infrastructure that power IRL’s data collection, processing, and quality workflows. This is a hybrid product and infrastructure role: you will design platform capabilities for application engineers while also owning the reliability, scalability, cost efficiency, and operation of the systems behind them. You will help shape the platform roadmap by understanding the needs of application engineers, prioritizing high-impact problems, and delivering simple, reliable, and scalable solutions. Who You Are - 5+ years of experience building and operating production software platforms. - Experience designing platform products, backend serv

pythonawsazure
View job →
M
Meltplan
📍 Bengaluru• Full-time
17 days ago

MeltPlan is building the “planning engine” for the $14 Tn construction industry, an AI system designed specifically to optimize decisions before construction begins. While design software optimizes use and aesthetics and construction software optimizes execution and control, MeltPlan is building the missing layer - software that optimizes decisions and tradeoffs upstream, before scope is locked, procurement begins, and change orders become inevitable. MeltPlan’s long-term goal is to help teams make construction “boring” by making planning more intense: surfacing constraints and tradeoffs early, aligning stakeholders before plans are frozen, and reducing the need for late-stage redlines, rework, and change orders. MeltPlan is founded by operators who have built at scale. Kanav previously co-founded Innovaccer, a $3Bn healthtech company focused on making US healthcare more affordable and accessible. He’s now applying that systems-level thinking to construction.He’s joined by Tanmaya Kala, former Project Executive at DPR Construction, who led large commercial, healthcare, and life sciences projects. We combine deep tech scale with real construction execution. What This Role Really is : We are seeking a detail-oriented and technically strong AI QA Engineer to ensure the quality, reliability, and performance of Large Language Model (LLM)-based systems. In this role, you will be responsible for designing and executing test strategies, validating model outputs, and building evaluation frameworks to enhance the accuracy, safety, and overall performance of AI-driven applications.We would particularly value candidates who have hands-on experience in developing evaluation frameworks (evals) for AI systems, along with strong expertise in comprehensive system testing and quality assurance practices.You are responsible for making MeltPlan work in the real world. What You'll Do: Design, develop, and execute evaluation frameworks (Evals) for Large Language Models (LLMs) and AI syst

pythonawsazure
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
17 days ago

Opportunity Overview: We’re looking for a senior-level automation engineer who will help raise the bar on release quality, environment reliability, and change safety across Cohere’s platform. You’ll partner closely with Product, Engineering, Platform, and SRE to build scalable automation, guardrails, and validation systems that reduce production risk while increasing delivery velocity. This is not a “test scripts only” role. You’ll shape automation strategy, embed quality into the SDLC, and help define how changes move safely from dev → staging → UAT → prod in a fast-moving healthcare platform. You’ll help define how quality scales as Cohere grows. This role has real influence over release safety, platform reliability, and how engineering teams ship software in a regulated, high-impact domain. You won’t just test features — you’ll shape how Cohere delivers them safely to production. What you’ll do: Own and evolve Cohere’s end-to-end test automation strategy across UI, API, config changes, and critical workflows Design and maintain scalable E2E automation frameworks for multi-tenant, payer-specific workflows Build automated validation for deployment guardrails, release readiness, and production change safety Partner with Platform/DevOps to integrate automation into CI/CD pipelines and deployment workflows Create automated coverage for high-risk paths (authorization flows, partner integrations, file pipelines, feature flags, config changes) Drive test reliability, flake reduction, and actionable failure signals Define and enforce quality gates for prod releases, blue/green and canary deployments, and config changes Collaborate with Product and Engineering to ensure business outcomes are testable, measurable, and observable Improve test data management and environment stability to enable reliable automation at scale Mentor engineers on testability, automation best practices, and quality-first development Partner with SRE and Security to ensure production readines

typescriptawsci/cd
View job →
M
Mongodb
📍 Alberta• Full-time• From $210K/yr
1mo ago

MongoDB’s Developer Productivity organization exists to help engineers build and deliver high-quality software through a highly effective software development process and a strong foundation of shared tools and services. We are looking for a Senior Director to lead our Pipeline team. This role is tasked with bringing together the major systems and experiences that power software delivery at MongoDB. The team’s mission is to provide a reliable, scalable, secure, and effective platform for ensuring fast software deployability, leveraging AI native approaches. We are open to in-office, flexible or remote hiring across Canada. The Team The Pipeline organization sits within Developer Productivity and is responsible for the systems, services, and user experiences that define MongoDB’s software delivery ecosystem. This is mission-critical infrastructure operating at substantial scale and supports a variety of software product delivery needs. Success in this role requires excellent product judgment for developer-facing experiences, strong systems and platform leadership, and the ability to align multiple teams around a cohesive strategy. Candidate Profile We’re looking for a senior engineering leader who can unify product-minded developer tooling with deep platform and operational excellence. The right candidate is passionate about developer productivity and has a track record of leading managers and teams through organizational growth, technical complexity, and cross-functional change. They should be comfortable owning a broad portfolio that spans developer experience, reliability and scale, release systems, telemetry, and operational health. They should also be able to work effectively with senior leaders and partners across engineering and product to set direction, allocate resources, and make trade-offs that balance near-term delivery with long-term platform function. The right candidate for this role will 12+ years of hands-on software engineering experience bui

mongodbawsazure
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime