Jobiba hiring network

Software Reliability Engineer Jobs

6,326 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

DU
18 days ago

About the Team DoorDash Labs is an independent team within DoorDash. We explore robotics and automation to transform last-mile logistics in the long term. If you have a passion for applying robotics solutions in a service used by millions of people, then we want to talk to you! About the Role We're looking for an experienced technical operator to lead live testing, deployment, and operational validation of cutting-edge autonomous technologies. This role sits at the intersection of engineering and operations, helping ensure new capabilities are safely deployed, thoroughly evaluated, and translated into actionable engineering feedback. You’re excited about this opportunity because you will… Lead and oversee live testing across transport, deployment, and safety validation. Partner closely with hardware, software, and business operations teams to validate new product capabilities while providing guidance and mentorship to junior team members. Conduct and document complex tests for autonomous technologies, evaluating robot behavior, identifying issues, and validating new features and requirements. Provide actionable technical feedback to engineering teams based on test outcomes. Exercise technical judgment during live testing by evaluating robot behavior, assessing operational risk, distinguishing expected behavior from product defects, and determining when engineering escalation or additional validation is required. Develop and implement testing processes, protocols, and checklists that improve the safety, efficiency, and reliability of new products. Utilize internal tools to analyze logs, investigate issues, document findings, and track issues through resolution. Mentor junior team members in structured debugging and documentation practices. Conduct detailed analyses and generate comprehensive reports that identify trends, summarize findings, and provide strategic recommendations to engineering and operations partners. We’re excited about you because… 2+ years of exper

awsgitlinux
View job →
O
Okta
📍 Toronto• Full-time• C$108K – C$135K/yr
22 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. This is a contract position through our staffing partner Magnit. 108,000.00 - 135,000.00 - 162,000.00 CAD Annual This role is not eligible for the Okta-sponsored benefits listed below. Magnit will provide any locally required benefits. Okta seeks a skilled Senior Recruiter to drive full-lifecycle recruitment and build strategic talent pipelines for our Engineering organization across North America. As part of our AMER Tech Recruiting team, you will be a trusted talent advisor responsible for sourcing, engaging, and delivering top-tier engineering talent while maintaining an "always recruiting" mindset in a fast-paced, high-growth environment. What You'll Be Doing Own full-lifecycle recruitment for Engineering and technical roles (Software Engineering, Site Reliability, Security, TPM, Product) across US & Canada, managing a flexible req load that scales with business priorities. Partner strategically with hiring managers and leadership to understand talent needs, define role scope, advise on talent gap mitigation, and challenge assumptions to ensure hiring decisions strengthen long-term organizational capability. Build and execute talent strategies that balance external hiring with internal mobility, creating sustainable pipelines that reflect commitment to diversity, inclusion, and high-performing engineering culture. Drive metrics-informed recruiting decisions by developing KPIs, analyzing recruiting data, and using insights to optimi

awsrestmachine learning
View job →
O
1mo ago

About the Team Training Runtime designs the core distributed runtime that powers everything from early research experiments to frontier-scale model runs. We work on building robust, scalable, high performance components to support our distributed training workloads. Our priorities are to maximize the productivity of our researchers and our hardware, with the goal of accelerating progress towards AGI. Within Training Runtime, the Process Management team develops the distributed OS responsible for launching, coordinating, and supervising the large numbers of processes that make up modern training workloads. Our runtime sits beneath training frameworks and on top of research infrastructure, ensuring jobs run reliably across massive clusters while maintaining performance, stability, and observability. Success for us is measured by both system reliability and researcher velocity - enabling ideas to scale from experiments to production training runs. About the Role As a Training Runtime: Process Management Engineer , you will work on the software that ties thousands of computers together and exposes them as a unified system. This system has to serve individual researchers running multiple parallel experiments, as well as our largest training runs spanning 100’s of thousands and even millions of machines and accelerators. This requires easy to use, introspectable systems that can promote a fast debugging and development cycle, as well as relentless optimization for scale while maintaining stability and performance throughout. You will work primarily in Rust , building high-performance asynchronous systems with a strong emphasis on performance, correctness, and scalability. Working at this scale and at the frontier of AI development poses novel challenges. Out-of-the-box approaches often don’t work. The problems you will be working on are highly ambiguous and require strong design judgment as well as proficient execution to advance the state of our infrastructure. We’re loo

pythonawslinux
View job →

NVIDIA has been redefining computer graphics, desktop gaming, and enhanced computing capabilities for more than 25 years. Today, we are tapping into the unlimited potential of AI to define the next era of computing. As a NVIDIAN, you will work on problems that sit at the boundary of architecture, silicon, firmware, software, and production, where strong judgment matters as much as technical depth. We're the Silicon Design for Productization (DFP) Team, within the broader Silicon Co-Design Group, and we turn power and thermal design into executable productization methodology. Power and thermal are among the most complicated problems we work on at NVIDIA because they sit at the intersection of architecture, workload behavior, silicon variation, firmware policy, platform constraints, and product goals. Small decisions here have an outsized impact on performance, efficiency, reliability, bring-up speed, and ultimately what the product can deliver in the field. We define how features move from concepts to bring-up, characterization, validation, and release. In this role, you will help us build that bridge. We're looking for an engineer who reasons from first principles, flourishes with ownership in a fast-paced environment, and uses AI with sound judgment. What you’ll be doing: Lead the effort across multi-functional teams to keep the program’s power and thermal productization strategy clear, executable, and on track. Create methodology and silicon test plan based controller designs and architecture, including characterization process, debug tools, fuse/firmware settings and lab requirements. Drive resolution for challenging silicon issues through structured hypotheses, measurement plans, and root-cause closure. Steward the Power and Thermal playbook when the existing productization methodology

CH
Cohere Health
📍 Hyderabad• Full-time
18 days ago

Opportunity Overview: We’re looking for a Manager, Platform Engineering that can lead and grow a high-performing engineering team focused on Developer Experience, DevOps, SRE, and Quality. You will own the systems and processes that enable teams to build, test, release, and operate software with high velocity and reliability, driving engineering efficiency and operational excellence across the organization. What you’ll do: Lead a fast-paced, autonomous team of engineers focused on platform engineering, developer experience, DevOps, SRE, and quality engineering Own and drive the internal developer platform strategy and roadmap, improving how engineering teams build, test, deploy, and operate services Create transparency into engineering efficiency and system health through meaningful metrics across delivery, reliability, and quality Enable teams to move faster by improving CI CD pipelines, environments, tooling, and overall developer workflows Provide technical leadership across platform, infrastructure, and reliability, helping teams build scalable and resilient systems Ensure strong engineering practices across release processes, testing, quality, reliability, and security Define and enforce release guardrails, validation standards, and rollback mechanisms to improve production safety Improve environment stability and consistency across development, QA, and pre production environments Drive test strategy and automation maturity to improve overall product quality and confidence in releases Define and implement observability, monitoring, and alerting standards across systems Improve incident detection, response, and RCA practices, ensuring learnings translate into platform and system improvements Drive cloud infrastructure best practices across AWS, containers, and infrastructure as code Foster a culture of ownership, reliability, and continuous improvement within the team Provide innovative solutions for attracting, developing, and retaining top engineering talent I

CH
18 days ago

Opportunity Overview: We’re seeking a strategic and execution-focused Technical Program Manager to drive large, cross-functional initiatives from concept through launch. In this role, you will lead end-to-end delivery of complex, multi-team programs—such as platform migrations, architectural modernization, reliability improvements, and major product launches—while defining clear milestones, success metrics, and sequencing plans. You’ll partner deeply with Engineering to proactively manage dependencies, mitigate risk, and ensure technical and business alignment. This role also strengthens our execution systems by improving program transparency, standardizing launch rigor, and translating technical complexity into clear executive-level insights. What you’ll do: Lead end-to-end delivery of multi-team initiatives (e.g., platform migrations, architectural modernization, reliability improvements, major product launches) Define milestones, critical paths, and measurable success criteria Break ambiguous initiatives into executable phases with clear sequencing Identify and manage cross-team technical dependencies Proactively surface risks and tradeoffs with mitigation plans Run structured risk reviews and escalation processes when needed Participate in technical design discussions to understand architecture and constraints Ensure design reviews, capacity planning, and rollout strategies happen early Collaborate with EMs and engineers leaders to align on realistic timelines Establish program tracking mechanisms that create transparency without bureaucracy Standardize launch readiness, milestone reviews, and postmortem follow-ups Identify systemic delivery bottlenecks and drive continuous improvement Provide concise executive updates with clear status, risks, and asks Maintain decision logs and documentation for key initiatives Translate technical complexity into business impact What you’ll need: Must-haves 6+ years of experience in technical program management, engineering pr

awsazuregcp
View job →
P
Particle41
📍 India• Remote
13 days ago

UI/UX Designer Projects often need the technical expertise of more than one person. Particle41 provides expert teams that embed directly into businesses for immediate impact and execution towards developing outstanding software. Our clients range from startups to small and medium-sized enterprises. As a UI/UX designer, you’ll work on a handful of digital products across the web and mobile. You’ll communicate with clients to help them understand their customers’ needs, and make design recommendations based on your findings. In This Role, You Will: Apply human-centered design principles to software creation. Develop concepts and prototypes for stakeholders. Iterate quickly on your designs in a lean, agile environment. Articulate design decisions and diplomatically navigate feedback. Conduct user research and analyze data to inform design decisions. Ensure brand and product cohesion by leveraging design systems. Requirements Gathering and Analysis Collaborate with designers, product managers, and other stakeholders to gather requirements and translate them into technical solutions. Participate in requirement analysis sessions to understand business needs and user requirements. Provide technical insights and recommendations during the requirements-gathering process. Agile Development Participate in Agile development processes, including sprint planning, daily stand-ups, and sprint reviews. Work closely with Agile teams to deliver software solutions on time and within scope. Adapt to changing priorities and requirements in a fast-paced Agile environment. Testing and Debugging Conduct thorough testing and debugging to ensure the reliability, security, and performance of applications. Write unit tests and validate the functionality of developed features and individual elements. Writing integration tests to ensure different elements within a given application function as intended and meet desired requirements. Identify and resolve software defects, c

REMOTEawsazure
View job →
TI
TEGNA India
📍 Chennai• Full-time
18 days ago

TEGNA Inc. helps people thrive in their local communities by providing the trusted local news and services that matter most. With 64 television stations in 51 U.S. markets, TEGNA reaches more than 100 million people monthly across web, mobile apps, streaming, and linear television, while also maintaining a strong global presence in India with offices in Bangalore and Chennai that support technology, product, and business operations initiatives. Together, we are building a sustainable future for local news. Senior Backend Java Developer Position Summary TEGNA is looking for an experienced Senior Backend Java Developer to join our highly collaborative and agile engineering team. This role is ideal for someone who enjoys solving complex technical challenges and building scalable backend platforms in high-performance environments. You will work on designing and developing cloud-native applications and distributed systems while collaborating across engineering teams to deliver reliable and innovative solutions. What You’ll Do • Design, develop, and implement scalable backend services using Java, Spring Boot, and AWS. • Build and optimize microservices-based architectures to support high-performance applications. • Develop reliable and scalable solutions for distributed systems operating in high-traffic environments. • Partner with cross-functional teams to understand requirements and deliver quality solutions. • Improve application performance, resiliency, scalability, and system reliability. • Participate in code reviews and contribute to engineering best practices and quality standards. • Drive continuous improvements through emerging technologies and modern development practices. • Support cloud deployment, monitoring, and operational excellence initiatives. What You’ll Bring • Bachelor’s or master’s degree in computer science, Software Engineering, or

javaawsci/cd
View job →
A
13 days ago

AI Teammates are Asana’s flagship AI innovation - autonomous agents embedded directly into enterprise workflows that reason, act, and proactively drive work forward. The AI Teammates Platform (AITP) team owns the execution engine and platform layer that makes these AI Teammates effective, reliable, and scalable. We build the core systems behind agentic execution: model evaluation and rollouts, proactive detection of blocked work, execution quality infrastructure, tool orchestration, and the platform APIs that power AI across all of Asana engineering. Rather than building a monolithic feature, we build a composable foundation. If a Teammate reasons, acts, or autonomously improves - that’s us. We’re looking for an experienced Technical Lead to drive the technical strategy, architecture, and execution for the AI Teammates Platform. You will bridge the gap between frontier AI capability and enterprise-grade software engineering, operating in close partnership with frontier model providers while building systems that scale to millions of users. If you are passionate about applied AI, agentic orchestration, high-reliability backend systems, and growing high-performing engineering teams, we’d love to hear from you! This role is based in our San Francisco office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What You’ll Achieve Set the technical strategy for the AI Teammates execution engine, tool orchestration, and platform APIs within Asana, advocating for engineering-driven investments with a vision for keeping our systems flexible, reliable, and maintainable to meet customer needs now and in the future Drive the team to continually and holistically

aiproject management
View job →
E
18 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Portworx team to build and deliver our highest-quality product suite. In this role, you will write clean, scalable code with a strong focus on quality, reliability, and user-centric design. You will directly contribute to building a new SaaS platform that delivers a secure, consistent, and best-in-class experience for customers purchasing and managing Portworx offerings. As a core developer, you will take ownership of designing and implementing critical features across the entire Portworx portfolio. WHAT YOU’LL DO Design & Scale SaaS Microservices: Develop, test, and integrate high-performance microservices and features into the Portworx product suite, ensuring high availability in distributed systems. Drive End-to-End Delivery: Lead software lifecycle activities including architectural design, code reviews, unit/functional testing, documentation, and continuous integration and deployment (CI/CD). Partner Across Teams: Collaborate with product managers, cross-functional engineering peers, and early-adopter customers to transform requirements into production-ready software. Own Product Quality & Iteration: Take full ownership of feature stability by proactively incorporating customer feedback and rapidly resolving issues identified during testing and deployment. Innovate & Experiment: Research emerging technologies and cloud infrastructure tools to push performance boundaries and continuously i

javaawskubernetes
View job →
E
18 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Portworx team to build and deliver our highest-quality product suite. In this role, you will write clean, scalable code with a strong focus on quality, reliability, and user-centric design. You will directly contribute to building a new SaaS platform that delivers a secure, consistent, and best-in-class experience for customers purchasing and managing Portworx offerings. As a core developer, you will take ownership of designing and implementing critical features across the entire Portworx portfolio. WHAT YOU’LL DO Design & Scale SaaS Microservices: Develop, test, and integrate high-performance microservices and features into the Portworx product suite, ensuring high availability in distributed systems. Drive End-to-End Delivery: Lead software lifecycle activities including architectural design, code reviews, unit/functional testing, documentation, and continuous integration and deployment (CI/CD). Partner Across Teams: Collaborate with product managers, cross-functional engineering peers, and early-adopter customers to transform requirements into production-ready software. Own Product Quality & Iteration: Take full ownership of feature stability by proactively incorporating customer feedback and rapidly resolving issues identified during testing and deployment. Innovate & Experiment: Research emerging technologies and cloud infrastructure tools to push performance boundaries and continuously i

javaawskubernetes
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As the Team Lead for Initiator & Protocol Engineering, you will spearhead the critical bridge between our industry-leading FlashArray and the Linux/VMWare ecosystems. You will drive the performance and reliability of our storage protocol stacks—spanning NVMe over Fabrics and Fibre Channel—ensuring Pure Storage remains the gold standard for enterprise connectivity. Collaborating closely with cross-functional hardware and software teams, you’ll mentor a high-caliber engineering squad to solve complex kernel-level challenges and influence the global Linux upstream community. WHAT YOU'LL DO Own the Protocol Lifecycle: Lead the development, maintenance, and optimization of Linux and VMWare initiator stacks (NVMeoF, FC-SCSI, iSCSI) and target drivers to ensure seamless, high-performance integration with Pure FlashArray. Drive System Resilience: Architect enhancements for Fibre Channel and NIC driver stacks that improve RAS (Reliability, Availability, and Serviceability), specifically focusing on multipathing logic and link health monitoring. Technical Leadership & Mentorship: Guide a team of senior and junior engineers through complex project deliveries, conducting deep-dive code reviews and setting the technical bar for C/C++ and Python development within the kernel space. Solve the Impossible: Act as the final escalation point for the most challenging system-level bugs found in the field or internal testing, u

pythonawslinux
View job →
E
Everpure
📍 Bengaluru• Full-time
18 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE The FlashArray team builds an innovative, high-performance, and highly available portfolio of products designed for demanding, mission-critical applications. While we deliver a hardware storage array, over 90% of our engineering team focuses on software engineering. Our customers value FlashArray for its simplicity of management, continuous feature upgrades, and ability to stay on the cutting edge without downtime. Recently, we extended these capabilities into the public cloud with CloudSnap and Cloud Block Store for AWS, enabling customers to leverage cloud agility for both traditional IT and cloud-native applications. WHAT YOU’LL DO Design & Implement: Create innovative algorithms and technologies for high-performance systems targeting six-nines (99.9999%) reliability. End-to-End Ownership: Lead feature innovation from initial concept through to shipped product. Problem Solving: Analyze and resolve complex technical challenges through persistent problem-solving and technical insight. Collaborate & Deliver: Partner closely with smart, collaborative peers to deliver features that directly enhance customer experience and satisfaction. Learn & Grow: Continuously expand domain expertise in systems software within a supportive, knowledge-sharing environment. WHAT YOU BRING Software Development Experience: 3+ years of professional development experience using C, C++, Python, Go, Java, or related prog

pythonjavaaws
View job →
E
18 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the FlashArray team to build high-performance, resilient storage software that powers mission-critical applications worldwide. As a core member of an agile engineering group, you will design and deliver zero-downtime algorithms that directly shape our enterprise storage and public cloud offerings. In this role, you will partner closely with global systems engineers and product teams to translate complex technical challenges into scalable, real-world solutions. Your work will directly impact how thousands of global enterprises manage, protect, and scale their data seamlessly. WHAT YOU’LL DO Design & Deliver Resilient Systems: Architect and implement high-performance algorithms for enterprise storage products, ensuring platform reliability, end-to-end delivery from concept to release, and six-nines availability. Expand Cloud Architecture: Extend core platform capabilities into public cloud environments (such as AWS), driving performance and agility for both traditional IT and cloud-native applications. Drive Problem Solving & Quality: Analyze and resolve complex systems software challenges, optimizing storage internals and contributing to continuous continuous platform upgrades. Collaborate & Mentor: Partner with multidisciplinary engineering peers to review code, refine architecture, and maintain high engineering standards across distributed software projects. WHAT YOU BRING Systems Programming Exp

pythonjavaaws
View job →
D
18 days ago

DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. What you’ll do: Manage end-to-end release cadences for both DataHub Core (OSS) and our Managed Cloud Coordinate cross-functional teams including engineering, product, QA, DevOps, documentation, and customer success Define Quality Gates: Establish the "Go/No-Go" criteria that ensure every release is rock-solid Build Automation: Design the dashboards and runbooks that turn manual release chaos into a streamlined machine Risk Mitigation: Identify breaking changes and dependency conflicts before they hit production Serve as the central point of contact for all release-related questions and updates Facilitate release retrospectives and drive continuous improvement initiatives Required Qualifications 5+ years of experience in technical program management or release management roles Proven track record of managing complex software releases in fast-paced startup environments Deep understanding of software development lifecycles, CI/CD practices, and DevOps principles Experience with both open-source and enterprise software release models is a strong plus Strong technical background with ability to understand architecture, dependencies, and technical trade-offs Excellent project management skills with proficiency in tools like Jira, Linear, or similar Outstanding commun

dockerkubernetesci/cd
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime