About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
Jobiba hiring network
Software Reliability Engineer Jobs
6,326 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software, and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers, and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Validation leadership team, the Principal Execution and Quality Validation Engineer will be responsible for driving validation execution, automation, product quality, and release readiness across Graphcore silicon and platform technologies. The role combines deep expertise in hardware validation, test automation, and quality engineering with a strong focus on execution excellence. Working closely with architecture, design, verification, firmware, software, systems engineering, and validation teams, the successful candidate will develop scalable validation methodologies, improve test coverage and execution efficiency, and ensure products meet the highest standards of functionality, reliability, and performance before customer deployment. As a recognized technical leader within the validation organisation, this role will influence validation strategies, automation roadmaps, and quality practices across multiple projects and engineering disciplines. The Team The Validation Execution and Quality team sits within the Validation organisation and is responsible for improving validation effectiveness, te
Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu
Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu
Senior -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to deb
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software, and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers, and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Validation leadership team, the Senior Execution and Quality Validation Engineer will be responsible for executing validation plans, developing automation solutions, and improving product quality across Graphcore silicon and platform technologies. Working closely with architecture, design, verification, firmware, software, systems engineering, and validation teams, the successful candidate will contribute to scalable validation methodologies, improve test coverage and execution efficiency, and help ensure products meet high standards of functionality, reliability, and performance before customer deployment. The role requires strong technical skills, attention to detail, and a passion for improving validation quality through effective execution, automation, and continuous improvement. The Team The Validation Execution and Quality team sits within the Validation organisation and is responsible for improving validation effectiveness, test execution efficiency, product quality, and release readiness across Graphcore silicon and platform products. The team develops validation methodologies, automation
About the Team pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model. About the Role We’re looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You’ll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production — directly improving how quickly and reliably research moves forward. In this role, you will: Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance. Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention. You might thrive in this role if you: Are excited about the potential of personal AGI and want to build the infrastructure that enables it. Have strong software engineering fundamentals and experience building or operating large-scal
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Company Description Okta is the leading independent provider of enterprise identity. The Okta Identity Cloud enables organizations to securely connect the right people to the right technologies at the right time. With over 6,500 pre-built integrations to applications and infrastructure providers, Okta customers can easily and securely use the best technologies for their business. Over 7,950 organizations, including 20th Century Fox, JetBlue, Nordstrom, Slack, Teach for America, and Twilio, trust Okta to help protect the identities of their workforces and customers. Position Description We are seeking an experienced Full Stack Senior Software Engineer to play a key role in building and scaling the Okta Recovery Vault (ORV). This team is responsible for Okta's enterprise-grade soft-delete and object recovery capability, designed to protect critical identity objects (Users and Groups) from accidental or malicious deletion. As a Senior Engineer, you will own the technical design, implementation, and operational reliability of critical components within our real-time, high-fidelity recovery system — spanning backend services and the admin-facing UI that customers use to review and restore their data. You will solve complex engineering problems around identity preservation (UUIDs) and relationship restoration—including group memberships, app assignments,
We are seeking an Engineering Manager to join our growing Gurugram Product & Technology team to provide technical direction, direct architecture, and implement core parts of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As an Engineering Manager on this new team, you will be responsible for leading and growing an engineering team, taking on challenging, high-visibility projects that improve and enhance the performance, scalability, and reliability of the distributed systems infrastructure for this new product. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Position Expectations Provide technical leadership and mentorship to a team of engineers, fostering a culture of innovation, quality, and continuous improvement Drive the architectural vision for the platform, ensuring it runs equally well on public clouds, private cloud environments and on-premise Work with product managers, program managers, design & analytics teams and other teams to define, prioritize and deliver new features that delight our users and drive platform improvements Take responsibility for the planning and execution of major features, raise delivery risks Own the monitoring, operations, and maintenance of the systems your team develops Enable the team to operate efficiently by removing technical obstacles, coordinating with other teams on dependencies, and prioritizing the team's overall well-being Contribute to planning for organizational growth, including allocation of engineering resources, participate in hiring and assignment of projects Qualifications 8+ years of experience of building distributed systems, and/or foundatio
About the Team OpenAI's People team helps hire, develop, and support the people building safe and beneficial AGI. Within that team, People Systems builds the technical foundation that enables our HR, recruiting, payroll, benefits, and performance operations to scale with quality, speed, and rigor. We work at the intersection of HR systems, software engineering, and internal tooling. Our goal is not just to keep core systems running, but to build durable technical leverage for the company. About the Role We're hiring a Workday Engineer to help design, build, and operate the systems that power critical people workflows at OpenAI. This is a highly technical role for someone who combines strong Workday expertise with real engineering fluency. You'll build reliable integrations, improve system architecture, automate complex workflows, and help connect Workday to internal tools, external platforms, and emerging AI-driven systems. You should be comfortable going beyond configuration work. We're looking for someone who can reason through ambiguous systems problems, write and debug technical solutions, work effectively in Git-based environments, and use modern developer workflows, including CLI-driven tooling, to build and operate with speed and discipline. You'll partner closely with cross-functional teams across People, Finance, Security, and Engineering, including our People Innovations team, to build systems that are secure, scalable, and practical. Some work will involve improving mature production infrastructure; some will involve building entirely new workflows and capabilities from scratch. In this role, you will Design, build, and maintain Workday integrations, applications, and workflow automations across domains such as payroll, benefits, recruiting, performance, and case management Improve the reliability, quality, and scalability of People systems through strong engineering, testing, and operational practices Build technical solutions that connect Workday with i
Interactive Design Engineer - HPIQ Description - About The Role As a Product Design Engineer at HP IQ, you’ll work at the intersection of creativity, engineering, and design to develop groundbreaking devices that redefine how people interact with technology. You’ll collaborate closely with industrial designers, hardware engineers, and interaction designers to transform ideas into functional prototypes and production-ready designs. From 3D CAD modeling and hands-on prototyping to engineering analysis and design validation, you’ll help solve complex mechanical challenges and drive designs from early concepts toward production. You’ll have the opportunity to take ownership of meaningful engineering work while learning from an experienced, multidisciplinary team and contributing to products that push the boundaries of what’s possible. What You Might Do Design mechanical parts, components, and assemblies for new consumer devices using 3D CAD (NX) Create detailed 2D engineering drawings, define specifications and tolerances, and work closely with overseas vendor partners to ensure accuracy and quality Design and run experiments, applying engineering analysis and test results to guide design decisions Evaluate materials, manufacturing processes, and design tradeoffs to develop robust and scalable solutions Work cross-functionally with hardware, software, industrial design, and other engineering teams to bring concepts from ideation through development Conceptualize, design, and build prototypes that quickly validate ideas Participate in design reviews, clearly communicating design decisions, technical analysis, and tradeoffs to both technical and non-technical audiences Build, test, troubleshoot, and iterate prototypes to identify issues and improve product performance, reliability, a
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. Today, we are tapping into the unlimited potential of AI to define the next era of computing. As an NVIDIAN, you will address challenges spanning architecture, silicon, firmware, software, and production — and excellent judgment matters as much as technical depth! We are the Silicon Power Team within the Silicon Co-Design Group. We architect and deliver groundbreaking solutions for productizing NVIDIA's chips across consumer, professional, server, embedded, mobile, and automotive markets. Silicon characterization, correlation to arch and design expectations, product spec finalization, and productization techniques and infrastructure are our day-to-day work — always on the bleeding edge of the industry. Small decisions here have outsized impact on performance, efficiency, reliability, bring-up speed, and ultimately what the product delivers in the field. We are hiring a Senior Silicon Power Engineer to own power-feature productization on a flagship silicon program. This is not a coordination role, and it is not a compliance role — it is the seat where power features either work at scale or become the reason a program slips. The two highest-leverage problems in this seat: Close the hardest multi-functional power failures before they gate a program. Take ambiguous, cross-boundary issues across architecture, firmware, validation, and platform to root-cause closure — with productized fixes and reusable methodology the next program can inherit. Build AI-enabled characterization as a real capability, not a demo. Every bring-up generates terabytes of characterization, shmoo, and telemetry data. Deploy AI workflows for data analysis, metric extraction, trend detection, and cross-bring-up correlation — with the guardrails and validation discipline to make them trustworthy enough to gate production decisions! <
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Senior Professional Services Engineer at GitLab, you'll work directly with customers to deliver installation, migration, training, and advisory services that help them adopt GitLab successfully. You'll lead engagements from single-node Omnibus installs to large reference architectures built with infrastructure as code (IaC) and configuration as code. You'll also guide migrations from other systems to GitLab SaaS or self-managed deployments, and help customers make practical decisions that improve reliability, security, and day-to-day workflows. In this role, you'll work closely with customers and Git
Job Title Automation Engineer - C# Job Description Automation Engineer - C# As an Automation engineer, you will ensure that the complete and integrated MR systems meet the requirements as defined in the System Requirements Specifications and that all features are implemented and verified correctly. To improve test efficiency and coverage, selected verification tests and regression test suites are automated using C# .NET. The Test Automation Engineer plays a key role in the development, execution, and maintenance of automated test suites, working in close collaboration with verification engineers and the test automation team located in Bangalore, India. Your Role: 5+ years of proven experience in software development and/or test automation within a complex, high‑tech environment. Effectively communicates with stakeholders, escalates or removes impediments, supports risk management, and drives continuous improvement. Keeps technical knowledge up to date and translates emerging trends (e.g., Model‑Based Testing, AI‑driven testing) into practical applications within a high‑tech environment. Coaches and mentors team members on test automation practices, tools, and processes. Strong expertise in software development, testing, and debugging, with a quality‑first mindset. Expertise in C#. Good understanding of modern test automation trends, frameworks, and best practices. Experience in setting up, evolving, and maintaining test automation infrastructure. Strong quality drive, with attention to robustness, reliability, and maintainability of test solutions. Experience working in global, multicultural teams, collaborating across sites and disciplines. Demonstrates a continuous improvement mindset and leads by example. Strong communication and documen
NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent engineering manager to own and deliver an end to end manageability stack for Data Center Systems. We are seeking an experienced manager who is deeply technical, hands-on, and has a wide system view. You will manage a team of experts, design & build OpenBMC based manageability software stack for NVIDIA’s next generation Data Center Compute Systems. We want to grow our teams with the smartest people in the world. If you're creative and autonomous, we want to hear from you! What you’ll be doing: Own and deliver OpenBMC based manageability stack for next generation Data Center Compute Systems. Own firmware delivered to data centers in terms of quality, reliability and telemetry performance. Manage and lead a distributed team of software engineers to deliver firmware stack with high quality. Work with data center architects and cloud customers for correct requirements and scope implementation to ensure speed of light product development. Work closely with cross functional teams to ensure scalable manageability architecture for all data centers products Drive efficiency, reliability and optimization in firmware architecture from a data center view point. Work closely with customers and internal teams to resolve issues at Speed of Light. What we need to see: BS, MS, or PhD in EE/CS or related field o
Get new software reliability engineer jobs by email
Daily job updates · Unsubscribe anytime