Forward was founded in 2013 by four Stanford Ph.D.s, building the industry's first network digital twin: a mathematically accurate model of the production network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change before it touches production. That founding instinct still defines how we work. We're accurate and evidence-driven, relentless about clarity, and we'd rather be certain than comfortable, building a groundbreaking platform that transforms how teams run and secure networks across every major cloud and vendor environment. Global leaders like Goldman Sachs, PayPal, S&P Global, IBM, and Dell trust Forward, alongside fast-growing enterprises and government agencies, realizing an average of $14.2 million in annual benefits, according to IDC. Backed by top-tier investors, including A. Capital, Andreessen Horowitz, Goldman Sachs, MSD Partners, Omega Venture Partners, Section 32, and Threshold Ventures, and headquartered in Santa Clara, we're most proud of our team: curious people who'd rather build what doesn't exist than accept how things have always been done. Forward is currently seeking a Senior Backend Software Engineer to work as part of our Platforms team. You will play a critical role in designing, developing, and scaling the core backend services and infrastructure that support our SaaS and on-prem deployments. Your contributions will have a direct impact on the stability, performance, and scalability of our platform, helping to ensure an exceptional experience for our customers. Responsibilities: Platform development: Contribute to the design and development of storage systems, job scheduling systems, data ingestion frameworks, monitoring frameworks etc to ensure high system performance and availability. Feature development: Build and maintain backend frameworks that support essential platform features Scalability & Reliability: Develop scalable, high-performing
Jobiba hiring network
Reliability Engineer Jobs
2,028 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Forward is transforming how the world’s most complex networks are managed and secured. Founded in 2013 by four Stanford Ph.D.s, we built the industry’s first network digital twin — a mathematically precise model of the production network that gives IT teams unmatched visibility, verification, and agility across every major cloud and vendor environment. Our customers include global leaders such as Goldman Sachs, PayPal, S&P Global, IBM, and Dell, as well as fast-growing enterprises and government agencies. According to IDC, Forward customers realize an average of $14.2 million in annual benefits through improved efficiency and security. Backed by world-class investors including Andreessen Horowitz, Goldman Sachs, MSD Partners, and Threshold Ventures, Forward offers a people-centric, innovative culture where brilliant minds are shaping the future of network reliability, security, and AI-ready operations. Forward is currently seeking experienced Java developers to work as part of our Network team. Responsibilities Help bring the best ideas from the software development world into the networking industry. Contribute to our code base, systems and software architecture as a member of our engineering team. Help create and optimize network device models for different device vendors and protocols. Help create infrastructure needed to configure, collect and test network devices. Work with peers who are experts in Networking, Distributed Systems, Big Data and Search. Requirements 5+ years of work experience in software development 3+ years of work experience with Java BS in Computer Science or related degree Solid software engineering experience with large code bases Basic understanding of networking and TCP/IP. Strong verbal and written communication skills. Nice to haves Working knowledge of how switches, routers, firewalls or load balancers work. Experience working with networking protocols such as BGP/OSPF/IS-IS, IPv4/IPv6, MPLS, VLAN, VXLAN, etc. This position is a re
A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO As Database Support Engineer, you’ll support various critical database platforms across Development, QA, UAT, and Production environments. The role partners closely with application teams, application support, and database engineers and operates within a Follow‑the‑Sun model to ensure availability, performance, and reliability of database services. Key responsibilities include: • Provide operational support for enterprise database platforms in both on-prem private cloud and public cloud • Monitor database health, capacity, performance, and availability, and respond to alerts, diagnose issues, and perform timely remediation • Perform routine maintenance activities (patching, upgrades, housekeeping etc) • Troubleshoot database‑related incidents and collaborate on root cause analysis • Work closely with application owners, application support teams, and DB Engineers • Provide guidance on database best practices and operational standards • Participate in cross‑team problem resolution and continuous improvement initiatives • Contribute to design, implementation and testing of automation and self service capabilities of DB platforms • Drive continuous improvement, identifying opportunities to reduce toil and increase platform efficiency. • Participate in a Follow‑the‑Sun operating model, including shift‑based coverage and handoffs WHAT’S REQUIRED • Bachelor’s degr
Here’s a summary of the role: Build software that matters, take real technical ownership, and use modern AI tooling to do your best work. This is a hands-on senior engineering role for someone who enjoys solving complex product problems, shaping robust solutions, and helping teams deliver reliable services at scale. You’ll work on secure, scalable microservices and APIs using TypeScript, AWS, and modern engineering practices. You’ll play a leading role within a collaborative product engineering team, owning complex features end to end, contributing to design and architectural decisions, supporting production systems, and helping raise the bar across backend, cloud, and AI-assisted development workflows. Here’s a breakdown of what you’ll do, not all of it, just the important stuff: Own and deliver complex backend services and APIs using Node.js, TypeScript, and AWS , from technical design through release and production support. Contribute to design and architecture discussions, making pragmatic decisions that balance delivery speed, maintainability, scalability, and security. Mentor and support less experienced engineers through code reviews, pairing, technical guidance, and day-to-day collaboration. Work closely with product managers, designers, and engineers across the team to turn requirements into practical, reliable solutions. Use AI tools to accelerate coding, debugging, testing, research, and documentation, while validating outputs carefully and applying sound judgment. Strengthen service reliability, observability, and engineering quality by improving monitoring, incident response, testing, and development practices. These are the essentials you’ll need to get an interview: 5 to 8 years of professional software engineering experience delivering production systems in an agile environment. Strong backend development s
Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human evaluation and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Platform Engineering team. In this role, you will lead the design and development of core data storage, streaming, caching, and indexing platforms and underlying systems. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the architecture, design, implementation, and reliability of our foundational data platforms and systems, working closely with stakeholders and internal customers to understand and refine requirements. Collaborate with cross-functional teams to define, design, and deliver new features. Proactively identify opportunities for, and driving improvements to, current programming practices, including process enhancements and tool upgrades. Present technical information to teams and stakeholders, providing
Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. About the General Agents Team The General Agents team, part of Scale’s Enterprise organization, builds robust general agents for customer use cases and applications. The team sits at the intersection of frontier agent development and real-world deployment, translating state-of-the-art reasoning and agentic capabilities into reliable, production-grade systems that drive real economic value. Our agents are scalable systems built around recurring enterprise problem domains, with a strong emphasis on generalization, extensibility, and deployment across many customers. About the Role As a Senior/Staff Machine Learning Engineer (MLE) on the General Agents team, you’ll play a critical role in designing, building, and deploying production-ready AI agents that solve high-impact enterprise problems. You will work across the full agent lifecycle—from model and system design to evaluation, deployment, and iteration—bridging cutting-edge agentic techniques with the constraints and requirements of real customer environments. You will: Design and implement end-to-end agent systems that combine LLM reasoning, tool use, memory, and control logic to solve recurring enterprise use cases. Build scalable, reliable agent architectures that can be deployed across many customers with varying data, tools, and constraints. Develop evaluation frameworks, datasets, environments, and metrics to measure agent performance, reliability, and business impact in production settings. Collaborate closely with product managers, customers, data annotators, and other engineering teams to translate enterprise requirements into robust agent designs. Productionize frontier agent techniques (e.g.,
OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric—with AI embedded across the automation lifecycle—Redwood accelerates the path to autonomous operations. Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT We are looking for a Software Engineer, Platform & Integrations . Working closely with senior and lead engineers, you will design, develop, and maintain high-quality features that power enterprise data exchange for more than 1,000 customers worldwide. This is an incredible opportunity to deepen your expertise in cloud-native architectures, enterprise security, and modern DevOps practices in a fast-growing product environment. Feature Development & Design: Write clean, maintainable, and well-tested code using Java and Spring Boot to deliver scalable backend services and microservices. Platform Reliability: Contribute to enhancing the monitoring, logging, and observability of our core platform to ensure high availability and performance. Security & Compliance: Implement secure coding practices to safeguard data exchange and maintain compliance across our cloud infrastructure. Collaborative Execution: Work within an agile team, collaborating closely with QA, Product, and fellow engineers to deliver high-qual
Manufacturing Test Engineering Manager Position Summary We are seeking an experienced Manufacturing Test Engineering Manager to lead the development and execution of the end-to-end manufacturing test strategy for next-generation AI server platforms and datacenter infrastructure. This role is responsible for defining and driving the manufacturing test architecture from L6 board assembly through L11 rack-level integration and final system validation , ensuring world-class product quality, manufacturability, and production scalability. This leader will manage a team of 3–5 Manufacturing Test Engineers while partnering closely with Hardware, Firmware, Platform, Validation, Quality, Operations, and Joint Design Manufacturing (JDM) partners. The role owns the manufacturing test strategy, test coverage, factory test infrastructure, manufacturing capacity planning, and continuous improvement of manufacturing quality. Key Responsibilities Manufacturing Test Strategy Define and own the end-to-end manufacturing test strategy from L6 board assembly through L11 rack integration . Develop standardized manufacturing test methodologies that optimize quality, throughput, cost of test, and scalability across multiple products and JDM sites. Establish manufacturing test standards, best practices, and engineering processes that support high-volume server manufacturing. Technical Leadership & People Management Lead, mentor, and develop a team of 3–5 Manufacturing Test Engineers supporting multiple hardware programs. Establish team priorities, allocate resources, and ensure successful execution of manufacturing test deliverables. Foster a culture of technical excellence, accountability, collaboration, and continuous improvement. Serve as the primary technical escalation point for manufacturing test and production issues. Cross-Functional Engineering Collaboration Partner with Hardware, Platform, Firmware, Validation, Reliability, Quality, and Operations teams to ensure manufac
1743 - This position is in Austin, Texas. Position Summary We are seeking an experienced Board-Level Hardware Validation Engineer to define and execute the validation and verification of complex electronic systems throughout the product lifecycle. This role is responsible for defining validation strategies, developing test plans, executing hands-on testing, analyzing failures, and working directly with ODM partners to ensure products meet performance, reliability, quality, and compliance requirements before mass production. The ideal candidate combines strong electrical engineering fundamentals with practical lab expertise and is comfortable personally performing validation activities while coordinating with cross-functional teams and manufacturing partners. Key Responsibilities Validation Strategy & Planning Define comprehensive board-level and inter-board validation plans based on product requirements, design specifications, and customer use cases. Develop validation methodologies covering functional, electrical, thermal, power, signal integrity, reliability, and stress testing. Establish test coverage, acceptance criteria, qualification requirements, and release gates. Review hardware architecture, schematics, component specifications, and interface topologies to identify validation risks early in the design cycle. Define incremental validation and regression coverage for component substitutions, design changes, and firmware updates. Hands-On Validation Execution Develop, automate, and execute validation tests on prototype and production-intent hardware. Perform board bring-up, functional verification, electrical characterization, and system-level integration testing. Validate communication interfaces, control signals, and timing requirements. Verify power sequencing, reset behavior, leakage current, and recovery across operating states. Execute temperature and voltage corner testing against approved operating limits. Use oscilloscopes, logic an
Job Summary Reporting to the Memory Validation leadership team, the Senior Silicon DDR/HBM Validation Engineer will be responsible for the bring-up, validation, characterization and debug of advanced memory subsystems used in next-generation AI compute platforms. The role will focus on DDR and HBM technologies, working closely with silicon design, firmware, characterization, platform and systems teams to ensure robust memory subsystem functionality, performance and reliability. The successful candidate will take ownership of significant validation activities, contribute to debug and root-cause analysis efforts, and help improve validation methodologies, automation and infrastructure. The Team The Memory Validation team sits within the Validation organisation and is responsible for the bring-up, validation, characterization and debug of memory subsystems across Graphcore silicon and platform products. The team supports DDR and HBM validation activities throughout the product lifecycle, from first silicon through production readiness. Engineers work closely with architecture, RTL, firmware, characterization, systems and platform teams to ensure memory technologies meet functionality, performance, reliability and performance objectives. Responsibilities and Duties Execute validation and bring-up activities for DDR and HBM memory subsystems Verify memory bring-up software, firmware and scripts against defined project requirements Debug firmware, hardware and system-level issues and contribute to root-cause analysis activities Analyse system logs, validation data and characterization results to identify failures and performance issues Perform PHY characterization and analog-level analysis during stress testing and validation activities Develop and execute functional, stress, performance and corner-case validation tests Perform signal integrity, voltage, frequency and timing measurements using laboratory instrumentation Char
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. Today, CLEAR is well-known as a leader in digital and biometric identification, reducing friction for our members wherever an ID check is needed. We’re looking for a Senior Software Engineer to establish our Observability framework and foundations. You will join us to accelerate building and scaling our innovative systems that support our growing identity platform. You will drive on Observability best practices to find and fix gaps in our observability and our overall systems. You will also lead practices such as load testing, capacity planning, game days, chaos testing, and incident post-mortems. What You Will Do: Embed within the Engineering pillar to deeply understand the product and implement observability across all key flows Facilitate and build load testing cases, ensuring we understand the limits and scaling factors of our services and systems Contribute to observability and support the design of new services and systems, ensuring highly reliable and scalable concepts are implemented Build and lead practices such as game days, chaos engineering, and failure analysis Build long-term capacity plans, with an eye toward reliability and cost-efficiency Who You Are: 6+ experience writing production-grade software in a modern language, such as Java and Python. Strong knowledge of distributed systems concepts (think CAP theorem), microservices architecture, and distributed tracing . Experience with modern observability systems such as Datadog. Experience with performance debugging tools and patterns. You should be able to read a f
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We're looking for Senior Fullstack Software Engineers to help build the next generation of CLEAR's identity platform. Beyond verifying identity, we're creating a secure, networked digital identity that enables seamless experiences across travel, enterprise, healthcare, financial services, and beyond. As a Senior Software Engineer, you'll own complex technical problems from design through deployment, partnering closely with Product, Design, Security, and Operations to deliver reliable, scalable solutions. We're looking for engineers with a strong builder mindset who thrive in ambiguity, take ownership, and enjoy turning ideas into production systems. Level and team matching (open roles across the three pillars that make up Technology at CLEAR: Core Identity, CLEAR1 , and CLEAR Travel ) will occur towards the end of our interview process. Tech stack overview: Java / React / Typescript What you’ll do: Design, build, test, and deploy scalable full-stack applications that power CLEAR's identity platform. Own projects end-to-end from technical discovery and architecture through implementation, rollout, and operational support. Partner closely with Product, Design, Data, Security, and Operations to translate business problems into simple, scalable technical solutions. Drive engineering excellence by improving system reliability, performance, testing, observability, and developer experience. Contribute to architectural decisions and continuously improve the scalability, security, and maintainability of our platform. Mentor teammates through tho
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We're looking for Senior Fullstack Software Engineers to help build the next generation of CLEAR's identity platform. Beyond verifying identity, we're creating a secure, networked digital identity that enables seamless experiences across travel, enterprise, healthcare, financial services, and beyond. As a Senior Software Engineer, you'll own complex technical problems from design through deployment, partnering closely with Product, Design, Security, and Operations to deliver reliable, scalable solutions. We're looking for engineers with a strong builder mindset who thrive in ambiguity, take ownership, and enjoy turning ideas into production systems. Level and team matching (open roles across the three pillars that make up Technology at CLEAR: Core Identity, CLEAR1 , and CLEAR Travel ) will occur towards the end of our interview process. Tech stack overview: Python / React / Typescript What you’ll do: Design, build, test, and deploy scalable full-stack applications that power CLEAR's identity platform. Own projects end-to-end from technical discovery and architecture through implementation, rollout, and operational support. Partner closely with Product, Design, Data, Security, and Operations to translate business problems into simple, scalable technical solutions. Drive engineering excellence by improving system reliability, performance, testing, observability, and developer experience. Contribute to architectural decisions and continuously improve the scalability, security, and maintainability of our platform. Mentor teammates through t
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Fullstack Software Engineer on CLEAR’s Healthcare team, you will build and scale secure, interoperable identity and data solutions that connect patients, providers, and partners. You’ll operate at the intersection of modern web platforms, healthcare interoperability standards, and high-assurance identity systems powering frictionless, trusted healthcare experiences nationwide. A brief highlight of our tech stack: Python / React / Typescript AWS cloud What you’ll do: Design and deliver secure, scalable fullstack solutions that integrate with enterprise EHR systems and national health information exchange frameworks Build and maintain healthcare data integrations leveraging FHIR (RESTful APIs/JSON) and HL7 v2 messaging to enable compliant, real-time data exchange Develop identity resolution and patient matching capabilities using identifiers such as MRNs and NPIs to ensure integrity across disparate clinical systems Partner with Engineering, Security, Product, and Health Information Management teams to implement compliant, audit-ready workflows for regulated healthcare processes Collaborate with external vendors (e.g., Epic Technical Services) to troubleshoot integration issues, manage deployments across TST/PRD environments, and ensure production reliability How you’ll measure success: Successful delivery and stability of FHIR/HL7 integrations across healthcare partners Reduction in data integrity issues related to patient matching and identity resolution High system uptime and successful production deployments across tiered
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We're looking for Senior Fullstack Software Engineers to help build the next generation of CLEAR's identity platform. Beyond verifying identity, we're creating a secure, networked digital identity that enables seamless experiences across travel, enterprise, healthcare, financial services, and beyond. As a Senior Software Engineer, you'll own complex technical problems from design through deployment, partnering closely with Product, Design, Security, and Operations to deliver reliable, scalable solutions. We're looking for engineers with a strong builder mindset who thrive in ambiguity, take ownership, and enjoy turning ideas into production systems. Level and team matching (open roles across the three pillars that make up Technology at CLEAR: Core Identity, CLEAR1 , and CLEAR Travel ) will occur towards the end of our interview process. Tech stack overview: Java / React / Typescript What you’ll do: Design, build, test, and deploy scalable full-stack applications that power CLEAR's identity platform. Own projects end-to-end from technical discovery and architecture through implementation, rollout, and operational support. Partner closely with Product, Design, Data, Security, and Operations to translate business problems into simple, scalable technical solutions. Drive engineering excellence by improving system reliability, performance, testing, observability, and developer experience. Contribute to architectural decisions and continuously improve the scalability, security, and maintainability of our platform. Mentor teammates through tho
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime