Jobiba hiring network

Reliability Engineer Jobs

2,028 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

T-
Tubi - Canada
📍 Toronto• Full-time• From C$1.4M/yr
18 days ago

About the Role: We're hiring Senior and Staff Data Platform Engineers to join the Data Infrastructure teams in Toronto. Together these teams own the infrastructure that processes billions of events per day: Spark-on-Kubernetes, Flink and Kinesis pipelines, a multi-petabyte Delta Lake, a large-scale MemoryDB feature store, Databricks multi-environment operations, and the catalog and lifecycle systems that govern it. The team is small and senior. Each engineer owns major platform components: you design it, build it, and support it in production. This is a hybrid-role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: Spark-on-Kubernetes — EKS-based compute platform for Spark workloads: cluster configuration, Pod Identity IAM, job environment setup, Kustomize overlays, and shadow canary validation Event ingestion — Rust services and Flink jobs processing billions of events per day over Kinesis; throughput, reliability, on-call response, and AI-assisted operational tooling to reduce toil Platform infrastructure — Terraform modules for environment provisioning, cross-account AWS IAM, ARC runner infrastructure, and CI/CD for data platform changes Feature store and ML compute — Flink-based real-time feature pipelines feeding a large-scale MemoryDB cluster; GPU capacity governance and Databricks multi-environment operations for ML training workloads Workflow orchestration and CDC — Airflow-based DAG deployment, change data capture pipeline operations, and data quality monitoring Your Background: 3+ years building and operating production data platform infrastructure at the cluster or platform level, across Spark, Flink, Kinesis, Kubernetes, or equivalent Deep experience in at least one of: Spark-on-K8s cluster operations, Rust-based data or systems engineering, Kubernetes platform engineering and IaC, or data catalog and governance tooling Production AWS experience or equivalent: EKS, S3, Kinesis, and mu

pythonjavaaws
View job →
T-
18 days ago

About the Role: The Machine Learning team at Tubi drives the innovation behind personalized user experiences. With the largest inventory in the industry and hundreds of millions of viewers, we tackle problems in the space of recommendations, search, content understanding, and ads optimization that shape the future of streaming. We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction and execution for scaling our machine learning capabilities while ensuring our distributed systems and infrastructure can support innovation at massive scale. You will combine technical depth with leadership excellence to guide teams that deliver both foundational ML systems and high-performance distributed services. This is a hybrid role for our Toronto office. What You'll Do: Lead and manage high-performing teams across ML engineering and ML infrastructure, fostering a culture of innovation, collaboration, and growth. Define and execute the strategic roadmap for ML systems, including recommendation, personalization, and ads optimization. Oversee the design, development, and deployment of scalable ML pipelines: data ingestion, feature engineering, model training, evaluation, and serving. Architect distributed systems to support ML workloads at scale, ensuring reliability, observability, and operational excellence. Partner closely with Product, Engineering, and Content teams to align on business goals and deliver impactful ML-driven experiences. Support best practices in experimentation, evaluation, and ML system monitoring. Ensure cost efficiency, scalability, and performance in ML infrastructure investments. Your Background: 10+ years of industry experience spanning machine learning engineering and distributed systems. 3+ years of leadership and management experience, with a proven ability to build and lead strong t

awsmachine learningai
View job →
F
FalconX
📍 Bengaluru• Full-time
18 days ago

Who are we? FalconX is a pioneering team of operators, investors, and builders committed to revolutionizing institutional access to the crypto markets. Operating at the intersection of traditional finance and cutting-edge technology, FalconX addresses the industry's foremost challenges: Navigating the digital asset market can be complex and fragmented, with limited products and services that support trading strategies, structures, and liquidity found in conventional financial markets. As a comprehensive solution for all digital asset strategies from start to scale, FalconX operates as the connective tissue empowering clients with seamless navigation through the ever- evolving cryptocurrency landscape. Senior Software Engineer - Risk Location: Bengaluru (Onsite) Impact You’ll architect and build the next-generation risk solutions for FalconX and its clients. FalconX is the institutional gateway to digital asset markets. FalconX processes more than $1T+ in trading volume across spot and derivative instruments, has financed $2.5B+ in loan originations and is the 1st CFTC-registered swap dealer focused on cryptocurrency derivatives. You will evolve FalconX’s risk technology to be best-in-class. You will work on cutting-edge technology with the right balance of speed, accuracy and reliability. Well-reputed financial institutions and crypto firms will trust and rely on your solutions for their day-to-day work. You will help unlock increased business potential and opportunities for FalconX and our clients. Responsibilities Develop, maintain and enhance FalconX’s proprietary risk management tools, infrastructure, risk data, and processes. Architect and build scalable, robust, performant risk solutions for institutional customers and internal teams. Work closely with multiple members of the risk team to develop, maintain and improve the risk management stack. Coach and mentor teammates on supporting and maintaining FalconX’s risk solutions. Requirements Degree in Compu

pythonawsgit
View job →
Z
Zocdoc
📍 Pune• Full-time
18 days ago

Our Mission You call. You wait. You call again. In every other part of your life, you book in seconds. In healthcare, you’re blocked. We’re here to give power to the patient. For nearly 20 years, we’ve built the leading healthcare marketplace - helping tens of millions of people find and book the care they need. Now, we’re going further: building our infrastructure beyond Zocdoc’s marketplace to power access to care wherever patients search, from provider websites and insurance directories to search engines, AI platforms, and more. Healthcare still lacks something every other major consumer industry takes for granted: a seamless way to go from seeking to getting . We don’t want to own the front door to care; there isn't one. We want to make sure all of those doors open when patients are knocking. Fixing healthcare starts with fixing access to it. And we're still just getting started. Your Impact on Our Mission We are looking for a Senior Software Engineer to join the teams that build Zocdoc's provider platform — Provider Infrastructure and Provider Roster Management. These teams create and scale the core systems that: Store and serve provider, locati on, an d roster data as a reliable source of truth. Power onboarding and ongoing management experiences for practices of all sizes, from solo providers to large health systems. Ensure provider changes propagate quickly and safely to downstream systems like Search, Availability, Analytics, and Integrations — so patients always see accurate information. We're building towards a future where Zocdoc can onboard and manage tens of thousands of providers under a single organization in days, not months, with predictable performance and reliability across the stack. Your work will directly improve how quickly providers get live on Zocdoc, how confidently clients manage their rosters, and how reliably downstream experiences behave as we scale. As a Senior Engineer here, you'll balance meaningful individu

reactawsgit
View job →
D
DevRev
📍 Austin• Full-time
18 days ago

About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. About the role We are looking for a Senior Data Engineer to help build and evolve the data platform that powers critical business decisions and customer-facing experiences. You will own significant parts of our data architecture that is main powerhouse of DevRev Computer’s memory for accurate and efficient Answers. As a part of data team, you will design and operate scalable data systems, and work closely with Software Engineering, AI Agent teams, Data Science, and Product teams to turn complex data requirements into reliable, high-quality data products. This role is ideal for an experienced engineer who enjoys solving challenging problems involving large-scale data, distributed systems, database architecture, and performance optimization. You will have significant technical ownership and the opportunity to influence the direction of our agentic data platform while helping raise the engineering bar across the team. Responsibilities Own data architecture for large-scale, high-impact projects, making thoughtful tradeoffs across scalability, reliability, performance, maintainability, and operational cost. Design, build, and operate scalable data pipelines and data systems that reliably ingest, transform, store, and serv

javascriptpythonjava
View job →
CC
CCL Confidential
📍 Vancouver• Full-time• C$125K – C$200K/yr
18 days ago

We deliver foundational systems that shape the future of how technology is used in a top performing quantitative equity fund that manages over $78+ billion USD in financial assets. This is a fantastic opportunity in the exciting intersection of finance and technology where investment decisions are made using technology. Quantitative equity funds use programmed investment strategies and as a result, our technology team is crucial to its success. The team is headquartered and deeply rooted in West Coast Vancouver. We place high value on maintaining an entrepreneurial spirit and creating a culture where each of us has opportunities to succeed. What You Will Do The technology infrastructure team plays an essential role through innovative technologies on our hybrid (on-premise and cloud based) platform: distributed computing, petabyte-scale data storage, containerization, non-traditional high-performance databases, process orchestration, monitoring, data visualization and DevOps. You own the entire technology infrastructure life cycle: Engineer and support software and systems infrastructure. Introduce new foundational technologies that advance our software engineering capabilities to the next level. Collaborate with our software development teams on support issues and improvements to our infrastructure tools, processes, and software. Act as a conduit between our application development teams, and IT, network security, and other stakeholders to align priorities and translate business requirements into technical designs. Improve systems infrastructure reliability. Gather and analyze metrics from operating systems and applications to assist in performance tuning, fault finding and business continuity planning. Design, plan and implement solutions in an entrepreneurial spirit. What You Bring Programming Knowledge – you have an undergraduate, graduate, or post-graduate degree in a computer-related field OR exceptional programming skills gain

typescriptpythonjava
View job →

Locations: South Jordan, UT Salary: $56,000 Launch Your Career in Technology Every app, website, payment, and digital service relies on technology running smoothly behind the scenes. When something goes wrong, Production Support Engineers are the people who investigate the issue, restore service, and help prevent it from happening again. If you're curious, analytical, and enjoy solving problems, this is an opportunity to build hands-on experience with cloud platforms, Linux, automation, databases, and large-scale enterprise systems from day one. What Is Production Support? Production Support Engineers keep business-critical applications running reliably in live environments. Think of it this way: Software Engineers build the platform. QA Engineers test the platform. Production Support Engineers keep the platform running when it matters most. Working at the intersection of technology and business, you'll troubleshoot issues, automate processes, and help improve the reliability and performance of systems used by thousands, or even millions, of people every day. If you enjoy solving puzzles, working under pressure, and understanding how large-scale systems work, this could be the perfect place to start your career. What You'll Do As part of a global production engineering team, you'll: Help support large-scale applications and platforms used by leading organizations around the world. Monitor business-critical applications and services to ensure high availability and performance. Investigate and resolve production incidents across applications, infrastructure, databases, and cloud environments. Analyse logs, alerts, and system metrics to identify root causes and prevent recurring issues. Partner with software engineers, infrastructure teams, and business stakeholders to improve system reliabilit

javascriptpythonjava
View job →
HI
HP IQ
📍 San Francisco• Full-time• $162K – $225K/yr
18 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Senior Software Engineer, Cloud Services, you will create scalable, reliable backend systems that support HP IQ's mission to transform the way people work. We value how we work as much as what we deliver. Our journey of continuous learning and evolution requires a flexible and resilient services platform that fosters innovation, experimentation, and adaptation. To enable this, we focus on designs and tools rooted in strong engineering principles like abstraction, composition, virtualization, automation, and iterative development cycles. What You Might Do Design, develop, and maintain backend services and RESTful APIs using Java and Spring Boot. Write clean, efficient, and well-tested code following established best practices. Collaborate with frontend developers, product managers, and other engineers to deliver end-to-end features. Integrate with databases and external services, ensuring performance, security, and reliability. Participate in code reviews, debugging, and performance optimization efforts. Contribute to CI/CD pipelines and support application deployment in cloud and

javasqlpostgresql
View job →
HI
HP IQ
📍 San Francisco• Full-time• $140K – $225K/yr
18 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As the Senior Software Engineer, Tooling and Development Infrastructure, you will play a critical role in shaping the developer productivity tools and automated testing strategy. You’ll collaborate closely with design, development, and quality teams to plan, design, and implement robust automated tools and services that ensure the quality and reliability of our AI software stack. You will be highly hands-on in your work and collaborate closely with stakeholders. This position offers a unique opportunity to influence the development of cutting-edge automation frameworks, foster a culture of quality, and contribute to the long-term success of the organization. What You Might Do Develop and implement automation frameworks and testing strategies that cover the entire software stack, from backend systems to user-facing features. Identify, evaluate, and integrate new tools that streamline development. This includes everything from code quality tools and to Infrastructure-as-Code (IaC) solutions. Lead continuous improvement efforts for our build, release, and test systems, ensuring a robust

pythonredisci/cd
View job →
HI
HP IQ
📍 San Francisco• Full-time• $179K – $252K/yr
18 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Lead Software Engineer, Cloud Services, you will create scalable, reliable backend systems that support HP IQ's mission to transform the way people work. We value how we work as much as what we deliver. Our journey of continuous learning and evolution requires a flexible and resilient services platform that fosters innovation, experimentation, and adaptation. To enable this, we focus on designs and tools rooted in strong engineering principles like abstraction, composition, virtualization, automation, and iterative development cycles. What You Might Do Lead technical strategy and execution across development and infrastructure, designing and implementing scalable, secure cloud-native systems while establishing architecture standards and engineering best practices. Own end-to-end delivery of platform and cloud services, including infrastructure buildout, application architecture, APIs, deployment pipelines, observability, reliability, and performance, while mentoring engineers and driving cross-functional alignment. Work with modern containerization and cloud technologies, includi

javaredisdocker
View job →
H
18 days ago

Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Manager, Electrical Engineering is responsible for the electrical systems of the KARNO generator, including high-voltage power electronics, battery systems, low- and high-voltage architecture, wiring harnesses, and the hardware that converts linear motion into electrical output. This is a working manager role: the position leads and develops a team of electrical engineers while remaining directly involved in technical execution, including circuit architecture, schematic review, and hardware bring-up in the lab. The Manager is accountable for the technical excellence, safety, and reliability of the electrical engineering function, and for establishing the design standards and review practices the team works to. The position plans team capacity, owns hiring and development for the electrical engineering staff, and partners with mechanical, controls, supply chain, and program management on system integration. KARNO systems are deployed in data center, military, and industrial applications. AI at Hyliion At Hyliion, AI is core to how we work. We equip every team member with leading AI tools and count on you to use them — to move faster, solve harder problems, and help us realize the full potential of KARNO technology for the world. Duties and Responsibilities Own critical electrical designs personally, including regular time at the bench and in the test cell, while leading the team as a practicing engineer. Lead the electrical engineering team in the design and development of KARNO generator electrical systems, including high-voltage power electronics, battery systems, linear generator power stages, and low-voltage controls hardwar

aigoexcel
View job →
S
StarRez
📍 Australia• Full-time
18 days ago

About StarRez StarRez is the global leader in student housing software, providing innovative solutions for on and off-campus housing management, resident wellness and experience, and revenue generation. Trusted by 1,400+ clients across 25+ countries, StarRez supports more than 4 million beds annually with its user-friendly, all-in-one platform, delivering seamless experiences for students and administrators. With offices in the United States, Australia, the UK, and India, StarRez blends the robust capabilities of a global organization with the personalized care and service of a trusted partner. The Role: We are looking for a Senior Quality Engineer - AI to help shape how StarRez evaluates, validates, and improves AI-powered product experiences. You will play a pivotal role in elevating our quality practices for AI-powered product experiences and tooling. You will be a product expert within our engineering organization, deeply understanding workflows and customer outcomes. This role will balance hands-on individual contributor responsibilities with leading, influencing, and coaching your peers. You’re someone who is passionate about seeing systems holistically and is eager to drive the highest standards of accuracy, safety, usefulness, and reliability. You will help define what "good" AI output means for StarRez, build repeatable evaluation systems, coach reviewers and subject matter experts, and turn subjective feedback into measurable product improvement. You will work at the intersection of AI engineering, product, and domain knowledge — partnering with engineers, product managers, and SMEs to raise the bar on how StarRez evaluates, monitors, and improves AI tools. Key Responsibilities: Lead or contributed to end-to-end AI quality and evaluation strategy for product experiences involving LLMs, RAG, prompts, tool-use, or agent workflows. Provided quality-focused input during ticket grooming, discovery, and feature discussions, with clear guidance on testability, ob

aigorust
View job →
W
Wellhub
📍 Brazil• Full-time• Remote
18 days ago

Your wellbeing, our mission. Join a company shaping a healthier world. GET TO KNOW US At Wellhub we're revolutionizing workplace wellness. Our platform connects employees worldwide to the best partners for fitness, mindfulness, therapy, nutrition, and sleep—all in one simple subscription. Headquartered in NYC with team members in Europe, North America and South America, we’re on a mission to make every company a wellness company. We believe work should be fulfilling, inspiring, and balanced. Here, you’ll find a team that values wellbeing, collaboration, and different perspectives, where passion and creativity push boundaries to create real impact. Your contributions will help shape a healthier, more balanced world for you and millions of people globally. Join us in redefining the future of wellbeing! THE OPPORTUNITY We are hiring a Staff Platform Engineer with a dedicated focus on our Observability ecosystem for our Platform area in Brazil ! This is a Remote – Brazil position, meaning you can work from anywhere within the country. Please note that this role is only open to candidates in Brazil. In an environment of rapid growth and high-scale distributed architecture, your mission is to transform Observability from a passive toolset into a strategic asset using open source standards. You will act as an architect of efficiency and reliability , building a global platform that empowers engineering teams to "own what they build" with confidence. We are moving beyond basic monitoring to build a comprehensive "Observability as a Service" ecosystem. You will be responsible for evolving a self-service platform that balances performance with cost-effectiveness, solving complex challenges related to high-cardinality metrics, log retention strategies, and distributed tracing. We strive to eliminate friction. You will design the "Golden Paths" that allow developers to instrument their code instantly and gain high-fidelity signals without operatio

REMOTEpythonawsazure
View job →
E
Everpure
📍 Prague• Full-time
18 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Everpure is expanding beyond traditional storage to help organizations understand, govern, and activate their data in the AI era. As part of Everpure’s Data Management business, 1touch brings capabilities in data discovery, classification, contextualization, enrichment, and security posture management across SaaS, on-premises, cloud, and hybrid environments. Together, we are building an intelligent data-management platform that helps enterprises turn complex, distributed data into trusted, governed, and AI-ready information. Joining this team means contributing to a high-impact transformation at the intersection of data, cloud, security, and artificial intelligence. As a Senior Golang Engineer , you’ll build core backend capabilities that enable enterprises to discover, understand, and protect sensitive data across complex environments. You’ll solve challenging engineering problems around scale, concurrency, performance, and reliability , with meaningful ownership over how our platform evolves. Working closely with Engineering, Product, Architecture, and DevOps/SRE, you’ll help expand the platform to new data sources and increasingly complex customer environments. WHAT YOU'LL DO Design and build high-performance backend systems in Go that operate reliably across complex enterprise environments. Solve challenging engineering problems around concurrency, distributed systems, scalability, performance, and fault

awsrestai
View job →
E
Everpure
📍 Bengaluru• Full-time
18 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Exa team and lead the charge in redefining enterprise storage by unifying block, file, and object protocols across hybrid-cloud environments. You will combine deep technical expertise in distributed systems with hands-on people leadership to guide architectural decisions and mentor high-impact engineers. This is a unique opportunity to build new engineering teams from the ground up and drive industry-leading innovation alongside Product and Architecture partners. Your work will directly impact how customers consume, scale, and operate mission-critical storage infrastructure. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management, validation, and support teams to align technical roadmaps, establish architectural standards, and drive enterprise

pythonjavaaws
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Top cities for Reliability Engineer

City links are canonicalized and require at least 20 current jobs.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.