Role: Engineering Manager – Communications Location: Remote Team size: 15 engineers (backend + frontend) Type: Full-time About the Role We’re looking for an Engineering Manager who thrives at the crossroads of leadership, hands-on engineering, and solving problems that don’t come with an instruction manual. You’ll lead a team of talented backend and frontend engineers who are building the bridges between systems that power the full customer journey for financial institutions. In this role, you won’t just be overseeing work, you’ll be rolling up your sleeves, writing and reviewing production code, guiding architecture decisions, and coaching engineers to deliver their best work. The systems you’ll help build will connect modern cloud APIs with decades-old banking platforms, bringing reliability and elegance to what often starts as messy complexity. Our integration platform touches everything—from communication stacks to core banking systems, lending and mortgage platforms, payment gateways, and AI Platform. Every integration is an opportunity to shape how our customers experience our products end-to-end. And because we take an AI-augmented approach to software development, you’ll be part of a team that uses AI tools to augment SDLC and write better code, automate testing, and ship faster without compromising quality. What You’ll Be Doing Lead by Example – Stay hands-on with coding, designing architectures and reviewing code while guiding the team toward engineering excellence. Own the Integration Layer – Architect and scale connections across diverse systems, from sleek modern APIs to finicky legacy protocols. Champion the Customer Experience – Partner with Product, Implementation, and customer teams to ensure integrations truly solve real-world challenges. Collaborate Without Boundaries – Work closely with other engineering leaders to ship features that feel seamless across products. Build for the Long Run – Keep systems observable, reliable, and perform
Jobiba hiring network
Reliability Engineer Jobs
2,049 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As the Engineering Manager for Drive Qualification, you will lead a high-performing Bangalore team dedicated to validating Everpure-developed SSDs across performance, reliability, and firmware maturity. In this impactful leadership position, you will own the end-to-end validation strategy for hyperscale and datastore programs, establishing a center of excellence for system-level robustness. Partnering closely with cross-functional firmware, hardware, and analytics teams globally, your mission is to deliver comprehensive qualification coverage that ensures our enterprise storage platforms launch with ultimate confidence and quality. WHAT YOU’LL DO Lead and Scale the Team: Coach, mentor, and grow a multi-level validation engineering team, building a culture of ownership, clear domain expertise, and continuous career development. Drive Validation Strategy: Own the execution roadmap for core qualification domains (including PCIe, NVMe/OCP, and power-loss robustness) across critical milestone gates from engineering samples to final product release. Foster Cross-Functional Alignment: Partner with global hardware, firmware, and program management stakeholders to align on test coverage, coordinate issue triage, and deliver clear risk assessments that dictate release readiness. Advance Automation and Infrastructure: Champion the expansion of Python and Linux-based automation frameworks, regression infrastructure, and data
Backend Engineer (Senior Level) - SDE IV We're looking for a Senior Backend Engineer to lead the architecture and evolution of backend services that deploy and serve machine learning models in production. You'll work closely with ML Engineers, Platform, and Product teams to build scalable, reliable systems and drive technical direction across multiple teams. What You’ll Do Design and drive the long-term architecture of backend services for biometrics and ML model serving. Collaborate with core platform and backend teams on organization-wide architectural initiatives. Partner with business and engineering teams to design and deliver cross-cutting platform capabilities. Lead architectural reviews, mentor engineers, and promote engineering best practices. Build and maintain backend services for deploying and serving ML models Monitor service reliability, performance, and scalability in production Deploy and operate services on AWS using ECS + Fargate, SageMaker, or EC2 + Kubernetes Support real-time and batch inference workflows Contribute to CI/CD pipelines and deployment automation What We’re Looking For Strong expertise in backend development using Java and working knowledge of Python. Experience mentoring engineers and driving architectural decisions. Working knowledge of Python, especially for ML-related workflows Hands-on experience with AWS (e.g., DynamoDB, ECS, EC2, Redis, S3, SageMaker) Familiarity with Terraform or other infrastructure-as-code tools, and experience with CI/CD and production monitoring Experience with observability tools (Datadog, New Relic, etc.) Experience with containers and orchestration (Docker, ECS, etc.) Understanding of how ML models are deployed and served in production Experience with Kubernetes Nice to Have Experience with MLOps or ML platform engineering. Experience with asynchronous programming and event-driven systems. Jumio Values: IDEAL: Integrity, Diversity, Empowerment, Accountability, Leading Innovation Equal Opportunities :
About the Role At Jumio, the Software Development Engineer IV - QA (SDE-IV, QA) is a senior technical role focused on ensuring the quality, performance, and reliability of highly scalable web portals and distributed backend systems. Our platform spans multiple Java Spring Boot microservices and customer-facing web portals , deployed across AWS ECS, EKS, and Lambda , and integrated through event-driven messaging using SNS/SQS . In this role you will design and drive the test automation strategy for both UI (Playwright) and API/service layers, set the quality bar for the team, and act as a force multiplier — mentoring other engineers and embedding quality earlier in the development lifecycle. You will work closely with development, product, and DevOps teams to ensure our products meet the highest standards of quality, scalability, and security. This is a hands-on senior IC role: you will write code, but you will also influence architecture, own cross-service test strategy, and make build-vs-buy decisions for testing tooling. T-Shaped Engineering Expectation As part of Jumio's engineering culture, you will adopt a T-shaped engineering approach. Beyond deep expertise in test automation and quality engineering, you will contribute across the development lifecycle — understanding software architecture, participating in design and API-contract discussions, reviewing application code, and ensuring our distributed systems are testable, observable, and resilient by design. Role Value This role is critical to ensuring the reliability, scalability, and security of Jumio's products. By architecting and maintaining automated testing frameworks across web, API, and event-driven layers, you will enable faster, higher-confidence releases and reduce production risk in a complex microservices environment. What You'll Do Test Architecture & Strategy Define and own the end-to-end automated test strategy across web portals and backend microservices, balancing UI, API, contract, integ
About THG Ingenuity THG Ingenuity is a fully integrated digital commerce ecosystem, designed to power brands without limits. Our global end-to-end tech platform is comprised of three products: THG Commerce, THG Studios, THG Fulfilment. Each represents a single, unified solution, overcoming challenges and taking brands direct-to-consumer. Our client portfolio includes globally recognised brands such as Coca-Cola, Nestle, Elemis, Homebase, and Proctor & Gamble. Shift Engineer - £46,000 - £47,000 + Shift allowance £5,702 About the Role Your objective, through modern maintenance systems, will be to maintain business as usual for the plant operations, to continuously improve the availability/reliability of the equipment and to ensure we dispatch quality products safely. The shift engineer will provide a technical support resource to the Fulfilment operation, delivering optimum asset reliability and performance through an effective asset maintenance and Continuous Improvement approach. You will have a good understanding of modern maintenance techniques as well as good working operational knowledge of a Fulfilment site. You should also be capable of ensuring a quick and effective reaction to daily business needs as they are presented. Responsibilities: Technical support resource for maintenance across plant to ensure a safe, timely and effective close out of live plant issues. Proactive fault prevention and continuous improvement activities to deliver key maintenance KPIs Work on and assist with developing the Plant PPM program, while using your RCA skills in response to Reactive tasks, to move the department towards a predictive Maintenance culture. Complete root cause analysis reports on major engineering stoppages across site. Support the close out of health and safety, Quality and CI actions lists. Reliability focused using modern techniques such as TPM, RCA, RCM and FMEA. Coaching and supporting
At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Financial Systems owns the data and reporting foundation for Accounting, and the operational reliability of the pipelines that power reporting, reconciliations, and automation. We are building a single source of truth for financial information using dbt and Snowflake to enable scalable BI and process automation across the org. We are hiring an Analytics Engineer focused on maintaining and optimizing our finance data platform , improving reliability, efficiency, and performance of our pipelines and core datasets. This role is ideal for someone who enjoys operational ownership, building strong foundations, and making data systems easier and safer to run at scale. What You’ll Do Build and maintain dbt models and core datasets that support Accounting reporting and downstream automation use cases. Improve platform reliability through strong testing patterns, alerting, and runbooks. Own and document data pipelines and lineage, ensuring changes are understandable and auditable. Identify, troubleshoot, and resolve production data issues, and drive root-cause fixes. Optimize performance and cost in Snowflake and dbt to support scaling needs. Partner with engineering and business stakeholders to translate requirements into durable, well-tested data assets. Contribute to the team’s best practices in version control, code review, documentation, and release hygiene (GitHub-based workflows). What We Look For 3+ years of experience in analytics engineering, data engineering, or similar roles working with production data systems. Strong SQL skills and hands-on experience building in dbt (modeling, testing, documentation). Experience operating in a modern development workflow: Git and pull-request based collaboration (GitHub preferred). Familiarity with standard IDEs and collaborative debugging practices. E
About the Role Sigma Computing is redefining business intelligence by making complex data analysis accessible through a high-performance platform built for the modern data stack. The Compiler Team plays a foundational role in this mission by transforming user-driven spreadsheet interactions into highly optimized SQL queries, enabling seamless exploratory analytics on cloud data warehouses. As a member of the Compiler Team, you will join a group of engineers dedicated to building the core systems and abstractions that power Sigma’s intuitive spreadsheet interface, ensuring speed, reliability, and scalability for all users. What You Will Be Doing Tackle core challenges at the intersection of data modeling, query compilation, and large-scale interactive analytics—making it possible for end-users to query data warehouses efficiently without deep technical knowledge Design, build, and maintain sophisticated compiler infrastructure and intermediate representations that translate spreadsheet operations into optimized query plans Apply advanced optimization strategies to improve performance and accuracy across a wide range of query workloads and data architectures Contribute to both backend (Rust) and key frontend foundations (TypeScript), evolving critical abstractions that enable end-to-end workflow optimizations and new features Debug, analyze, and resolve complex issues, ensuring robustness and maintainability in a rapidly evolving product Collaborate with engineers and product stakeholders to review designs and code, driving technical best practices and architectural decisions throughout the team and company Qualifications We Need 5+ years experience engineering high-quality software systems Demonstrated success building and maintaining complex infrastructure or core platform services Deep understanding of Computer Science fundamentals, particularly in compilers, algorithms, SQL Optimization Passion for teamwork, technical ownership, and continually
DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You’ll Do: Play a key role in ensuring product quality, reliability, and performance. Define test strategies, lead automation initiatives. Design and execute manual and automated tests for UI, API, and database layers. Develop and maintain test automation frameworks using Selenium, Playwright, or similar tools. Write, execute, and optimize test scripts integrated with CI/CD pipelines (e.g. Jenkins). Perform API and database validations, analyze logs, debug failures, and report defects clearly. Collaborate closely with developers, product managers, and DevOps for end-to-end quality assurance. Apply strong coding, testing, and problem-solving skills to identify risks and improve test coverage. Who You Are: Bachelor’s degree in engineering (CS / IT) or equivalent degree from a well-known Institute / University. 1+ years of hands-on experience with QA Automation Framework development & Design (Preferred language: Python). Strong understanding of testing methodologies. Experience with Python, Perl, Shell Scripting, Selenium, Test Automation (QA), and Software Testing (QA) (Preferred). Experience in Software Development, SDET (must have). Strong problem analysis, troubleshooting and debugging skills. Experience in databases, preferably MySQL. REST/API testing experience is a plus. Ability to integrate end-to-end tests with CI/CD pipelines and monitor and improve metrics around test coverage. Ability to work in a dynamic and agile de
Software Engineer Build technology where every nanosecond matters. At Graviton, software isn't just a tool that supports trading. It is the infrastructure behind every research breakthrough, every trading decision and every competitive advantage. As a Software Engineer , you'll work on systems where performance, reliability and precision matter at an extraordinary scale. You'll partner closely with software engineers and quantitative researchers to build technology that processes enormous volumes of market data, powers quantitative research and supports live trading. You'll take on problems that don't have obvious answers — from designing high-performance systems and distributed infrastructure to eliminating bottlenecks measured in microseconds and building tools that make our researchers and engineers faster. Your work will go into production, be measured against real-world performance and have a direct impact on how our trading systems operate. If you enjoy solving hard engineering problems, understanding systems at a deep level and pushing technology to its limits, you'll feel right at home. What You'll Work On You'll work across the engineering stack that powers our quantitative research and trading platforms. Depending on your team, your work may include: Designing and building high-performance, low-latency systems in modern C++. Building distributed systems that process and analyze massive volumes of market data . Designing systems where latency, throughput and reliability directly influence trading performance . Working on Linux systems, networking, concurrency and multithreaded applications. Profiling systems, identifying bottlenecks and optimizing performance at the hardware and software level. Building robust infrastructure that supports quantitative research and live trading. Debugging complex production systems and solving problems where correctness and reliability are critical. Designing internal platforms and developer tools that accelerate research an
Role Overview Build the software services that power products, platforms, and better business decisions. As a Software Engineer II, you’ll develop scalable backend applications, APIs, integrations, and AI-enabled features using Python and cloud technologies. You’ll contribute to solutions from design through production, helping improve reliability, performance, security, and developer productivity. This is an opportunity to solve meaningful engineering challenges, grow your technical ownership, and collaborate with experienced engineers across the development lifecycle. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design and build scalable backend services, REST APIs, integrations, and reusable software components using Python and, where relevant TypeScript. Develop data ingestion, transformation, service-to-service communication, and automation capabilities that support reliable product experiences. Contribute to AI-enabled features and use AI development tools responsibly to improve coding, testing, research, documentation, and delivery. Apply sound engineering practices across architecture, performance, security, testing, debugging, and maintainability. Deploy and operate services using AWS and CI/CD workflows, contributing to monitoring, troubleshooting, documentation, and continuous improvement. Partner with engineers and cross-functional colleagues through design discussions, code reviews, technical problem-solving, and knowledge sharing. These are the essentials you’ll need to get an interview 3–5 years of professional experience building and delivering production software in an agile environment. Strong hands-on experience with Python and backend development, including APIs, integrations, or service-oriented applications. Experience working with cloud platforms, preferably AWS, and familiarity with deployment or CI/CD practices. Working knowledge of software design principles, testing, debugging, performance optimization, an
Role Overview You’re a hands-on backend engineer who enjoys owning features end to end and working on real products that customers rely on every day. In this Software Engineer II role, you’ll help build and evolve a Third Party Risk Management SaaS platform using Laravel and PHP, designing scalable APIs and services that keep performance and reliability front and center. You’ll work in a product-focused team that owns its services from architecture and implementation through deployment, monitoring, and continuous improvement. You’ll mentor junior engineers, influence technical decisions, and use modern AI-powered tools thoughtfully to ship better code faster. If you’re looking for a mid-level role with real ownership, modern tooling, and the chance to grow your impact, this is for you. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design, build, and maintain backend features and RESTful APIs in Laravel within a modern TALL stack environment. Own well-defined stories from implementation through deployment, monitoring, and iteration, ensuring performance and reliability. Contribute to architectural discussions and technical decisions that shape the Third Party Risk Management platform. Review code, improve test coverage, and strengthen CI/CD and engineering standards across the team. Mentor Software Engineer I colleagues through code reviews, pairing, and knowledge sharing. Use AI tools (e.g. GitHub Copilot, ChatGPT) to accelerate coding, debugging, testing, and documentation—while critically validating outputs and ensuring safe, responsible use. These are the essentials you’ll need to get an interview 3–5 years of professional software engineering experience in an agile, fast-paced environment. Strong experience with PHP and Laravel, ideally within the TALL stack (Tailwind, Alpine.js, Laravel, Livewire). Solid understanding of relational databases (MySQL or MariaDB), including data modelling and query optimisation. Experience designin
MeltPlan is building the “planning engine” for the $14 Tn construction industry, an AI system designed specifically to optimize decisions before construction begins. While design software optimizes use and aesthetics and construction software optimizes execution and control, MeltPlan is building the missing layer - software that optimizes decisions and tradeoffs upstream, before scope is locked, procurement begins, and change orders become inevitable. MeltPlan’s long-term goal is to help teams make construction “boring” by making planning more intense: surfacing constraints and tradeoffs early, aligning stakeholders before plans are frozen, and reducing the need for late-stage redlines, rework, and change orders. MeltPlan is founded by operators who have built at scale. Kanav previously co-founded Innovaccer, a $3Bn healthtech company focused on making US healthcare more affordable and accessible. He’s now applying that systems-level thinking to construction.He’s joined by Tanmaya Kala, former Project Executive at DPR Construction, who led large commercial, healthcare, and life sciences projects. We combine deep tech scale with real construction execution. What This Role Really is : We are seeking a detail-oriented and technically strong AI QA Engineer to ensure the quality, reliability, and performance of Large Language Model (LLM)-based systems. In this role, you will be responsible for designing and executing test strategies, validating model outputs, and building evaluation frameworks to enhance the accuracy, safety, and overall performance of AI-driven applications.We would particularly value candidates who have hands-on experience in developing evaluation frameworks (evals) for AI systems, along with strong expertise in comprehensive system testing and quality assurance practices.You are responsible for making MeltPlan work in the real world. What You'll Do: Design, develop, and execute evaluation frameworks (Evals) for Large Language Models (LLMs) and AI syst
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We’re looking for an early career Software Engineer to join our Infrastructure team to accelerate building and scaling our innovative systems that support our growing identity platform. In this role, you will build the next-generation infrastructure that underpins all systems at CLEAR. The ideal candidate for this role will approach challenges with an eye toward reliability, simplicity, and scalability. What You'll Do: Develop and maintain a streamlined process for engineers to effortlessly build and deploy scalable and reliable software-defined networking solutions on AWS. Enhance our compute platform (Kubernetes) with new functionalities and features, focusing on AWS networking services and concepts such as VPCs, Route Tables, Security Groups (SGs), ALBs/ELBs, and Route53, optimize service communication and management. Collaborate across engineering teams to advocate for and implement best practices in observability, utilizing tools like Splunk or Datadog to ensure robust network monitoring. Act as a product owner for our infrastructure, collecting feedback and requirements from engineering teams to address pain points and develop solutions, particularly in the realm of AWS networking and cloud-native design principles. What you're great at: 0-2 years of experience in infrastructure and platform development and AWS cloud services. Proficient in Python, with understanding of Kubernetes and container orchestration tools like EKS and ECS. Understand AWS networking services, including VPC design, SGs, NATGWs, ALBs/ELBs, Rout
Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is growing quickly. You will play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features, while also balancing scalability, reusability, and performance. Role Overview: We're looking for a Staff Platform Engineer to serve as the technical backbone of our Engineering organization. You'll own the technical strategy, and delivery of our platform — spanning architecture, DevOps, SRE, security, Dev-ex. This is a hands-on staff level role: you'll set technical direction, drive cross-team alignment, and be the senior escalation point for platform challenges. What you’ll do: Drive platform reliability, scalability, security, and cost efficiency across all environments. Technical Leadership: Provide technical leadership for platform components, Influence the technical strategy and architecture of our cloud platform, from CI/CD pipelines to observability and incident response. Design and implement platform components and reusable integration patterns that minimize custom development efforts, reduce the time spent on repetitive tasks, and ensure that integrations scale across multiple healthcare systems Partner closely with Architecture, DevOps, SRE, and Security teams to deliver cohesive platform solutions Cross-Functional Collaboration: Work closely with product teams, and solutions architects to understand integration needs and ensure the platform meets current and future business requirements. Serve as a senior escalation point for infrastructure and platform incidents Establish frameworks for: AI governance and compliance. Observability of systems. Traceability of decisions and outputs. Ensure enterprise readiness with security, auditability, and reliability in production environments. Security & Compliance : Ensure all p
Opportunity Overview: We are seeking a Lead Software Engineer to join our Integrations team. In this role, you will be designing, developing, and scaling highly available healthcare integration systems supporting prior authorization workflows across providers, payers, and delegated entities. You'll direct a fast-paced, autonomous,agile team of software engineers in the design, development, and operational support of a growing enterprise integration platform. This is an opportunity to drive technical excellence at the intersection of healthcare interoperability and modern distributed systems. What you’ll do: Technical Leadership: Provide technical leadership across architecture, system design, platform scalability, reliability, and operational excellence. Platform Engineering: Design and build scalable, resilient, and high-performing systems that support critical business workflows and enterprise integrations. Integration Solutions: Lead the development and maintenance of secure integrations with internal and external platforms, partners, and third-party systems. Cloud & Automation: Drive cloud infrastructure, deployment automation, and software delivery practices that enable reliable and efficient releases. Distributed Systems: Design and support event-driven and distributed architectures that enable scalable and fault-tolerant processing. Operational Excellence: Establish monitoring, observability, and incident response practices to ensure system reliability, performance, and availability. Quality Engineering: Champion automated testing, quality assurance, and engineering best practices throughout the software development lifecycle. Production Support: Lead the resolution of complex production issues and drive continuous improvement in platform stability and operational efficiency. Cross-Functional Collaboration: Partner with product, operations, data, security, and business stakeholders to deliver solutions aligned with organizational goals. Agile Delive
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime