Jobiba hiring network

Reliability Engineer Jobs

2,028 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Staff Software Engineer - Testing & Automation Exceptional software engineering is challenging. Amplifying it to ensure that multiple teams can concurrently create and manage a vast, intricate product escalates the complexity. As a Staff Engineer within the Verification Platform team at Sumo Logic, you will drive the implementation and optimization for our verification platform as well as the modernization of our CI/CD pipelines. Your mission is to develop and sustain automated tooling for all testing, verification, and functional requirements, leveraging AI reasoning and machine learning models to predict and prevent delivery issues, while integrating advanced security validation and non-functional requirements into our delivery lifecycle. You will contribute significantly to establishing automated delivery pipelines, empowering autonomous teams to create independently deployable services, and progressing Sumo Logic’s internal Platform-as-a-Service. This role sits at the intersection of Platform Engineering, Quality Engineering, DevSecOps, and Developer Productivity, helping teams deliver secure, reliable, and independently deployable services at scale. Responsibilities Strategy & Leadership: Drive technical direction and design for a modern Quality Engineering platform, driving the adoption of AI reasoning for enhanced automation of all testing, verification, and functional requirements. Pipeline Modernization: Lead the modernization of CI/CD pipelines to include automated security validation, compliance checks, and other critical non-functional requirements, with a focus on integrating AI/ML for intelligent pipeline optimization and risk prediction. Framework Ownership: Own the delivery pipeline and release automation framework for all Sumo services, ensuring improvements in developer productivity, deployment frequency, and release reliability. Cross-Team Collaboration: Educate and collaborate with teams during design and development phases to ensur

pythonjavareact
View job →
Z
Zscaler
📍 Bellevue• Full-time• From $102.4K/yr
16 days ago

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Production Engineer to join our team. This role is available as a hybrid opportunity 3 days a week in San Jose, CA or Bellevue, WA reporting to Production Engineering in the Cloud Infrastructure & Operations department. Join Zscaler to be a force multiplier for the reliability of a global platform processing 200+ billion transactions daily across tens of millions of enterprise users. In this role, you will provide the technical vision and hands-on execution to drive an "automation-first" culture across the company. By maturing our observability and architectural standards, you will directly reduce our Mean Time to Mitigate (MTTM) and shape the scalability of our globally distributed, multi-cloud infrastructure. What you’ll do (Role Expectations) Implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments Drive an "automation-first" cult

pythonvueaws
View job →
T
Toradex
📍 Bengaluru• Full-time
16 days ago

Toradex is a global company strongly focused on engineering & technology. We’re powered by a diverse & uniquely gifted workforce. We pursue the best people to propel our innovative vision of embedded computing and IoT. If you’re interested in being a driving force at an agile technology company, engineering clever computing solutions & helping other companies bring their products to life, we should talk. Description We are looking for a DevOps Engineer to strengthen our cloud operations and engineering practices, with a focus on reliable website delivery, secure AWS foundations, and fast but controlled delivery of new services. The position combines AWS operations, infrastructure as code, CI/CD, automation, and pragmatic software engineering. The person should be confident working with services for edge delivery, compute, storage, databases, DNS, security, and observability without relying on manual console changes as the default operating model. The role also supports on-premises to cloud migration, global service optimization, and practical responses to increasing AI-driven traffic. We value candidates who can use modern AI-assisted development effectively to spin up proof-of-concept projects quickly, while still applying disciplined Git, review, security, and deployment practices. About you You enjoy building stable, secure, and maintainable infrastructure that supports business-critical services. You can work independently and take ownership of cloud environments, deployments, and operational improvements. You are comfortable balancing speed, reliability, cost, and security when making technical decisions. You communicate clearly with technical and non-technical stakeholders and explain trade-offs in a practical way. You document your work well and create clear runbooks and support material for future maintenance. You are methodical when troubleshooting incidents and stay calm when systems are under pressure. You are curious about modern traffic patt

javascripttypescriptpython
View job →

DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You’ll Do: We are looking for a Senior Software Engineer – Platform Operations based in Pune, India, who will play a key role in ensuring the reliability, performance, and operational excellence of DeepIntent's platform and data ecosystem. This role requires a strong engineering mindset with the ability to troubleshoot complex technical issues, understand distributed data architectures, and collaborate across Engineering, Product, Analytics, and Customer-facing teams to deliver timely and effective solutions. As part of the Operations organization, you will work closely with Engineering to support production systems, improve operational processes, and drive platform stability. The ideal candidate is a self-motivated problem solver who is passionate about learning new technologies, improving system reliability, and delivering exceptional customer outcomes through engineering excellence. Serve as the engineering interface between Customer-facing teams, Analytics, Product, and Engineering organizations. Partner with Platform Support, Client Success, and other customer-facing teams to investigate and resolve complex platform-related issues. Analyze application, API, and data pipeline issues to identify root causes and drive timely resolution. Develop and standardize operational tools, and interfaces to support analytical and operational use cases. Monitor data pipeline executions, investigate failures, and implement corrective and pre

pythonjavasql
View job →
D
16 days ago

DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You’ll Do: We are looking for a Software Engineer to help build and scale our core backend systems and data infrastructure. In this role, you will work hands-on to develop the foundational data pipelines, storage solutions, and robust architectures that drive our healthcare advertising solutions and support our core products, reporting APIs, and analytics initiatives. This is an excellent opportunity for a growth-oriented engineer to work with massive datasets, modern cloud technologies, and cross-functional teams to deliver high-performance, fault-tolerant solutions. Build & Operate: Develop, test, and maintain highly reliable, scalable, and cost-optimized distributed systems and data architectures. Enable Self-Service Data: Create automated ingestion, storage, and transformation pipelines that make it simple for downstream users to access and utilize new datasets. Empower Machine Learning: Design and operate data pipelines specifically tailored to support the complex workflows of our Data Scientists and Machine Learning Engineers. Drive Operational Excellence: Help implement and champion DataOps and DevOps practices across the team to ensure system reliability and smooth deployments. Contribute to Best Practices: Play an active role in establishing and refining formal data practices, architectures, and engineering standards for the organization. Cross-Functional Collaboration: Partner effectively with business stakeholders,

pythonjavasql
View job →
P
16 days ago

A Career with Point72's Technology Team As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. What you'll do Lead the design, development, and operation of scalable, enterprise-grade AI/ML architectures and systems with a strong emphasis on reliability, availability, and performance. Lead and mentor a team of engineers, driving technical direction, code quality, and iterative delivery of large-scale solutions. Partner closely with data scientists, engineers, product teams, and compliance to integrate AI/ML solutions into existing and new products. Own the end-to-end lifecycle of GenAI services, including LLM inference, model serving, and proxy/gateway layers that support multiple downstream applications. Define and uphold engineering best practices around observability, scalability, security, and cost efficiency for AI/ML platforms. Evaluate tools, technologies, and processes to ensure the highest quality and performance of AI/ML systems. Stay abreast of the latest advancements in AI/ML technologies and methodologies, and translate them into pragmatic solutions for the business. Ensure compliance with industry standards and best practices in AI/ML. What's required Bachelor's or Master's degree in Computer Science, Engineering, or a related field. 10+ years of experience in software/AI/ML engineering, with a proven track record of successful delivery of complex, production-grade systems. Demonstrated experience building large-scale enterprise-grade services with high reliability, availability, and observability (SLO/SLA-driven en

pythonjavaaws
View job →
P
Point72
📍 Bengaluru• Full-time
17 days ago

JOB TITLE Observability Engineer A CAREER WITH POINT72'S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. WHAT YOU’LL DO Design observability capabilities that give engineering teams clear insight into application health, platform performance, and issues affecting users Build scalable collection pipelines for metrics, logs, and traces across cloud-based and on-premises environments Develop actionable alerting standards that reduce noise, shorten incident response, and highlight the most important signals Partner with application and infrastructure teams to define service health indicators and improve operational readiness before production launches Automate monitoring configuration, dashboard deployment, and reliability checks to support consistent observability across the technology environment Analyze production incidents to identify telemetry gaps and improve detection, diagnosis, and recovery Create dashboards and reporting views that help teams understand trends, capacity risks, and reliability outcomes Establish practical observability

pythonagileai
View job →

You’ll shape the future of a business‑critical platform as the technical lead across both product engineering and cloud infrastructure. You’ll modernize a mature .NET application running on AWS today, while steering its evolution toward a cloud‑native, React/Node.js, AI‑enabled architecture. If you enjoy owning architecture end‑to‑end, from backend and frontend through CI/CD, DevOps, and AWS infrastructure, this role gives you real influence at Staff Engineer level and the opportunity to set engineering standards that others follow. You’ll spend your time leading complex .NET and React features, designing scalable AWS infrastructure with Infrastructure as Code, and building automation that makes releases fast, safe, and repeatable. You’ll work on performance, reliability, and modernization in equal measure—fixing what’s slowing the platform down today and designing what it will look like in the next generation. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture and development of enterprise .NET services and APIs that power a business‑critical platform. Design and operate AWS infrastructure (using AWS CDK in TypeScript) to support secure, scalable, multi‑environment deployments. Build and optimize CI/CD pipelines (AWS CodePipeline, CodeBuild, Windows build agents) to make shipping .NET and React changes fast and reliable. Drive modernization initiatives across the stack, including clean architecture, refactoring legacy components, and reducing technical debt. Design and tune PostgreSQL and MSSQL database solutions for performance, scalability, and reliability. Mentor engineers and influence engineering practices across teams, raising the bar on cloud, DevOps, and software design. These are the essentials you’ll need to get an interview Significant experience (typically 8+ years) delivering and operating scalable enterprise software, owning both application code and cloud infrastructure. Deep hands‑on expertise with C

typescriptreactnode.js
View job →

Here’s a summary of the role: Build cloud software that matters, grow your technical depth, and use modern AI tooling to do your best work. This is a hands-on engineering role for someone who enjoys solving product problems, writing clean code, and helping services run reliably at scale. You’ll work on secure, scalable microservices and APIs using TypeScript, AWS , and modern engineering practices. You’ll be part of a collaborative product engineering team where you can own features, contribute to design discussions, support production systems, and keep growing across backend, cloud, and AI-assisted development workflows. Here’s a breakdown of what you’ll do, not all of it, just the important stuff: Design, build, test, and improve backend services and APIs using Node.js, TypeScript, and AWS . Take ownership of well-defined features from planning through release, including code quality, deployment, and production support . Work closely with product managers, designers, and other engineers to turn requirements into practical, reliable solutions. Contribute to technical design conversations, code reviews, and engineering standards that keep the team moving well. Use AI tools to speed up research, coding, debugging, testing, and documentation, while checking outputs carefully and applying sound judgment. Help keep systems secure, observable, and maintainable by improving monitoring, reliability, and day-to-day development practices. These are the essentials you’ll need to get an interview: 3 to 5 years of professional software engineering experience building production applications in an agile environment. Strong backend development skills with Node.js and TypeScript, including experience building APIs or microservices. Experience with React or Angular in a product engineering environment. Hands-on experience with

typescriptreactnode.js
View job →
DC
17 days ago

Role Overview Build reliable software services that power products, platforms, and business decisions. As a Senior Software Developer, you’ll design and deliver scalable applications, backend services, and integrations that perform well in production and evolve with changing business needs. You’ll apply strong software engineering practices across APIs, data-intensive applications, cloud services, AI-enabled solutions, and deployment pipelines. You’ll help shape technical solutions, improve system reliability, and contribute to a high-quality engineering culture. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design and develop scalable backend services and applications using Python or TypeScript. Lead the development of APIs, integrations, reusable software components, and AI-enabled features. Build reliable solutions for data ingestion, manipulation, service-to-service communication, and intelligent automation. Apply AI technologies and modern software engineering practices to improve product capabilities, developer productivity, and operational efficiency. Make sound technical decisions around architecture, performance, security, scalability, and maintainability. Deploy and operate applications using AWS services and CI/CD practices while improving testing, monitoring, documentation, and delivery standards. These are the essentials you’ll need to get an interview 5+ years of professional experience developing and delivering production software. Strong hands-on experience with Python; TypeScript or similar languages is also valuable. Proven experience building backend services, APIs, integrations, and service-oriented applications. Experience applying AI technologies, such as generative AI, machine learning services, intelligent automation, or AI-enabled application features. Strong understanding of software design principles, testing, debugging, performance optimization, and secure development. Experience working with cloud platfor

typescriptpythonaws
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $252K/yr
17 days ago

The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Staff Software Engineer, you will orchestrate the implementation of vertical features and horizontal capabilities to include mentoring other engineers on defining requirements with stakeholders and communication tradeoffs of technical implementations on feature and capabilities until they are accepted by the stakeholders. You will: Orchestrate feature implementation across the Federal engineering team to ensure architectural consistency. Define technical strategy for agentic guardrails, explainability, and fleet orchestration. Ensure system reliability and performance across multiple security classifications and network types. Mentor engineers in the process of defining requirements with stakeholders and gathering acceptance. Communicate high-level technical trade-offs and implementation strategies to senior government stakeholders and Scale C-Suite members. Influence the long-term product strategy and technical roadmap for the Federal business unit. Consult on the architecture of AI-powered solutions for large-scale federal contracts. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in

awsazuregcp
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $252K/yr
17 days ago

Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale our product in a fast-paced environment. The ideal candidate will have a strong understanding of software engineering principles and practices, as well as experience with large-scale distributed systems. You will be responsible for owning large new areas within our product, working across backend, frontend, and interacting with LLMs and ML models. You will solve hard engineering problems in scalability and reliability. You will: Own large new areas within our product Work across backend, frontend, and interacting with LLMs and ML models Deliver experiments at a high velocity and level of quality to engage our customers Work across the entire product lifecycle from conceptualization through production Be able, and willing, to multi-task and learn new technologies quickly Ideally you'd have: 7+ years of full-time engineering experience, post-graduation Experience scaling products at hyper growth startups Experience tinkering with or productizing LLMs, vector databases, and the other latest AI technologies Proficient in Python or Javascript/Typescript, and SQL Experience with Kubernetes Experience with major cloud providers (AWS, Azure, GCP) Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval

javascripttypescriptpython
View job →
SA
Scale AI
📍 Argentina• Full-time
17 days ago

Software Engineer Argentina; Uruguay Software Engineer - Robotics & Autonomous Systems Scale's Robotics business unit is dedicated to solving the data bottleneck in Physical AI across Robotics, Autonomous Vehicles, and Computer Vision. In this role, you'll be a key contributor building production systems for robotics data collection, model training pipelines, and evaluation infrastructure. You'll have the opportunity to own critical parts of our robotics platform, work directly with cutting-edge robotics and AV customers, and shape the future of embodied AI systems. You Will: Own and architect large-scale data processing pipelines for robotics and autonomous vehicle datasets Build ML training and fine-tuning pipelines using Scale's robotics data Work across backend (Python, Node.js , C++), and frontend (React, TypeScript) stacks to build end-to-end solutions Develop tools and real-time systems for robotics data collection, teleoperation, model evaluation, data curation, and data annotation Interact directly with robotics and AV stakeholders to understand their technical needs and drive product development Design comprehensive monitoring and evaluation frameworks for robotics models and data quality Solving complex, late-stage industry challenges in concurrent and real-time robotic systems, with strict attention to timing constraints and data integrity. This often involves deep investigation, reviewing academic papers, and direct collaboration with robotics vendors Collaborate with ML engineers and researchers to bring robotics research into production Deliver features at high velocity while maintaining system reliability and performance Ideally, You Have: At least 6 years of high-proficiency software engineering experience, with a strong background in complex systems and the ability to independently research, analyze, and unblock hard technical problems. Strong programming skills in Python and TypeScript/Node.js for production systems Experience with React and m

typescriptpythonreact
View job →

Scale’s rapidly growing Global Public Sector team is focused on using AI to address critical challenges facing the public sector around the world. Our core work consists of: Creating custom AI applications that will impact millions of citizens Generating high-quality training data for custom LLMs Upskilling and advisory services to spread the impact of AI As a Full Stack Software Engineer (Forward Deployed), you’ll collaborate directly with public sector counterparts to quickly build full-stack, AI applications, to solve their most pressing challenges and achieve meaningful impact for citizens. At Scale, we’re not just building AI solutions—we’re enabling the public sector to transform their operations and better serve citizens through cutting-edge technology. If you’re ready to shape the future of AI in the public sector and be a founding member of our team, we’d love to hear from you. You will: Partner with public sector clients to scope, collect feedback and implement solutions for complex problems, including spending up to two weeks per month in client offices for feedback and delivery. Architect production-grade applications that integrate AI models with full-stack frameworks, managing everything from interactive UIs to backend APIs and systems. Deploy and manage infrastructure within cloud environments, ensuring the highest levels of system integrity, security, scalability, and long-term reliability. Contribute to core platform features designed to be reused across diverse international client use cases. Partner with design, product, and data teams to build robust applications aligned with the broader technical architecture. Ideally you’d have: Bachelor’s degree in Computer Science or a related quantitative field 5+ years of post-graduation, full-stack engineering experience with demonstrated proficiency in React (required), TypeScript, Next.js, Python, Node.js, PostgreSQL or MongoDB plus hands-on experience with Docker, Kubernetes, and Azure

typescriptpythonreact
View job →
SA
17 days ago

At Scale, our mission is to develop reliable AI systems for the world's most important decisions. For 10 years, Scale has provided the high-quality data and full-stack technologies that power the world's leading models, and has helped enterprises and governments build, deploy, and oversee AI applications that deliver real impact. We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force. Scale's internship is not a side project. Interns own real, shipped work on the same roadmaps as full-time engineers, with mentorship from world-class talent and a culture that values ownership, speed, and truth-seeking. Many of our interns return as full-time Scaliens. Example Projects Build reinforcement learning and post-training data pipelines that power frontier model development Develop evaluation infrastructure that measures model reliability for enterprise and public sector customers Ship agentic AI applications and the tooling that makes them observable, testable, and safe to deploy Ship tools that accelerate the growth of new qualified contributors on Scale's platform Build fraud-detection systems that remove bad actors and keep Scale's contributor base safe and trusted Use models to estimate the quality of tasks and contributors, and guarantee quality on requests at large scale Devise advanced matching algorithms that pair contributors to customers for optimal turnaround and accuracy Create optimized and efficient UI/UX tooling, in combination with ML algorithms, for 100k+ contributors completing billions of complex tasks Develop new AI infrastructure products to visualize, query, and explore Scale data Requirements A graduation date in Fall 2027 or Spring 2028 with a Bachelor's degree (or equivalent) in a relevant field (Computer Science, EECS, Computer Engineering, Statistics) Available for a Summer 2027 internship (May/June start dates) in San Franci

typescriptpythonreact
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.