Jobiba hiring network

Reliability Engineer Jobs

2,028 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. As a Senior Technical Support Engineer, you will be a trusted advisor and technical resource for our customers. This is a hands-on role for someone who thrives in dynamic environments, loves troubleshooting complex technical issues, and is passionate about delivering exceptional support experiences. You’ll be responsible for resolving high-impact technical issues, driving customer success, and collaborating closely with product, engineering, and customer su

pythonsqlaws
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Equipment Ownership & Stability Management Act as the Equipment Owner for assigned process tools or modules Ensure equipment availability, reliability, and performance to meet manufacturing requirements Provide technical support to the manufacturing equipment repair Equipment Abnormality Analysis & Problem Solving Conduct systematic analysis of equipment alarms, failures, and recurring issues Identify and define clear root causes and failure mechanisms Develop, validate, and implement improvement actions based on data analysis and experiments Create and improve equipment maintenance procedures, checklists, and troubleshooting BKMs. PM / CM & Component Lifetime Optimization Increase tool uptime through systematic problem-solving, quality workmanship, and routine review of equipment documentation. Define, optimize, and execute Preventive Maintenance (PM) strategies, including scope and frequency Support PM release, change

artificial intelligenceairecruitment
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Responsibilities and Tasks Equipment Ownership & Stability Management Act as the Equipment Owner for assigned process tools or modules Ensure equipment availability, reliability, and performance to meet manufacturing requirements Provide technical support to the manufacturing equipment repair Equipment Abnormality Analysis & Problem Solving Conduct systematic analysis of equipment alarms, failures, and recurring issues Identify and define clear root causes and failure mechanisms Develop, validate, and implement improvement actions based on data analysis and experiments Create and improve equipment maintenance procedures, checklists, and troubleshooting BKMs. PM / CM & Component Lifetime Optimization Increase tool uptime through systematic problem-solving, quality workmanship, and routine review of equipment documentation. Define, optimize, and execute Preventive Maintenance (PM) strategies, including scope and frequency <s

artificial intelligenceairecruitment
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Responsible for understanding the processes used to assess and mitigate risks related to the introduction of new package technologies so that you can improve efficiency and effectiveness by implementing automation, data analysis and machine learning/AI solutions. Set up data bases, develop data ingestion pipelines, and implement automated data analysis and reporting tools. Support problem solving and risk assessment for internal or customer quality issue by pulling and analyzing data. Contribute to the advancement of technology at Micron through mentoring, publishing technical papers (internal and external), and developing innovative solutions to challenging problems. Implement Automation, Data Analysis, and AI Solutions. Collaborate with Engineering teams to Map Package DDQA processes and data streams. Set up and optimize databases and develop solutions to improve efficiency and effectiveness. Understand the needs of internal customers and develop solutions. Support Problem Solving and Risk Assessment for Quality Issues. Pull relevant product information, manufacturing data, and reliability data based on given problem statements. Determine the appropriate dataset and treatment required to answer questions posed by problem solving teams. This could include producing data visualizations, machine learning models, statistical inferences, and web applications. Provide recommendations about root cause findings and product risk, based on data analysis. Collaboratively Communicate Findings and Best Practices. Share best practices with global teams to enable a cultur

javascriptpythonjava
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron is advancing a historic $15 billion investment in semiconductor manufacturing in Boise, Idaho, with DRAM production planned for the second half of the decade. As a global leader in memory and storage solutions, Micron is investing more than $150 billion worldwide over the next decade to expand groundbreaking manufacturing capabilities and drive innovation across the semiconductor industry. Join the Boise expansion team and help build the future of semiconductor manufacturing. You will work alongside experienced engineers to develop advanced processes, solve complex manufacturing challenges, and support the production of technologies that power everything from artificial intelligence to next-generation computing systems. As a Dry Etch Process Engineer, you will contribute to the development, optimization, and sustainment of semiconductor manufacturing processes. You will gain hands-on experience in a brand new fabrication environment, working with multi-functional teams to improve yield, quality, reliability, and productivity. This is an excellent opportunity to apply engineering fundamentals, develop technical expertise, and grow your career in semiconductor manufacturing. Responsibilities Support the development and optimization of dry etch processes to improve product performance, yield, and manufacturing efficiency. Analyze process and manufacturing data to identify trends, solve issues, and implement corrective actions. Partner with process integration, equipment,

pythonsqlartificial intelligence
View job →

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. Service Levels is one of New Relic's most commercially critical products — it powers SLO compliance and reliability measurement for thousands of customers. We're looking for an experienced Engineering Manager to lead a senior, high-performing team building the next generation of service level management at scale. You'll lead a team that includes lead-level engineers with deep domain expertise, and your value will come from enabling their best work — not directing it. You'll own delivery, quality, and team health while partnering closely with product and design to ship features that directly impact New Relic's commercial momentum. What you'll do Lead a full-stack engineering team of 6-8 across backend (Java/Spring Boot) and frontend (React/TypeScript) Own end-to-end delivery — sprint planning, execution, quality, and release Set clear expectations, manage performance equitably, and develop engineers at every level Partner with Product Manager and XD to translate requirements into technically sound, deliverable plans Drive architectural discussions and hold the team to engineering excellence standards Identify

typescriptjavareact
View job →

Work Flexibility: Onsite Stryker is seeking a Staff Advanced Manufacturing, Automation/Software Engineer to join our Advanced Operations team, supporting the Medical Division - Acute Care Business Unit. In this role, you will lead the design, development, and deployment of advanced automation and production test systems for new product introductions. This role is critical to ensuring reliable, scalable, and compliant manufacturing for patient support and patient environment products. You will serve as both a technical owner and a supplier-facing leader, architecting systems internally while guiding external partners to deliver high‑quality automation solutions. You will have the opportunity to work on the design transfer of new products from research through development and into production. This is a hybrid role based out of Portage, MI. The team works onsite 4-5 days per week to support collaboration and project needs. What you will do: Automation System Architecture & Development Lead the full lifecycle of industrial automation and test systems from requirements, architecture, and design through implementation, validation, and release. Define system-level requirements encompassing mechanical, electrical, software, controls, and safety considerations. Develop and integrate control software, embedded interfaces, test sequences, and operator interfaces (HMI/SCADA). Ensure robust performance, maintainability, reliability, and alignment with design intent and manufacturing needs. Troubleshoot complex processes, software, and equipment issues; optimize system performance and uptime. Supplier & Equipment Vendor Leadership Manage automation and equipment suppliers, including capability assessments, technical reviews, process monitoring, and on‑site visits. Create clear, comprehensive

pythonlinuxc#
View job →

Observability Pipelines (OP) is Datadog's on-premise, vendor-agnostic telemetry pipeline product. As an Engineering Manager on the team, you'll own people management and engineering execution for one of OP's core missions, spanning areas like Integrations (ingesting from and routing to the many source and destination systems customers rely on), streaming insights, cost control, or pipeline capabilities, reliability and scalability. You'll partner directly with Product to help shape the roadmap, and work closely with your peer EMs and senior ICs to define how OP operates and grows. This is an opportunity to build your management craft while having real influence over the technical direction of a fast-growing product area. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Own people management and engineering execution Establish a strong operating rhythm for the team Drive high standards for on-call rotations and incident response Partner with Product on the roadmap, balancing product priorities with technical realities Lead, coach, and grow the careers of engineers on your team Who You Are: Experienced managing engineers directly, comfortable owning a team’s operating rhythm end-to-end, from planning through execution and stakeholder communication to incident and on-call ownership Have a technical background in distributed systems and data infrastructure Have experience with on-premises or customer-installed software concepts A product-minded partner to have on the team — you enjoy working with Product on strategy Experience with high-performance or Rust-based data pipeline systems Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications o

aigorust
View job →
D
Datadog
📍 New York• Full-time• From $280K/yr
1mo ago

The Detection Platform organization is responsible for helping customers identify, understand, and act on issues across their environments through alerting, event intelligence, and autonomous detection capabilities. As Director, Detection Platform, you will lead a group of engineering managers and teams responsible for foundational alerting infrastructure, event management, monitor creation experiences, and AI-powered detection systems. This role sits at the center of Datadog’s efforts to evolve how customers detect, investigate, and respond to operational issues at massive scale. You will partner closely with Product Management, Applied Science, Design, and Engineering leaders to shape the future of detection and observability experiences for Datadog customers while leading a growing organization of engineers. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead a multi-team engineering organization responsible for alerting, event management, monitor creation experiences, and autonomous detection capabilities. Define and execute the technical and organizational strategy for the Detection Platform while aligning stakeholders across Engineering, Product, Design, and Applied Science. Drive innovation in AI-powered detection, anomaly identification, and signal generation that helps customers proactively identify and resolve issues. Scale highly available platform systems that process hundreds of millions of evaluations while maintaining reliability, performance, and operational excellence. Develop and mentor engineering managers and technical leaders, fostering a culture of execution, collaboration, and technical rigor. Champion customer-centric product thinking by balancing platform investments with intuitive user experiences and measurable customer

machine learningaigo
View job →
O
1mo ago

About the Team Training Runtime designs the core distributed runtime that powers everything from early research experiments to frontier-scale model runs. We work on building robust, scalable, high performance components to support our distributed training workloads. Our priorities are to maximize the productivity of our researchers and our hardware, with the goal of accelerating progress towards AGI. Within Training Runtime, the Process Management team develops the distributed OS responsible for launching, coordinating, and supervising the large numbers of processes that make up modern training workloads. Our runtime sits beneath training frameworks and on top of research infrastructure, ensuring jobs run reliably across massive clusters while maintaining performance, stability, and observability. Success for us is measured by both system reliability and researcher velocity - enabling ideas to scale from experiments to production training runs. About the Role As a Training Runtime: Process Management Engineer , you will work on the software that ties thousands of computers together and exposes them as a unified system. This system has to serve individual researchers running multiple parallel experiments, as well as our largest training runs spanning 100’s of thousands and even millions of machines and accelerators. This requires easy to use, introspectable systems that can promote a fast debugging and development cycle, as well as relentless optimization for scale while maintaining stability and performance throughout. You will work primarily in Rust , building high-performance asynchronous systems with a strong emphasis on performance, correctness, and scalability. Working at this scale and at the frontier of AI development poses novel challenges. Out-of-the-box approaches often don’t work. The problems you will be working on are highly ambiguous and require strong design judgment as well as proficient execution to advance the state of our infrastructure. We’re loo

pythonawslinux
View job →
L
Litmos
📍 India• ₹35L – ₹45L/yr
14 days ago

Are you looking for an opportunity to help solve one of today's biggest business challenges? AI is changing the pace of business, and organizations everywhere are struggling to help their workforce, partners, and customers keep up. At Litmos, we're building the Learning Acceleration Platform that helps organizations build human capability faster—and we're looking for people who are passionate about making a meaningful impact for customers while growing alongside a collaborative, people-first team. Litmos is the Learning Acceleration Platform that helps organizations build capability faster, adapt at the speed business changes, and scale learning to anyone, anywhere. Combining an intuitive platform, AI-powered capabilities, trusted content, expert services, and a broad ecosystem of integrations, Litmos helps organizations accelerate workforce productivity, improve customer adoption and retention, enable high-performing partners, and reduce organizational risk through continuous learning. Organizations such as Hewlett Packard Enterprise, Graco, Sabre, and Russell Mineral Equipment trust Litmos to accelerate learning across their workforce, partners, and customers. Today, more than 11K customers with 30 million learners across 150 countries and 37 languages use Litmos to build the capabilities their organizations need to succeed. Backed by Francisco Partners, one of the world's leading technology investment firms, we're investing in the future of learning—and the people who are building it. Learn more at www.litmos.com . We are seeking a Manager, Cloud Engineering to lead and develop a team of cloud engineers while serving as a key technical authority for the architecture, reliability, and evolution of our cloud platform. This role combines hands-on technical leadership with team development, driving infrastructure modernization, automation, and operational excellence across our global SaaS environment while building and mentoring a high-performing enginee

pythonawsazure
View job →

Our Purpose Mastercard powers economies and empowers people in 200&#43; countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Manager, Software Engineering Overview We are seeking a hands-on Manager, Software Engineering to lead the development of scalable data platforms, backend systems, and data pipelines supporting Mastercard's Portfolio Intelligence products. This role will initially focus on leading a team of Data Engineering contractors while providing strong technical leadership across architecture, delivery, and engineering excellence. The ideal candidate is a recent practitioner who has built and operated modern data platforms and backend services, can confidently review architecture proposals and code, and leverages AI-powered engineering tools to improve productivity, quality, and delivery speed. As the team evolves, this leader will play a key role in building and developing a high-performing organization of full-time engineers. What You'll Do Technical Leadership Provide technical leadership for backend services, data pipelines, and platform capabilities that support analytics, reporting, and AI-driven products. Review architecture and design documents to ensure solutions are scalable, maintainable, secure, and aligned with long-term platform strategy. Conduct and oversee code reviews, promoting engineering best practices, reliability, security, and operational excellence. <br

pythonjavaaws
View job →
E
Everpure
📍 Bengaluru• Full-time
18 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As an Escalation Engineer on our FlashBlade team, you will serve as the premier technical authority driving customer trust and operational stability across complex, enterprise-scale storage environments . You will collaborate closely with front-line Support, Engineering, and product leaders to rapidly resolve high-impact technical challenges and transform complex system failures into long-term product reliability. By bridging real-world customer insights with engineering solutions, you will elevate team performance and ensure our enterprise customers achieve flawless platform availability. WHAT YOU'LL DO Drive High-Stakes Escalation Resolution: Take end-to-end ownership of critical, multi-platform system issues—evaluating hardware, software, networking, and environmental factors—to rapidly restore service, perform root-cause analysis, and protect customer business continuity. Elevate Engineering Talent & Knowledge: Mentor and coach support team members through joint case triage, structured technical trainings, and internal documentation, accelerating technical capabilities and resolution velocity across the organization. Bridge Product Engineering & Customer Insights: Partner directly with Product Engineering to relay real-world system behavior, ensuring critical customer feedback, feature enhancements, and bug fixes trickle back into core product design. Lead Strategic Customer Communications: Facilitate

awslinuxrest
View job →
E
Everpure
📍 Bengaluru• Full-time
18 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As an Escalation Engineer on our FlashBlade team, you will serve as the premier technical authority driving customer trust and operational stability across complex, enterprise-scale storage environments . You will collaborate closely with front-line Support, Engineering, and product leaders to rapidly resolve high-impact technical challenges and transform complex system failures into long-term product reliability. By bridging real-world customer insights with engineering solutions, you will elevate team performance and ensure our enterprise customers achieve flawless platform availability. WHAT YOU'LL DO Drive High-Stakes Escalation Resolution: Take end-to-end ownership of critical, multi-platform system issues—evaluating hardware, software, networking, and environmental factors—to rapidly restore service, perform root-cause analysis, and protect customer business continuity. Elevate Engineering Talent & Knowledge: Mentor and coach support team members through joint case triage, structured technical trainings, and internal documentation, accelerating technical capabilities and resolution velocity across the organization. Bridge Product Engineering & Customer Insights: Partner directly with Product Engineering to relay real-world system behavior, ensuring critical customer feedback, feature enhancements, and bug fixes trickle back into core product design. Lead Strategic Customer Communications: Facilitate

awslinuxrest
View job →
L
Litmos
📍 India• Full-time• ₹35L – ₹45L/yr
18 days ago

Are you looking for an opportunity to help solve one of today's biggest business challenges? AI is changing the pace of business, and organizations everywhere are struggling to help their workforce, partners, and customers keep up. At Litmos, we're building the Learning Acceleration Platform that helps organizations build human capability faster—and we're looking for people who are passionate about making a meaningful impact for customers while growing alongside a collaborative, people-first team. Litmos is the Learning Acceleration Platform that helps organizations build capability faster, adapt at the speed business changes, and scale learning to anyone, anywhere. Combining an intuitive platform, AI-powered capabilities, trusted content, expert services, and a broad ecosystem of integrations, Litmos helps organizations accelerate workforce productivity, improve customer adoption and retention, enable high-performing partners, and reduce organizational risk through continuous learning. Organizations such as Hewlett Packard Enterprise, Graco, Sabre, and Russell Mineral Equipment trust Litmos to accelerate learning across their workforce, partners, and customers. Today, more than 11K customers with 30 million learners across 150 countries and 37 languages use Litmos to build the capabilities their organizations need to succeed. Backed by Francisco Partners, one of the world's leading technology investment firms, we're investing in the future of learning—and the people who are building it. Learn more at www.litmos.com . We are seeking a Manager, Cloud Engineering to lead and develop a team of cloud engineers while serving as a key technical authority for the architecture, reliability, and evolution of our cloud platform. This role combines hands-on technical leadership with team development, driving infrastructure modernization, automation, and operational excellence across our global SaaS environment while building and mentoring a high-performing enginee

pythonawsazure
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Top cities for Reliability Engineer

City links are canonicalized and require at least 20 current jobs.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.