Job Details: Job Description: The Role and Impact As a Product Failure Analysis Engineer at Intel, you will play a vital role in identifying and resolving product issues to enhance performance, quality, and reliability. You will conduct comprehensive failure analyses and collaborate with cross-functional teams to develop innovative solutions that improve product yield and reliability. Through your efforts, you will directly impact Intel's ability to deliver world-class technologies. Business group The role is part of Intel Corporation's Product Quality and Reliability organization, which focuses on ensuring Intel's products meet the highest standards of quality and reliability. This group supports Intel's mission by driving improvements in product performance, enhancing manufacturing processes, and advancing technology innovation. Key Responsibilities Perform failure analysis and root cause investigation on Intel products, including CPUs, SoCs, and AI interface graphic platforms. Conduct failure analyses to identify defects and recommend corrective actions to design, fab, assembly, and test teams to improve product reliability. Develop and refine failure analysis methodologies, tools, and processes to enhance efficiency and accuracy. Collaborate with cross-functional teams to implement process improvements and innovative solutions for future technologies. Maintain and optimize Failure Mode and Effects Analysis (FMEA) and control plans throughout product lifecycles. Support new product launches, enabling technology transfers and toolset improvements across Intel sites. Qualifications: Minimum Qualifications Bachelor's or Mas
Jobiba hiring network
Reliability Engineer Jobs
2,028 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Responsibilities include, but not limited to: Key Responsibilities Develop systems to optimize operational performance across manufacturing platforms. Design, build, and deploy predictive models related to equipment utilization and performance enhancement. Analyze manufacturing data to identify bottlenecks and optimization opportunities for both short-term and long-term planning. Build and deploy LLM-based AI solutions to support fab operations, enabling smarter decision-making and improved operational efficiency. Collaborate closely with PMs, IT, and Core AI teams to drive cross-functional AI initiatives. Build interactive dashboards and visualizations to support data-driven decision-making. Work within the GCP environment for model training, deployment, and maintenance. Explore and implement MLOps, AutoML, and other advanced techniques to enhance model performance and reliability. Required Qualifications Master’s degree or above in Computer Science, Statistics, Engineering, or related fields 3+ years of experience in data science Proficient in Python, SQL, TensorFlow or other deep learning frameworks Skilled in data visualization tools Strong communication skills and experience in cross-functional collaboration </li
About DataCamp Data and AI skills are critical for thriving today, and DataCamp is the platform that empowers everyone to learn them. We help individuals and Fortune 1000 companies close the data and AI skills gap through world-class learning, hands-on training, and a global community of expert instructors. In this role, you'll collaborate with instructors and teams across curriculum, engineering, product, and marketing to build and scale the assessments and certifications that prove those skills—helping millions worldwide validate what they've learned and signal it to employers. About the role This is an individual contributor role. You will own the design and production of DataCamp's assessment and certification content, working with subject-matter experts and building AI systems that generate, review, and continuously refresh high-quality, defensible questions at scale. Here's what your day-to-day will look like: Own the end-to-end lifecycle for assessment and certification content—from blueprint and item design through review, publication, and ongoing maintenance—and manage deadlines across it. Design psychometrically sound assessments: write and review items, define scoring and passing standards, and safeguard validity, reliability, and fairness. Build and operate AI workflows—prompting, pipelines, and evaluation—that draft, grade, and quality-check items, so content production scales without sacrificing rigor. Source and collaborate with subject-matter experts to keep certifications current with the fast-moving data and AI landscape. Map certifications to the real skills hiring managers screen for, so a DataCamp credential is a credible signal to employers. Benchmark against the wider certification landscape—competing programs, industry standards, and emerging role definitions—and set the strategy for where our certifications go next. Continuously assess certification performance using pass rates, item statistics, and learner and employer feedback to dri
About Us: Sauce Labs is the world’s largest full-lifecycle, test automation platform, and the company behind Selenium. Trusted by 80% of the world’s top ten largest financial institutions and over 300,000 enterprise users, Sauce Labs provides the only AI platform capable of turning business intent into autonomous testing and quality assurance. With a proprietary dataset of 8.7 billion test runs, Sauce Labs empowers the Fortune 2000 to bridge the gap between AI-driven code generation and enterprise-grade software quality. Learn more at saucelabs.com . The Role: We are seeking an innovative and experienced AI Architect to join our engineering leadership team. This is a strategic role that will be instrumental in designing and building the next generation of AI-powered features for our continuous testing platform. You will be responsible for architecting scalable and robust AI solutions that transform how our customers gain insights from their test data and production environments, and how they create tests. Responsibilities: Define AI Architecture: Lead the design and architecture of cutting-edge AI/ML solutions for new product offerings, ensuring scalability, performance, quality and reliability within a cloud-native environment. AI-Powered Insights (Test & Production): Architect AI systems to derive actionable insights from vast quantities of test run logs and analytics data. This includes identifying patterns, anomalies, and performance trends. Production Error Reporting Integration: Design AI solutions that integrate with our existing error reporting product to analyze production issues for mobile and web applications, providing deeper understanding and predictive capabilities. Unified Data Intelligence: Develop architectures for combining insights from both test runs and production data, creating a holistic view of application quality and user experience. Automated Failure Analysis & Remediation: Architect AI models and systems t
About the Team The Growth Platforms team builds the systems and operating foundations that help OpenAI grow responsibly. We partner across the product portfolio to connect customer signals, identity and consent, campaign workflows, measurement, and product experiences into an AI-enabled growth engine. Our work helps teams launch, learn, and scale with a high bar for data quality, privacy, reliability, and customer trust. About the Role We’re looking for an experienced marketing technology and operations leader to drive cross-functional work at the intersection of growth, measurement, data, and automation. Your mission will be to turn fragmented tools, signals, and workflows into reliable, measurable, AI-enabled capabilities that teams can use safely at scale. You’ll work across Growth, Marketing Operations, Product, Engineering, Data Engineering, Data Science, Security, Privacy, Legal, and Revenue Operations, as well as external advertising platforms, measurement providers, and implementation partners. You’ll translate business requirements and privacy constraints into data contracts, integration designs, rollout plans, and reliable first-party data systems. This is a hands-on, high-impact role for someone who brings structure to ambiguity and moves from event schemas, APIs, and data quality assurance to operating cadences, partner enablement, and executive updates. This role is based in San Francisco or New York City with a hybrid office expectation. In this role, you will: Own the operating model for Growth’s marketing technology stack across identity, consent, audiences, activation, measurement, and experimentation. Own and operate the complete paid-media tracking and measurement system, including website pixels, server-to-server conversion events, mobile measurement integrations, identity and consent controls, attribution methods, and timely signal delivery to advertising platforms. Design and implement event schemas, data mappings, APIs, and integrations; valid
NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services. What you’ll be doing: Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale. Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation. Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data. Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems. Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation. What we need to see: BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10+ years operating production infrastructure or platform services. Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges. <l
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Fab16N Equipment Manager, you will lead and manage a team of equipment engineers to oversees the installation, modification, upgrade and maintenance of manufacturing equipment. Study equipment performance, reliability, upgrades and safety issues. Establish programs and solutions for increasing uptime and for equipment problems that affect the manufacturing process, and provide technical support to the manufacturing equipment repair and process engineering organizations. Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Join our world-class team at Micron, where brand new ideas and collaboration drive groundbreaking advancements in memory and storage solutions. Responsibilities: Guide local equipment owners through the installation and qualification of Diffusion tools, ensuring flawless performance. Complete all required safety training and ensure all work follows strict safety policies. Partner with Process Engineering, Operations, and IE teams to develop and execute strategi
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Fab16N Equipment Manager, you will lead and manage a team of equipment engineers to oversees the installation, modification, upgrade and maintenance of manufacturing equipment. Study equipment performance, reliability, upgrades and safety issues. Establish programs and solutions for increasing uptime and for equipment problems that affect the manufacturing process, and provide technical support to the manufacturing equipment repair and process engineering organizations. Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Join our world-class team at Micron, where brand new ideas and collaboration drive groundbreaking advancements in memory and storage solutions. Responsibilities: Guide local equipment owners through the installation and qualification of Diffusion tools, ensuring flawless performance. Complete all required safety training and ensure all work follows strict safety policies. Partner with Process Engineering, Operations, and IE teams to develop and execute strategic plans f
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Fab16N Equipment Manager, you will lead and manage a team of equipment engineers to oversees the installation, modification, upgrade and maintenance of manufacturing equipment. Study equipment performance, reliability, upgrades and safety issues. Establish programs and solutions for increasing uptime and for equipment problems that affect the manufacturing process, and provide technical support to the manufacturing equipment repair and process engineering organizations. Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Join our world-class team at Micron, where brand new ideas and collaboration drive groundbreaking advancements in memory and storage solutions. Responsibilities: Guide local equipment owners through the installation and qualification of PVD tools, ensuring flawless performance. Complete all required safety training and ensure all work follows strict safety policies. Partner with Process Engineering, Operations, and IE teams to develop and execute strategic plans fo
NVIDIA pioneers computer graphics, gaming, AI, and accelerated computing. We are looking for a Senior Solution Architect with full-stack software engineering experience to join our team and play an important role in developing Sales AI applications. This position offers the opportunity to design, build, and evolve solutions that bring generative AI and intelligent workflows into everyday sales experiences. You will work across the application stack and collaborate with Product, AI and machine learning, Data, Security, Solution Architecture, and Engineering teams to deliver secure, reliable, and scalable solutions used globally. What you’ll be doing: Collaborate with application teams to design, develop, and maintain scalable full-stack solutions for enterprise sales workflows. Guide technical solutions across front-end, back-end, APIs, data services, integrations, and cloud infrastructure. Translate product requirements and business needs into secure, maintainable solutions and intuitive user experiences. Integrate generative AI models, AI services, APIs, retrieval systems, and agentic workflows into production applications. Design application architectures that support performance, availability, observability, security, scalability, and long-term maintainability. Lead technical design discussions, compare implementation approaches, make informed architecture decisions, and evaluate emerging technologies. Improve engineering practices for testing, code quality, continuous integration and delivery, monitoring, documentation, and production readiness. Investigate complex issues and develop solutions that improve reliability and user experience. Mentor engineers, share technical knowledge, and contribute to engineering standards and collaborative team practices. What we need to see: <
Job Title Lead Software Technologist I - UI Job Description Minimum required Education: Bachelor's / Master's Degree in Computer Science, Software Engineering, Information Technology or equivalent. Job title: Lead Software Technologist I Your role: Lead development of intuitive clinical user interface applications used by clinicians in high-acuity environments. Collaborate with UX teams and clinicians to ensure applications meet usability and safety requirements. Provide technical leadership across modern UI technologies and supporting C++ layers for embedded devices. Lead and mentor Platform development teams on software architecture, design patterns, and embedded systems best practices through daily hands-on collaboration. Lead and mentor application development teams on application services for physiological data visualization, and clinical decision support features. Establish and enforce quality standards, development methodologies, and coding practices that drive continuous improvement in software reliability and performance. Conduct rigorous code reviews and provide constructive technical feedback to ensure adherence to medical device software standards (IEC 62304, FDA regulations) Optimize application performance by identifying and resolving bottlenecks in resource-constrained embedded Linux environments. Drive the adoption of AI-enabled development tools and demonstrate measurable productivity improvements across the team. Collaborating with cross-functional teams including Product Management, QA, Regulatory, and other engineering leads to define and deliver features. Support software lifecycle management activities including sustaining engineering, defect resolution, and platform evolution.
Assistant Manager - Instrumentation at Adani’s Copper Manufacturing Complex in Mundra, you will play a pivotal role in ensuring the optimal performance, reliability, and safety of all instrumentation systems critical to copper production processes. Your expertise in instrumentation engineering will drive continuous improvement initiatives, maintenance strategies, and technology integration aligned with our commitment to operational excellence and sustainability. You will collaborate closely with cross-functional teams to support production goals, uphold quality standards, and embrace Adani’s values of innovation, integrity, and community responsibility. This position offers a unique opportunity to contribute meaningfully to one of the world’s leading copper manufacturing facilities while advancing your professional growth within a dynamic, inclusive, and high-performance work environment. Source: Adani Group | Job ID: 48032
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We lead complex technical programs that help Plaid scale its engineering platform. We partner across Engineering, Infrastructure, Data, Security, ML, Legal, and Product to deliver company-wide technical initiatives that improve reliability, scalability, and developer productivity. You'll lead strategic technical programs from planning through execution. You'll partner with engineering leaders to align stakeholders, manage dependencies, drive decisions, and ensure successful delivery of complex initiatives. You'll work across a variety of technical domains, adapting quickly to new challenges and helping teams execute effectively. As a Technical Program Manager, you will lead high-impact, cross-functional initiatives. As a generalist, you may work on a variety of programs. An example is one that strengthens Plaid's data and machine learning platforms. You will partner with engineering, product, data, legal, privacy, and business stakeholders to drive complex technical programs from planning through execution. Your work will help improve data governance, modernize machine learning infrastructure, and accelerate the adoption of trusted, high-quality datasets that power analytics, artificial intelligence
About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Through strategic partnerships and self-built campuses, we are scaling one of the world's fastest-growing AI infrastructure platforms. The Supply Chain organization ensures critical infrastructure components—from compute systems and networking equipment to integrated rack solutions—are sourced, manufactured, qualified, and delivered with the speed and reliability required to support frontier AI development. We partner closely with Hardware Engineering, Manufacturing Quality Engineering, Infrastructure Delivery, Hardware Operations, Finance, and suppliers worldwide to build a resilient, scalable supply chain capable of supporting rapid infrastructure expansion. As Industrial Compute continues to grow, Supply Chain serves as the operational bridge between engineering innovation and large-scale infrastructure deployment. About the Role We are seeking a Supply Chain Manager to lead strategic execution across sourcing, supplier operations, manufacturing quality, and infrastructure delivery for OpenAI's AI infrastructure portfolio. This role will oversee a multidisciplinary team responsible for strategic sourcing, manufacturing quality engineering, and technical program management while partnering closely with engineering, finance, hardware operations, and deployment teams. You will drive supplier strategy, manufacturing readiness, production planning, quality performance, and operational execution across the full hardware lifecycle. Success requires balancing long-term supplier strategy with day-to-day execution. You'll establish scalable operating mechanisms, strengthen supplier partnerships, manage complex cross-functional programs, and ensure OpenAI can rapidly deploy AI infrastructure without compromising quality, cost, or reliability. This is a people leadership role responsible for developing a high-performing organization while driving operati
About the Team OpenAI’s Platform team powers how millions of developers and enterprises build with our models. We provide APIs and agentic solutions used by global startups and fortune 500s. We work closely with product, engineering, design, and go-to-market to build a world-class platform that pushes the frontier of AI capabilities. About the Role As a Data Scientist on the Platform team, you will drive a data-driven culture for OpenAI’s API and B2B solutions. You’ll define the metrics that matter for developer success and enterprise value, measure the impact of new models and features, and partner with PMs and engineers to improve model quality, reliability, latency, and cost. Your work will shape how thousands of products adopt agentic AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will Embed with the Platform product team as a trusted partner, uncovering ways to improve developer experience, reliability, and usage growth Define north-star metrics across the developer funnel (activation, retention, growth), as well as latency/cost guardrails for new features and models Design and interpret A/B tests and controlled rollouts (e.g., new model versions, pricing/limits, new API features, new B2B products) Build source-of-truth dashboards and self-serve data tools for product, engineering, and go-to-market teams Translate product learnings into actionable feedback for Research (e.g., failure modes, eval gaps, model response quality) You might thrive in this role if you have 5+ years in a quantitative role in ambiguous, high-growth environments (platforms, APIs, or B2B products a plus) Depth in SQL and Python, with a track record proposing, designing, and running rigorous experiments Experience defining and operationalizing metrics from scratch (including reliability/latency/cost and safety) Strong cross-functional communication with PMs, enginee
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime