About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a Rack Power Engineer with deep expertise in high-power conversion and distribution to design, qualify, and support power systems for AI supercomputers. You will own rack power solutions—including power shelves, AC/DC rectifiers, power supply units (PSUs), power management controllers (PMCs), and high-current distribution—from requirements and supplier development through deployment. You will also monitor fleet rack power health, lead debugging and root-cause investigations, and drive improvements into hardware, firmware, and qualification coverage. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own rack power architecture and requirements for high-power AI supercomputing systems, including power budgets, AC input interfaces, DC distribution, redundancy, efficiency, serviceability, and integration with data center infrastructure. Drive the design and supplier development of power shelves, rectifiers, PSUs, PMCs, busbars, connectors, and protection circuits. Review electrical designs and control behavior, and evaluate performance, cost, reliability, and availability trade-offs. Define and execute component, shelf, and rack qualification plans covering load transients, current sharing, hot-swap, startup and shutdown, redundancy failover, fault protection and recovery, thermal limits, and AC disturbances and ride-through
Jobiba hiring network
Reliability Engineer Jobs
2,028 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. POSITION SUMMARY CVS Health is seeking a highly skilled Staff Data Engineer, Observability Engineering to join the Enterprise Observability Platform organization and help advance the next generation of observability, infrastructure, and security data capabilities. The Staff Data Engineer, Observability Engineering will play a critical role in designing, building, and operating scalable data pipelines and data products that power enterprise observability, operational intelligence, and security analytics across the organization. The Staff Data Engineer, Observability Engineering is a senior individual contributor responsible for developing and optimizing Databricks-based data engineering solutions that ingest, transform, govern, and deliver high-volume telemetry, infrastructure, application, and security data. This role combines deep hands-on technical execution with ownership of engineering excellence, operational reliability, performance optimization, and data platform best practices. Working closely with Observability Engineering, Security Engineering, Infrastructure Engineering, and Data Platform teams, the Staff Data Engineer, Observability Engineering will contribute to the evolution of the enterprise observability lakehouse by building resilient ingestion frameworks, establishing data quality standards, enhancing governance controls, and driving efficient, scalable data processing patterns. The id
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary: We are looking for a Staff Systems Engineer - Digital to join our team, the foundational layer that powers access, governance, and intelligence across our digital products. You will work horizontally across engineering, product, and AI teams, leading, owning, and evolving the shared infrastructure that every team at the company depends on. You will be the primary authority on how information is modeled, governed, and served across operational, analytical, and AI workloads - driving quality, compliance, and reliability at scale. If you thrive in a role where your architecture decisions multiply the productivity and capability of entire teams, this is the opportunity for you. Key Responsibilities: Data Architecture & Platform Ownership: Define and own the enterprise data architecture strategy across operational, analytical, and AI/ML workloads Design and govern data models, data contracts, and canonical schemas used across product and platform teams Evaluate and standardize data platform tooling — data lakes, warehouses, streaming, and serving layers (GCP BigQuery, Pub/Sub, Dataflow, or equivalent) Serve as the primary point of contact and SME for shared data platform concerns acro
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Responsibilities Analyze inline/param/probe/DOE data to identify yield detractors and drive continuous improvement. Apply semiconductor device physics, process knowledge, and statistical tools to troubleshoot yield issues. Collaborate with module engineering teams to diagnose process/tool‑related yield variation and ensure timely resolution. Lead or participate in cross‑functional task forces to solve complex yield, defect, or process integration challenges. Perform root‑cause analysis using FMEA, 8D, SPC, and other structured methodologies. Publish clear Pareto analyses and maintain dashboards for assigned product lines. Support new technology transfer, process baseline setup, and qualification activities. Partner with equipment engineering, shift engineering, and quality teams to address long‑term defect or excursion issues. Conduct material, process, and equipment evaluations and recommend optimization strategies. Ensure product performance meets design and reliability requirements and propose corrective actions when gaps exist. Leverage AI-Enabled and AI-Assisted solutions, including Copilot, YMS Genie, Agentic AI, AI Agents, and Large Language Models (LLMs), to accelerate yield analysis, automate engineering workflows, summarize insights, and enhance diagnostic efficiency. <spa
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. In this position, you will be joining a brilliant team to collaborate with Micron’s various design and verification teams all over the world, support the efforts of design verification flow development, optimization, and implementation and work with cross-functional team to assure the best cost, quality, reliability, time-to-market, and customer satisfaction We are seeking a talented CAD senior engineer specialized in analog and digital simulations to join our high-achieving engineering team. As a CAD senior engineer, you will play a pivotal role in optimizing and simulating complex analog and digital circuits, contributing to the development of state-of-the-art products that shape our industry. You will get a chance to develop the industry's first and unique flows to enhance productivity and reduce pre-silicon failures. Responsibilities: Work closely with memory design teams and solve their daily challenges and provide complete solutions for the future. Proactively identify problem areas for improvement, propose, and develop innovative solutions. Develop highly scalable, and clean software systems. Continuously evaluate and implement new tools and technologies to improve the current software flows. Work with EDA vendors to evaluate and integrate E
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Are you excited to drive Quality initiatives on new product from Design through Qualification? Collaborating across multi-functional teams to deep dive and actively drive resolution and ensure customer satisfaction in a fast paced and dynamic environment? Innovate and develop Quality and Reliability Strategies and enhancements? All while being part of a fun and supportive team? As a Development Quality Engineer in DEMQRA DQA Design Team at Micron Technology, Inc., your responsibilities will include the following: Work with key development partners in Design, Product Engineering, Technology Development, Quality Assurance and Manufacturing to identify and mitigate potential quality and reliability issues during the Design and pre-qualification phase of new product Identify Design specific risks for new High Performance Memory product and collaborate with the Design, PE, TD and QA teams to ensure correct verification, validation and detection specs are in place to ensure visibility to the status and success of mitigation efforts Drive deep-dive problem solving and mitigation/resolution alignment across the portfolio for DRAM related quality and reliability issues Act as the Subject Matter Expert on Design related quality and Live Die reliability issues and advise Global Quality leadership on risk and mitigation strategies Partner with Design, PE, and QA to align and recommend changes to verification and/or validation flows to ensure they are optimized to meet Micron’s QR nee
Leidos is seeking a Test and Integration Engineer to lead cutting-edge testing and validation efforts within the Undersea Systems Division (USD) . This role is a unique opportunity to drive innovation in testing underwater vehicle systems, maritime sensors, subsea telemetry, and ISR solutions that support critical defense and national security missions. You'll contribute to a multidisciplinary team focused on testing, integrating, and validating advanced maritime technologies, ensuring system reliability and performance from prototype design to full-scale system deployment for ongoing Navy missions. Leidos’ Undersea Systems Division is a recognized leader in C4ISR technologies, delivering innovative, mission-critical solutions across sensor networks, unmanned systems, and tactical platforms . We’re known for achieving “industry firsts” in the most challenging maritime domains. Join us and be part of a world-class team delivering unmatched solutions for today's most pressing maritime missions. Why Join Us? Make an Impact : Your work will directly support U.S. maritime dominance and national security. Lead Innovation : Be at the forefront of applying innovative technology and autonomy to real-world maritime systems. Work with Experts : Collaborate with a top-tier team of engineers, scientists, and technicians located across the U.S. Shape the Future : Influence both the strategic and tactical direction of next-generation subsea technologies. What You’ll Do Test and Validate: Write, develop and execute comprehensive test plans, procedures, and protocols to ensure system functionality, reliability, and compliance with requirements. Integrate systems and conduct hands-on testing: Write, develop, and execute integration plans to bring complex systems together. Perform field</b
The SCG Architecture team is hiring a Senior Power Integrity Co-Design Engineer to architect and deliver di/dt mitigation across silicon, package, board, and platform. This role bridges architecture, silicon, and platform — translating product noise targets into shipped specifications, and feeding silicon findings back into the next generation's build. Success in this role requires strong systems thinking and a willingness to accept ambiguity. It also requires the ability to apply AI as a force multiplier while maintaining rigorous engineering judgment. What you'll be doing: Architect voltage-noise mitigation across the full stack — silicon, package, board, platform — and own the codesign trade-offs between them. Co-design noise features with Speed, Power, Reliability, Circuit Design , Power-Arch, ASIC, and platform teams. You're the connective tissue across the codesign web. Work with other team members to define product-level voltage noise targets, drive them to closure, and sign them off at shipment. Build and take ownership of the Sim-to-Si correlation methodology for noise. You know when a model is lying and when silicon is. Model and prototype next-gen noise features — transient sense, droop response, mitigation IP, and codify them so every future program inherits them. Lead show-stopper noise bugs during bringup. The critical issues stop with you. Drive architecture-level codesign tradeoffs across V/F Power Noise Reliability Thermal (Noise-Variation) and (Noise-to-Closure) boundary work, where the highest-leverage innovation lives. What we need to see: BS / MS / PhD in EE, CE, or related (or equivalent experience). 5+ years in silicon power integrity, voltage noise, or PDN. Deep expertise in at least one of
SCG sits at the crossroads of design, architecture, marketing, and productization—owning the journey from the architecture stage through final product definition across Gaming, Datacenter, Automotive, and Embedded markets. As a System Verification CoDesign Engineer, you will work on system-level speed features, develop the verification collaterals and automation infrastructure to characterize and validate them, and lead debug of the complex silicon issues that stand between a program and on-time shipment. This is a hands-on role for an engineer who combines deep technical craft with the drive to compress cycle time using modern tooling—including AI—without losing rigor. What You’ll Be Doing: Collaborate cross-functionally with system architects, hardware, firmware/software, process/reliability, and operations teams to co-design system-level speed features and deliver industry-defining products. Understand system level behavior and speed reliability margins, bounding box constraints and identify solutions that optimize margins . Translate hardware features and architectural requirements into verification techniques that achieve full coverage across testing flows. Perform closed loop validation by correlat ing silicon behavior against timing simulation and design expectations; provide actionable feedback to improve future designs. Define, prototype, and refine pre- and post-silicon bring-up flows to ensure
We anticipate the application window for this opening will close on - 30 Sep 2026 Careers that change lives start here. Medtronic is a global leader in healthcare technology with a Mission to alleviate pain, restore health, and extend life. Our 95,000 employees work across more than 150 countries to put patients first — developing innovative medical technologies that improve the lives of 72+ million patients each year. Your unique talents will help shape the future of healthcare while building a career grounded in purpose, growth, and impact. A Day in the Life The Acute Care & Monitoring (ACM) business develops technologies that help clinicians monitor, assess, and manage patients across the continuum of care. Within ACM, the Test Engineering function supports product development, verification, reliability, and manufacturing readiness through the design and execution of test methods, fixtures, and equipment used to evaluate product performance and system functionality. This team partners closely with systems, mechanical, electrical, software, quality, and manufacturing engineering teams to support the development and sustainment of medical technologies. This position is based in Lafayette, Colorado and is an onsite role. Travel up to 10% may be required to support testing activities, cross-functional collaboration, and business needs. As a Mechanical R&D Engineer II, you will support system-level, reliability, mechanical, and electrical testing activities for Acute Care & Monitoring products. This role combines hands-on laboratory testing, fixture development, data analysis, and cross-functional collaboration to evaluate product performance, investigate issues, and support verification activities throughout the product lifecycle. Primary Responsibilities <li
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity We are a global team of innovators shaping the future of observability. Our intelligent platform gives customers real-time insight into complex systems so they can innovate faster and operate reliably in an AI-first world. If you’re excited by high-throughput distributed systems and want to contribute to one of the largest and fastest-growing observability platforms, we’d love to hear from you. Join a backend engineering team focused on building and operating JVM-based services that ingest, process, and serve massive volumes of telemetry data. You’ll work on high-scale, low-latency systems that power mission-critical observability features used by engineers worldwide. What you'll do Design, build, and operate JVM-based microservices (primarily Java and Kotlin) with a focus on performance, scalability, and reliability. Own services end-to-end: architecture, implementation, deployment, monitoring, on-call participation, and continuous improvement. Apply strong concurrency and performance practices: asynchronous programming, backpressure, efficient I/O, memory management, and GC tuning. Build and evolve event-driven systems; work with Kafka for streaming, partitioning, consumer groups, and schema evolution.Instrument services for deep observability (metrics, logs, traces), define SLIs/SLOs, and use e Experience with Kafka or similar streaming technologies (topic/partition strategy, consumer lag, idempotency, schema compatibility) strongly preferred. Proficiency w
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are looking for an experienced Staff Software Engineer in Test to join our Identity Management Engineering (IDM) team serving the Privileged Access Team (PAM). This team is passionate about delivering large-scale, mission-critical software in a fast-paced Agile environment. In this role you'll be working with a team of highly-skilled and talented engineers, responsible for delivering sophisticated backend solutions that help Okta reliably operate at large scale and be highly available. As part of the team, you’ll be ensuring projects are completed with the highest quality and reliability using automation at every level for fast, robust and secure releases. Job Duties and Responsibilities: Review requirements and design specs to develop relative test plans and test cases Automate API tests, end-to-end tests, reliability/scale tests Work with engineering management to scope and plan engineering efforts Communicate and document QE plans for scrum teams to review Review application code, identify bug and other areas of weakness, architect tools for future coverage Automate all critical features to maintain zero-debt cadence Release features with solid quality Respond to production issues/alerts and customer issues during on-call rotation Be a strong customer advocate with a strong quality DNA. Requirements: 5+ years of QE experience preferably in an enterprise SaaS company 3+ years experience in quality engineering for enterprise level software. 5+ yea
Job Title Product Support Engineer (Open) Job Description As a Product Support Engineer, you will be responsible for providing advanced technical support for Philips healthcare solutions and medical informatics platforms. The role focuses on diagnosing and resolving complex technical issues, supporting customer escalations, and contributing to continuous improvements in product support processes to enhance customer satisfaction and operational performance. Your role: Provide advanced technical support for healthcare products and solutions, diagnosing and troubleshooting software, hardware, network, and system-related issues. Act as a subject matter expert for internal teams, field engineers, and customers, supporting complex technical investigations and escalations. Collaborate with global teams, including R&D, Product Management, Service, and Field Service organizations to drive issue resolution and product improvements. Advocate for customer needs throughout the product lifecycle while promoting service excellence and customer satisfaction. Lead technical troubleshooting efforts and provide Level 3 escalation support for critical customer issues. Support service process improvements, reliability initiatives, technical communications, and knowledge-sharing activities. Participate in the development of training materials and technical documentation for internal and external stakeholders. Drive continuous improvement initiatives focused on product performance, supportability, and customer experience. You're the right fit if: You have experience in Healthcare IT, Biomedical Engineering, Medical Informatics, Clinical Informatics, Technical Product Support, or related fields. You have experience diagnosing and troubleshooting complex hardware, software, networking, or system integration issues. Yo
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. Our Core Data Engineering team is responsible for designing and building all foundational datasets used across Robinhood and within our operational business areas. We build the foundational data models and core data layers that are consumed by downstream users, ensuring high data quality and reliability. We also develop internal data and AI tooling that enables teams across the company to scale their data development workflows efficiently. Our team collaborates with engineering, product, brokerage, crypto, and marketing teams to expand and grow Robinhood's products around the world ! We also partner with Machine Learning teams to build robust training datasets that power intelligent product features. As the Engineering Manager for our Toronto Data Engineering team, you will lead a team of exceptional engineers and drive the execution of key data initiatives. In this role, you will balance technical leadership with people management, dedicating approximately 60% of your time coaching and 40% to hands-on technical contributions, such as code and architecture reviews. You will drive roadmap planning and establish clear goals for the team, particularly as we expand into new mark
About Ubiquiti At Ubiquiti Inc., we create technology platforms for Businesses, Smart Homes, and Internet Service Providers, driven by our goal to connect everyone, everywhere. To date, Ubiquiti has shipped over 100 million devices worldwide, from ISP networking products to next generation of IT solutions. Our growth is made possible by the dedicated team of hundreds behind the scenes. From software developers and product managers to designers and strategists, Team UI is driven to achieve our common goal: Rethinking IT. At Ubiquiti, you’ll heighten your potential and broaden your horizons - all while shaping the future of connectivity. Join forces with us on our mission to build a better IT industry. We are currently looking for a highly skilled Backend Software Engineer (Node.js) to join our team in Stockholm, Sweden. Please note that applicants must live in Sweden and hold a valid work permit at the time of application to be considered for this role. Team: You will join the UniFi Talk team that builds Ubiquiti's VoIP/phone system inside the UniFi ecosystem. UniFi Talk gives users desk phones plus the UniFi Talk application running on compatible UniFi consoles. This role is well suited to an engineer who enjoys solving complex product and platform problems, improving reliability, and working across backend, embedded, cloud, and application boundaries. Responsibilities: Design, build, and maintain backend services in Node.js and TypeScript. Develop secure, scalable, and maintainable APIs and service workflows. Contribute to architecture and code design decisions across the team. Review code and help maintain strong engineering standards, release quality, and development workflows. Investigate production issues, debug complex failures, and ship robust fixes. Work closely with embedded, web, mobile, product, and design teams to deliver platform capabilities used across the UniFi ecosystem. Improve the reliability, observability, and developer experience of t
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime