We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. The Data team within Plaid’s Fraud organization builds the machine learning systems that power Plaid’s fraud detection products, leveraging Plaid’s unique network data to identify and stop fraud before it happens. The team owns the full ML lifecycle—from feature pipelines and model training to production serving and monitoring—building reliable, scalable systems that deliver high-quality fraud detection as we grow to support hundreds of customers. As a Senior Machine Learning Engineer, you will own the development of high-performance feature computation and online inference pipelines that power production machine learning systems at scale. You’ll build robust observability, monitoring, and automated debugging capabilities, while leveraging AI-assisted tools to investigate complex system behavior and maintain high reliability. You’ll partner closely with ML Infrastructure, Data Science, and Product teams to execute critical technical initiatives and deliver scalable, high-impact ML solutions. Responsibilities: Build and scale machine learning systems that power a rapidly growing fraud detection product in a fast-paced environment. Solve complex technical challenges at the intersect
Jobs in United States
Reliability Engineer Iii in United States
655 active opportunities · Updated October 2026
Showing
15 jobs
Explore current reliability engineer iii jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The ChatGPT organization at OpenAI supports our mission by bringing advanced AI capabilities to hundreds of millions of users worldwide. The Image Generation team is responsible for one of the fastest-growing experiences in ChatGPT, enabling users to create, edit, and transform images through natural language. Recent advances in our multimodal image models have dramatically improved image quality, instruction following, editing precision, consistency, and text rendering, unlocking entirely new creative and professional workflows. We work at the intersection of research, infrastructure, and product to build the systems that power image generation at global scale. Our team partners closely with researchers, product engineers, designers, and platform teams to bring state-of-the-art image capabilities to millions of users while continuously pushing the boundaries of what AI-powered creation can do. About the Role We are looking for an experienced Backend Engineer to join the Image Generation team and help build the systems that power image creation and editing across ChatGPT. You'll work on the core backend infrastructure that enables users to generate, edit, and iterate on visual content using cutting-edge multimodal AI models. This includes building highly scalable services, orchestration systems, APIs, storage platforms, and distributed infrastructure that support billions of image generations and editing workflows. You'll partner closely with product, research, and mobile teams to transform breakthrough AI capabilities into reliable, performant experiences used by millions around the world. In this role, you will: Design, build, and operate backend systems that power image generation and image editing experiences in ChatGPT. Develop scalable APIs, services, and infrastructure that support multimodal AI workflows. Optimize reliability, latency, throughput, and cost across large-scale distributed systems. Partner with researchers to productionize new im
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake runs large scale cloud infrastructure to deliver its own service — production and internal deployments, Kubernetes fleets, CI/CD, etc. Our cloud spend is in billions of dollars per year. The Cloud Efficiency team builds a unified, self-serve cloud efficiency platform along with AI skills and agents that makes spend observable, attributable, governable while driving recommendations and optimization of our cloud spend. AS A SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Design, develop, and maintain scalable platform for resource ownership registry, usage attribution, utilization measurement, and cost modeling. Build AI agents, tools and automation to enhance system monitoring, alerting, and root cause analysis. Improve and optimize data ingestion, storage, and query efficiency for cloud utilization, cost and efficiency data at scale. Collaborate with teams across Snowflake to understand attribution and observability needs and implement solutions that improve operational visibility. Contribute to open-source and industry best practices in monitoring and distributed systems monitoring. Ensure high availability, reliability, and performance of team-managed platforms by participating in on-call rotations and incident management. Partner with Finance, Product and Engineering
About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Where Data Does More. Join the Snowflake team. Join our ML Feature Store team where we're building cutting-edge product capabilities that power complex feature transformations and low latency feature serving. We're revolutionizing machine learning feature management and serving capabilities as part of the Snowflake ML suite of products. In the era of GenAI and agents, our team delivers high-quality, fresh feature solutions that make a real difference for our customers. IN THIS ROLE AT SNOWFLAKE, YOU WILL: Help define and own the roadmap for Snowflake Feature Store, working collaboratively with senior architects and ML team leadership Build and execute a vision for incorporating new advances in machine learning Ensure operational excellence of services and meet reliability, availability, and performance commitments Collaborate across ML partner teams to improve development velocity and capabilities Support team members in delivering high technical quality WE WOULD LOVE TO HEAR FROM YOU IF YOU HAVE: 10+ years of experience in designing and building data serving infrastructure and/or machine learning platforms. Strong track record working with machine learning systems and platforms. Strong understanding of computer science fundamentals. B.Sc . in Computer Science Fluency in Ja
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. The Integrations Operations Engineering (IOE) team strengthens Plaid's network with financial institutions by directly resolving integration issues to increase the reliability and quality of data, and by building new integrations to grow our financial network. We sit at the nexus of engineering, a deep understanding of Plaid's products and customers, and direct financial-institution relationships: we write and ship the code that keeps connectivity healthy, we use data to focus on the issues with the greatest customer impact, and we work directly with data partners (the financial institutions and platforms themselves) to resolve the problems that can't be fixed from our side alone. The team plays a mission-critical role in ensuring industry-leading connectivity so our customers can meet ever-expanding financial-services use cases and reach as many users as possible. What You'll Do Investigate and resolve the highest-impact integration issues by writing maintainable, tested code and deploying it to production, then monitor for regression or degradation after your changes ship. Prioritize by customer impact. The team runs a business-value-based prioritization model that automatically
As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. To provide substantive overlap with the team, this position must be in Eastern Time. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute.
$107.4K – $155.7K/yr
Interactive Design Engineer - HPIQ Description - About The Role As a Product Design Engineer at HP IQ, you’ll work at the intersection of creativity, engineering, and design to develop groundbreaking devices that redefine how people interact with technology. You’ll collaborate closely with industrial designers, hardware engineers, and interaction designers to transform ideas into functional prototypes and production-ready designs. From 3D CAD modeling and hands-on prototyping to engineering analysis and design validation, you’ll help solve complex mechanical challenges and drive designs from early concepts toward production. You’ll have the opportunity to take ownership of meaningful engineering work while learning from an experienced, multidisciplinary team and contributing to products that push the boundaries of what’s possible. What You Might Do Design mechanical parts, components, and assemblies for new consumer devices using 3D CAD (NX) Create detailed 2D engineering drawings, define specifications and tolerances, and work closely with overseas vendor partners to ensure accuracy and quality Design and run experiments, applying engineering analysis and test results to guide design decisions Evaluate materials, manufacturing processes, and design tradeoffs to develop robust and scalable solutions Work cross-functionally with hardware, software, industrial design, and other engineering teams to bring concepts from ideation through development Conceptualize, design, and build prototypes that quickly validate ideas Participate in design reviews, clearly communicating design decisions, technical analysis, and tradeoffs to both technical and non-technical audiences Build, test, troubleshoot, and iterate prototypes to identify issues and improve product performance, reliability, a
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview Join a team responsible for building the authentication and security solutions that help protect digital interactions across Mastercard's Identity Solutions platform. As a Senior Software Engineer, you will design, develop, and support secure, scalable, and high-performing applications that enable trusted identity verification, authentication, and fraud prevention capabilities. In this role, you will take ownership of complex technical challenges, contribute to software design and architecture decisions, and partner closely with product, security, and platform teams to deliver reliable, production-ready solutions. You'll play a key role in advancing engineering excellence through secure development practices, system reliability, automation, and continuous improvement while mentoring other engineers and influencing technical direction across the team. Role •Design, build, test, deploy, and maintain scalable, cloud-native applications and microservices •Develop REST APIs using Java and Spring Boot, focusing on performance, scalability, and reliability •Translate requirements into well-structured designs and architecture, ensuring maintainability and security •Lead and contribute to system design discussions, aligning with architectural standards and best practices<
We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Sr Design Engineer, you will work on design, simulation, and validation of next ‑ generation High ‑ Bandwidth Memory (HBM) architectures and circuit blocks. HBM requires advanced DRAM design knowledge combined with deep understanding of 3D stacked architecture, TSV signaling, wide I/O interfaces, PHY timing, power integrity, and system co ‑ optimization with GPUs/accelerators. This role sits at the intersection of DRAM design and high ‑ performance computing, enabling future AI/ML, HPC, and advanced graphics products. In this position, you will collaborate with Micron’s various design and verification teams all over the world and support the efforts of groups such as Product Engineering, Test, Probe, Process Integration, Assembly and Marketing to proactively design products that optimize all manufacturing functions and assure the best cost, quality, reliability, time-to-market, and customer satisfaction.
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. We are seeking a Principal Analog Design Engineer to lead the design, integration, and delivery of advanced analog and mixed-signal IPs for High Bandwidth Memory (HBM) products. This role is critical in developing and integrating high-performance analog subsystems within the HBM logic die, enabling industry-leading bandwidth, power efficiency, and reliability. As a principal engineer, you will provide deep technical leadership across analog design, IP integration, system alignment, and silicon execution, driving end-to-end success of HBM solutions. Responsibilities will include, but are not limited to: Lead the d esign and own critical HBM analog circuits, including: High ‑ speed transmitters and receivers Clock generation and distribution (PLLs, DLLs, CDRs) SerDes ‑ related analog blocks Biasing, reference, and calibration circuits
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron’s DRAM Design Engineering Group (DDEG) is where innovation meets excellence. We are advancing memory and storage technologies through collaborative engineering and creative problem-solving. Our team works at the forefront of semiconductor design, developing solutions that shape the future of memory products in a fast-paced, learning-focused environment. As a Design Verification Engineer, you will help develop next-generation memory technologies by verifying and optimizing digital and analog circuit designs. In this role, you will work closely with global multi-functional teams across the product lifecycle to deliver high-quality, manufacturable memory solutions that meet performance, reliability, cost, and customer requirements. Your work will directly contribute to bringing advanced memory products from concept to production. Responsibilities: Verify circuit functionality, reliability, power, and compliance with product specifications Drive verification planning, coverage closure, circuit debug, and design improvements Perform circuit modeling and simulation using industry-standard tools Support silicon validation, reticle experiments, and tape-out activities Partner with engineering, manufacturing, and product teams to deliver manufacturable designs Minimum Qualifications: Bachelor’s degree in Electrical Engineering or a related field 4+ years of semiconductor design, verifi
Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: Grow your career while continuing Exact Sciences’ inspiring work. Changing roles within the company allows you to develop your skills while changing the future of cancer in a new way. Schedule - Monday - Friday 8 am -4:30/5 PST Position Overview This role is responsible for diagnosing and resolving failures in enterprise laboratory instruments and automation systems, conducting root cause investigations, and escalating service issues to minimize downtime. It includes performing routine preventive maintenance and calibrations to ensure equipment reliability and compliance with service standards. The position requires accurate documentation aligned with Good Documentation Practices (GDP) and regulatory requirements such as OSHA, FDA, ISO, and CLIA. Success in this role involves developing technical expertise, training others, and collaborating across teams and vendors to support troubleshooting and project execution. Flexibility, commitment to quality, and a focus on process improvement—including SOP development and workflow optimization—are essential. Key Accountabilities: Include, but are not limited to, the following: Troubleshooting & Repair: Under general supervision, diagnose, repair, and resolve failures following established procedures on instrumentation and automation systems within the laboratory by applying technical expertise to restore functionality. Escalate unresolved or complex issues to senior s
Job Details: Job Description: This position is in the Intel Mask Operations within the Logic Technology Development team working in one of the most advanced semiconductor process technologies in the world. In this position the engineer will be an integral contributor to the ongoing production of photolithography masks for Intel's leading-edge silicon manufacturing solutions. In this position, you will work on-site, in a dynamic and collaborative environment solving complex and challenging technical problems on sophisticated manufacturing processes and equipment. As IMO module Engineer the responsibilities may include but not limited to: Process equipment installation and development. Equipment maintenance, management of troubleshooting activities, regular monitoring of process performance, defect analysis and reduction. Drives improvements on quality, reliability, cost, yield, process stability/capability, productivity, and safety/ergonomics. Working with cross functional teams to solve technical process and defect issues. Plans and conducts experiments to fully characterize the process throughout the development cycle. Establishes control systems to optimize and sustain production performance. Develops strategies to resolve difficult problems and establishes systems to manage these problems in the future. An ideal candidate should exhibit the below behavioral traits: Experience with rapid analysis of complex process issues and identification of a solution path Willing to work Independently with minimal direction, to be self-directing and show initiative Communication skills and demonstrated ability to summarize complex d
Other cities to consider
More places hiring for this role
Get new reliability engineer iii jobs in United States by email
Daily job updates · Unsubscribe anytime