Jobiba hiring network

Reliability Engineer Jobs

2,028 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

This is where your work makes a difference. At Baxter, we believe every person—regardless of who they are or where they are from—deserves a chance to live a healthy life. It was our founding belief in 1931 and continues to be our guiding principle. We are redefining healthcare delivery to make a greater impact today, tomorrow, and beyond. Our Baxter colleagues are united by our Mission to Save and Sustain Lives. Together, our community is driven by a culture of courage, trust, and collaboration. Every individual is empowered to take ownership and make a meaningful impact. We strive for efficient and effective operations, and we hold each other accountable for delivering exceptional results. Here, you will find more than just a job—you will find purpose and pride. Your role at Baxter The Principal Systems Engineer will serve as a Product Design Owner (PDO) responsible for technical owner for the design, risk and integration of infusion pump systems and/or projects, which combine electro-mechanical hardware, embedded software, and user interface components, as well as the interface with other related EM and/or digital products. This role drives operational excellence and predictable, consistent execution in our products, and design and risk integrity and may serve as Risk Owner on some projects. The engineer is accountable for product safety, performance, reliability, usability and regulatory compliance, as well as risk. The PDO drives design decisions, manages design control activities, may be accountable for a product risk file, and collaborates across engineering disciplines and cross-functions to deliver robust, reliable, safe and innovative infusion therapy solutions. What you will be doing: Leads interdisciplinary design and development of medical products in compliance with FDA, EU MDR,

recruitment
View job →
N
Nvidia
📍 Santa Clara, United States
1mo ago

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent engineering manager to own and deliver an end to end manageability stack for Data Center Systems. We are seeking an experienced manager who is deeply technical, hands-on, and has a wide system view. You will manage a team of experts, design & build OpenBMC based manageability software stack for NVIDIA’s next generation Data Center Compute Systems. We want to grow our teams with the smartest people in the world. If you're creative and autonomous, we want to hear from you! What you’ll be doing: Own and deliver OpenBMC based manageability stack for next generation Data Center Compute Systems. Own firmware delivered to data centers in terms of quality, reliability and telemetry performance. Manage and lead a distributed team of software engineers to deliver firmware stack with high quality. Work with data center architects and cloud customers for correct requirements and scope implementation to ensure speed of light product development. Work closely with cross functional teams to ensure scalable manageability architecture for all data centers products Drive efficiency, reliability and optimization in firmware architecture from a data center view point. Work closely with customers and internal teams to resolve issues at Speed of Light. What we need to see: BS, MS, or PhD in EE/CS or related field o

pythongitai
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
1mo ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will work on the systems software strategy and execution that brings new AI silicon from first power-on to a fully integrated system running production-representative models at expected functionality and performance. You will define how software exercises and validates compute, memory, interconnect, and I/O subsystems, then build the diagnostics, automation, and observability needed to find issues quickly. This role sits at the center of silicon, firmware, platform, systems, and workload teams. You will turn hardware specifications and performance targets into an end-to-end bringup plan, drive cross-functional debug, and establish the stress and regression infrastructure that makes each new platform reliable across operating environments. In this role, you will: Contribute to the end-to-end software bringup and validation strategy for new silicon and first-party systems. Define software-driven test coverage across compute, memory, interconnect, I/O, and their system-level interactions. Build diagnostics, test automation, telemetry, and regression infrastructure that accelerate first-silicon learning and issue isolation. Lead bringup from initial silicon arrival through board and system integration, docking, runtime enablement, and model execution. Design stress tests that characterize reliability, performance, and stability across workloads and operating conditions. Translate architecture specifications and performance models into measurable acceptance crit

REMOTEpythonawsrest
View job →
S
Sofi
📍 United States• Full-time
1mo ago

Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. Role Summary: We are looking for a talented and detail-oriented Sr Data Engineer to tackle data challenges. You will design, build, and maintain critical data pipelines and datasets, supporting areas like recruiting, compensation, talent management, and learning and development. Your work will enhance data accessibility and empower the People Team and business leaders to make informed decisions with high-quality, reliable data. Key Responsibilities: Develop and maintain robust data pipelines and datasets. Build foundational data products for key business areas. Enhance self-service data capabilities for the People Team. Ensure high standards in ETL/ELT operations, data quality, and pipeline reliability. Join us to drive impactful change and support SoFi's mission of fostering a thriving workplace through data excellence. What you’ll do: Design and build production dbt models in Snowflake that integrate Workday and other People systems into well-modeled, documented datasets, including slowly changing dimensions for People history. Build and operate Airflow DAGs that ingest People systems data and orchestrate dbt runs, keeping loads reliable and re-runnable. Own data quality and observability: dbt tests, freshness checks, row-count validation, and monitoring so issues are caught before stakeholders see them.

pythonsqlgcp
View job →

Job Details: Job Description: Shape the Future of Confidential Computing! Join Intel's elite engineering team to validate TDX (Trusted Domain Extensions) - a breakthrough technology redefining security and virtualization. Work on one of the industry's most complex and impactful products, where functional validation meets deep technical innovation. Why You'll Love This Role: High Impact: Ensure the reliability of TDX, a cornerstone of Intel's confidential compute. Deep Technical Challenge: Validate across the stack-firmware, OS, drivers, middleware, SDKs, and apps. Code-Driven: Build frameworks, automation, and reference implementations. Innovation at Scale: Operate at the intersection of security, virtualization, and cloud. Career Growth: Collaborate with world-class engineers on advanced architectures. Qualifications: Minimum Qualifications: Education BSc or MSc in Computer Science, Computer Engineering, or Electrical Engineering. Experience 5-7 years in software/hardware validation or verification. 5-7 years coding in higher-level languages: C or C&#43;&#43;. Skills Proficient in English. Excellent communication and teamwork skills. Strong analytical and problem-solving abilities. Preferred Qualifications: Operating Systems and Virtualization Pre/Post-Silicon validation Specman experience What we offer: At Intel, employees share in successes, enjoy comprehensive rewards and are inspired by an innovative & inclusive workplace. What can you expect when there is a match between us? <li

N
1mo ago

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform & node designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each bringing together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We’re searching for a highly motivated, technical leader to design, drive, and operationalize rack-scale factory and deployment flows for next-generation data center products. The ideal candidate will combine deep systems expertise, decisive technical leadership, and a passion for building reliable, debuggable, and scalable manufacturing and deployment solutions. What you’ll be doing: Lead and drive rack-scale/L11 flows for factory and initial data center deployment. Design and implement end-to-end factory workflows, including firmware flashing sequences, security provisioning, and deployment of software mitigations. Collaborate with data center architects, ODMs, and OEMs to define factory and data center requirements that ensure efficient and reliable production ramp. Champion reliability, debuggability an

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Rider Loyalty team is where riders become members. We build the membership, rewards, and benefits products that give people a reason to choose Lyft on every trip, and we make sure the value a rider has earned shows up at the moment it matters. Loyalty sits inside the Rider Loyalty, Partnerships, and Rider Pay (PLP) group. You will lead a team of engineers across iOS, Android, and Server. You will own the membership and rewards platform end to end and work daily with Product, Design, Data Science, and Partnerships. Responsibilities: Own the Loyalty roadmap from strategy through delivery. Turn goals like member growth and retention into an engineering plan, and manage the dependencies that run through Partnerships and Rider Pay. Build and scale the systems behind membership, rewards earning and redemption, and benefit delivery. Hold a high technical bar through architecture reviews, tech debt management, observability, reliability, and on-call. Grow engineers by matching people to the right opportunities, setting clear expectations, and giving feedback early. Experience: 5+ years building software professionally, including 2+ years directly managing engineers. You have managed a team that shipped both mobile and backend work, and you can still read and review code in at least one of those areas. You have owned a consumer product used by millions of people each month. You use AI tools in your own work and have a clear view of where they help and where they do not. BS/MS in Computer Science, Computer Engineering, or a related field, or equivalent practical experience. Benefits: Extended health and dental coverage options, along with life insurance and disability benefits Mental health benefits Family building benefits Child care and pet benefits Access to a Lyft funded Health Care Savings Accou

artificial intelligenceai
View job →

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team The Developer Productivity group is responsible for making Stripe’s developers happy and productive. We work on tools, processes, and code refactoring to accelerate Stripe engineering as Stripe scales. We’re looking for people with an interest in building the tools to improve the day to day experience of Engineers in Stripe. The ideal candidate will have a passion for solving developer experience problems, and a pragmatic ability to ship results iteratively—powered by a mix of technical expertise across some or all of: language processing tools; version control systems; build systems; and distributed systems engineering. You’ll be working on a mix of engineer-facing systems, CI infrastructure platforms, and big-data engineering. What you’ll do We have a ton of important work to do, which is why we’re hiring! Our active projects change all the time, but here are a few examples of recent projects so you can get an idea of the types of work we do: Build and manage systems to handle CI at massive scale—including batching, speculative stacking, merge-race inhibition, and more. Build and manage systems to handle our enormous CI test suite–identifying and managing flaky tests, assessing and reproducing flakiness, and optimizing for test effectiveness. Enhance our CI systems for reliability, including adaptive response to available capacity, resilience and self-recovery from outages, and highly-leveraged observability. Detect and isolate code breaka

D
Datadog
📍 Massachusetts• Full-time• From $192K/yr
1mo ago

About Datadog: We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—allowing for seamless collaboration and problem-solving among Dev, Ops and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team: The Revenue Data Engineering Teams designs, builds and runs the data pipelines and helper systems to accurately and in a timely manner quantify our customers’ usage across all Datadog products. This team is at the leading edge of any new product we release. The Revenue Data Processing team builds and operates the data pipelines that does billing, and cost attribution for all Datadog products. We process terabytes of data daily to power revenue-critical systems and are at the center of every new product launch at Datadog. As a Senior Software Engineer, you will own meaningful parts of a large-scale, mission-critical processing platform — driving architectural improvements, building new billing capabilities, and maintaining the high reliability bar our downstream consumers depend on. You Will: Design and build high-throughput data pipelines for billing and cost attribution Drive platform improvements — latency reduction, Spark optimization, sharding, and cross-datacenter reliability Own root-cause investigations on billing accuracy issues in collaboration with Finance and Product teams Contribute to new billing features Work across Python and Scala, with technologies including Spark, Airflow, Trino, and Apache Iceberg Participate in on-call rotation and maintain a high reliability bar for production systems Contribute to engineering standards and help grow the technical culture of the team You Are: You have significant experience building and operating production data pipelines at scale using Spark and Airflow

pythonaigo
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $207.5K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Team/ Role Paragraph: Coinbase is building the future of institutional trading as part of the Everything Exchange, powering the systems that let the world's largest financial institutions trade across markets with speed, depth, and trust. As a Market Data Engineer on the Markets team, you'll architect, build and own the systems that generate, capture, normalize, and distribute real-time market data across Coinbase at low latency - the ticker plant and distribution backbone of a unified trading platform - while leading a lean, high-impact team. It's a rare chance to architect market data infrastructure from the ground up, with the ownership of a startup and the reach of Coinbase. What you’ll do: Build and operate market data components: feed handlers, normalization, distribution, and venue connectivity. Design low-latency, high-throughput systems for real-time market data used by trading systems. Contribute to reliability and performance, participating in observability, on-call, and incident response. Write high-quality, well-tested code and help raise the quality bar on a lean team. Partner with engineers, product, and cross-functional stakeholders to deliver market data capabilities. Required skills and experience: 5+ years of backend engineering, with experience building and maintaining production low-latency systems. Hands-on experience with market data syst

REMOTEjavaawsai
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $218K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Team/ Role: Coinbase is building the future of institutional trading as part of the Institutional Everything Exchange, powering the systems that let the world's largest financial institutions trade across markets with speed, depth, and trust. As an Infrastructure Lead on the Exchange team, you'll own the infrastructure, deployment, and operational tooling behind a unified trading platform that runs across cloud and on-prem/colocated venues, keeping latency-sensitive, multi-node trading environments fast, reliable, and easy to build on. It's a rare chance to shape trading infrastructure from the ground up, with the ownership of a startup and the reach of Coinbase. What you’ll do: Own infrastructure, deployment, and operational tooling for latency-sensitive, multi-node trading environments across cloud and on-prem. Drive reliability and developer velocity, owning observability, deploy safety, and incident response for systems that run 24/7. Set the operational bar through standards, reviews, and automation, and reduce single-point-of-failure risk across the trading stack. Partner with engineers building the trading platform to make latency-sensitive systems operable and performant. Mentor engineers and build team resilience across a lean, high-impact environment. Partner cross-functionally with Product, Institutional Markets, and SRE to turn platform needs into a

REMOTEawslinuxai
View job →
S
Stripe
📍 San Francisco• Full-time
1mo ago

Who we are About Bridge Bridge is building a new payments platform powered by stablecoins to make global money movement simpler, faster, and more accessible. Through our APIs, businesses can send and receive funds across borders, give customers access to dollars through virtual accounts, and disburse USD globally. We believe stablecoin rails will become a foundational part of financial infrastructure, moving and settling trillions of dollars worldwide. Bridge is helping bring that future forward. Our team has built financial infrastructure at companies including Coinbase, Stripe, Square, Brex, Upstart, DoorDash, and Airbnb. We are united by the conviction that stablecoins can materially improve how money moves around the world. What you’ll do Build and ship end-to-end product experiences across frontend applications, backend services, APIs, and data systems. Design reliable, performant, and intuitive interfaces for developers, operations teams, and end users interacting with Bridge products. Develop and maintain the APIs and backend systems that power global payments, virtual accounts, and payouts. Own projects from technical design through launch and iteration, often operating with substantial autonomy and without dedicated product management support. Debug and resolve production issues across the stack, with a focus on reliability, performance, and customer impact. Make thoughtful tradeoffs among business priorities, user experience, speed of execution, and long-term technical quality. Who you are We’re looking for someone who meets the minimum requirements to be considered for the role. If you meet these requirements, you are encouraged to apply. The preferred qualifications are a bonus, not a requirement. Minimum requirements 6+ years of professional software-engineering experience. We are open to candidates across a range of seniority, from experienced individual contributors through senior technical leaders. Experience building and shipping production software

sqlrestai
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $186.1K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . We're hiring a Senior Software Engineer to join the CB Node team within the Platform organization. This team runs all the blockchain nodes that power every asset Coinbase offers to customers, currently spanning 55 unique protocols across Ethereum, Bitcoin, Solana, and beyond. You'll focus on performance optimization and reliability across the blockchain platform stack, diagnosing root causes instead of over-provisioning, reducing infrastructure spend, and building the observability and health-check systems that keep our nodes reliable at scale. What you'll do: Own deep performance analysis and optimization across the blockchain node stack, diagnosing root causes of resource consumption and implementing system-level fixes that reduce infrastructure spend while maintaining reliability SLOs. Build observability, health-check, and automated failover capabilities for blockchain nodes, ensuring the team can detect degradation and respond proactively rather than reactively scaling resources. Drive a right-sizing initiative across the node fleet, profiling workloads, identifying over-provisioned instances, and establishing capacity models that balance cost efficiency with headroom for reliability. Partner with protocol integration engineers, wallet teams, and indexer teams to extend performance improvements across the full blockchain platform stack, not just the node la

REMOTEreactawsai
View job →
C
Coder
📍 United States• Full-time• Remote
1mo ago

As a Senior Backend Engineer on Coder's Enterprise Experience team, you'll build the systems that help large organizations run Coder in production with confidence. You'll improve how Coder scales, how it's upgraded, and how reliably it performs in regulated, air-gapped, and enterprise environments. You'll work on a genuinely cross-functional team of backend, platform, and QA engineers who own the end-to-end experience for Coder operators. From designing new features to evolving Coder's architecture, you'll partner across engineering and product to solve complex problems and ship software that operators trust. What you'll do here Design and build new features end to end, from technical design through production rollout. Design and implement backend architecture changes that support Coder's long-term scalability goals. Investigate and resolve scalability bottlenecks under production-like load, from database access patterns to concurrency handling in coderd. Improve database migration safety and upgrade reliability through schema compatibility, background migrations, and safe rollback strategies. Own the backend side of issues surfaced by Coder operators and administrators. Document the design, implementation, and operational tradeoffs of the systems you build. Participate in code reviews, RFC-style design discussions, and on-call rotations for the services you own. What we're looking for 5+ years of professional software engineering experience, including significant production experience with Go. Deep understanding of Go's concurrency model, including goroutines, channels, the sync package, and debugging race conditions under real-world load. Experience designing and operating relational databases in production, including schema design, migrations, and transactions. Strong verbal and written communication skills. Exceptional debugging and troubleshooting skills, with the persistence to drive complex problems to resolution. A self-motivated, analytical engineer who enj

REMOTEawsgcpkubernetes
View job →
A
Amplitude
📍 Remote• Full-time• $198K – $299K/yr
1mo ago

The Developer Experience (DX) team at Amplitude builds and maintains the foundations that power how developers integrate, extend, and trust Amplitude across platforms. Our mission is to make Amplitude’s SDKs reliable, easy to adopt, and a joy to build on, so customers can confidently instrument their products and unlock insights at scale. We’re looking for a Staff Software Engineer, iOS to play a key technical leadership role on our DevEx team. In this role, you will lead the design and development of Amplitude’s core iOS SDKs, including Analytics and Session Replay , and serve as the iOS platform expert that other SDK teams, such as Statsig, Guides, and Surveys rely on. As a Staff Engineer, you’ll operate with a wide scope and high impact: setting technical direction for the iOS platform, driving cross-SDK architecture, improving performance and reliability, and raising the bar for developer experience across Amplitude’s mobile ecosystem. As a Staff Software Engineer, you will: Lead the technical direction, architecture, and long-term evolution of Amplitude’s iOS SDK platform. Own and drive development of core iOS SDKs, including Analytics and Session Replay, with a strong focus on performance, reliability, and ease of use. Act as the iOS platform expert and trusted partner for other SDK teams (Experiment, Guides, Surveys), enabling them to build on shared foundations safely and efficiently. Design and evolve shared infrastructure, APIs, and abstractions that scale across multiple iOS SDKs. Collaborate closely with Product, and Customer Support to ensure SDKs meet real customer needs. Lead cross-team technical discussions, reviews, and architectural decisions that span multiple SDKs. Improve developer experience through better APIs, documentation, tooling, testing strategies, and sample apps. Mentor senior and mid-level engineers, raising the overall quality and effectiveness of the team. You'll be a great addition to the team if you have: A strong focus on develop

aiswiftrust
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Top cities for Reliability Engineer

City links are canonicalized and require at least 20 current jobs.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.