Jobiba hiring network

Reliability Engineer Jobs

2,028 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Do you get excited about complex distributed systems? Do you love data? We’re looking for an experienced technology leader to lead and grow our global engineering team and help our customers build better software by building the next generation of Application Performance Monitoring (APM) by leading our Agent teams. New Relic is committed to giving our customers valuable insights into their systems, and New Relic’s APM agents are used by tens of thousands of companies to evaluate and improve the performance of their most important business applications. Opportunity to work from a remote office may be available depending on the applicant's location. What you'll do Hands on with data analysis and technical problem solving Leverage open source tools like OpenTelemetry to acquire data Help design and build our APM agents that run in our customer environments and give engineers deep insight into application performance, and business line owners actionable intelligence on their business performance Build processes that ensure reliability, scalability, team growth and execution Work across teams to build engineering plans and execute, getting things done in a distributed, large-scale, and fast-growing environment This role requires 4+ years of experience leading people and teams 10+ years of experience in Software Engineering Proven track record of leading and scaling strong technical teams Experience handling high performing self-directed remote teams Capable of div

javaairecruitment
View job →
P
Pinterest
📍 CA, United States• Full-time• Remote• From $1.7M/yr
1mo ago

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . The Production Engineering organization at Pinterest is accountable for ensuring overall Pinterest availability as well as enhancing Engineering teams' capability to design, build and operate robust systems at scale. Pinterest's applications and infrastructure handle billions of monthly page views and petabytes of data as Pinterest continues to grow and scale. As a Senior Production Engineer on Solutions Engineering, you will design and build AI agents, platforms, tools, frameworks and methodologies to assure the reliability of our large-scale distributed systems serving hundreds of millions of monthly active users, handling hundreds of thousands of requests per second, and managing tens of petabytes of data. You'll lead infrastructure modernization initiatives, build intelligent automation that eliminates operational toil and amplifies engineer

REMOTEpythonsqlmysql
View job →
O
1mo ago

About the Team The Finance & Supply Chain Engineering organization includes two complementary teams. Software Engineering builds internal full-stack applications, durable agentic workflows, plugins, MCPs, and measurable AI-enabled engineering practices. Data Engineering builds trusted analytics data assets for Finance and Supply Chain. The teams have distinct charters, with important shared dependencies and broad cross team partnerships across Engineering, Applications, Finance, and Supply Chain. About the Role We are looking for a hands-on senior technical leader who will report alongside the Software Engineering and Data Engineering managers. This is an individual-contributor role with no immediate people-management responsibility. The Tech Lead will raise the technical bar across both teams, participate in important cross-team or high-risk design decisions, and directly own and ship high-impact work. The role should improve team judgment and autonomy rather than act as a floating architect or universal approval gate. In this role, you will: Partner with the Software Engineering and Data Engineering managers as a peer technical leader; managers retain accountability for people, staffing, priorities, performance, and delivery commitments. Directly own the architecture, implementation, launch, and operation of one or more high-impact initiatives, remaining accountable for real outcomes rather than advisory output alone. Guide important design decisions that are cross-team, difficult to reverse, or material to security, financial controls, reliability, data quality, or long-term cost of ownership. Establish pragmatic engineering standards across architecture, APIs and data contracts, testing, security, reliability, observability, lineage, data quality, and operational ownership. Advance engineering standards for building with AI, including agentic workflows, evaluation, telemetry, adoption, and outcome measurement. Work across backend, full-stack, data, and big-d

awsrestai
View job →

NVIDIA has been redefining computer graphics, desktop gaming, and enhanced computing capabilities for more than 25 years. Today, we are tapping into the unlimited potential of AI to define the next era of computing. As a NVIDIAN, you will work on problems that sit at the boundary of architecture, silicon, firmware, software, and production, where strong judgment matters as much as technical depth. We're the Silicon Design for Productization (DFP) Team, within the broader Silicon Co-Design Group, and we turn power and thermal design into executable productization methodology. Power and thermal are among the most complicated problems we work on at NVIDIA because they sit at the intersection of architecture, workload behavior, silicon variation, firmware policy, platform constraints, and product goals. Small decisions here have an outsized impact on performance, efficiency, reliability, bring-up speed, and ultimately what the product can deliver in the field. We define how features move from concepts to bring-up, characterization, validation, and release. In this role, you will help us build that bridge. We're looking for an engineer who reasons from first principles, flourishes with ownership in a fast-paced environment, and uses AI with sound judgment. What you’ll be doing: Lead the effort across multi-functional teams to keep the program’s power and thermal productization strategy clear, executable, and on track. Create methodology and silicon test plan based controller designs and architecture, including characterization process, debug tools, fuse/firmware settings and lab requirements. Drive resolution for challenging silicon issues through structured hypotheses, measurement plans, and root-cause closure. Steward the Power and Thermal playbook when the existing productization methodology

C
Coinbase
📍 - USA• Full-time• Remote• From $243.9K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . The Core Infrastructure team within Coinbase's Platform product group builds the foundational systems that keep Coinbase online, secure, and scalable, owning the compute and networking platforms that power every product and service across the company. As the Group Product Manager for Core Infrastructure & Reliability, you'll own the product vision and multi-year strategy for Coinbase's cloud infrastructure, driving the design, operation, and scaling of the systems that underpin hundreds of billions of dollars in annual transaction volume. You'll partner deeply with Engineering, SRE, Security, and Finance to ensure Coinbase's infrastructure is reliable, cost-efficient, and resilient across multiple cloud environments and regions. What you’ll do: Own the product strategy and roadmap for Core Infrastructure, spanning compute, networking, multi-region and multi-cloud architecture, and platform reliability. Strengthen infrastructure reliability and resilience programs, defining platform-level SLOs, capacity planning, failover capabilities, and incident reduction targets to meet the uptime demands of a global financial platform. Lead evaluation and adoption of cloud infrastructure technologies (Kubernetes, service mesh, distributed storage, observability, infrastructure-as-code), making build-vs-buy decisions that balance cost, speed, and long-term scalability. A

REMOTEawskubernetesai
View job →

Job Requisition ID # 26WD100930 L'affichage de poste en français suivra / The French job posting follows. 26WD100930, Customer Advocate, Design & Manufacturing, Customer Reliability Position overview Autodesk Customer Technical Success is looking for an experienced Design & Manufacturing customer advocate to identify the systemic customer problems that matter most and turn those signals into clear decisions and measurable outcomes. The Customer Advocate brings together customer advocacy and product intelligence skills into one role. You will bring Design & Manufacturing product and workflow knowledge together with customer conversations, support trends, escalations, product data, sentiment, and business context to understand what is really happening, determine where Autodesk should invest attention, and build a clear case for action. Your value comes from applying technical expertise and judgment to customer signals, developing a point of view, validating it with customers and internal partners, and driving systemic issues toward a clear solution path. You will work closely with Product Management, Engineering, Technical Support, Technical Account Management, Support Readiness, Customer Success, and other teams across Autodesk. Responsibilities Identify, validate, and prioritize systemic customer pain using Support dat

E
18 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As a key contributor within the Drive Qualification Center of Excellence, you will ensure the performance and reliability of Everpure-developed SSDs for hyperscale and enterprise environments. You will partner with firmware and hardware teams to validate mission-critical storage components, transforming complex technical requirements into robust validation frameworks. Your mission is to guarantee that our storage solutions exceed global customer expectations through rigorous system-level testing and data-driven analysis. WHAT YOU’LL DO Design and deliver automated validation suites that stress-test PCIe, NVMe, and OCP compliance, ensuring firmware maturity and hardware robustness for the Everpure Platform. Drive technical root-cause analysis for complex failures across firmware and system layers, utilizing telemetry and logs to resolve performance or data integrity bottlenecks. Own and scale the regression infrastructure , improving test repeatability and coverage to accelerate the qualification cycle for next-generation NAND technologies. Collaborate with cross-functional engineering teams to provide clear risk assessments and quality metrics, directly influencing product readiness and release timelines. Develop custom validation tools and scripts in Python to automate the characterization of drive-level behavior under power-loss, snapshot, and high-volume scenarios. WHAT YOU BRING Deep Technical Expertise: Profi

pythonawsrest
View job →
E
18 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As a key contributor within the Drive Qualification Center of Excellence, you will ensure the performance and reliability of Everpure-developed SSDs for hyperscale and enterprise environments. You will partner with firmware and hardware teams to validate mission-critical storage components, transforming complex technical requirements into robust validation frameworks. Your mission is to guarantee that our storage solutions exceed global customer expectations through rigorous system-level testing and data-driven analysis. WHAT YOU’LL DO Design and deliver automated validation suites that stress-test PCIe, NVMe, and OCP compliance, ensuring firmware maturity and hardware robustness for the Everpure Platform. Drive technical root-cause analysis for complex failures across firmware and system layers, utilizing telemetry and logs to resolve performance or data integrity bottlenecks. Own and scale the regression infrastructure , improving test repeatability and coverage to accelerate the qualification cycle for next-generation NAND technologies. Collaborate with cross-functional engineering teams to provide clear risk assessments and quality metrics, directly influencing product readiness and release timelines. Develop custom validation tools and scripts in Python to automate the characterization of drive-level behavior under power-loss, snapshot, and high-volume scenarios. WHAT YOU BRING Deep Technical Expertise: Profi

pythonawsrest
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
19 days ago

Opportunity Overview: We’re looking for a Manager, Platform Engineering that can lead and grow a high-performing engineering team focused on Developer Experience, DevOps, SRE, and Quality. You will own the systems and processes that enable teams to build, test, release, and operate software with high velocity and reliability, driving engineering efficiency and operational excellence across the organization. What you’ll do: Lead a fast-paced, autonomous team of engineers focused on platform engineering, developer experience, DevOps, SRE, and quality engineering Own and drive the internal developer platform strategy and roadmap, improving how engineering teams build, test, deploy, and operate services Create transparency into engineering efficiency and system health through meaningful metrics across delivery, reliability, and quality Enable teams to move faster by improving CI CD pipelines, environments, tooling, and overall developer workflows Provide technical leadership across platform, infrastructure, and reliability, helping teams build scalable and resilient systems Ensure strong engineering practices across release processes, testing, quality, reliability, and security Define and enforce release guardrails, validation standards, and rollback mechanisms to improve production safety Improve environment stability and consistency across development, QA, and pre production environments Drive test strategy and automation maturity to improve overall product quality and confidence in releases Define and implement observability, monitoring, and alerting standards across systems Improve incident detection, response, and RCA practices, ensuring learnings translate into platform and system improvements Drive cloud infrastructure best practices across AWS, containers, and infrastructure as code Foster a culture of ownership, reliability, and continuous improvement within the team Provide innovative solutions for attracting, developing, and retaining top engineering talent I

CH
19 days ago

Opportunity Overview: We’re seeking a strategic and execution-focused Technical Program Manager to drive large, cross-functional initiatives from concept through launch. In this role, you will lead end-to-end delivery of complex, multi-team programs—such as platform migrations, architectural modernization, reliability improvements, and major product launches—while defining clear milestones, success metrics, and sequencing plans. You’ll partner deeply with Engineering to proactively manage dependencies, mitigate risk, and ensure technical and business alignment. This role also strengthens our execution systems by improving program transparency, standardizing launch rigor, and translating technical complexity into clear executive-level insights. What you’ll do: Lead end-to-end delivery of multi-team initiatives (e.g., platform migrations, architectural modernization, reliability improvements, major product launches) Define milestones, critical paths, and measurable success criteria Break ambiguous initiatives into executable phases with clear sequencing Identify and manage cross-team technical dependencies Proactively surface risks and tradeoffs with mitigation plans Run structured risk reviews and escalation processes when needed Participate in technical design discussions to understand architecture and constraints Ensure design reviews, capacity planning, and rollout strategies happen early Collaborate with EMs and engineers leaders to align on realistic timelines Establish program tracking mechanisms that create transparency without bureaucracy Standardize launch readiness, milestone reviews, and postmortem follow-ups Identify systemic delivery bottlenecks and drive continuous improvement Provide concise executive updates with clear status, risks, and asks Maintain decision logs and documentation for key initiatives Translate technical complexity into business impact What you’ll need: Must-haves 6+ years of experience in technical program management, engineering pr

awsazuregcp
View job →
E
ElevenLabs
📍 New York• Full-time• Remote
13 days ago

About ElevenLabs ElevenLabs is an AI research and product company transforming how we interact with technology. We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including Andreessen Horowitz, ICONIQ Growth and Sequoia. We've raised $781M in funding and our last valuation was $11B - multiples of 11, always. We have expanded from voice into three main platforms: ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale. ElevenCreative empowers creators and marketers to generate and edit speech, music, image, and video across 70+ languages. ElevenAPI gives developers access to our leading AI audio foundational models. Everything we do is the result of the creativity and commitment of our team - builders doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to work hard and create lasting positive impact, we want to hear from you. How we work High-velocity: Rapid experimentation, lean autonomous teams, and minimal bureaucracy. Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you. AI first: We use AI to move faster with higher-quality results. We do this across the whole company—from engineering to growth to operations. Excellence everywhere: Everything we do should match the quality of our AI models. Global team: We prioritize your talent, not your location. What we offer Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible. Growth paths: Joining ElevenLabs means joining a

REMOTEpythonai
View job →
P
Particle41
📍 India• Remote
13 days ago

UI/UX Designer Projects often need the technical expertise of more than one person. Particle41 provides expert teams that embed directly into businesses for immediate impact and execution towards developing outstanding software. Our clients range from startups to small and medium-sized enterprises. As a UI/UX designer, you’ll work on a handful of digital products across the web and mobile. You’ll communicate with clients to help them understand their customers’ needs, and make design recommendations based on your findings. In This Role, You Will: Apply human-centered design principles to software creation. Develop concepts and prototypes for stakeholders. Iterate quickly on your designs in a lean, agile environment. Articulate design decisions and diplomatically navigate feedback. Conduct user research and analyze data to inform design decisions. Ensure brand and product cohesion by leveraging design systems. Requirements Gathering and Analysis Collaborate with designers, product managers, and other stakeholders to gather requirements and translate them into technical solutions. Participate in requirement analysis sessions to understand business needs and user requirements. Provide technical insights and recommendations during the requirements-gathering process. Agile Development Participate in Agile development processes, including sprint planning, daily stand-ups, and sprint reviews. Work closely with Agile teams to deliver software solutions on time and within scope. Adapt to changing priorities and requirements in a fast-paced Agile environment. Testing and Debugging Conduct thorough testing and debugging to ensure the reliability, security, and performance of applications. Write unit tests and validate the functionality of developed features and individual elements. Writing integration tests to ensure different elements within a given application function as intended and meet desired requirements. Identify and resolve software defects, c

REMOTEawsazure
View job →
A
14 days ago

AI Teammates are Asana’s flagship AI innovation - autonomous agents embedded directly into enterprise workflows that reason, act, and proactively drive work forward. The AI Teammates Platform (AITP) team owns the execution engine and platform layer that makes these AI Teammates effective, reliable, and scalable. We build the core systems behind agentic execution: model evaluation and rollouts, proactive detection of blocked work, execution quality infrastructure, tool orchestration, and the platform APIs that power AI across all of Asana engineering. Rather than building a monolithic feature, we build a composable foundation. If a Teammate reasons, acts, or autonomously improves - that’s us. We’re looking for an experienced Technical Lead to drive the technical strategy, architecture, and execution for the AI Teammates Platform. You will bridge the gap between frontier AI capability and enterprise-grade software engineering, operating in close partnership with frontier model providers while building systems that scale to millions of users. If you are passionate about applied AI, agentic orchestration, high-reliability backend systems, and growing high-performing engineering teams, we’d love to hear from you! This role is based in our San Francisco office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What You’ll Achieve Set the technical strategy for the AI Teammates execution engine, tool orchestration, and platform APIs within Asana, advocating for engineering-driven investments with a vision for keeping our systems flexible, reliable, and maintainable to meet customer needs now and in the future Drive the team to continually and holistically

aiproject management
View job →
E
18 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Portworx team to build and deliver our highest-quality product suite. In this role, you will write clean, scalable code with a strong focus on quality, reliability, and user-centric design. You will directly contribute to building a new SaaS platform that delivers a secure, consistent, and best-in-class experience for customers purchasing and managing Portworx offerings. As a core developer, you will take ownership of designing and implementing critical features across the entire Portworx portfolio. WHAT YOU’LL DO Design & Scale SaaS Microservices: Develop, test, and integrate high-performance microservices and features into the Portworx product suite, ensuring high availability in distributed systems. Drive End-to-End Delivery: Lead software lifecycle activities including architectural design, code reviews, unit/functional testing, documentation, and continuous integration and deployment (CI/CD). Partner Across Teams: Collaborate with product managers, cross-functional engineering peers, and early-adopter customers to transform requirements into production-ready software. Own Product Quality & Iteration: Take full ownership of feature stability by proactively incorporating customer feedback and rapidly resolving issues identified during testing and deployment. Innovate & Experiment: Research emerging technologies and cloud infrastructure tools to push performance boundaries and continuously i

javaawskubernetes
View job →
E
18 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Portworx team to build and deliver our highest-quality product suite. In this role, you will write clean, scalable code with a strong focus on quality, reliability, and user-centric design. You will directly contribute to building a new SaaS platform that delivers a secure, consistent, and best-in-class experience for customers purchasing and managing Portworx offerings. As a core developer, you will take ownership of designing and implementing critical features across the entire Portworx portfolio. WHAT YOU’LL DO Design & Scale SaaS Microservices: Develop, test, and integrate high-performance microservices and features into the Portworx product suite, ensuring high availability in distributed systems. Drive End-to-End Delivery: Lead software lifecycle activities including architectural design, code reviews, unit/functional testing, documentation, and continuous integration and deployment (CI/CD). Partner Across Teams: Collaborate with product managers, cross-functional engineering peers, and early-adopter customers to transform requirements into production-ready software. Own Product Quality & Iteration: Take full ownership of feature stability by proactively incorporating customer feedback and rapidly resolving issues identified during testing and deployment. Innovate & Experiment: Research emerging technologies and cloud infrastructure tools to push performance boundaries and continuously i

javaawskubernetes
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Top cities for Reliability Engineer

City links are canonicalized and require at least 20 current jobs.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.