About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Quality leadership within Manufacturing Operations, the Senior Reliability Scientist is responsible for leading reliability activities across complex, high-performance systems. Working closely with established reliability experts and cross-functional teams, this role uses experimental data and advanced modelling to inform design decisions, validate product reliability and optimise serviceability strategies, including spares provisioning. The Team The Quality team within Manufacturing Operations is responsible for ensuring product robustness, reliability and lifecycle performance across Graphcore’s hardware portfolio. The team includes experienced reliability specialists and works closely with technology research, chip, board, system design, platform and operations teams to translate reliability insights into actionable improvements across the product lifecycle. Responsibilities and Duties: · Define and refine reliability requirements across silicon, board and system levels, working in partnership with research and design teams · Apply ad
Jobiba hiring network
Reliability Engineer Jobs
2,049 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Who are we? FalconX is a pioneering team of operators, investors, and builders committed to revolutionizing institutional access to the crypto markets. Operating at the intersection of traditional finance and cutting-edge technology, FalconX addresses the industry's foremost challenges: Navigating the digital asset market can be complex and fragmented, with limited products and services that support trading strategies, structures, and liquidity found in conventional financial markets. As a comprehensive solution for all digital asset strategies from start to scale, FalconX operates as the connective tissue empowering clients with seamless navigation through the ever- evolving cryptocurrency landscape. Responsibilities Be part of a trading systems engineering team, dedicated to building out the core trading platforms. Work closely with cross functional teams to improve the system reliability, scalability and security. Engage in and improve the quality supporting the platform. Build and manage systems, infrastructure and applications through automation. Provide operational support to internal teams working on the platform. Work on improvements to bring in high efficiency, reduce latency, deploy systems faster. Practice sustainable incident response and blameless postmortems. Together with your engineering team, you will share an on-call rotation and be an escalation contact for service incidents. Implement and maintain rigorous security best practices across all infrastructure, with a focus on minimizing attack surface and ensuring data integrity. Monitor system health and performance with a keen eye for identifying and resolving issues before they affect trading activity. Manage user queries and service requests (often requiring in depth analysis of the technical and/or business logic of our systems). Proactive approach to problem analysis and resolution of production incidents. Manage Issue tracking and prioritisation of day to day production incidents. Manage platf
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About the Team The Professional Services R&D team is a new, dynamic group at the forefront of innovation within Okta. Our mission is to design and build reusable, scalable assets and tools that empower our delivery teams and partners. By making customer engagements more efficient, streamlined, and cost-effective, we directly contribute to our customers' success and accelerate their time-to-value with Okta. This is a unique opportunity to join a strategic team from the ground up and shape the future of Okta's professional services. Position Summary As the DevOps Engineer for the R&D team, you will build and own the infrastructure and automation that enables us to develop and release software with speed and confidence. You will be responsible for creating and managing our CI/CD pipelines, defining our infrastructure as code, and ensuring our deployed assets are scalable, secure, and observable. You will be a key enabler of the team's agility, implementing the tools and processes that allow us to innovate and iterate quickly while maintaining a high bar for quality and reliability. Responsibilities Design, build, and maintain the team's CI/CD pipelines to automate the build, test, and deployment of our software assets. Manage and provision cloud infrastructure using Infrastructure as Code (IaC) principles and tools (e.g., Terraform, CloudFormation). Implement and manage monitoring, logging, and alerting solutions to ensure the health and performance of
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lake using open formats like Apache Iceberg, delivering deep correlation and long-term analytics at dramatically lower cost. A dynamic Knowledge Graph and chat-based AI SRE provide rich context and guided workflows so teams can move from detection to root cause and resolution significantly faster. The Infrastructure team at Observe by Snowflake is responsible for building, scaling, and operating the development and production environments that power our observability platform. We are a small, highly collaborative team with a broad scope, focused on delivering reliable infrastructure while continuously improving the systems that support our engineers and customers. What You’ll Do Design, build, and operate scalable cloud infrastructure in AWS supporting a high-scale observability platform. Improve system reliability, performance, and operational visibility across development and production environments. Develop and maintain CI/CD pipelines and internal tooling to improve developer productivity and deployment safety. Identify and mitigate security risks, and help maintain intern
Firmware Engineer BIOS/UEFI Description - The Firmware Engineer is responsible for ensuring that embedded firmware meets the highest standards of quality, robustness, and long‑term reliability. This role works closely with firmware developers, hardware engineers, architecture teams, QA, and cross‑functional partners to define quality metrics, validate system behavior, improve defect detection, and build processes that prevent regressions. Responsibilities Develop, execute, and maintain comprehensive firmware test plans, including functional, regression, stress, corner-case, and long-duration reliability tests. • • • Evaluates the firmware architecture, design and development plus testing methodologies to create strategy and implement methods to bring increased quality and reliability to the firmware. Reviews firmware code and test plans including functional, regression, stress, corner-case and long-duration reliability tests. Perform root-cause analysis for firmware defects, drive containment, corrective actions, and verification of fixes. Create and manage quality dashboards, KPI, failure rates, and reliability metrics. Evaluate failure modes exposed during design, development, manufacturing and field use; propose design or process improvements to prevent recurrence. Works with product architects, firmware teams, product managers and provides critical guidance, system-level debugging and troubleshooting to various teams, as a Subject Matter Expert. Reviews and evaluates designs and project activities for compliance with systems design and development guidelines and standards; provides tangible feedback to improve product quality and mitigate failure risk. Create and manage quality dashboards , KPIs, failure rates, and reliability metrics. Validate firmware integration with hardware, BIOS/UEFI, EC, micro
About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability Team Focus Areas Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely. Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads. Databases: Building high performance, distributed database systems that power all of OpenAI's product stack. In this
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking a strategic and results-driven engineering leader who is passionate about cloud agnostic infrastructure, operational excellence, and enabling engineering teams to operate autonomously and build with confidence. As Head of Infrastructure, you'll lead a talented and geographically distributed team of engineers across the SF Bay Area, India, and Europe, fostering a culture of collaboration, ownership, and continuous improvement. You'll own the infrastructure that underpins one of the world's most widely used API platforms, an environment handling ~80,000 requests per second at the front door, and be responsible for its reliability, scalability, and evolution. In addition to infrastructure, you'll own the Site Reliability Engineering (SRE) function at Postman, setting the standards and practices that keep the platform reliable at scale. You'll work closely with engineering managers, product managers, and platform teams to drive the technical roadmap for our cloud agnostic infrastructure and reliability practices, ensuring we can support a large and rapidly growing engineering organization. If you're p
About the Role: We are looking for a Senior DevOps Engineer to join our DevOps team at K Health. You will own and evolve the infrastructure underpinning a healthcare AI platform serving patients and enterprise health system partners. This is a high-ownership role: you will architect and operate cloud environments across K Health and its enterprise partners, lead complex infrastructure migrations, drive disaster recovery programs, and help build the next generation of AI-powered operations tooling. You will also mentor junior engineers and collaborate closely with product and engineering teams across the company. This is a hybrid role based in New York City (4 days/week in office) and includes participation in a daytime on-call rotation. What you will do: Own the design, implementation, and evolution of our GKE-based Kubernetes infrastructure across K Health and enterprise partner environments. Build and maintain our Terraform modular infrastructure library, including reusable modules with automated testing, across GCP, Cloudflare, and AWS. Architect, build, and maintain GitLab CI/CD shared pipeline templates used by all engineering teams (build, test, security scanning, deployment). Own and maintain self-hosted infrastructure software running in-cluster, including GitLab, ArgoCD, Langfuse, DependencyTrack, NGINX Ingress, and others. Implement and support security and compliance controls across infrastructure and the software supply chain - secrets management, pipeline secret detection, container scanning, SOC2 and HIPAA. Drive disaster recovery readiness: design failover scenarios, author runbooks, and lead periodic DR tests. Lead development of AI-powered operations tooling and agentic infrastructure. Monitor, troubleshoot, and improve production system reliability; respond to incidents during on-call shifts. Mentor junior DevOps engineers and establish team-wide engineering standards. What we are looking for: 5+ years of experience in DevOps, platform engineering,
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role: Join our Infrastructure Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Infrastructure Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff Infrastructure Engineers who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. You Will: Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios. Optimize Performance and Infrastructure: Collaborate with core infrastructure and product teams to performance tune and optimize our cloud deployments (Kubernetes, Docker, GCP). Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions. Elevate Developer Experience: Design and implement improvements to our build, test, and deployment systems to make software delivery faster, safer,
About the Team This team builds and operates the systems that enable OpenAI researchers to run reliable, scalable, and efficient research workflows. The team sits close to research and works across infrastructure, systems, and automation to make sure researchers have the tools and environments they need to move quickly. The work spans software engineering, infrastructure, systems administration, cluster operations, and reliability engineering. As OpenAI’s infrastructure evolves from bespoke bare-metal systems toward more standard, scalable platforms, the team needs engineers who can understand how systems work end-to-end and build the right abstractions without reinventing the wheel. About the Role As a Software Engineer on this team, you will build and operate the infrastructure that supports frontier research and critical research-facing systems. You will work on systems that sit close to the metal, but the role is not limited to classic operations or sysadmin work. We are looking for someone who can reason about networking, bootstrapping, Kubernetes, scalability, automation, and reliability - while also writing software to make these systems better over time. This role is a strong fit for an independent, high-ownership engineer who enjoys reliability-heavy infrastructure work but still wants to build. You do not need to come in as a kernel expert or highly algorithmic optimization engineer, but you should be deeply curious about infrastructure, comfortable debugging complex systems, and excited to support researchers doing novel work. We expect you to: Build and operate reliable infrastructure for research workloads and research-facing services. Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services. Improve cluster bootstrapping, provisioning, automation, and deployment workflows. Debug issues across networking, compute, storage, orchestration, and service reliability layers.
About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. Our team builds the software, tooling, and operational systems that help manage this fleet at scale. We work across production engineering, distributed systems, capacity management, and operational automation to improve reliability, reduce manual work, and make better use of available compute. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will develop the systems that help manage the GPU fleet powering ChatGPT, including tooling for fleet health, capacity planning, operational automation, and incident response. You will work closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and compute utilization. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference. Develop systems for capacity planning, fleet health monitoring, and resource utilization. Automate operational workflows, including incident detection, diagnosis, and response. Identify and address bottlenecks affecting fleet reliability, scalability, and performance. Partner with infrastructure, research, and product engineering teams to improve the compute platform. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems. Have a background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering. Have built software that automates operational workflows and reduces manual work. Have worked with distributed infrastructure, cluster orchestration, or large-scale int
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron's Global Supplier Quality organization is seeking a Controller Quality Principal Engineer to lead the quality strategy, qualification, and continuous improvement of storage and memory controller (ASIC/SoC) manufacturers supporting Micron's SSD, embedded, and storage solutions portfolio. This is a senior technical leadership role responsible for driving controller supplier quality performance from design qualification through mass production and field support The successful candidate will collaborate across functions with ASIC Development Team, Compose Engineering, Product Engineering, Dependability, Manufacturing, and Commodity Management, as well as directly with controller IC vendors, foundries, third party reliability labs and OSAT (outsourced assembly and test) partners, to ensure controller quality, reliability, and supply continuity meet Micron's standards. This role can be based in Taiwan, Hyderabad, or San Jose and will work extensively across time zones with global partners and suppliers! Develops, evaluates, revises, and applies technical quality assurance protocols/methods to inspect and test in-process raw materials, production equipment, and finished products. Ensures activities and items are in compliance with both company quality assurance standards and applicable government regulations. Performs analysis and identifies trends in the inspection of finished products, in-process materials and bulk raw materials, and recommends corrective actions when vital. Ensures that established manufacturing inspection, sampling and statisti
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron's Global Supplier Quality organization is seeking a Controller Quality Principal Engineer to lead the quality strategy, qualification, and continuous improvement of storage and memory controller (ASIC/SoC) manufacturers supporting Micron's SSD, embedded, and storage solutions portfolio. This is a senior technical leadership role responsible for driving controller supplier quality performance from design qualification through mass production and field support The successful candidate will collaborate across functions with ASIC Development Team, Compose Engineering, Product Engineering, Dependability, Manufacturing, and Commodity Management, as well as directly with controller IC vendors, foundries, third party reliability labs and OSAT (outsourced assembly and test) partners, to ensure controller quality, reliability, and supply continuity meet Micron's standards. This role can be based in Taiwan, Hyderabad, or San Jose and will work extensively across time zones with global partners and suppliers! Develops, evaluates, revises, and applies technical quality assurance protocols/methods to inspect and test in-process raw materials, production equipment, and finished products. Ensures activities and items are in compliance with both company quality assurance standards and applicable government regulations. Performs analysis and identifies trends in the inspection of finished products, in-process materials and bulk raw materials, and recommends corrective actions when vital. Ensures that established manufacturing inspection, sampling and statisti
As a Senior Sales Engineer , you will be the primary technical resource for our Account Executive team. You will share your product and technical expertise through presentations, product demonstrations, and technical evaluations (Proof Of Values). As the technical expert, you will work with clients to understand their requirements and pain points, then design the right solution for their business needs. During the sales cycle you will guide clients through trials and POVs, demonstrating Sumo Logic’s ability to meet and exceed their requirements and building a positive relationship that will provide continuous value to our customers. Finally, you will have the opportunity to work cross-functionally with our Product Management and Engineering teams to share your knowledge and experiences to ultimately improve our business and our customers’ success. We seek talent who wants to leverage their technical and people skills to help deliver solutions to clients directly and become a trusted advisor in the process. Above all else, you should be highly self-motivated and extremely curious to learn more about Sumo Logic and the vast problems that it can solve. Responsibilities Partner with the Account Executives to understand customer challenges and mains, and articulate Sumo Logic’s value proposition, vision, and strategy to customers Technically close complex opportunities through advanced competitive knowledge, technical skill, and credibility Understand and help orchestrate all phases of the sales cycle, including leading technical validations during the Proof of Value phase Be successful working with all levels of an organization, from executives down to individual developers and Site Reliability Engineers Deliver product and technical demonstrations of the Sumo Logic service Work cross functionally with Product Management and Engineering to improve the Sumo Logic service based on your experience with customers Requirements B.S. in Computer Science, Engineerin
As a Senior Solutions Engineer, you will be the primary technical resource for our ANZ Accounts team. You will share your product and technical expertise across security and observability use cases, through presentations, product demonstrations, and technical evaluations (Proof Of Values). As the technical expert, you will work with clients to understand their requirements and pain points, then design the right solution for their business needs. During the sales cycle you will guide customers through POVs, demonstrating Sumo Logic’s ability to meet and exceed their requirements and building a positive relationship that will provide continuous value to our customers. Finally, you will have the opportunity to work cross-functionally with our Product Management and Engineering teams to share your knowledge and experiences to ultimately improve our business and our customers’ success. We seek talent who wants to leverage their technical and people skills to help deliver solutions to clients directly and become a trusted advisor in the process. Above all else, you should be highly self-motivated and extremely curious to learn more Sumo Logic and the vast problems that it can solve. Responsibilities Partner with the Account Executives / Account Managers to understand customer challenges and mains, and articulate Sumo Logic’s value proposition, vision, and strategy to customers Technically close complex opportunities through advanced competitive knowledge, technical skill, and credibility Understand and help orchestrate all phases of the sales cycle, including leading technical validations during the Proof of Value phase Be successful working with all levels of an organization, from executives down to individual developers and Site Reliability Engineers Deliver product and technical demonstrations of the Sumo Logic service Work cross functionally with Product Management and Engineering to improve the Sumo Logic service based on your experience with customers Requirem
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime