About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.
Jobiba hiring network
Reliability Engineer Jobs
2,049 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.
Job Details: Job Description: Intel's Design Quality and Reliability organization is seeking an AI Platform Engineer to architect and build an enterprise-grade AI platform for mission-critical engineering work. This platform will enable Intel engineers to analyze complex design, qualification, and reliability data; automate engineering workflows; access organizational knowledge; and make faster, evidence-based decisions throughout the product lifecycle. The successful candidate will combine strong software engineering fundamentals with expertise in AI-native and agentic development. They will be highly proficient with Agentic AI coding assistants and able to use these tools responsibly to accelerate architecture, implementation, testing, debugging, and documentation. This role requires close collaboration with Design, Quality and Reliability, Product Engineering, Manufacturing, IT, Information Security, and other Intel stakeholders. Responsibilities 1. Architect and develop Intel's reusable AI platform for Design Quality and Reliability. 2. Build AI agents and workflows for engineering data analysis, qualification planning, risk assessment, knowledge retrieval, reporting, and process automation. 3. Apply Agentic AI coding assistants to accelerate software development while maintaining rigorous engineering review and validation. 4. Integrate AI capabilities with Intel engineering databases, quality-management systems, internal APIs, spreadsheets, documentation repositories, and workflow tools. 5. Develop production-grade backend services, APIs, data pipelines, model gateways, and agent-orchestration components. 6. Establish shared platform capabilities for identity, access control, tool authorization, memory, observability, evaluation, and auditability. 7. Implement human approval, deterministic validation, and rollback controls for consequential engineering actions. 8.
Job Details: Job Description: As a System Simulation Engineer, you will play a critical role in developing and optimizing simulation models to support the design and validation of cutting-edge systems. Your day-to-day activities will include creating, refining, and analyzing system-level simulations to assess performance, reliability, and functionality. Your expertise will directly contribute to the development of innovative solutions that drive Intel's success in delivering high-quality products to market. Intel Corporation is a world leader in designing and manufacturing advanced technologies that power the future of computing. The organization you will join focuses on enabling innovative engineering solutions that support Intel's broader goals of advancing technology leadership. By fostering collaboration across teams and disciplines, this group drives impactful contributions to Intel's product portfolio and technological growth. The primary responsibilities for this role will include, but are not limited to: Develop and maintain system-level simulation models to evaluate performance, reliability, and functionality. Collaborate with cross-functional teams to gather requirements and ensure alignment with design and validation goals. Analyze simulation results and provide actionable insights to improve system design and architecture. Optimize simulation frameworks for accuracy and efficiency, ensuring alignment with project timelines. Document processes, methodologies, and findings to facilitate knowledge sharing and collaboration. Qualifications: Minimum Qualifications : Minimum qualifications are required to be initially considered for this
About the Team The Plugin Ecosystem team builds the platform and product experiences that let people extend ChatGPT and Codex. We work on plugins, skills, connectors, interactive apps, and open standards like the Model Context Protocol (MCP). We make plugins easy to discover, install, and use, ensure they’re invoked at the right time, and help people find new ways to get value from them. We want anyone to be able to turn a useful workflow into a plugin, share it, and have other people use it. A plugin can package instructions and skills with connections to the tools and data it needs. Our work spans creation and publishing, reliable execution across our products, clear permissions and approvals, and the controls admins need to bring plugins to their organizations. We work closely with research to improve plugin quality as models evolve. About the Role We’re looking for product-minded engineers to build the systems behind plugins and improve how models use them. Depending on your focus, you may scale generalist infrastructure and identity-related integrations across products, or improve plugin quality at the intersection of backend engineering and applied AI or work on the product experience itself to drive plugin usage. You’ll work across teams and own problems from diagnosis and design through implementation and release. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and ship APIs, SDKs, and services that developers use to extend ChatGPT and Codex. Build intuitive experiences that help users discover, install, and use plugins to get more done. Make plugins easier to create, test, publish, update, and share. Improve when and how models use plugins, from choosing the right plugin to completing a task. Work with Research to diagnose failures and measure improvements as models evolve. Improve plugin reliability and interaction quality acros
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: Notion’s Data Foundations team builds and operates the batch and streaming infrastructure behind our product features, analytics, search, and AI experiences. We’re looking for a hands-on technical leader to shape the next generation of this platform as Notion serves larger customers, expands globally, and supports more data-intensive products. You’ll identify the highest-leverage problems, set direction, build and develop a high-performing team, and lead multi-quarter initiatives across our data lake, streaming, distributed-compute, governance, and reliability systems. You’ll stay close to critical technical decisions while creating clear ownership, growing engineers and technical leaders, and helping the team execute as one—partnering closely with Data Engineering, Data Product, Search, AI, Infrastructure, and Security. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You
We are seeking a Staff Engineer to join our growing team to provide technical direction and implement core parts of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Staff Engineer on this new team, you will be responsible for providing technical leadership to teams developing cutting edge technologies related to enabling deployment at scale of AI applications. You will take on challenging, high-visibility projects that improve and enhance the performance, scalability, and reliability of the distributed systems infrastructure for this new product. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day. We value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We're looking to speak with candidates based in the New York City area for our hybrid or in-office working models. Position Expectations Work closely with product management, product engineering, product design peers as well as other teams within the company to define the first version and future evolution of the service Design, build and deliver well-tested core pieces of the platform in collaboration with other vested parties Contribute to shaping architecture, code reviews and development practices, developer experience as the teams and product grow Mentor fellow engineers and assume ownership and accountability of projects Qualifications Strong background in building core components for high scale compute and data distributed systems 8+ years experience of building distributed systems, and/or foundational cloud services at scale and an interest in working with Python, Go and Java Proven success in designing, writing, testing, debugging, performance tuning, possessing a strong grip on the foundational materials of computer science and maintaining distributed and/or highly concurrent software s
Position Overview We are looking for a Software Engineer II to build and deliver scalable software solutions across our products. You will work on modern web applications and cloud-based services using Node.js, React, TypeScript, AWS, PostgreSQL, MSSQL, and Docker, while contributing to AI-enabled features and integrations. You will collaborate closely with other engineers, product managers, and cross-functional teams to develop reliable, maintainable, and production-ready solutions. This role provides an opportunity to work with modern AI technologies including Python, AWS Bedrock, MCP, RAG, and agentic AI workflows while developing strong expertise in cloud-native software engineering. What You'll Do Develop and maintain scalable backend services and APIs using Node.js, TypeScript, and JavaScript. Build responsive and maintainable frontend applications using React. Design and implement integrations with AWS services and contribute to cloud-native application development. Develop and maintain applications using PostgreSQL and MSSQL, including writing efficient queries and working with database schemas. Build, test, and deploy applications using Docker and modern CI/CD practices. Contribute to AI-enabled product features using Python, AWS Bedrock, RAG, MCP, and AI integration patterns. Work with the team to integrate LLM capabilities, APIs, tools, and data sources into production applications. Write clean, maintainable, and well-tested code following established engineering practices. Participate in code reviews, technical discussions, debugging, and production issue resolution. Develop unit and integration tests and contribute to improving application quality and reliability. Monitor application performance and troubleshoot issues across development and production environments. Collaborate with senior engineers and architects to implement technical solutions aligned with product and engineering requirements. Stay current with emerging technologies, particularly in
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Job Title: Software Engineer II, Data Analytics and Engineering Intro: We’re looking for a Software Engineer II, Data Analytics and Engineering to improve the quality, reliability and velocity of data science and product development at Pinterest. You’ll build scalable data foundations, analytics tooling and analysis pipelines that enable trusted, self-service access to datasets, insights and metric investigations across cross-functional teams. What you’ll do: Develop and document practical instrumentation and experimentation standards, then partner with product engineering teams to apply them to priority product development work. Build and improve scalable analysis pipelines and tooling that produce reliable insights at scale and strengthen understanding of key data structures and metrics. Create tools and processes that enable Data Scientists a
The Lead EMS/SCADA will be responsible for the design, development, integration, testing, and deployment of Energy Management Systems (EMS) and SCADA solutions for utility-scale Battery Energy Storage System (BESS) projects. The role will drive system architecture, control strategies, monitoring solutions, and communication interfaces to ensure optimal performance, reliability, and grid compliance. Source: Adani Group | Job ID: 58677
About the Team OpenAI’s User Operations team shepherds our customers’ adoption of AI and ensures that our customers' product experience is nothing short of exceptional. We are building the very first post-AGI support team. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others, to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. Within Premium Support, Dedicated Support Engineers combine deep technical troubleshooting with an enduring understanding of our most strategic customers’ architectures, critical workloads, and business priorities. Through proactive reliability work, ownership during incidents, and AI-powered support capabilities, we help customers operate successfully as their use of OpenAI grows. About the Role We’re looking for a senior leader to build and scale our Dedicated Support Engineering function globally. You will define its strategy, build the team, and establish how we deliver technically rigorous, proactive support for customers running some of the most complex and consequential workloads on OpenAI. This role combines organizational leadership, technical judgment, and executive customer engagement. You will establish a model in which DSEs develop deep customer context, independently advance difficult investigations, anticipate operational risks, and drive issues through resolution. You will also turn what the team learns into improvements that benefit customers across OpenAI. You should bring experience building technical organizations that maintain long-term accountability for enterprise customers. Leadership in Technical Account Management, enterprise Support Engineering, or a comparable technical customer function is particularly rel
About the Role: We are looking for a talented Automation Engineer to join our Automation Engineering team in Toronto. In this role, you will be responsible for designing and implementing automated tests for Mobile development. You will collaborate closely with QA engineers and developers to build scalable test frameworks, improve automation coverage, and contribute to the efficiency of our multi-platform release process. You will also design data-driven end-to-end checks around playback and ad insertion , integrate them into CI/CD pipelines as quality gates, and operate a reliable device lab to prevent regressions from shipping. Your work will directly accelerate testing and release velocity while improving revenue-critical reliability across Tubi’s Android and IOS apps. This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: Design, implement, and maintain automated tests for mobile development (Android & iOS) Contribute to the development and optimization of cross-platform automation frameworks. Write and maintain test scripts in JavaScript/TypeScript , using frameworks such as Puppeteer, Appium, WebDriverIO, Selenium. Ensure test cases are integrated into CI/CD pipelines and provide reliable feedback on product quality. Help identify flaky tests, investigate root causes, and improve test stability. Collaborate with developers and QA engineers to clarify requirements and improve test strategies. Participate in code reviews and follow best practices for test automation . Your Background: Bachelor’s degree or above in a technical field (e.g., Computer Science, Engineering, Mathematics), or equivalent industry experience. 3+ years of hands-on experience in automation testing for mobile devices Strong programming skills in JavaScript/TypeScript (preferred), or Python/Java. Experience with automation frameworks (e.g. Puppeteer, Appium,, WebDriverIO, Selenium, Playwright, T
About the Role: Tubi is one of the largest free streaming platforms in the US, serving a large-scale streaming audience across Web, iOS, Android, Roku, Fire TV, Apple TV, and game consoles. Quality at this scale isn't a checkbox — it's a competitive advantage. We're looking for an Automation Engineering Manager to lead the team responsible for building and operating Tubi's multi-platform test automation infrastructure. You will own the strategy, tooling, and execution quality across our client surfaces — from video playback and ad delivery to content discovery and onboarding. This role is for a hands-on technical leader who can set direction, influence cross-functional roadmaps, and stay close enough to the code to guide architecture, review critical implementation decisions, and unblock complex technical issues. You will build a team that ships reliable automation at speed — and you will help the team move toward AI-native automation practices: fluent in AI tooling, proactive about applying it, and disciplined about using it responsibly. This is a hybrid role based out of either our San Francisco or Toronto office. You must be willing to travel to either location at least 2 days a week. What You'll Do: Test Strategy & Quality Planning Define and own Tubi's multi-platform automation strategy — covering Web, iOS, Android, CTV (Roku, Fire TV, Apple TV, Smart TVs, game consoles), and API layers. Establish testing standards, coverage targets, and quality gate policies across the CI/CD pipeline to protect release confidence and production reliability. Design specialized test strategies for business-critical scenarios: video playback (HLS/DASH), ad insertion, content recommendation surfaces, and user authentication flows. Use AI-assisted analysis (e.g., failure pattern clustering, test gap detection) to continuously improve test strategy based on real production signal and defect trends — not gut instinct. Automation Framework & Infrastructure Lead the
About Wolt At Wolt, we create technology that brings joy, simplicity and earnings to the neighborhoods of the world. In 2014 we started with delivery of restaurant food. Now we’re building the delivery of (almost) everything and you’ll find us in over 500 cities in 30 countries around the world. In 2022 we joined forces with DoorDash and together we keep on dreaming big and expanding across the globe. Working at Wolt isn’t always easy, but it’s definitely exciting. Here you’ll learn more, build more, and ship more than in most other companies. You’ll be challenged a lot, but also have a lot of fun on the way. So, if you’re a self-starter with drive and entrepreneurial spirit, this could be the ride of your life. What you’ll do: Build and maintain high-throughput backend services using Go . Collaborate with product managers, designers, and frontend developers to ship features that support internal support agents across the globe. Design systems that are scalable , resilient , and easy to maintain. Lead and contribute to architectural discussions and technical decision-making. Write well-tested code and help the team maintain high code quality standards. Our humble expectations: 7+ years of professional software engineering experience, with a proven track record of building and scaling complex systems. 2+ years of production experience in Golang , with the ability to mentor others and drive best practices across the team. Strong hands-on experience with both SQL and NoSQL databases — especially Cassandra. Solid understanding of designing and operating low-latency, high-throughput distributed systems . Nice to have Background in Node.js or other backend languages. Familiarity with cloud infrastructure (AWS, GCP) and event-driven architectures. Previous on-call experience , with a pragmatic approach to reliability and incident management. What we value A product-oriented mindset — you think beyond the ticket, understand the “why” behind the work, and aim to create real
Who are we? FalconX is a pioneering team of operators, investors, and builders committed to revolutionizing institutional access to the crypto markets. Operating at the intersection of traditional finance and cutting-edge technology, FalconX addresses the industry's foremost challenges: Navigating the digital asset market can be complex and fragmented, with limited products and services that support trading strategies, structures, and liquidity found in conventional financial markets. As a comprehensive solution for all digital asset strategies from start to scale, FalconX operates as the connective tissue empowering clients with seamless navigation through the ever- evolving cryptocurrency landscape. Role Overview As an Engineering Manager for the Credit team, you will lead the core engine driving our credit systems, including the margin notification engine and loan booking system. Your primary focus will be on maximizing data accuracy, system reliability, and architectural integrity. You will lead a high-performing group of senior engineers, stream-lining technical processes, preventing over-engineering, and maintaining a high standard of delivery alongside key stakeholders. Role Split & Focus Areas Technical Leadership & Architecture (40%): Drive long-term system maintenance, architectural refactoring, and code quality. Ensure technical debt is addressed without blocking impactful business features. Process & Stakeholder Management (20–30%): Partner with product and business stakeholders to prioritize deliverables. Make decisive trade-offs and cut unnecessary overhead to maintain lean execution. People Management & Mentorship (20–30%): Manage, mentor, and guide senior technical talent across performance cycles, career growth, and talent acquisition. Key Responsibilities Oversee and maintain the reliability, precision, and efficiency of core credit applications (Margin Notification Engine, Loan Booking Systems). Drive architectural roadmap decision
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime