Jobiba hiring network

Senior Software Reliability Engineer Jobs

7,292 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

A
Airbnb
📍 United States• Full-time• From $248K/yr
1mo ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Media Foundation team builds the core platforms and infrastructure that power photo and video capture, upload, processing, storage, and delivery across Airbnb's products, at a massive global scale. We partner closely with Product, Design, Data Science, Trust, and Infrastructure teams to ensure every image and video guests and Hosts see is fast, high-quality, and reliable, from the moment a Host uploads a photo to the moment a Guest views it in search results. The Difference You Will Make: As the Senior Engineering Manager for Media Foundation, you will lead a team of engineers to build and operate Airbnb's media infrastructure, the foundation powering every photo, video, and document that users interact with on the platform. Capabilities include uploads, AI/ML transformations, distribution, and presentation within the product. You will shape the team’s vision and help steer the organization towards our larger goals. This includes day-to-day operations, technical deep dives, community engagement with internal engineering teams, and strategic planning. As the technical and organizational leader for Media Foundation, you own the reliability, scalability, and evolution of the systems that power every media interaction on the Airbnb platform from the moment a host uploads a photo to the instant a guest streams a listing video. Your experience in both media technology and people management will be essential to the maintenance, modernization, and innovation of Airbnb's media platform. You set the bar for engineering excellence, define what great looks like for your team, and crea

aigorust
View job →
J
16 days ago

Backend Engineer (Senior Level) - SDE IV We're looking for a Senior Backend Engineer to lead the architecture and evolution of backend services that deploy and serve machine learning models in production. You'll work closely with ML Engineers, Platform, and Product teams to build scalable, reliable systems and drive technical direction across multiple teams. What You’ll Do Design and drive the long-term architecture of backend services for biometrics and ML model serving. Collaborate with core platform and backend teams on organization-wide architectural initiatives. Partner with business and engineering teams to design and deliver cross-cutting platform capabilities. Lead architectural reviews, mentor engineers, and promote engineering best practices. Build and maintain backend services for deploying and serving ML models Monitor service reliability, performance, and scalability in production Deploy and operate services on AWS using ECS + Fargate, SageMaker, or EC2 + Kubernetes Support real-time and batch inference workflows Contribute to CI/CD pipelines and deployment automation What We’re Looking For Strong expertise in backend development using Java and working knowledge of Python. Experience mentoring engineers and driving architectural decisions. Working knowledge of Python, especially for ML-related workflows Hands-on experience with AWS (e.g., DynamoDB, ECS, EC2, Redis, S3, SageMaker) Familiarity with Terraform or other infrastructure-as-code tools, and experience with CI/CD and production monitoring Experience with observability tools (Datadog, New Relic, etc.) Experience with containers and orchestration (Docker, ECS, etc.) Understanding of how ML models are deployed and served in production Experience with Kubernetes Nice to Have Experience with MLOps or ML platform engineering. Experience with asynchronous programming and event-driven systems. Jumio Values: IDEAL: Integrity, Diversity, Empowerment, Accountability, Leading Innovation Equal Opportunities :

REMOTEpythonjavaredis
View job →
S
StarRez
📍 Australia• Full-time
16 days ago

About StarRez StarRez is the global leader in student housing software, providing innovative solutions for on and off-campus housing management, resident wellness and experience, and revenue generation. Trusted by 1,400+ clients across 25+ countries, StarRez supports more than 4 million beds annually with its user-friendly, all-in-one platform, delivering seamless experiences for students and administrators. With offices in the United States, Australia, the UK, and India, StarRez blends the robust capabilities of a global organization with the personalized care and service of a trusted partner. The Role: We are looking for a Senior Quality Engineer - AI to help shape how StarRez evaluates, validates, and improves AI-powered product experiences. You will play a pivotal role in elevating our quality practices for AI-powered product experiences and tooling. You will be a product expert within our engineering organization, deeply understanding workflows and customer outcomes. This role will balance hands-on individual contributor responsibilities with leading, influencing, and coaching your peers. You’re someone who is passionate about seeing systems holistically and is eager to drive the highest standards of accuracy, safety, usefulness, and reliability. You will help define what "good" AI output means for StarRez, build repeatable evaluation systems, coach reviewers and subject matter experts, and turn subjective feedback into measurable product improvement. You will work at the intersection of AI engineering, product, and domain knowledge — partnering with engineers, product managers, and SMEs to raise the bar on how StarRez evaluates, monitors, and improves AI tools. Key Responsibilities: Lead or contributed to end-to-end AI quality and evaluation strategy for product experiences involving LLMs, RAG, prompts, tool-use, or agent workflows. Provided quality-focused input during ticket grooming, discovery, and feature discussions, with clear guidance on testability, ob

aigorust
View job →
M
Mongodb
📍 Alberta• Full-time• From $210K/yr
1mo ago

MongoDB’s Developer Productivity organization exists to help engineers build and deliver high-quality software through a highly effective software development process and a strong foundation of shared tools and services. We are looking for a Senior Director to lead our Pipeline team. This role is tasked with bringing together the major systems and experiences that power software delivery at MongoDB. The team’s mission is to provide a reliable, scalable, secure, and effective platform for ensuring fast software deployability, leveraging AI native approaches. We are open to in-office, flexible or remote hiring across Canada. The Team The Pipeline organization sits within Developer Productivity and is responsible for the systems, services, and user experiences that define MongoDB’s software delivery ecosystem. This is mission-critical infrastructure operating at substantial scale and supports a variety of software product delivery needs. Success in this role requires excellent product judgment for developer-facing experiences, strong systems and platform leadership, and the ability to align multiple teams around a cohesive strategy. Candidate Profile We’re looking for a senior engineering leader who can unify product-minded developer tooling with deep platform and operational excellence. The right candidate is passionate about developer productivity and has a track record of leading managers and teams through organizational growth, technical complexity, and cross-functional change. They should be comfortable owning a broad portfolio that spans developer experience, reliability and scale, release systems, telemetry, and operational health. They should also be able to work effectively with senior leaders and partners across engineering and product to set direction, allocate resources, and make trade-offs that balance near-term delivery with long-term platform function. The right candidate for this role will 12+ years of hands-on software engineering experience bui

mongodbawsazure
View job →
M
Mongodb
📍 United States• Full-time• From $168K/yr
1mo ago

MongoDB’s Developer Productivity organization exists to help engineers build and deliver high-quality software through a highly effective software development process and a strong foundation of shared tools and services. We are looking for a Senior Director to lead our Pipeline team. This role is tasked with bringing together the major systems and experiences that power software delivery at MongoDB. The team’s mission is to provide a reliable, scalable, secure, and effective platform for ensuring fast software deployability, leveraging AI native approaches. We are open to in-office, flexible or remote hiring across the US. The Team The Pipeline organization sits within Developer Productivity and is responsible for the systems, services, and user experiences that define MongoDB’s software delivery ecosystem. This is mission-critical infrastructure operating at substantial scale and supports a variety of software product delivery needs. Success in this role requires excellent product judgment for developer-facing experiences, strong systems and platform leadership, and the ability to align multiple teams around a cohesive strategy. Candidate Profile We’re looking for a senior engineering leader who can unify product-minded developer tooling with deep platform and operational excellence. The right candidate is passionate about developer productivity and has a track record of leading managers and teams through organizational growth, technical complexity, and cross-functional change. They should be comfortable owning a broad portfolio that spans developer experience, reliability and scale, release systems, telemetry, and operational health. They should also be able to work effectively with senior leaders and partners across engineering and product to set direction, allocate resources, and make trade-offs that balance near-term delivery with long-term platform function. The right candidate for this role will 12+ years of hands-on software engineering experience bui

mongodbawsazure
View job →

Senior Machine Learning Engineer Description - We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability. Key Responsibilities MLOps Platform and Architecture Design and implement a scalable MLOps platform using AWS and Databricks. Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models. Build standardized workflows that move models from development and validation into staging and production. Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure. Establish clear separation between development, testing, staging, and production environments. Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives. CI/CD and

pythonawsazure
View job →
I
12 days ago

Job Details: Job Description: The Role and Impact As a Systems and Solutions Engineer, you will drive the design, development, and integration of systems that combine software, firmware, board, and silicon/SoC components to meet specific customer needs. In this role, you will play a key part in defining, implementing, and optimizing solutions to ensure high performance, reliability, and quality across the system lifecycle. Your work will directly impact the seamless functionality and user experience of cutting-edge technologies, enhancing Intel's position in delivering innovative systems to global customers. Business Group You will be joining the Silicon and Platform Engineering Group (SPE), an organization committed to advancing Intel's mission of delivering world-class silicon and platform solutions. The group focuses on developing integrated systems that align with customer needs and support Intel's broader goals of leadership in technology innovation. SPE collaborates across diverse domains to ensure Intel platforms meet performance, reliability, and scalability requirements. Key Responsibilities - Design and develop software, firmware, and hardware solutions that integrate seamlessly across system components. - Lead the definition and implementation of system architecture, translating business opportunities into technical requirements and use cases. - Evaluate technical risk and optimize systems for ease of use, reliability, security, availability, and sustainability. - Drive technical solutions to address customer challenges, deploying systems and conducting benchmarks to validate performance. - Collaborate with cross-functional teams to influence next-generation requirements and solutions, guiding research and academic collaborations as needed. - Conduct lab experiments to simulate real-life environments, analyze prototype performance, and refine system

recruitment
View job →
D
DevRev
📍 Chennai• Full-time
16 days ago

About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. About the role: We are looking for a Senior Data Engineer to help build and evolve the data platform that powers critical business decisions and customer-facing experiences. You will own significant parts of our data architecture that is main powerhouse of DevRev Computer’s memory for accurate and efficient Answers. As a part of data team, you will design and operate scalable data systems, and work closely with Software Engineering, AI Agent teams, Data Science, and Product teams to turn complex data requirements into reliable, high-quality data products. This role is ideal for an experienced engineer who enjoys solving challenging problems involving large-scale data, distributed systems, database architecture, and performance optimization . You will have significant technical ownership and the opportunity to influence the direction of our agentic data platform while helping raise the engineering bar across the team. What you'll do: Own data architecture for large-scale, high-impact projects, making thoughtful tradeoffs across scalability, reliability, performance, maintainability, and operational cost. Design, build, and operate scalable data pipelines and data systems that reliably ingest, transform, store, and se

javascriptpythonjava
View job →
J
Jumio
📍 India• Full-time• Remote
16 days ago

About the Role At Jumio, the Software Development Engineer IV - QA (SDE-IV, QA) is a senior technical role focused on ensuring the quality, performance, and reliability of highly scalable web portals and distributed backend systems. Our platform spans multiple Java Spring Boot microservices and customer-facing web portals , deployed across AWS ECS, EKS, and Lambda , and integrated through event-driven messaging using SNS/SQS . In this role you will design and drive the test automation strategy for both UI (Playwright) and API/service layers, set the quality bar for the team, and act as a force multiplier — mentoring other engineers and embedding quality earlier in the development lifecycle. You will work closely with development, product, and DevOps teams to ensure our products meet the highest standards of quality, scalability, and security. This is a hands-on senior IC role: you will write code, but you will also influence architecture, own cross-service test strategy, and make build-vs-buy decisions for testing tooling. T-Shaped Engineering Expectation As part of Jumio's engineering culture, you will adopt a T-shaped engineering approach. Beyond deep expertise in test automation and quality engineering, you will contribute across the development lifecycle — understanding software architecture, participating in design and API-contract discussions, reviewing application code, and ensuring our distributed systems are testable, observable, and resilient by design. Role Value This role is critical to ensuring the reliability, scalability, and security of Jumio's products. By architecting and maintaining automated testing frameworks across web, API, and event-driven layers, you will enable faster, higher-confidence releases and reduce production risk in a complex microservices environment. What You'll Do Test Architecture & Strategy Define and own the end-to-end automated test strategy across web portals and backend microservices, balancing UI, API, contract, integ

REMOTEjavascripttypescriptjava
View job →

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

pythonlinuxartificial intelligence
View job →

About Inspira Education Inspira Education Group is one of the fastest-growing edtech startups in the US. We started with a simple mission to democratize access to high-quality coaching so that every student in the world has an equal opportunity to access the best opportunities. As the world’s leading network of top admissions coaches in medical, legal, business, and college studies, we’re building software and services in one place—disrupting long-entrenched application processes with products and experiences that strive to provide an equal platform for candidates from diverse backgrounds worldwide. As one of the fastest-growing edtech firms in the world, we are backed by some of the leading venture capital firms and investors in the world, including Zeev Ventures, Quiet Capital, Craft Ventures and Jeff Fluhr (Founder of Stubhub). About the role We’re looking for a strong full-stack engineer who can own the complete product development process: understand a business problem, define the solution, design the user experience, build the software, and improve it after launch. You’ll work closely with leadership and business teams, combining hands-on engineering with product management and design responsibilities. You should be highly effective with AI coding tools and have the technical depth to independently review, debug, secure, and maintain everything you ship. This is an in-person role requiring 5 day/week in our NYC office. What you’ll own Translate business needs and user feedback into product requirements, user flows, prototypes, and prioritized development plans. Design and build polished applications across the front end, back end, database, and integrations. Make architecture decisions and scope releases that balance speed, reliability, and future maintainability. Use AI tools throughout development to accelerate implementation, testing, debugging, and documentation. Own deployment, production monitoring, incident resolution, and ongoing improvemen

javascripttypescriptpython
View job →
G
12 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

pythonartificial intelligenceai
View job →
G
16 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

pythonlinuxai
View job →
G
Graphcore
📍 Austin• Full-time
16 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

pythonaiexcel
View job →
G
16 days ago

Senior -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to deb

pythonlinuxai
View job →
🔔

Get new senior software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime