Jobiba hiring network

Senior Software Reliability Engineer Jobs

7,292 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OneTrust
📍 Atlanta• Full-time• From $116.5K/yr
16 days ago

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge We're looking for a Senior Software Engineer that will report to the Development Manager / R&D Head. In this role, you will part of the R&D Team that works on mission-critical applications. Your Mission Engage and partner with various Engineering, Operations, and Product teams to design, deliver, and maintain a highly available and performant application platform. Build and implement application observability and platform monitoring tools to continuously improve the customer experience Eliminate toil by automating processes, tuning alerts, and improving code where it is most needed Frequently evaluate new ideas and trends to identify potentially useful tools and techniques Collaborate with different functional groups to identify gaps, prioritize, and resolve issues Defining, implementing, and maintaining SLIs and SLOs aligned with customer experience. Design and instrument SLIs such as latency, error rates, and availability across critical services Manage and enforce error budgets to balance system reliability with product feature v

pythonjavasql
View job →
O
OpenAI
📍 San Francisco• Full-time• $295K – $380K/yr
1mo ago

About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Senior Software Engineer, ML Systems & Training Infrastructure, you will be a deeply hands-on engineering force multiplier for the robotics team. You will help keep the training framework and surrounding infrastructure healthy, review and improve code quickly, debug failures across ML systems and infrastructure, and unblock researchers and engineers when the path from idea to working training job gets rough. We’re looking for people who love writing, reading, reviewing, and fixing code; who can get productive quickly in unfamiliar systems; and who bring strong practical judgment without a lot of ego or process overhead. This role will be based in San Francisco, CA and be expected in office 5 days per week and offer relocation assistance to new employees. In this role, you will: Review, improve, and clean up code across training frameworks and adjacent infrastructure. Identify risky or low-quality changes before they land, and raise the code quality bar without slowing the team down. Debug issues across ML training systems, GPUs, clusters, networking, and related infrastructure. Help researchers and engineers unblock broken training jobs, flaky workflows, and brittle internal tooling. Improve the reliability, maintainability, and usability of the robotics team’s training framework. Move quickly on practical engineering problems that directly affect team velocity. You might thrive in this role if you: Have strong software engineering fundamentals and excellent code review judgment. Have experience with ML systems, training fr

awsrestai
View job →
C
Cohere
📍 European Union• Full-time• From £180K/yr
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? North is Cohere's cutting-edge AI workspace platform, designed for enterprise users. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools, data and applications. The North for Finance team builds specialized AI solutions for finance teams while accelerating adoption of agentic workflows for their mission-critical operations. You'll work at the intersection of AI innovation and financial services, developing features that integrate domain-specific knowledge with enterprise-grade reliability. As an engineer on this team you will help transform complex financial workflows through intelligent automation. You'll extend core agentic platform capabilities with built-in data isolation, structured-data manipulation and collaborative governance, positioning North as the leading AI workspace for finance. As a Senior Software Engineer, you will: Design, build, ship, and maintain customer facing workflows and agentic automations

M
Mindbody
📍 United States• Full-time
1mo ago

At Playlist, life's richest moments happen when people step away from screens to move, connect, explore, and play. We're building the definitive platform for intentional living, connecting people with inspiring experiences in fitness, wellness, and beyond. With popular brands like Mindbody and ClassPass, Playlist empowers businesses and individuals, making it effortless for aspirations to become actions. Join us in reshaping technology's role to foster meaningful, real-world connections. Mindbody equips wellness entrepreneurs with technology to support thriving businesses and create exceptional experiences. Innovation and curiosity drive our culture, connecting businesses and individuals through cutting-edge solutions. Join us if you're passionate about enhancing wellness through technology. The Role You'll Play: At Playlist, we're reimagining how technology can foster meaningful, real-world connections. As a Senior Software Engineer on the SmartDesk team, you'll be at the forefront of building an AI-powered front desk assistant that transforms how wellness businesses communicate and operate. Crafting end-to-end AI-powered experiences that seamlessly connect wellness businesses with their clients Designing and implementing robust web and messaging services that power conversational and workflow automation Integrating large language models (LLMs) and generative AI frameworks with a laser focus on safety, reliability, and user trust Collaborating closely with product, design, and applied AI teams to transform complex challenges into intuitive solutions Architecting scalable systems across frontend, backend, and integration layers Mentoring teammates through thoughtful code reviews and

javascripttypescriptpython
View job →
G
16 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Quality leadership within Manufacturing Operations, the Senior Reliability Scientist is responsible for leading reliability activities across complex, high-performance systems. Working closely with established reliability experts and cross-functional teams, this role uses experimental data and advanced modelling to inform design decisions, validate product reliability and optimise serviceability strategies, including spares provisioning. The Team The Quality team within Manufacturing Operations is responsible for ensuring product robustness, reliability and lifecycle performance across Graphcore’s hardware portfolio. The team includes experienced reliability specialists and works closely with technology research, chip, board, system design, platform and operations teams to translate reliability insights into actionable improvements across the product lifecycle. Responsibilities and Duties: · Define and refine reliability requirements across silicon, board and system levels, working in partnership with research and design teams · Apply ad

aigoexcel
View job →

The Engineering Lead Analyst – SonarQube & Code Quality Engineering is a senior-level engineering role responsible for leading static code analysis, automated code quality governance, security vulnerability remediation, and AI-augmented developer enablement across enterprise software delivery pipelines. In this role, you will champion software reliability, maintainability, clean-coding standards, and automated quality gates. You will partner with development teams, system architects, and platform engineering to integrate and manage enterprise-scale code quality platforms (such as SonarQube) both on-premises and in cloud/SaaS environments. Additionally, you will drive modern engineering practices by embedding Behavior-Driven Development (BDD) within your own software delivery and leveraging Agentic AI workers and Model Context Protocol (MCP) architectures to optimize developer experience, streamline code governance, and boost engineering velocity. Key Responsibilities 1. Code Quality & Static Analysis Platform Ownership Lead the architecture, deployment, administration, and continuous enhancement of enterprise Static Application Security Testing (SAST) and Code Quality platforms (e.g., SonarQube , DeepSource, Codacy, Semgrep). Configure, calibrate, and enforce automated Quality Gates, code rulesets, technical debt calculation models, and code-coverage baselines across multi-language enterprise repositories. Oversee version upgrades, patching, high availability, and operational maintenance for on-premises and SaaS/cloud-hosted code quality infrastructure. 2. CI/CD & Pipeline Integration <li style=

javascripttypescriptpython
View job →
S
Smartsheet
📍 Bangalore, INDIA• Full-time• Hybrid
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Automation is the key to creating highly reliable and secure large-scale software systems. Are you someone who engineers solutions to problems rather than simply fixing the same thing over and over again? Can you protect Smartsheet against attackers? We are looking for a Senior DevSecOps Engineer to join our global Security Operations team. In this critical role, you will be a leader in maturing our security and reliability posture by treating both as software engineering challenges. You will engineer and operate a highly reliable, scalable, and defensible production environment, directly impacting our ability to deliver a world-class service to our customers 24/7. This is a unique opportunity to blend deep expertise in Site Reliability Engineering (SRE) and modern Security Operations, working at the intersection of infrastructure, automation, and security to build a platform that is resilient and secure by design. You Will: Engineer Secure and Resilient Infrastructure: Design, build, maintain, and improve secure, scalable, and highly available infrastructure in our multi-cloud environment (primarily AWS) using Infrastructure as Code (IaC) principles with tools like Terraform, Kubernetes, and Helm. Automate Proactive Security: Engineer an

pythonawskubernetes
View job →

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge As a Senior Staff Software Engineer, you will serve as a technical leader for OneTrust’s AI Governance (AIG) platform, driving the design, scalability, and reliability of systems that enable enterprises to deploy and govern AI and LLM-powered applications responsibly. You will deeply understand how customers build, deploy, and operate AI systems, and translate those needs into secure, compliant, and observable platform capabilities. Your Mission Development Lead the design and development of Java/Python microservices and shared libraries integrating with AI platforms for OneTrust’s AI Governance product. Design, build, and test cloud-native applications deployed on Microsoft Azure using Core Java, REST, and the Spring ecosystem. Lead the architecture and development of reusable AIG reporting and dashboard capabilities that integrate governance data from SQL databases and analytical platforms with runtime observability signals. Design reusable semantic-layer and metric-abstraction capabilities, including dataset contracts, metric defini

pythonjavasql
View job →

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge As a Senior Staff Software Engineer, you will serve as a technical leader for OneTrust’s AI Governance (AIG) platform, driving the design, scalability, and reliability of systems that enable enterprises to deploy and govern AI and LLM-powered applications responsibly. You will deeply understand how customers build, deploy, and operate AI systems, and translate those needs into secure, compliant, and observable platform capabilities. Your Mission Development Lead the design and development of Java/Python microservices and shared libraries integrating with AI platforms for OneTrust’s AI Governance product. Design, build, and test cloud-native applications deployed on Microsoft Azure using Core Java, REST, and the Spring ecosystem. Lead the architecture and development of reusable AIG reporting and dashboard capabilities that integrate governance data from SQL databases and analytical platforms with runtime observability signals. Design reusable semantic-layer and metric-abstraction capabilities, including dataset contracts, metric defini

pythonjavasql
View job →
P
1mo ago

About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: Join a team that builds robust, real-time distributed systems for a cutting-edge database. We care about performance, reliability, scalability, and most of all learning and having fun together. Whether you’re a seasoned coder or just getting started, if you’re passionate about technology and eager to learn, you’ll fit right in. Who we are: We show up to work, ready to collaborate and build technologies that make a difference, with people who genuinely care. We chase improvements such as tail latencies, bytes throughput, cache hit rate, and operational cost efficiency. We believe learning is ongoing and that even the most complex problems can have simple solutions. What You’ll Do: Collaborate with teammates to design and build database features that power AI applications. Learn how to tune performance and support reliability in distributed systems (don’t worry, we’ll guide you). Help Pinecone run smoothly on popular cloud providers. Take ownership of your work and grow your skills every day. Have fun. Who You Are: 5+ years of work experience - programming in Rust, Go, C++, or a comparable language. You’re genuinely curious about distributed systems and eager to dive deep into technical challenges. You approach problems with creativity and persistence, and you’re comfortable asking thoughtful questions or seeking feedback. You’re excited to learn, value constructive feedback, and appreciate mentorship. Bonus Points: You have hands-on experience with cloud platforms (AWS, GCP, Azure) or have demonstrated an ability to pick u

awsazuregcp
View job →
E
12 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE We are seeking an experienced Platform Software Engineer for our Systems Software Team. You will be working as part of a dynamic team and will be responsible for designing, developing, and testing system software functionality for Pure’s upcoming platforms. The work spans the gamut of Systems software and you will have the opportunity to work in a wide range of areas and features ranging from Platform drivers to networking and storage layers. WHAT YOU'LL DO Plan and influence the lifecycle of new Hardware Platforms. Work on problems ranging from design, bring up, to deployment, upgrades and fleet level reliability. Participate in the full lifecycle of new hardware platforms from early bring up through manufacturing release. Work closely with peer teams to debug complex HW/FW of new server hardware, including CPUs, chipsets, and peripheral components. Debug complex HW/FW issues across x86, PCIe, NVMe, and networking using lab tools (oscilloscope, logic analyzer, JTAG) and kernel/driver traces. Design, implement and improve remote server management capabilities (e.g., using standards like Redfish) and enhance Reliability, Availability, and Serviceability (RAS) features. Design, write and maintain software components in C/C++, Python, Golang and RUST. Collaborate with vendors on requirements specification and follow through to system delivery. Work closely with hardware engineers, system architects, and o

pythonlinuxai
View job →

Are you looking for an opportunity to help solve one of today's biggest business challenges? AI is changing the pace of business, and organizations everywhere are struggling to help their workforce, partners, and customers keep up. At Litmos, we're building the Learning Acceleration Platform that helps organizations build human capability faster—and we're looking for people who are passionate about making a meaningful impact for customers while growing alongside a collaborative, people-first team. Litmos is the Learning Acceleration Platform that helps organizations build capability faster, adapt at the speed business changes, and scale learning to anyone, anywhere. Combining an intuitive platform, AI-powered capabilities, trusted content, expert services, and a broad ecosystem of integrations, Litmos helps organizations accelerate workforce productivity, improve customer adoption and retention, enable high-performing partners, and reduce organizational risk through continuous learning. Organizations such as Hewlett Packard Enterprise, Graco, Sabre, and Russell Mineral Equipment trust Litmos to accelerate learning across their workforce, partners, and customers. Today, more than 11K customers with 30 million learners across 150 countries and 37 languages use Litmos to build the capabilities their organizations need to succeed. Backed by Francisco Partners, one of the world's leading technology investment firms, we're investing in the future of learning—and the people who are building it. Learn more at www.litmos.com . We are looking for a Senior Full Stack Engineer to build and evolve our platform using .NET, React, and SQL Server-based systems. You will work across services, APIs, and front-end systems with a focus on scalability, reliability, and maintainability. You will operate within a globally distributed Agile team (Scrum and ShapeUp), where engineers own delivery, from design through deployment, while collaboratin

reactsqlazure
View job →

The NVIDIA PerfTech team is looking for a talented C&#43;&#43; Software Engineer to help build the next generation of AI-powered developer tools. You will apply strong C&#43;&#43; and software-engineering fundamentals while gaining hands-on experience with agentic workflows, retrieval systems, and AI services. In this role, you will contribute to Genie, NVIDIA’s company-wide AI knowledge and developer-productivity service. You will work across C&#43;&#43; tools and AI services to help engineers find information, understand complex systems, and work more effectively. What You’ll Be Doing: Develop production-quality C&#43;&#43; components, APIs, and integrations for NVIDIA’s AI-powered developer-tools ecosystem. Build capabilities connecting native C&#43;&#43; tools with Genie’s retrieval and agentic features. Contribute to agentic workflows, retrieval systems, ingestion pipelines, MCP tools, APIs, and enterprise integrations. Build benchmarks and improve retrieval quality, reliability, performance, and resource usage. Own features from investigation and design through implementation, testing, and delivery. Collaborate with graphics, software, and hardware teams developing performance-analysis and developer tools. What We Need to See: Bachelor’s or Master’s degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience. 5&#43; years of modern C&#43;&#43; programming skills gained through professional experience, internships, or substantial technical projects. Good understanding of data structures, algorithms, object-oriented design, multithreading, debugging, and testing. Ability and motivation to work across C&#43;&#43; systems and Python-based AI services. Familiarity with AI-powered applications, agentic workflows, retrieval systems, or related technologies. Abil

pythonaic++
View job →
HI
HP IQ
📍 San Francisco• Full-time• $140K – $225K/yr
16 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role HP IQ’s Connectivity team is seeking a Software Engineer with deep expertise in device software development. The ideal candidate will bring strong knowledge of connectivity stack and hands-on experience developing, integrating, and optimizing device software to deliver industry-leading user experiences. You will work at the intersection of Wi-Fi, Bluetooth, and emerging device-to-device transport technologies, tackling complex challenges in performance, reliability, and low-latency communication. This role offers the opportunity to contribute to cutting-edge innovations that are redefining how people and devices seamlessly connect across the modern enterprise. What You Might Do Design, develop, and integrate connectivity software features across Android, Windows, and embedded platforms, including SDKs, frameworks, and system services. Implement, optimize, and tune wireless networking protocols and sensing algorithms with a focus on enterprise-scale architectures and deployments. Design, implement, and troubleshoot peer-to-peer technologies to deliver secure, reliable, and low-latency

javaredislinux
View job →
L
Litmos
📍 India• Full-time• ₹2.6Cr – ₹3.4Cr/yr
16 days ago

Are you looking for an opportunity to help solve one of today's biggest business challenges? AI is changing the pace of business, and organizations everywhere are struggling to help their workforce, partners, and customers keep up. At Litmos, we're building the Learning Acceleration Platform that helps organizations build human capability faster—and we're looking for people who are passionate about making a meaningful impact for customers while growing alongside a collaborative, people-first team. Litmos is the Learning Acceleration Platform that helps organizations build capability faster, adapt at the speed business changes, and scale learning to anyone, anywhere. Combining an intuitive platform, AI-powered capabilities, trusted content, expert services, and a broad ecosystem of integrations, Litmos helps organizations accelerate workforce productivity, improve customer adoption and retention, enable high-performing partners, and reduce organizational risk through continuous learning. Organizations such as Hewlett Packard Enterprise, Graco, Sabre, and Russell Mineral Equipment trust Litmos to accelerate learning across their workforce, partners, and customers. Today, more than 11K customers with 30 million learners across 150 countries and 37 languages use Litmos to build the capabilities their organizations need to succeed. Backed by Francisco Partners, one of the world's leading technology investment firms, we're investing in the future of learning—and the people who are building it. Learn more at www.litmos.com . We are looking for a Senior Full Stack Engineer to build and evolve our platform using .NET, React, and SQL Server-based systems. You will work across services, APIs, and front-end systems with a focus on scalability, reliability, and maintainability. You will operate within a globally distributed Agile team (Scrum and ShapeUp), where engineers own delivery, from design through deployment, while collaboratin

reactsqlazure
View job →
🔔

Get new senior software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime