Jobiba hiring network

Software Reliability Engineer Jobs

6,326 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. In this role you will: Develop interactive, data-rich user interfaces using React, TypeScript, and Vega, with a focus on integrating LLM-driven features (e.g., natural language querying, generative UI, and AI-assisted data storytelling). Lead the end-to-end delivery of substantial product features, ensuring AI outputs are presented with high reliability and low latency. Work closely with PMs, UX designers, and AI/ML engineers to bridge t

javascripttypescriptjava
View job →
PE
Private Employer
📍 Dublin• Full-time• Hybrid
1mo ago

Shape the Future with Dun & Bradstreet At Dun & Bradstreet, we believe data has the power to create a better tomorrow. As a global leader in business decisioning data and analytics, we help companies worldwide grow, manage risk, and innovate. Since 1841, businesses have trusted us to turn uncertainty into opportunity. We’re a diverse, global team that values creativity, collaboration, and bold ideas. Are you ready to make an impact and help shape what’s next? Join us! Explore opportunities at dnb.com/careers. The Software Engineer II is responsible for designing, developing, and delivering scalable data platform solutions that support high-performance data processing and analytics capabilities. This role operates with increased autonomy, contributing to system design, driving delivery of platform capabilities, and ensuring the reliability and scalability of the Dun & Bradstreet data ecosystem. The role plays a key part in modernizing and optimizing data platforms, including identity resolution and Match systems, to meet growing business demands.

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. The Role You will work on our Metrics platform - enabling our users to query, visualize, and alert on billions of time series quickly and effectively. You'll own meaningful parts of that stack, drive performance and scalability improvements, and contribute to architectural decisions that shape where the platform goes next. This isn't a maintenance role, the metrics backend is being actively evolved, and you'll be a core part of that. Thi

aigorust
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. The Role You'll be our dedicated expert in query execution and query performance. That means owning the query execution service end-to-end: working on caching strategies, incremental execution, query rewrites, and other optimizations that directly affect the speed and cost of running Observe at scale. You'll also be the go-to resource when query latency issues arise during customer evaluations and new deal cycles, diagnosing root causes

aigorust
View job →
N
Nvidia
📍 Remote, Poland• Remote
12 days ago

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, passionate, and self-motivated, we want to hear from you! We are looking for an experienced networking software engineer. An awesome candidate is highly technical who is also comfortable with dealing with enterprise customers. You will join a team of Solution Engineers focused on the Mellanox Networking, DGX Platforms, Container Orchestrators, Deep Learning containers, and other Enterprise related system software. SW Solution Engineers spend approximately 50% of their time helping customers with their most complex problems and 50% of their time doing R&D related work. This individual should have proven grasp of datacenter and networking technologies, to provide comprehensive solutions for complex installations, maintenance, or operations for a broad scope of leading-edge networking products. What you'll be doing: Take ownership and drive customer issues with Ethernet or InfiniBand network adapter/DPU deployments from inception to resolution. Develop features and tools as part of solution engineering efforts to support all Enterprise Service offerings including but not limited to Networking products. Work with NVIDIA Enterprise customers and internal users to improve the availability, reliability, and overall experience of working with NVIDIA Networking products. Bring independent analysis, communication, and problem-solving to customer experience. Collaborate with engineering to document, recreate and solve issues. What we need to see: BSc in Computer Science, Electrical Engineering, Computer Engineering, or related field (or equivalent experience). 8+ years system software developm

REMOTEkuberneteslinux
View job →
N
12 days ago

We are seeking a highly skilled and hard-working Senior Test Developer / test engineer to join our multifaceted Enterprise Software QA team. This role offers an outstanding opportunity to leave your mark on the design, construction, optimization and testing of large-scale infrastructure for various foundational NVIDIA unified cloud services and data center offerings. If you are a dedicated engineer with strong expertise in cloud infrastructure and distributed systems and want to apply your skills with AI tools, this role could fit you perfectly. You will thrive in an exciting, innovative environment. What you'll be doing: Work with development teams on test plans for all layers of SW stack for cloud infrastructure, execution, reviews, failure analysis and assessing overall quality and risk. Work with customer PMs on software issues including technical feedback from OEMs and CSPs. Develop key benchmarks to track execution and deploy process improvements to improve efficiency Leverage AI skills to expedite the test scope, test plan, execution and automation workflows. Lead NVIDIA Cloud and Data Center bring up activities which will involve validation, reporting, working with engineering to debug issues, providing design input at times, adding coverage in different areas. Design, develop and maintain CI/CD pipelines for continuous testing in cloud environments when needed. Perform performance, scalability, and reliability testing of cloud services. Implement and maintain test environments in cloud platforms such as AWS, Azure, or Google Cloud. Supervise the infrastructure to alert on significant events, ensuring the highest level of system performance and reliability. Work with various different partner teams to ensure availability of clusters to test on and take the lead in resolve all issues. Working with tea

awsazuredocker
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join a world-class team of engineers building the next generation of enterprise storage solutions. As a key contributor, you'll be at the forefront of innovation, developing and optimizing the Linux kernel to push the boundaries of performance and reliability. You'll play a vital role in shaping the future of our products, collaborating with a brilliant team to solve complex challenges and deliver groundbreaking results. WHAT YOU'LL DO Develop new features within the Linux kernel in support of Pure’s enterprise storage products. Maintain and patch existing Linux code to resolve difficult problems, including customer issues. Optimize performance of the kernel within Pure’s arrays to meet customer requirements Work cross-functionally and with partners and vendors, to diagnose and resolve problems at the boundary of hardware and software Lead the architecture and development of software from initial concept to release, ensuring high-quality, resilient, and high-performance outcomes. Collaborate and share knowledge with peers, providing mentorship as necessary. Participate in code reviews and collaborate with cross-functional teams to define requirements for upcoming enterprise storage server projects. WHAT YOU BRING Deep, hands-on experience in Linux kernel/Unix and device driver development, with a proven ability to ship high-performance, resilient products. Minimum 5 years experience preferred A

linuxaic++
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

REMOTEkubernetesaigo
View job →
C
1mo ago

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary We are seeking a highly experienced Principal / Director-level Full-Stack Software Development Engineer to join our Digital Caremark organization and lead the architecture, design, and delivery of next-generation digital solutions. This role spans AI-enabled applications, scalable digital platforms, and enterprise integrations that power critical healthcare and pharmacy experiences. This is a senior technical leadership role for a hands-on engineer who can operate across the full stack—from intuitive front-end applications to resilient backend services—while setting architectural direction, influencing engineering standards, and mentoring teams. The ideal candidate combines deep technical expertise, platform thinking, and strong collaboration skills to build secure, scalable, API-first solutions in a highly regulated environment. Key Responsibilities: Architecture & System Design • Define and drive architecture for large-scale distributed systems and digital platforms • Lead design reviews and set architecture standards and best practices • Champion API-first, microservices, and event-driven architecture patterns • Ensure systems meet scalability, reliability, security, and compliance requirements • Balance performance, cost, and speed in technical decision-making Fu

typescriptpythonjava
View job →

About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: Join a team that builds robust, real-time distributed systems for a cutting-edge database. We care about performance, reliability, scalability, and most of all learning and having fun together. Whether you’re a seasoned coder or just getting started, if you’re passionate about technology and eager to learn, you’ll fit right in. Who we are: We show up to work, ready to collaborate and build technologies that make a difference, with people who genuinely care. We chase improvements such as tail latencies, bytes throughput, cache hit rate, and operational cost efficiency. We believe learning is ongoing and that even the most complex problems can have simple solutions. What You’ll Do: Collaborate with teammates to design and build database features that power AI applications. Learn how to tune performance and support reliability in distributed systems (don’t worry, we’ll guide you). Help Pinecone run smoothly on popular cloud providers. Take ownership of your work and grow your skills every day. Have fun. Who You Are: 5+ years of work experience - programming in Rust, Go, C++, or a comparable language. You’re genuinely curious about distributed systems and eager to dive deep into technical challenges. You approach problems with creativity and persistence, and you’re comfortable asking thoughtful questions or seeking feedback. You’re excited to learn, value constructive feedback, and appreciate mentorship. Bonus Points: You have hands-on experience with cloud platforms (AWS, GCP, Azure) or have demonstrated an ability to pick u

awsazuregcp
View job →
E
ElevenLabs
📍 United Arab Emirates• Full-time
1mo ago

About ElevenLabs ElevenLabs is an AI research and product company transforming how we interact with technology. We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including Andreessen Horowitz, ICONIQ Growth and Sequoia. We've raised $781M in funding and our last valuation was $11B - multiples of 11, always. We have expanded from voice into three main platforms: ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale. ElevenCreative empowers creators and marketers to generate and edit speech, music, image, and video across 70+ languages. ElevenAPI gives developers access to our leading AI audio foundational models. Everything we do is the result of the creativity and commitment of our team - builders doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to work hard and create lasting positive impact, we want to hear from you. How we work High-velocity: Rapid experimentation, lean autonomous teams, and minimal bureaucracy. Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you. AI first: We use AI to move faster with higher-quality results. We do this across the whole company—from engineering to growth to operations. Excellence everywhere: Everything we do should match the quality of our AI models. Global team: We prioritize your talent, not your location. What we offer Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible. Growth paths: Joining ElevenLabs means joining a

pythonrestai
View job →
J
Jamf
📍 Us Remote• Full-time• Remote• From $113.3K/yr
1mo ago

At Jamf, we believe in an open, flexible culture based on respect and trust. Our track record and thriving work environment all stem from the freedom we grant ourselves to get the job done right. We take pride in helping tens of thousands of customers around the globe succeed with Apple. The secret to our success lies in our connectivity, while operating with a high degree of flexibility. Work-life balance remains our priority while feeling connected is important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: The Senior Software Engineer is responsible for building the tools required to help organizations succeed with Apple. Lead others on the agile team to break down problems and apply the appropriate designs and practices to build Jamf products. Subject matter expert in various Jamf components and product offerings. Mentor and coach others while delivering new components and features with high quality and reliability. You may be required to work periodically at a Jamf office or collaborative work location with other Jamf employees in your area for certain events or moments that matter. What you can expect to do in this role : Break down customer problems into work you and the team can execute on. Independently complete tasks from start to finish with high quality. Ability to communicate technical concepts to stakeholders. Use your knowledge of Engineering best practices to ask the right questions, solve problems and build great software with a high level of quality. Produce designs for new and existing features. Clearly communicate technical concepts with others in the organization (Technical Communication, Support, Product and Cloud). Performs all job responsibilities in alignment with the core values, mission and purpose of the organization. Adheres to the highest moral, ethical and legal standards to deliver and environment that promotes respect, innovation and creativity

REMOTErestagileai
View job →
A
Asana
📍 Warsaw• Full-time• $372K – $432K/yr
1mo ago

We're looking for a Senior Infrastructure Engineer who brings strong software engineering skills and a deep understanding of production systems. This role is a good fit for someone who enjoys building systems that make infrastructure more scalable, reliable, and easy to operate – using code, not runbooks. You'll work with a highly collaborative team to design and build the internal platforms that power all of Asana, from product features to AI systems to offline analytics. Our tech stack includes: AWS, Kubernetes (EKS), MySQL (RDS), OpenSearch, DynamoDB, Redis, Terraform, Datadog, TypeScript, Scala, Go, and Python. We’re especially interested in people who think like backend engineers but care deeply about systems – things like failure modes, operational cost, debuggability, and performance. This role is based in our Warsaw office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do, and your recruiter can share more about the in-office requirements. We offer a Contract of Employment (UoP) for our employees in Poland. What you’ll achieve: Design and build frameworks, tools, and services that improve the reliability, observability, and scalability of Asana’s infrastructure. Lead end-to-end projects, from scoping and design through to rollout, across multiple systems and teams. Improve the operability of stateful infrastructure like MySQL, OpenSearch, and DynamoDB – and help drive Asana’s long-term vision for storage reliability. Debug production issues across the stack. Yes, there’s an on-call rotation – but this isn’t a pager monkey role. You’re here to fix things properly and make sure they don’t break again. Partner with product teams to shape a service-oriented architecture that enables fast, reliable development. Share knowledge through code reviews, design discussions, and mentorship. Abou

typescriptpythonsql
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

OpenAI’s charter calls on us to ensure the benefits of AI are distributed broadly and safely. Our Health AI team focuses on expanding access to high-quality medical expertise and aims to set a high standard for deploying AI responsibly in high-stakes domains. Improving health will be one of the defining impacts of AGI. Today, millions of people lack access to reliable medical information, and clinicians around the world face increasing time and resource constraints. We are building AI systems that support patients, clinicians, and health workers, while meeting the highest standards for safety, reliability, and privacy. We are seeking full stack software engineers to help build and scale products used by consumers and care providers globally. You will work closely with product, design, and research teams to ship real systems in a fast-moving, high-impact environment. In this role, you will: Design and build scalable fullstack systems for consumer and enterprise health. Own end-to-end feature development—from early design and implementation through deployment, monitoring, and iteration. Build and maintain data pipelines and services that meet strict privacy, security, and compliance requirements (e.g., HIPAA). Collaborate closely with researchers and safety teams to integrate reliability, evaluation, and guardrails into production systems. Debug, optimize, and harden systems to support high availability, performance, and global scale. Take ownership of ambiguous problems and drive them to practical, high-quality solutions. You might thrive in this role if you: Are deeply motivated by improving health outcomes and expanding access to medical expertise. Are a strong engineer who enjoys building durable, well-designed systems. Have 5+ years of experience writing maintainable, production-quality code. Can operate with high agency—owning problems end-to-end with minimal supervision. Enjoy working in fast-moving, cross-functional teams with engineers, product managers, desi

awsgitrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI's Industrial Compute organization is responsible for planning, delivering, operating, and optimizing the compute infrastructure that powers frontier AI. As OpenAI scales toward becoming an intelligence utility, Industrial Compute coordinates a complex lifecycle spanning infrastructure strategy, capacity planning, provider partnerships, fleet operations, product demand, and financial planning. The organization manages one of the largest and fastest-growing compute footprints in the world, where decisions around capacity allocation, deployment readiness, utilization, reliability, and product demand directly impact product availability, customer experience, and business performance. The Capacity Systems team builds the software platforms, data systems, and automation frameworks that connect these functions into a shared operating model. We transform fragmented planning workflows into scalable systems that enable teams to understand what compute was contracted, delivered, healthy, allocated, and ultimately converted into business and research outcomes. About the Role We are seeking a Capacity Systems Software Engineer to build the platforms and services that power Industrial Compute planning, forecasting, optimization, and operational decision-making. In this role, you will design and develop software systems that connect infrastructure delivery, fleet health, capacity allocation, demand forecasting, deployment readiness, financial planning, and product consumption into a unified system of record. Your work will help OpenAI make better decisions about where compute should be deployed, how capacity should be allocated, and how infrastructure investments translate into business value. You will partner closely with Capacity Planning, Fleet Operations, Infrastructure Engineering, Product, Finance, Supply Chain, and Strategic Sourcing teams to replace spreadsheet-driven workflows with scalable software systems that enable visibility, automation, and dec

typescriptpythonjava
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime