For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. The Grid Service and Platform Engineering team is looking for a highly motivated and collaborative Software Engineering Manager. This role involves leading the engineering of mission-critical, tier 0 service infrastructure, the foundational data platform that powers Smartsheet at scale. You will oversee services that handle millions requests per day, operate at 99.999% availability, and deliver low-latency, high-throughput performance for millions of customers worldwide. We are an agile team that operates iteratively, focused on building high-quality software and adhering to rigorous operational best practices across complex, cross-functional distributed systems. This full-time position reports to the Director, Engineering and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Manage one or more related teams of 6–10+ software engineers, driving development of tier 0 grid services and platform infrastructure that millions of customers depend on daily. Own and uphold 99.999% service availability targets across critical platform services, embedding reliability engineering, incident management, and on-call rigor into team culture. Help architect and guide technical vision to evolve low-latency, high-throughput service platforms capable of sustaining millions requests per day with predictable, consistent performance under load. Guide and mentor engineers on distributed systems architecture, scalability patterns, and platform best pr
Jobiba hiring network
Software Reliability Engineer Jobs
6,428 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. About the Role & Team Everpure acquired Portworx to create the industry's most complete Kubernetes Data Services Platform for cloud-native applications. In this role, you will be supporting multi-cloud data services for Kubernetes and containerized workloads, enabling leading enterprises to run mission-critical data applications smoothly across public and private clouds. WHAT YOU'LL DO Analyze & Support: Troubleshoot and support large-scale customer deployments in public/private clouds across all severity levels. Technical Expertise: Provide hands-on guidance during all deployment phases, POCs, pre-sales calls, and production environments for key accounts. Cross-Functional Collaboration: Partner with engineering teams to analyze logs, reproduce complex customer issues, and develop long-term fixes. End-to-End Ownership: Track customer support cases end-to-end, triage multi-layer software stack issues, and escalate to core engineering when needed. Knowledge Sharing: Author and maintain KB articles, FAQs, and technical documentation for internal teams and customers. WHAT YOU BRING Experience: 2 - 4+ years in customer-facing technical support or Site Reliability Engineering (SRE). Containers & Orchestration: Solid working knowledge of Kubernetes, OpenShift, Tanzu, or VMware container solutions ( CKA certification is a plus ). Cloud Platforms: Hands-on experience with AWS, Azure, GCP, or related cloud technol
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Automation is the key to creating highly reliable and secure large-scale software systems. Are you someone who engineers solutions to problems rather than simply fixing the same thing over and over again? Can you protect Smartsheet against attackers? We are looking for a Senior DevSecOps Engineer to join our global Security Operations team. In this critical role, you will be a leader in maturing our security and reliability posture by treating both as software engineering challenges. You will engineer and operate a highly reliable, scalable, and defensible production environment, directly impacting our ability to deliver a world-class service to our customers 24/7. This is a unique opportunity to blend deep expertise in Site Reliability Engineering (SRE) and modern Security Operations, working at the intersection of infrastructure, automation, and security to build a platform that is resilient and secure by design. You Will: Engineer Secure and Resilient Infrastructure: Design, build, maintain, and improve secure, scalable, and highly available infrastructure in our multi-cloud environment (primarily AWS) using Infrastructure as Code (IaC) principles with tools like Terraform, Kubernetes, and Helm. Automate Proactive Security: Engineer an
We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices
We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, http://incident.io , Rootly, Jira Cloud Platform, Cortex, or similar and know how to navigate operational challenges of different
We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, incident.io, Rootly, Jira Cloud Platform, Cortex, or similar and know how to navigate operational challenges of different s
About BlockTech BlockTech is a fast-paced algorithmic trading firm facilitating global cryptocurrency derivatives and spot trading while expanding into new markets. As we continue to grow rapidly, we are looking for a Software Engineer to join our Foundation team amid our exciting scale-up phase! You will Build & optimize: Design, develop, and maintain high-reliability, low-latency, and high-throughput foundational systems that enable our trading and technology teams to scale efficiently. Ingest & aggregate: Collect trading business data with minimal latency impact and ingest both public and private exchange information into our trading system. Store & stream: Develop and maintain infrastructure for real-time data aggregation and long-term storage, as well as our Kafka-based messaging systems. Collaborate & support: Work closely with multiple teams, assisting them in integrating with and making the most of our foundational systems. Innovate: Drive projects from concept to deployment with full ownership, and explore new tools, frameworks, and approaches to keep our infrastructure best-in-class. The Foundation team develops core software infrastructure (libraries, frameworks, and systems) for BlockTech, solving common problems and lending its expertise to enable other teams to stay focused on their respective domains. They own, develop, and configure a wide variety of critical, high-reliability software, ranging from low-level ultra-low-latency shared memory IPC libraries to high-throughput data buses and data aggregation systems including Kafka, NATS, PostgreSQL, and Iceberg. They work primarily in Rust, but also use Python and SQL. If you thrive on low-level problem-solving, building robust frameworks from scratch, and enabling others to move faster, this role is for you. What We're Looking For Essential: 5+ years of experience as a Software Engineer, with a strong focus on systems-level optimisation and awareness of hardware constraints Proficiency
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Coinbase's Developer Infrastructure - Test team exists for one reason: every Coinbase engineer should get fast, reliable test signals so the company can test and ship faster. As AI accelerates the pace of code generation, test infrastructure is becoming a critical path for how quickly Coinbase delivers value to customers. As a Staff Software Engineer on the Platform team, you'll set the technical direction for how Coinbase tests and ships software, owning the systems that turn testing into a speed advantage instead of a bottleneck. What you'll do: Define and own the technical strategy for test infrastructure across Coinbase engineering, with feedback speed as a core design constraint. Build and operate core test infrastructure services, including test orchestration, smart test selection, sharding, flaky-test detection, and test result analysis. Drive measurable improvements in test feedback speed and signal reliability so engineers can ship with confidence and without reruns. Own systems end to end, including architecture, observability, SLOs, and on-call operations. Partner with engineering teams across Coinbase to identify bottlenecks and turn them into platform improvements. Mentor engineers, raise technical standards, and shape how the organization approaches test infrastructure. Required Skills and Experience: 10+ years building and operating production so
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role Build AI-powered developer experience that makes engineers enabled. You’ll join our Developer Experience team, united by the mission of “Engineers can do their work quickly, easily, comfortably, and safely.” You’ll be part of a team spanning the US and India that turns ambiguous goals into shippable projects that meaningfully improve developer throughput and reliability. Your team owns AI experiences, remote and local dev environments, builds and CI, deploys, testing, and verification. You’ll work on a relatively broad scope within a well-resourced team. This role is based in Hyderabad, India. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve You build an AI harness to auto-resolve bugs reported for the Notion app. This is high-leverage work for the engineering org and requires solid engineering to improve resolution accuracy. You improve CI and build reliability so engineers get early feedback when changes break CI. We’re
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Software Engineer, Core Reliability on the Infra Reliability team within Platform , you'll help Coinbase scale 50x by improving reliability, security, and deployment safety across our production environment. This team owns the systems that secure service configurations and secrets, reduce customer-facing incidents, and strengthen deployment infrastructure supporting thousands of services and hundreds of daily releases. You'll lead high-impact reliability projects that make our entire service environment more resilient and safer for customers. What you'll do: Own the design and delivery of reliability projects and features that improve resiliency across Coinbase's service environment in partnership with other engineering teams. Partner with critical T0/T1 services to understand architecture, improve scalability, and reduce operational toil. Build and enhance systems that securely manage service configurations and secrets at scale. Improve canary-based release systems and expand deployment capabilities to support thousands of services and hundreds of daily deployments with fewer incidents. Drive reliability best practices and strengthen reliability culture across engineering teams at Coinbase. Required Skills and Experience: 5+ years of software engineering experience designing, building, and maintaining production services in service-oriented architectures
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Why Reliability? Roblox serves over 100 million people every day across a platform that is constantly evolving — and behind every experience is infrastructure that has to work, every time, at massive scale. The Reliability team at Roblox operates at the depth and breadth of the Roblox stack. Availability of the platform is a key company goal. We are hiring our first Senior Machine Learning engineer within our team. As a Senior Machine Learning Engineer within Reliability, you will help set the direction for how machine learning systems/practices can be leveraged to improve the reliability of the overall Roblox platform. You will own the architectural and execution roadmap of leveraging massive data across - logs, traces, metrics, production changes, to proactively detect issues before they become real problems (MTTD) and/or reduce time to resolve incidents (MTTR). You will have the opportunity to cross functionally collaborate with other similar teams at Roblox to define best practices and software. You will: Help define the roadmap for leveraging Machine Learning Engineering to improve Production Systems Reliability at Roblox. Improve r
At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. The Collections team, as part of the Repayments area, is on a mission to build a robust platform that will maximize the recovery of delinquent users by sending the right message to the right people at the right time, while monitoring and detecting problems and ensuring high reliability of the engineering systems. Collaborating closely with our product managers, backbook risk teams and other engineering teams , you will effectively manage loans throughout the delinquency phase of their lifecycle and develop and implement recovery strategies. We are looking for a highly motivated software engineer to build the next generation platform solutions that will allow us to manage our constantly growing volume while maintaining high availability and reliability of the systems. You will work closely with your team mates in the Collections team and Product to build robust collections solutions which will enable us to keep our loans portfolio healthy and help our customers pay on time and recover from any delays. What You'll Do With the support of your team, you will work on tasks that contribute to the team's projects and goals. You will work collaboratively and proactively with your team and stakeholders, bringing them along for your work and helping to create visibility and dialog regarding the risks and trade-offs related to your work. You will strike the right balance of speed and quality in your work, ensuring that we hit our business goals while protecting our systems from downtime. You will contribute to a sense of community on your team by engaging in growth and development activities What We Look For You have previous work or internship experience designing, developing and launching backend systems at scale and are experienced using one of Python or Kotlin. You are familiar with the building b
At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. The Collections team, as part of the Repayments area, is on a mission to build a robust platform that will maximize the recovery of delinquent users by sending the right message to the right people at the right time, while monitoring and detecting problems and ensuring high reliability of the engineering systems. Collaborating closely with our product managers, backbook risk teams and other engineering teams , you will effectively manage loans throughout the delinquency phase of their lifecycle and develop and implement recovery strategies. We are looking for a highly motivated software engineer to build the next generation platform solutions that will allow us to manage our constantly growing volume while maintaining high availability and reliability of the systems. You will work closely with your team mates in the Collections team and Product to build robust collections solutions which will enable us to keep our loans portfolio healthy and help our customers pay on time and recover from any delays. What You'll Do · With the support of your team’s tech lead and manager, you will break down larger projects into individual tasks, deliver them in multiple phases, and collaborate with others to ensure timely delivery of your work. · You will support your peers and stakeholders in the product development lifecycle by collaborating with product management, design & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. · You will support the operations and availability of your team’s artifacts by creating and monitoring metrics, escalating when needed, and supporting “keep the lights on” & on-call efforts. · You will contribute to a sense of community on your team by engaging in growth and devel
At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. The Collections team, as part of the Repayments area, is on a mission to build a robust platform that will maximize the recovery of delinquent users by sending the right message to the right people at the right time, while monitoring and detecting problems and ensuring high reliability of the engineering systems. Collaborating closely with our product managers, backbook risk teams and other engineering teams , you will effectively manage loans throughout the delinquency phase of their lifecycle and develop and implement recovery strategies. We are looking for a highly motivated software engineer to build the next generation platform solutions that will allow us to manage our constantly growing volume while maintaining high availability and reliability of the systems. You will work closely with your team mates in the Collections team and Product to build robust collections solutions which will enable us to keep our loans portfolio healthy and help our customers pay on time and recover from any delays. What You'll Do With the support of your team, you will work on tasks that contribute to the team's projects and goals. On-Call Rotation - There would be an on-call rotation for this role as a requirement You will work collaboratively and proactively with your team and stakeholders, bringing them along for your work and helping to create visibility and dialog regarding the risks and trade-offs related to your work. You will strike the right balance of speed and quality in your work, ensuring that we hit our business goals while protecting our systems from downtime. You will contribute to a sense of community on your team by engaging in growth and development activities What We Look For You have previous work or internship experience designing, developing and launching backend systems at scale an
At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. The Collections team, as part of the Repayments area, is on a mission to build a robust platform that will maximize the recovery of delinquent users by sending the right message to the right people at the right time, while monitoring and detecting problems and ensuring high reliability of the engineering systems. Collaborating closely with our product managers, backbook risk teams and other engineering teams , you will effectively manage loans throughout the delinquency phase of their lifecycle and develop and implement recovery strategies. We are looking for a highly motivated software engineer to build the next generation platform solutions that will allow us to manage our constantly growing volume while maintaining high availability and reliability of the systems. You will work closely with your team mates in the Collections team and Product to build robust collections solutions which will enable us to keep our loans portfolio healthy and help our customers pay on time and recover from any delays. What You'll Do · With the support of your team’s tech lead and manager, you will break down larger projects into individual tasks, deliver them in multiple phases, and collaborate with others to ensure timely delivery of your work. · You will support your peers and stakeholders in the product development lifecycle by collaborating with product management, design & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. · You will support the operations and availability of your team’s artifacts by creating and monitoring metrics, escalating when needed, and supporting “keep the lights on” & on-call efforts. · On-Call Rotation - There would be an on-call rotation for this role as a requirement · Y
Get new software reliability engineer jobs by email
Daily job updates · Unsubscribe anytime