Jobiba hiring network

Software Reliability Engineer Jobs

6,326 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

P
Pendo
📍 New York• Full-time• $300K – $325K/yr
1mo ago

The Team + The Role Our Emerging Team is focused on building AI Products for our product experience (PX) platform. We build from the ground up to explore, prototype, and ship AI-native experiences that change how software teams understand and serve their users. This is not an AI layer added to existing product; it is a deliberate bet on what product intelligence looks like next. The team operates with high autonomy, moves quickly, and builds products without clear precedents. As a Staff Software Engineer (AI), you will sit at the intersection of deep technical capability and strong product judgment. You will design and build production-grade AI systems, including RAG pipelines, agentic workflows, and LLM-powered features, while making clear tradeoffs across prompting, fine-tuning, architecture, evaluation, and deployment. You will also partner closely with product, design, and engineering stakeholders to frame the right problems and communicate technical decisions clearly. This role is based in our New York office. What this looks like day-to-day Applied AI systems: Design and build AI-native systems, including RAG pipelines, agentic workflows, and LLM-powered product features. You will take ideas from prototype through production and ensure they can support real users. Model strategy: Make principled decisions about when to prompt, when to fine-tune, and when to use a different technical approach entirely. You will explain those tradeoffs clearly to engineers and non-engineers. Evaluation and guardrails: Instrument and evaluate model outputs rigorously by defining evaluation frameworks and identifying hallucinations early. You will implement guardrails that hold up under real-world usage and load. Productionize AI ownership: Own model deployment, monitoring, latency optimization, cost management, and reliability at scale. You will ensure AI systems are observable, performant, and production-ready. Full-stack delivery: Contribute across the stack when needed to get

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. We are looking for a skilled and motivated Principal Software Engineer who is passionate about continuous learning and eager to grow along with us in a fast-paced, innovative environment. You will work remotely from Bulgaria and will be reporting to an engineering leader located in Bulgaria. You Will: Lead the design and implementation of Smartsheet's next-generation architecture, ensuring scalability, security, and performance for millions of global users. Define and drive architecture strategy, making key technical decisions that shape the future of the platform. Review and guide technical project designs, providing feedback during design review presentations to ensure system resilience and scalability. Take ownership of cross-functional technical initiatives, aligning teams around common architectural goals while driving large-scale projects to completion. Foster strong technical leadership, mentoring senior engineers and influencing best practices across multiple engineering teams. Lead deployment reviews for high-impact projects, ensuring they meet scalability, performance, and security requirements. Collaborate closely with product management and other business stakeholders to balance market needs with technical constraints, driving innovation while maintaining technical rigor. Advocate for quality and operational excellence, ensuring systems are monitored, tested, and maintained to meet the highest reliability standards. Perform other duties as assigned. You Have: Proven experience in system architecture and the d

REMOTEpythonjavasql
View job →
S
Smartsheet
📍 -REMOTE, USA-• Full-time• Remote
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. The App Core team is looking for a curious, growth-minded Software Engineer II to work on the core infrastructure of Smartsheet, focusing on the reliability, stability, and scale of our foundational systems. This role is a perfect fit for a developer who has the software engineering basics down and wants to fast-track their skills at the intersection of generalist backend engineering and Infrastructure as Code (IaC). This team practices mob programming for the majority of the workday. We operate through collective code ownership, meaning you'll be collaborating in real-time with teammates most of the time—not working solo on isolated tasks. We focus on building high-quality software while adhering to rigorous operational best practices. If you are passionate about continuous learning and ready to dive into the core foundation of our platform, we want to hear from you. You Will: Contribute to Core Infrastructure: Help build and maintain the core infrastructure that serves as the backbone for Smartsheet. Contribute to a robust environment that ensures the foundational reliability, stability, and performance expected by all of our users. Collaborate via Mob Programming: Work closely with the team every day in a re

REMOTEpythonjavasql
View job →
S
Smartsheet
📍 Bellevue• Full-time• From $1.7M/yr
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Corporate Systems Engineering builds and operates the software platforms, integrations, and automations that power Smartsheet’s core business functions across Finance, Sales/GTM, and People & Culture. Our team owns mission-critical systems and workflows that enable how the company hires, sells, bills, pays, reports, and scales. We operate at the intersection of software engineering, enterprise platforms, and business-critical data, treating internal systems with the same rigor, reliability, and product mindset as customer-facing software. The Automation team builds human-to-system and system-to-system automations that reduce manual effort and friction across the business. We combine cloud-native services, agentic AI, and workflow orchestration to enable employees to interact with enterprise systems through intelligent, secure, and auditable automation. As a Senior Software Engineer I (Automation), you will lead the design, build, and operation of systems and workflows that directly support business execution at scale. You will own complex technical initiatives, partner with Product Managers and stakeholders on technical roadmaps, and mentor junior engineers. This full-time position reports to the Sr. Director, Development and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Architect AI Agents: Take a leading role in designing Agentic Workflows using AWS Step Functions and Bedrock Agents that reason

javascripttypescriptpython
View job →
S
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Corporate Systems Engineering builds and operates the software platforms, integrations, and automations that power Smartsheet’s core business functions across Finance, Sales/GTM, and People & Culture. Our team owns mission-critical systems and workflows that enable how the company hires, sells, bills, pays, reports, and scales. We operate at the intersection of software engineering, enterprise platforms, and business-critical data, treating internal systems with the same rigor, reliability, and product mindset as customer-facing software. The Finance Systems team engineers and operates the platforms that support financial operations, including ERP, procurement, billing, and compliance. We work across configuration, extensibility, and integration to ensure systems are scalable, auditable, and resilient, treating code, configurations, and controls with the same rigor as software. As a Senior Software Engineer I (Finance Systems), you will lead the design, build, and operation of systems and workflows that directly support business execution at scale. You will own complex technical initiatives, partner with Product Managers and stakeholders on technical roadmaps, and mentor junior engineers. You will report into a Manager, Enterprise Systems, and can be based in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Systems Architecture & Optimization: Engineer and lead the end-to-end lifecycle—analysis, prioritization, and tec

REMOTEvueawsrest
View job →
S
Sofi
📍 San Francisco• Full-time
1mo ago

Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. The Role: We are seeking a highly skilled and experienced Staff Software Engineer to join our Test Platform team. In this role, you will have the opportunity to directly impact the design and architecture of our Developer Platform and tooling that enables SoFi engineers to create and deliver high-quality solutions. You will collaborate and partner with a curious team of engineers to design and deliver solutions that raise the testing and reliability standards for our backend and web applications. This new team will be focused on building a green field project to enable autonomous testing for our AI driven SDLC, leveraging cutting edge technologies to deliver a highly resilient and thorough platform. If you are a seasoned Staff Software Engineer with a passion for building products that just work and enabling developers to build reliable services, and a strong background in distributed systems, we invite you to apply for this exciting and new opportunity. What You’ll Do: Provide technical leadership for initiatives in Testing and Reliability, with a focus on integrating AI-driven automation and autonomous testing practices. Collaborate with product engineering teams to understand requirements and design platform capabilities that are efficient, robust, and developer-friendly. Architect and implement solution

pythonjavaaws
View job →
R
Roblox
📍 San Mateo• Full-time• From $153.1K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. We're looking for an outstanding engineer to join the Roblox Datasets team. We develop and maintain highly leveraged core datasets, frameworks, and tooling that support the growing demand for analytics across the company. Our team sits within Foundation AI, and we partner closely with Product, Engineering, Analytics, and Data Science. We help these teams move faster by making core data more reliable, better modeled, easier to discover, and easier to use. This team sits at an important intersection between platform and product. We partner broadly across major Roblox pillars, including Infra, Economy, Creator, Engine, Ads, Discovery, Growth, Compliance, Safety, Apps, and Social. That means you will have broad exposure to both technical and business problems. You Will: Build and improve the data pipelines, datasets, and internal tools that power decision-making across Roblox Partner with engineers, data scientists, and product teams to turn messy, ambiguous data needs into reliable, reusable data assets Help improve data quality, observability, reliability, and discoverability across the stack Contribute to foundational work that supports multiple product pillars, not just a single surface are

pythonsqlaws
View job →
T
Twilio
📍 - Ireland• Full-time• Remote
1mo ago

Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Senior Software Engineer. About the job This position is needed to design, build, and optimize the core signalling infrastructure that powers real-time video communications for our customers. You will play a key role in ensuring high performance, reliability, and scalability of our video platform, enabling seamless and secure video experiences. Responsibilities In this role, you’ll: Design, implement, and maintain video signalling protocols and server components for real-time video calls (e.g., WebRTC, SIP, RTCP/RTP) in a highly scalable distributed system. Collaborate with cross-functional distributed teams and various stakeholders to deliver high-performance, low-latency media experiences. Ensure secure transmission and compliance with industry best practices (e.g., end-to-end encryption, privacy standards). Contribute to architectural decisions and code reviews, mentoring junior engineers as needed. Stay current with

REMOTEjavaawsazure
View job →
T
Twilio
📍 - US• Full-time• Remote• $138.7K – $173.4K/yr
1mo ago

Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Software Engineer, Platform Engineering (L3) About the job This position is a critical engineering role within Twilio Platform Engineering, requiring a hands-on engineer capable of developing, deploying, and managing highly available, massive-scale distributed systems. Our systems regularly process more than 12 billion emails during peak events like Black Friday, and our throughput requirements continue to scale rapidly. As an L3 engineer, you will build and operate resilient backend services at scale and contribute to the design and reliability of our dual-cloud infrastructure span across Amazon Web Services (AWS) and Microsoft Azure. You'll run Kubernetes beyond the boundaries of managed services, automate infrastructure with Terraform, and write production code to help keep distributed systems healthy under real production load while using modern AI-assisted tooling to move faster. Responsibilities In this role

REMOTEawsazuredocker
View job →
T
Twilio
📍 - India• Full-time• Remote
1mo ago

Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as our next Staff engineer (L4), Twilio’s Segment team. About the job As a Staff Engineer on the Twilio Segment Data platform/ pipelines team, you’ll build and scale systems that process several hundred thousands of data points per second. You will lead the development of high-scale ingestion and data processing systems You'll be designing, operating and maintaining complex distributed systems, ensuring reliability, performance, and cost-efficiency while querying petabytes of data for our customer data platform (CDP). Responsibilities In this role, you’ll: Design and deliver robust, high-scale routing experiences for the Data platform/ pipelines team for Twilio Segment. Ship features that opt for high availability and throughput with eventual consistency Collaborate with engineering and product leads, as well as teams across Twilio Segment Support the reliability and security of the platform Build and optimize globally available and high

REMOTEjavaawsgcp
View job →
T
Twilio
📍 - US• Full-time• Remote• $138.7K – $173.4K/yr
1mo ago

Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Software Engineer, Email Platform (L3) About the job This position is a critical engineering role within Twilio SendGrid, requiring a hands-on engineer capable of developing, deploying, and managing highly available, massive-scale distributed systems. Our systems regularly process more than 12 billion emails during peak events like Black Friday, and our throughput requirements continue to scale rapidly. As an L3 engineer, you will act as a key driver of execution within our core services. You will be heavily involved in modernizing our backend systems, optimizing our extensive Go-based microservices, and contributing to the design and reliability of our dual-cloud infrastructure span across Amazon Web Services (AWS) and Microsoft Azure. Responsibilities In this role, you’ll: WEAR THE CUSTOMER’S SHOES: Architect and ship reliable, high-velocity features that handle critical traffic with low end-to-end latency. Part

REMOTEpythonjavasql
View job →
A
Asana
📍 New York City• Full-time• $248K – $282K/yr
1mo ago

We’re looking for a Staff Software Engineer to integrate and improve our AI developer experience so that engineers at Asana can use AI to increase their velocity. As part of the AI Developer Productivity team, you’ll set technical direction and build the next generation of AI-powered developer tools across editors, IDEs, CLIs, code review, and cloud and local coding agents. This role is based in our New York City office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Design and build AI-augmented workflows and systems that help coding agents understand the codebase, follow engineering practices, and enable Asana engineers to complete software development tasks faster and with more confidence. Design and scale autonomous cloud agents that take on complex, multi-step engineering tasks to reduce toil and enable engineers to focus on higher-leverage work. Build and refine IDE, editor, and CLI integrations that make AI-assisted development feel intuitive in engineers' daily work. Create reusable agent skills, tools, context, and integrations that teams across Asana can build on rather than reinvent. Improve AI-assisted code review workflows so engineers get faster, higher-quality feedback before and during review. Drive adoption of AI developer tools across engineering through usability improvements, measurement, documentation, and enablement. Set technical direction for the team, balancing experimentation with reliability, maintainability, and long-term platform thinking. Partner cross-functionally with teams across the engineering organization to understand engineering needs, identify workflow friction, and scale high-impact solutions

D
Datadog
📍 New York• Full-time• From $244K/yr
1mo ago

Role Summary: Datadog is seeking a Staff Software Engineer to help shape the future of our Bring Your Own Cloud (BYOC) Logs offering by unifying observability pipelines with log management software that customers deploy and manage in their own infrastructure. This role will focus on building and scaling systems that process, route, and store high-volume observability data within customer-managed infrastructure. You will operate as a hands-on technical leader, driving architecture, cross-team delivery, and product direction across a complex and evolving space. This is a high-impact opportunity to influence product strategy, mentor engineers, and solve deeply technical challenges at scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Make customer-controlled deployments feel like a managed Datadog product: deployment, upgrades, configuration, observability, diagnostics, reliability, and secure operation across diverse customer cloud environments Build and scale high-throughput systems for log processing, routing, and transformation across distributed environments Lead cross-team initiatives, aligning engineers, product managers, and stakeholders to deliver complex, multi-team projects Design and implement software that runs reliably that customers deploy and operate within their own cloud infrastructure. Improve system performance, scalability, and cost efficiency through thoughtful trade-off analysis and capacity planning Contribute hands-on to critical code paths, debugging, and deployment challenges in customer environments Who You Are: You have significant experience building software that is installed, deployed, and operated in customer environments rather than only as a fully managed SaaS service. You have strong expertise in distributed systems,

awsazuregcp
View job →
D
Datadog
📍 Portugal• Full-time• Remote
1mo ago

This role is part of Datadog’s Security Agent team, which powers critical security capabilities across Workload Protection, Vulnerability Management, Cloud Security products, and other emerging security offerings. As a Staff Software Engineer, you will lead the design and development of low-level Linux instrumentation and runtime security technologies that help customers detect threats, monitor system activity, and protect cloud-native workloads at scale. You will work on complex technical challenges involving eBPF, Linux kernel internals, performance-sensitive systems, and large-scale data collection while influencing technical direction across multiple product teams. This role offers significant ownership, broad organizational impact, and the opportunity to shape the future of Datadog’s security platform. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the architecture and development of security agent capabilities that power runtime threat detection and workload protection across Datadog Security products. Design and build reusable eBPF-based monitoring functionality for process, file, and network visibility within Linux environments. Drive end-to-end delivery of new features, from technical strategy and design through implementation, testing, and rollout. Establish and evolve testing methodologies that improve platform coverage, detection quality, reliability, and performance. Partner with product, security, infrastructure, and engineering teams to deliver shared platform capabilities used across multiple Datadog products. Provide technical leadership by influencing engineering direction, mentoring peers, and helping resolve complex cross-functional challenges. Who You Are: You have significant experience building software in Linux environments,

REMOTElinuxaigo
View job →
D
Datadog
📍 Spain; Paris, France• Full-time
1mo ago

We are looking for a Senior Software Engineer to help us take REDAPL, our Referential Data Platform, to the next level. REDAPL is Datadog's main platform for tracking our customers' infrastructure resources and relationships. The platform enables products where customers can understand, keep track of, and gain insights into their infrastructure related to performance, cost, security, and more. Many Datadog products use REDAPL today such Cloud Security Posture Management, Resource Catalog, Cloud Cost Management, and Service Catalog and others - REDAPL ingests more than 4.5mil updates/second. As a Senior Engineer, you will drive, lead and collaborate on projects both inside and outside the platform. You can expect to contribute to key technical decisions relating to our data ingestion, processing, and query pipelines. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Build a query engine that supports efficient relationship traversals for our most demanding workloads. Contribute to design and drive high-priority, high-visibility projects to increase the platform's value, resilience, and scalability across multiple teams. Lead and guide other engineers through architectural platform decisions Identify potential system risks and trends in reliability and design solutions to address them Provide input on prioritizing engineering-led initiatives in short- and long-term planning and roadmaps Collaborate with internal product teams to understand their requirements and how we plan for their product growth as they integrate and depend on REDAPL Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering or related scientific field or equivalent experience You have worked extensively with multiple types of data stores You have contributed to in

aigorust
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime