Jobiba hiring network

Senior Software Reliability Engineer Jobs

7,101 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

A
Asana
📍 New York• Full-time• $202K – $223K/yr
1mo ago

The Product team builds full-stack features end-to-end. From designing our data models to implementing the subtle interaction behaviors that differentiate good software from great software. We work closely with UI designers and are supported by our infrastructure team. We aim to delight users with both large new features and smaller, daily product enhancements—thanks to our continuous deployment architecture. We want to create a superlative user experience, down to the smallest details. This role is based in our New York City office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve You will work full-stack to develop new features used by millions of Asana users Participate in every step of the product development process with an open and curious team Develop clean, beautiful code and leave it better than you found it Experience growth and development by being paired with a mentor who will support and guide you through opportunities to stretch and learn About you 5+ years of experience working within large, well maintained codebases Experience using Typescript and React Excellent communication skills for collaborating with other teams Sound judgment when balancing moving quickly with producing quality code and long-term code maintainability Passionate about creating a superlative user experience and attentive to details Appreciate productivity and care deeply about helping teams collaborate more effectively and efficiently, including your own Demonstrates curiosity about AI tools and emerging technologies, with a willingness to learn and leverage them to enhance productivity, collaboration, or decision-making At Asana, we're committed

typescriptreactrest
View job →

The Incident Insights & Readiness SRE team at Datadog fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. Our users are Datadog engineers, and we build the software, tooling, and operational frameworks that help them prepare for, respond to, and learn from incidents. We work closely with engineering teams across Datadog to analyze incidents and turn those insights into better tools, stronger incident response, and organizational learning. Our efforts empower Datadog to navigate unexpected failures confidently, efficiently, and with a commitment to continuous learning and systems improvement. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. The Incident Insights & Readiness SRE team at Datadog fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. Our users are Datadog engineers, and we build the software, tooling, and operational frameworks that help them prepare for, respond to, and learn from incidents. We work closely with engineering teams across Datadog to analyze incidents and turn those insights into better tools, stronger incident response, and organizational learning. Our efforts empower Datadog to navigate unexpected failures confidently, efficiently, and with a commitment to continuous learning and systems improvement. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own and improve the on-call experience for the company by establishing best practices and building platforms to support on-call rotations and compensation

typescriptpythonkubernetes
View job →

About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale with trillions of data points per day, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams for tens of thousands of companies globally. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team The Datadog Security Libraries team owns the customer-side integrations behind our run-time security products App & API Protection , Workload Protection , and Code Security . Our libraries let customers automatically manage application security risk with continuous, real-time monitoring of vulnerabilities and threats against their web applications, serverless applications, and APIs, in production. Automatically integrated with Application Performance Monitoring (APM) distributed tracing and code-level context, our software empowers development, operations, and security teams to build and run secure applications. As a polyglot team we ship and maintain the security capabilities of Datadog's tracing libraries across .NET , Java , Go , Node.js , Python , Ruby , and PHP , on top of a shared C++ core and a set of HTTP proxy integrations (primarily Envoy, NGINX, and HAProxy). Our code runs inside thousands of production applications around the world. Recent work spans exploit prevention (RASP) and WAF detections, API Security, code security (IAST and SCA), and AI-assisted ("agentic") onboarding, always measured by real product outcomes and operational telemetry. The Opportunity We're looking for a senior, polyglot engineer to contribute across several of our security libraries, with .NET or Java expertise. You'll design and build security integrations and detection features, take them from prototype to production-hardened, and own them operationally as they instrument thousands of applications. As a se

pythonjavanode.js
View job →

About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale with trillions of data points per day, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams for tens of thousands of companies globally. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team The Datadog Security Libraries team owns the customer-side integrations behind our run-time security products App & API Protection , Workload Protection , and Code Security . Our libraries let customers automatically manage application security risk with continuous, real-time monitoring of vulnerabilities and threats against their web applications, serverless applications, and APIs, in production. Automatically integrated with Application Performance Monitoring (APM) distributed tracing and code-level context, our software empowers development, operations, and security teams to build and run secure applications. As a polyglot team we ship and maintain the security capabilities of Datadog's tracing libraries across .NET , Java , Go , Node.js , Python , Ruby , and PHP , on top of a shared C++ core and a set of HTTP proxy integrations (primarily Envoy, NGINX, and HAProxy). Our code runs inside thousands of production applications around the world. Recent work spans exploit prevention (RASP) and WAF detections, API Security, code security (IAST and SCA), and AI-assisted ("agentic") onboarding, always measured by real product outcomes and operational telemetry. The Opportunity We're looking for a senior, polyglot engineer to contribute across several of our security libraries, with .NET or Java expertise. You'll design and build security integrations and detection features, take them from prototype to production-hardened, and own them operationally as they instrument thousands of applications. As a se

pythonjavanode.js
View job →
D
Datadog
📍 Tel Aviv• Full-time
1mo ago

The eBPF APM team is building a zero-instrumentation observability solution that automatically discovers services on every host, supports both plaintext and TLS-encrypted traffic, classifies Layer 7 protocols, decodes service-level traffic, and reports RED (requests, errors, duration) metrics. Leveraging deep expertise in eBPF, the team operates across a wide range of Linux kernel versions, distributions, and complex customer environments. In addition to low-level networking, the team solves challenges related to protocol versioning, TLS detection across diverse languages and runtimes, and resilient performance in production systems We’re looking for a senior engineer with strong systems-level thinking and a good understanding of Linux. You should be comfortable working close to the kernel, ideally with experience in eBPF, or with a strong desire to dive into it. Proficiency in C/C++/ Go is essential, and familiarity with networking protocols, TLS internals, or distributed tracing is a strong advantage. You’ll join a high-impact team tackling ambitious technical challenges—like decoding traffic across multiple protocols, and ensuring high-fidelity metrics in complex, real-world environments. You’ll be expected to lead design and implementation efforts, contribute to roadmap planning, and collaborate across teams to ensure our solution remains robust, scalable, and frictionless for our users. This role is a great fit for engineers who thrive on low-level, performance-sensitive problems, and want to shape the future of observability through cutting-edge kernel technology. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Design and build core components of our zero-instrumentation APM product using eBPF and Go Develop systems to aut

linuxrestai
View job →
D
Datadog
📍 Massachusetts• Full-time• From $192K/yr
1mo ago

About Datadog: We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. You will: Solve a scaling bottleneck in a critical service Deploy a new feature to production, progressively rolling it out with feature flags Investigate and fix a production issue from a service your team owns Design a way to scale up a service for more traffic With your team, plan the most important projects to work on next Who You Are: You have significant experience in one or more languages You value code simplicity and performance You can design architecture to solve problems at high scale You have a BS/MS/PhD in a scientific field or equivalent experience You want to work in a fast, high-growth startup environment that respects its engineers and customers You have demonstrated ability to use AI coding tools in day-to-day workflows and build, validate, and refine AI-generated output in products You can design AI Backend systems, with awareness of quality, cost, and latency tradeoffs 6+ years of experience Bonus points: You've worked at high scale with systems like Redis, Cassandra, Kafka You wrote your own data pipelines once or twice before You have a strong background in statistics You have significant experience with Go, C, or Python You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how You’re motivated to push the boundaries of how AI can improve software engineering best practices and contribute to building AI-enabled products Datadog values people from all walks of life. We understand not everyone will meet all the above qualificat

pythonredisai
View job →
D
Datadog
📍 Georgia• Full-time• From $244K/yr
1mo ago

The Language Tools team enables ~1,500 Datadog developers to build, test, and package millions of lines of Go, Python, Java, Rust, and TypeScript in our backend monorepo. Our success is measured by their productivity and satisfaction. They use the tools that we develop and support several times a day, in both development and CI environments. We use the Bazel open source build system as a foundation. The team is growing rapidly, both with Datadog and as we absorb other repositories into the monorepo. As a senior software engineer on the team, you will own projects from start to finish, both greenfield and brownfield. You will gain first-hand understanding of what Datadog developers need, and inform our roadmap. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Invent build, test and packaging tools that are simpler and more reliable to use. Push performance and cost efficiency at scale, raising cache hit rates and cutting CI times and compute spend across millions of targets. Treat CI like SREs treat prod, making sure our pipelines are green and fast. Prepare, run, and finish complex migrations. Contribute back to the Bazel ecosystem, upstreaming fixes and shaping features we depend on. Who You Are: An expert in Bazel and/or one of the languages listed above. A well-rounded engineer. You must broadly understand the various types of software projects that are built, tested, and packaged with our tools. Both careful and fearless. The changes we make impact the velocity of hundreds of engineers. They are risky but necessary. User-focused. We help Datadog engineers to use the tools that we develop, and continuously improve their usability, so they don’t need our help the next time. Ideally, you have ex

typescriptpythonjava
View job →
D
Datadog
📍 Massachusetts• Full-time• From $192K/yr
1mo ago

Distributed Systems engineers at Datadog design, implement and run in production the foundational platforms powering our applications. Your data pipelines will ingest, store, analyze and query in real-time billions of events per second from companies all over the globe. The platforms are optimized for durability, high availability, low latency, internet-scale footprint and operability. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Build fault-tolerant, horizontally scalable solutions running in multi-tenant environments Write in Go, Java Rust or C++, amongst other languages Use Kafka, Redis, Cassandra, Elasticsearch and other open-source components Own meaningful parts of our service, have an impact, grow with the company Who You Are: 6+ years of experience You have a BS/MS/PhD in a scientific field or equivalent experience You have significant backend programming experience in one or more languages (Go, Java, Rust, C++) You have been exposed to working on problems (high durability / low latency /…) You can get down to the low-level when needed You care about simple designs and performance You want to work in a fast, high-growth startup environment that respects its engineers and customers You have demonstrated ability to use AI coding tools in day-to-day workflows and validate, critique, and refine AI-generated output. Bonus: you’re motivated to push the boundaries of how AI can improve software engineering best practices and contribute to building AI-enabled products. This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. Datadog values peo

javaredisai
View job →
D
1mo ago

We are building the best platform in the world for engineers to understand, scale, and protect their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, application tracing, and security insights for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way . At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Solve a scaling bottleneck in a critical service Deploy a new feature to production, progressively rolling it out with feature flags Investigate and fix a production issue from a service your team owns Design a way to scale up a service for more traffic With your team, plan the most important projects to work on next Who You Are: You have significant experience in one or more languages You value code simplicity and performance You can design architecture to solve problems at high scale You have a BS/MS/PhD in a scientific field or equivalent experience You want to work in a fast, high-growth startup environment that respects its engineers and customers You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how You have demonstrated ability to use AI coding tools in day-to-day workflows and build, validate, and refine AI-generated output in products You can design AI Backend systems, with awareness of quality, cost, and latency tradeoffs Bonus: You’re motivated to push the boundaries of how AI can improve software engineering best practices and contribute to building AI-enabled products Datadog values people from all walks of life. We understand not everyone will meet all

aigorust
View job →

Come join the Server Ingress Security team, where we are rearchitecting MongoDB Server’s ingress networking to make MongoDB clusters even more secure. This new team is building the Atlas Network Protection layer, a set of performant, security-critical services that harden MongoDB's pre-authentication attack surface and provides the ability to respond rapidly to emergent threats. We are looking for talented Senior Engineers to join the team and be founding members, where you will play a crucial role in our multi-year roadmap. Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies security and systems engineering fundamentals to protect a popular database at scale, join us! We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. Candidate Profile 5+ years of experience building production-quality systems software Experience with large backend/compiled codebases and performance-sensitive software, preferably in Rust Bonus points for experience working hands-on in security-sensitive or networking-adjacent domains Strong systems fundamentals, including multi-threaded programming and performance profiling. Bonus points for: Understanding of network protocols, TLS, and connection lifecycle management Familiarity with security concepts such as attack surface reduction, input validation, memory safety, and defense-in-depth architectures Excellent verbal and written technical communication skills, with a strong desire to collaborate with colleagues Strong time management skills and the ability to realistically assess project complexity B.Sc. in Computer Science or a related field, or equivalent practical experience, with strong competencies in data structures, algorithms, and software design/architecture. Interest in the theory and practice of high-availability, security-critical systems Position Expectations Design, implement, and operate production

mongodbawsazure
View job →
M
Mongodb
📍 San Francisco• Full-time• From $126K/yr
1mo ago

Join the Atlas Search Query team to design and develop the next generation of Search query architecture, optimization, and execution. Atlas Search is a growing cloud service that allows users to execute complex search and vector search queries using the MongoDB Query Language. Our users can focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team is building a cloud-based distributed system responsible for the core components of search including data ingestion, performance, query language, query execution, for both relevance-based search and vector search. Our product is being adopted quickly and there are many interesting projects. This is a technical role where you will be responsible for the success of complex Search Query feature development. We are looking to speak to candidates who are based in San Francisco, CA for our hybrid working model. What You’ll Do Lead complex projects across the MongoDB ecosystem, for instance, development of a new Search aggregation framework within the MongoDB aggregation framework. Set project level strategy, architect features, and lead projects to successful execution Identify, design, and implement features enhancing our query language, performance, and operability Perform code reviews with peers and make recommendations on how to improve our software development processes Influence and grow team members through active mentoring and leading by example What We Look For 5+ years experience in data management/search systems, ideally with a strong query processing and optimization background Experienced in the development and maintenance of stateful distributed systems Eager to shape the technological direction of a complex system and have the ability to lead initiatives through collaboration with others Experienced in debugging and profiling multithreaded applications written in Java and Rust Bonus: experience with designing high-volume query engines, such as a datab

javamongodbaws
View job →
M
Mongodb
📍 New York• Full-time• From $126K/yr
1mo ago

The MongoDB Cloud Services Team is a diverse group of contributors working together to help our users manage MongoDB at global scale. The Cloud Team is responsible for MongoDB Atlas: our database as a service offering, and fastest growing product, which allows users to deploy fault-tolerant, globally distributed MongoDB clusters in just minutes. The Backup Team delivers essential infrastructure to help our customers in their hour of need - providing the ability to quickly restore a massive, distributed database to any point in time at the click of a button. The Backup Team’s mission is to make MongoDB backup more reliable, faster, and also cheaper. This team is responsible for the Backup Agent (Go), the extensive server-side infrastructure (Java) which manages 100s of TB of data and processes billions of operations per day, and the user interface (Javascript) that customers use to manage their backups. Common project themes are performance, scaling, and ease of use. We are looking to speak to candidates who are based in New York for our hybrid working model. We're looking for someone who is Skilled at writing large-scale, distributed backend systems in a compiled language (Java, C#, Go, etc.) Fond of chasing down tough problems in a distributed systems environment Cool under pressure - has wrangled production crises, and secretly finds this a little fun Experienced with Linux, and able to correlate application performance problems with underlying hardware limits Comfortable working across the stack of a modern web application Always striving to expand their knowledge Curious, collaborative and intellectually honest Responsibilities Work closely with product teams, considering the user’s perspective while helping the team achieve success Collaborate with team members over best practices and core concepts Hold yourself accountable to your actions, maintaining the balance between accomplishing goals with research & development Own our

javascriptjavamongodb
View job →

MongoDB is looking for an outstanding person to join our newly created Forward Deployed Engineering team and take on a key role in our extended R&D organization. Forward Deployed Engineering is linking the work of teams engaged on application modernization roles with our Product and Engineering teams. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Many organizations have built up large estates of legacy applications. Lack of scalability and resilience, long development times, operating cost, and inability to run on cloud are common issues with these applications. To address these issues, organizations are engaging in large transformational Application Modernisation programs. MongoDB is recognized as the developer data platform of choice for transactional systems that provide the best scalability, resiliency and developer experience in the cloud as well as on premises. Organizations are continuously migrating workloads from these legacy applications to new platforms, often based on microservices, using MongoDB. Such transformations are time intensive and often risky. Tooling based on generative AI promises to accelerate these transformations in a way never seen before. Forward Deployed Engineering is responsible for exploring the possibilities of Generative AI technologies and providing invaluable feedback to MongoDB’s R&D teams to drive future capabilities of MongoDB and the MongoDB ecosystem. Application Modernization Engineers will work alongside project teams that are executing Application Modernisation projects with customers. The successful candidate will be responsible for evaluation, build, and applying tools in modernization projects, facilitating the usage of such tools and processes across the different project teams, identifying opportunities for tooling deployment, selecting potential 3rd party tools, contributing to the development of tooling prototypes and helping to shape the produ

javasqlpostgresql
View job →
M
1mo ago

The MongoDB Query Execution Team is hiring software engineers who want to join us in developing a high performing, reliable and modular distributed query system. Our engineers work on implementing and maintaining execution algorithms, building new query language features, tuning database performance, and more to power our customers' critical workloads. This role can be based out of our Dublin office or remotely in Ireland. Relocation can be supported. Position Expectations Understand and improve current functionality of the MongoDB query engine Contribute high quality C++ code and give and solicit feedback in code reviews Identify, design, implement, test, and support new features related to query performance and robustness, query language enhancements, diagnostics for query performance problems, and integration with other products and tools Work constructively with peers to deliver excellent technical solutions Candidate Profile 5+ years of experience in systems programming Experience in databases and/or data management systems is a huge plus, but not a requirement Hands-on experience building industrial-strength software Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Experience with large code bases, preferably in C++, C, Rust or a similar compiled language B.Sc in Computer Science or similar field, or equivalent practical experience Interest in the theory and practice of database query engines. Hands-on experience or M.Sc./Ph.D in the domain is a plus Success Measures In three months you’ll have contributed to the development of a project slated for the next major version, as well as fixed a few bugs in a minor version of our latest stable release series In six months, you’ll have taken on code review responsibilities and are independently delivering complex functionality and squashing bugs independently In twelve months, you’re leading the development of a new major feature and are h

mongodbawsazure
View job →
M
Mongodb
📍 New York• Full-time• From $126K/yr
1mo ago

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. We are looking for a talented Senior Software Engineer to join our team as we execute on a multi-year roadmap and prepare to launch and rapidly scale services that handle petabytes of data. Come do some of the most interesting work of your career as we Think Big and Go Far for our customers! This role can be based out of our New York City office (hybrid working model) or remotely in the North America region. What you’ll do Design, build, and operate control plane services powering an elastic and multi-tenant storage layer for thousands of database instances. Solve problems around maintaining high availability and performance during load spikes, hardware failures, cloud provider outages, and other disruptions. Contribute to a culture of operational excellence through dashboards, playbooks, and on-call improvements. Lead complex technical projects from planning through successful deployment with clear stakeholder updates. Partner closely with peers across database, cloud, and infrastructure engineering teams as well as project management to investigate incidents and develop long-term roadmaps. Mentor junior engineers and foster a collaborative team environment. We’re looking for someone with 5+ years of professional software development experience building, deploying, and operating multi-tenant cloud services with a focus on operational excellence. Experience with large backend/compiled codebases, such as Rust or C/C++. Experience with containerization and orchestration platforms (e.g. Kubernetes). Experience with observability tooling (e.g. time series metrics, dashboards). Experience with distributed systems

mongodbawsazure
View job →
🔔

Get new senior software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime