Clear all

Jobiba hiring network

Senior Distributed Systems Engineer Data Platform Analytics And Alerts Jobs

4,932 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current senior distributed systems engineer data platform analytics and alerts jobs. Use filters to narrow by work mode, employment type, experience and date posted.

C
Coinbase
📍 - USAFull-timeRemoteFrom $186.1K/yr
29 days ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Senior Software Engineer, AI Transformation You'll join a high-performing team of engineers driving AI transformation at Coinbase as a Senior Software Engineer on the IT Operations team within ESTO. This team builds custom products through full-stack engineering and scales the infrastructure powering Coinbase's AI products, with direct exposure to senior leadership in a fast-paced, incubator-style environment. You'll own the reliability and automation of critical AI infrastructure, ensuring our systems are resilient, observable, and secure at scale. What you'll do: Own end-to-end delivery of AI products by building production-grade distributed systems, including serving infrastructure, data pipelines, and deployment orchestration across the full stack throughout the SDLC. Drive platform adoption by designing clean APIs, abstractions, and developer-facing tooling that enable product teams to integrate AI capabilities without bespoke infrastructure requests. Partner with engineering and product leadership across Platform and other product groups to align infrastructure requirements, resolve cross-team technical dependencies, and define shared platform contracts. Shape engineering standards and technical culture by establishing architectural patterns, mentoring engineers, and raising the bar on code quality, observability, and operational excellence. Build f

REMOTEpythonawsdocker
View job →
N
Newrelic
📍 IndiaFull-time
28 days ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. The Infrastructure product organization develops New Relic infrastructure instrumentation agents, next generation data processing and management services, vulnerability management, and security testing capabilities for on-prem and cloud customers. We work with data at a scale using a diverse tech stack (Go, Java, JavaScript, React GraphQL, Kubernetes, many public cloud web services, and more). As a senior backend engineer, you will help us build and extend next generation solutions such as a control plane for customers to manage their data pipelines at scale. New Relic is looking for engineers who are interested in building a brand-new observability experience. This high-impact engineering position is a phenomenal opportunity to own and build a set of next generation services and capabilities for the company. We are searching for a motivated engineer who is ready for a career-defining role in their next opportunity. We look forward to talking with you! What you'll do ● Design, Build, maintain, and scale back-end services and their support tools. ● Participate in architectural definitions with a high degr

javascriptjavareact
View job →
P
Pagerduty
📍 LisbonFull-time
27 days ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. About the role PagerDuty’s Operations Cloud runs on a platform that ingests billions of signals and turns them into real-time action for thousands of customers. We’re looking for an early-career AI/ML Engineer who is excited to grow at the intersection of two disciplines: large-scale distributed systems and machine learning. In this role you will help build and ship AI systems that run in production at PagerDuty’s scale — powering Incident Management AI Agents, event intelligence, and the LLM-powered capabilities embedded across our platform. You’ll work alongside senior engineers on real production problems, learning how AI features go from a prototype to something that serves reliably at scale. We are looking for a candidate who is genuinely excited about building with modern AI — LLMs, agents, and retrieval — eager to learn how resilient, high-throughput systems are built, and motivated to grow into an engineer who is strong in both. What you’ll do Contribute to AI-powered features — LLM agents, retrieval, and event intelligence — that operate on high-volume, real-time data, with support and guidanc

awsazuregcp
View job →
M
29 days ago

The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our Toronto or Vancouver offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design primitives of at least one of AWS, Azur

mongodbawsazure
View job →
M
Mongodb
📍 United StatesFull-timeFrom $127K/yr
29 days ago

The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim

mongodbawsazure
View job →
P
Pinterest
📍 IlFull-timeRemoteFrom $1.7M/yr
29 days ago

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . The Production Engineering organization at Pinterest is accountable for ensuring overall Pinterest availability as well as enhancing Engineering teams' capability to design, build and operate robust systems at scale. Pinterest's applications and infrastructure handle billions of monthly page views and petabytes of data as Pinterest continues to grow and scale. As a Senior Production Engineer on Solutions Engineering, you will design and build AI agents, platforms, tools, frameworks and methodologies to assure the reliability of our large-scale distributed systems serving hundreds of millions of monthly active users, handling hundreds of thousands of requests per second, and managing tens of petabytes of data. You'll lead infrastructure modernization initiatives, build intelligent automation that eliminates operational toil and amplifies engineer

REMOTEpythonsqlmysql
View job →
G
Godaddy
📍 United StatesFull-timeFrom $154K/yr
29 days ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the world, delivering the object, block, and file storage platforms that power GoDaddy's hosting infrastructure, internal services, OpenStack environments, and next-generation AI/HPC workloads. If you're passionate about distributed systems, storage architecture, and solving failure scenarios at massive scale, this is an opportunity to work on infrastructure few engineers will experience in their careers. Ceph is a strategic platform at GoDaddy — not an ancillary service. Our global footprint includes 80+ production clusters, 20,000+ OSDs, 1,830 storage nodes, 300 PB of raw capacity, and 69 billion objects spanning five datacenters across three continents. The platform supports RBD, RGW (S3/Swift), and CephFS workloads through more than 1,550 pools, 574,000 placement groups, and 900+ MDS daemons, creating engineering challenges that demand deep expertise in storage architecture, data durability, performance optimization, automation, and observability. As a Lead Senior Site Reliability Engineer, you'll serve as one of the principal technical leaders for GoDaddy's Ceph platform. You'll design the next generation of storage clusters, lead major platform upgrades, drive capacity and hardware strategy, and establish the standards that govern how the platform scales. You'll be the engineer the team turns to for the most complex s

pythonkubernetesai
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Staff / Lead Software Engineer to architect, design, and scale our External Observability Platform . In this role, you will lead the technical strategy for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex cross-functional architecture initiatives, and skilled at mentoring senior engineers while holding your own with the brightest technical minds in the industry. Key Responsibilities Architect & Scale Distributed Infr

javavueaws
View job →
C
Coinbase
📍 - USAFull-timeRemoteFrom $218K/yr
29 days ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Staff Software Engineer, Core AI Infrastructure You'll join a high-performing team of engineers driving AI transformation at Coinbase as a Senior Software Engineer on the IT Operations team within ESTO. This team builds custom products through full-stack engineering and scales the infrastructure powering Coinbase's AI products, with direct exposure to senior leadership in a fast-paced, incubator-style environment. You'll own the reliability and automation of critical AI infrastructure, ensuring our systems are resilient, observable, and secure at scale. What you'll do: Own end-to-end delivery of AI products by building production-grade distributed systems, including serving infrastructure, data pipelines, and deployment orchestration across the full stack throughout the SDLC. Drive platform adoption by designing clean APIs, abstractions, and developer-facing tooling that enable product teams to integrate AI capabilities without bespoke infrastructure requests. Partner with engineering and product leadership across Platform and other product groups to align infrastructure requirements, resolve cross-team technical dependencies, and define shared platform contracts. Shape engineering standards and technical culture by establishing architectural patterns, mentoring engineers, and raising the bar on code quality, observability, and operational excellence. Build

REMOTEpythonawsdocker
View job →
N
Nuro
📍 Mountain ViewFull-timeFrom $193.9K/yr
29 days ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro takes a machine-learning-first approach to autonomous driving, and the ML Infrastructure team builds and operates the infrastructure that makes that possible. We own the systems that train the models at the core of the Nuro Driver™ - from distributed GPU training and closed-loop reinforcement learning, to the workflows, orchestration, observability, and cost management that keep the fleet running efficiently. Our work sits directly on the critical path of autonomy development. When a training run stalls, when a pipeline silently regresses, or when GPU utilization slips, it shows up in how fast the rest of the company can ship. We care as much about reliability and operational maturity as we do about raw scale. About the Work Contribute to Nuro’s training infrastructure, spanning multi-generation accelerators, and multi-cluster scheduling and orchestration. Design and operate large-scale data pipelines - batch and strea

pythongcpkubernetes
View job →
D
Datadog
📍 New YorkFull-timeFrom $272K/yr
29 days ago

About Datadog: We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Opportunity: Datadog’s Senior Staff Engineers are technical leaders operating at the forefront of large-scale systems design, building the infrastructure that will support our next five years of growth and beyond. They do this in three major ways: As individual contributors, they bring world-class technical depth to build industry-leading systems in areas such as observability data platforms, distributed query engines, and real-time event streaming at global scale. As technical leaders, they apply broad architectural perspective and deep systems thinking to align design decisions across teams and domains. They work across complex, multi-team problem spaces to define long-term technical direction, drive large-scale initiatives forward, and ensure consistent execution. As engineering stewards, they play a key role in evolving our systems and engineering culture. They actively participate in Datadog’s senior technical community, bringing external insights and internal experience to elevate engineering standards and mentor the next generation of technical leaders. Examples of projects a Senior Staff Engineer may lead include designing and launching a new distributed data storage engine capable of handling hundreds of millions of records per second, building the real-time infrastructure behind a new observability product, or re-architecting a core service to support exponential growth in throughput and complexity. What You’ll Do: Be the technical owner of multiple critical systems or architecture areas, often spanning several t

aigorust
View job →
D
Datadog
📍 France; Sophia Antipolis, FranceFull-time
29 days ago

About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale with trillions of data points per day, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams for tens of thousands of companies globally. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team The Datadog Security Libraries team owns the customer-side integrations behind our run-time security products App & API Protection , Workload Protection , and Code Security . Our libraries let customers automatically manage application security risk with continuous, real-time monitoring of vulnerabilities and threats against their web applications, serverless applications, and APIs, in production. Automatically integrated with Application Performance Monitoring (APM) distributed tracing and code-level context, our software empowers development, operations, and security teams to build and run secure applications. As a polyglot team we ship and maintain the security capabilities of Datadog's tracing libraries across .NET , Java , Go , Node.js , Python , Ruby , and PHP , on top of a shared C++ core and a set of HTTP proxy integrations (primarily Envoy, NGINX, and HAProxy). Our code runs inside thousands of production applications around the world. Recent work spans exploit prevention (RASP) and WAF detections, API Security, code security (IAST and SCA), and AI-assisted ("agentic") onboarding, always measured by real product outcomes and operational telemetry. The Opportunity We're looking for a senior, polyglot engineer to contribute across several of our security libraries, with .NET or Java expertise. You'll design and build security integrations and detection features, take them from prototype to production-hardened, and own them operationally as they instrument thousands of applications. As a se

pythonjavanode.js
View job →
D
29 days ago

About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale with trillions of data points per day, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams for tens of thousands of companies globally. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team The Datadog Security Libraries team owns the customer-side integrations behind our run-time security products App & API Protection , Workload Protection , and Code Security . Our libraries let customers automatically manage application security risk with continuous, real-time monitoring of vulnerabilities and threats against their web applications, serverless applications, and APIs, in production. Automatically integrated with Application Performance Monitoring (APM) distributed tracing and code-level context, our software empowers development, operations, and security teams to build and run secure applications. As a polyglot team we ship and maintain the security capabilities of Datadog's tracing libraries across .NET , Java , Go , Node.js , Python , Ruby , and PHP , on top of a shared C++ core and a set of HTTP proxy integrations (primarily Envoy, NGINX, and HAProxy). Our code runs inside thousands of production applications around the world. Recent work spans exploit prevention (RASP) and WAF detections, API Security, code security (IAST and SCA), and AI-assisted ("agentic") onboarding, always measured by real product outcomes and operational telemetry. The Opportunity We're looking for a senior, polyglot engineer to contribute across several of our security libraries, with .NET or Java expertise. You'll design and build security integrations and detection features, take them from prototype to production-hardened, and own them operationally as they instrument thousands of applications. As a se

pythonjavanode.js
View job →
G
29 days ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Senior Backend Engineer on the Plan: Spec-Driven Development team, you'll help build GitLab's intent-to-code loop: agentic workflows that turn a stated intent into a refined work item, an implementation plan, and verified, mergeable code. You'll own early-stage backend work across the flows, application programming interfaces, data models, and evaluation systems behind this experience, using Ruby on Rails, Python, PostgreSQL, GitLab Duo Agent Platform, large language model application programming interfaces, and an artificial intelligence gateway. You'll join a small, distributed team working on one

REMOTEpythonsqlpostgresql
View job →
P
Postman
📍 IndiaFull-time
1mo ago

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Search Team at Postman is responsible for enabling users to quickly find and get started with the APIs that they are looking for. Postman is growing at a rapid pace, and this manifests into an ever-increasing volume of data that users create and consume, within their teams and in the Public API Network. We focus on improving discovery and ease of consumption over this data. We are looking for a Senior Engineer with 6+ years of experience deep backend expertise on search and ETL systems and a strong product mindset, to lead core initiatives on our search platform. In this role, you'll work at the intersection of infrastructure, relevance, and developer experience—designing systems that power search across the platform. You’ll bring a bias for action, a strong backend foundation, and the curiosity to explore beyond traditional boundaries, including areas like high performance web services, high volume data pipelines, machine learning, and relevance tuning. What You'll Do Own end to end architecture and roadmap of search platform consisting of distributed indexing pipelines, storage infra and high performance web serve

javascriptpythonjava
View job →
🔔

Get new senior distributed systems engineer data platform analytics and alerts jobs by email

Daily job updates · Unsubscribe anytime