Jobiba hiring network

Staff Platform Engineer Observability Jobs

15 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current staff platform engineer observability jobs. Use filters to narrow by work mode, employment type, experience and date posted.

W
Wellhub
📍 BrazilFull-timeRemote
3 days ago

Your wellbeing, our mission. Join a company shaping a healthier world. GET TO KNOW US At Wellhub we're revolutionizing workplace wellness. Our platform connects employees worldwide to the best partners for fitness, mindfulness, therapy, nutrition, and sleep—all in one simple subscription. Headquartered in NYC with team members in Europe, North America and South America, we’re on a mission to make every company a wellness company. We believe work should be fulfilling, inspiring, and balanced. Here, you’ll find a team that values wellbeing, collaboration, and different perspectives, where passion and creativity push boundaries to create real impact. Your contributions will help shape a healthier, more balanced world for you and millions of people globally. Join us in redefining the future of wellbeing! THE OPPORTUNITY We are hiring a Staff Platform Engineer with a dedicated focus on our Observability ecosystem for our Platform area in Brazil ! This is a Remote – Brazil position, meaning you can work from anywhere within the country. Please note that this role is only open to candidates in Brazil. In an environment of rapid growth and high-scale distributed architecture, your mission is to transform Observability from a passive toolset into a strategic asset using open source standards. You will act as an architect of efficiency and reliability , building a global platform that empowers engineering teams to "own what they build" with confidence. We are moving beyond basic monitoring to build a comprehensive "Observability as a Service" ecosystem. You will be responsible for evolving a self-service platform that balances performance with cost-effectiveness, solving complex challenges related to high-cardinality metrics, log retention strategies, and distributed tracing. We strive to eliminate friction. You will design the "Golden Paths" that allow developers to instrument their code instantly and gain high-fidelity signals without operatio

REMOTEpythonawsazure
View job →
CH
Cohere Health
📍 HyderabadFull-time₹2K – ₹2K/yr
3 days ago

Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is growing quickly. You will play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features, while also balancing scalability, reusability, and performance. Role Overview: We're looking for a Staff Platform Engineer to serve as the technical backbone of our Engineering organization. You'll own the technical strategy, and delivery of our platform — spanning architecture, DevOps, SRE, security, Dev-ex. This is a hands-on staff level role: you'll set technical direction, drive cross-team alignment, and be the senior escalation point for platform challenges. What you’ll do: Drive platform reliability, scalability, security, and cost efficiency across all environments. Technical Leadership: Provide technical leadership for platform components, Influence the technical strategy and architecture of our cloud platform, from CI/CD pipelines to observability and incident response. Design and implement platform components and reusable integration patterns that minimize custom development efforts, reduce the time spent on repetitive tasks, and ensure that integrations scale across multiple healthcare systems Partner closely with Architecture, DevOps, SRE, and Security teams to deliver cohesive platform solutions Cross-Functional Collaboration: Work closely with product teams, and solutions architects to understand integration needs and ensure the platform meets current and future business requirements. Serve as a senior escalation point for infrastructure and platform incidents Establish frameworks for: AI governance and compliance. Observability of systems. Traceability of decisions and outputs. Ensure enterprise readiness with security, auditability, and reliability in production environments. Security & Compliance : Ensure all p

awsci/cdgit
View job →

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity We are looking for a Staff Engineer to join our Observability team. This role is ideal for a highly technical engineer who thrives on uncovering the truth behind complex system behaviors, diagnosing difficult production challenges, and driving platform-wide improvements. As a Staff Engineer, you will operate as a force multiplier across engineering teams, helping Postman build world-class observability capabilities while improving reliability, performance, and developer productivity. You will partner closely with Infrastructure, Platform, Product, Security, and Data teams to identify systemic issues, establish operational excellence, and ensure engineering teams have the visibility they need to operate at scale. What You'll Do Drive the technical vision and architecture for Postman's observability platform. Design and build scalable solutions for metrics, logging, tracing, alerting, and operational analytics. Investigate complex production issues, identify root causes, and drive long-term corrective actions. Partner with engineering teams to improve service reliability, availability, performance, and operational mat

pythonjavanode.js
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Staff / Lead Software Engineer to architect, design, and scale our External Observability Platform . In this role, you will lead the technical strategy for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex cross-functional architecture initiatives, and skilled at mentoring senior engineers while holding your own with the brightest technical minds in the industry. Key Responsibilities Architect & Scale Distributed Infr

javavueaws
View job →
G
18 days ago

Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team... GEDE (Global Edge and Domains Engineering) keeps GoDaddy's domains, edge, and aftermarket platforms running for millions of customers. Within GEDE, the Reliability Engineering team (Domains Production Engineering) is the group that gets the call when something breaks — and, more importantly, builds the systems that mean it breaks less often. We own the monitoring, compliance, and patching toolchain for the whole Domains infrastructure footprint, provide advanced incident support, and are actively modernizing how the org detects, diagnoses, and even auto-remediates issues with AI-assisted tooling! What you'll get to do... Build and evolve observability using Prometheus/Mimir, the Grafana LGTM stack, Elastic/OTEL, and Site24x7 — closing gaps across the org. Steer our cloud migration journey and bolster our efforts to keep the services reliable and performant. Own patching compliance and vulnerability remediation at scale across a mixed on-prem + AWS fleet, hitting hard SLA targets. Operate and extend our multi-tenant Kubernetes/ArgoCD platform, including the migration of core services. Contribute to our AI/automation initiatives: auto-generating runbooks from Prometheus alerts, ServiceNow change-risk scoring, and other tooling that reduces toil for the whole team Consult with partner dev teams on metrics, alert thresholds, and monitoring standards — this is a platform-enablement role, not just a ticket queue. Mentor other engineers on the team and help mature our operational practices. Your experience should include

pythonsqlpostgresql
View job →
C
Coinbase
📍 - USAFull-timeRemoteFrom $218K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Senior Software Engineer, Access & Authorization As a Staff Software Engineer on the Access & Authorization team within the Platform group, you'll help shape the systems customers use to securely access Coinbase. This team builds and operates distributed services at the intersection of security, identity, developer infrastructure, and customer experience. You'll design foundational authentication and authorization capabilities that balance strong security with a simple experience, creating leverage across every Coinbase product and partner integration. What you'll do: Own end-to-end design, implementation, and operation of capabilities across authentication, OAuth, sessions, tokens, 2FA, account recovery, and edge authorization. Build secure, low-latency Go and gRPC services used across Coinbase's web, mobile, API, and partner experiences, with strong observability, failure testing, and measurable SLOs. Lead technical designs for complex initiatives such as signed authorization tokens, passkeys, biometric recovery, device-aware authentication, and mobile OAuth. Partner with product, security, fraud, and infrastructure teams to balance customer conversion, account protection, regulatory requirements, and engineering constraints. Simplify integration patterns and build self-service tooling that enables product teams to adopt access and authorization capabilitie

REMOTEredisawsai
View job →
A
Amplitude
📍 RemoteFull-time$165K – $247K/yr
1mo ago

About the Role Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Senior Platform Engineer, you'll own medium-to-high-complexity platform projects end-to-end and help shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll partner with Staff engineers and product teams to make Kubernetes effortless across the engineering org, building self-service automation and scalable AWS infrastructure that lets product teams ship faster, safer, and with less cognitive load. If you're excited about building the systems that other engineers will rely on every day, this role is for you. Key Responsibilities Lead high-impact platform projects — design and ship capabilities that move the needle on developer experience, reliability, or security, and set the bar for quality, testing, and safe deployment practices. Build the AI-augmented platform. Design tooling and workflows that help engineers get more out of AI-assisted development — think infra primitives that are easy to reason about, automated review, and policy-as-code that keeps the guardrails strong as AI shifts how code gets written. Own Infrastructure-as-Code for Kubernetes, AWS, and GCP using Terraform, Helm, Kustomize, and emerging tooling — and make it consumable enough that an LLM can safely PR against it. Evolve our CI/CD backbone (Argo CD / Workflows / Rollouts, GitHub Actions) to make deploys faster, safer, and easier to reason about. Instrument and operate. Drive observability with Datadog and Amplitude, own dashboards and SLOs, and use the data to push reliability forward. Participate in on-call, lead incident response when needed, and turn postmortems into durable platform improvements. Reduce toil and tech debt with pragmatic remediation

pythonawsgcp
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Overview We are looking for a strong software and platform engineer to join our Production Engineering team in Bangalore as an individual contributor in FA ProductionEng APJ. This role will help build and operate internal platforms that improve how we provision, observe, govern, and troubleshoot engineering infrastructure at scale. The fleet management use cases that give teams a single place to understand and operate the test infrastructure. If you enjoy building internal platforms that remove friction, improve visibility, and make engineering teams faster and more effective, this role is for you. Why This Role Is Unique This is not a typical application development role.You will work on internal platforms that directly shape how engineering teams consume and manage shared infrastructure. The role spans platform engineering, workflow automation, observability, API-driven services, and infrastructure lifecycle management. The right candidate will work on systems such as: Self-serviceability workflows and lease-based testbed governance. Developer Platform dashboards and APIs used for triage, visibility, and product trend observation. Testbed and workflow orchestration across fleet management domains. Impact This role is a high-leverage engineering investment. The work will improve how engineering teams provision testbeds, understand failures, operate shared infrastructure, and move faster with less friction. Better

awskuberneteslinux
View job →
D
1mo ago

As a Staff Engineer on the Data Platform Experience team, you'll help shape how Datadog engineering teams build, operate, and evolve products on the Observability Data Platform. You'll lead the design and delivery of shared platform capabilities that reduce developer friction, improve operational visibility, and enable engineering teams to move faster with confidence. This role combines deep distributed systems expertise with technical leadership across multiple teams, influencing platform strategy while remaining hands-on in the code. You'll have the opportunity to solve company-wide challenges spanning cost intelligence, operational tooling, platform health, and developer experience. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead strategic engineering initiatives that improve how product teams build, operate, and evolve services on the Observability Data Platform. Design and build scalable platform capabilities for cost intelligence, including cloud cost allocation, trend analysis, and optimization recommendations. Develop operational intelligence and self-service tooling that helps engineering teams understand platform health, troubleshoot incidents, and improve operational efficiency. Drive reusable platform services and developer workflows that increase engineering autonomy while reducing operational complexity across multiple products. Provide technical leadership across teams by influencing architecture, mentoring engineers, and raising engineering standards through hands-on technical contributions. Participate in the team's on-call rotation and continuously improve platform reliability, observability, and operational excellence. Who You Are: You have experience designing and building large-scale SaaS or cloud platforms with deep expertise i

javakubernetesai
View job →
D
Datadog
📍 Spain; Paris, FranceFull-time
1mo ago

As a Staff Engineer on the Data Platform Experience team, you'll help shape how Datadog engineering teams build, operate, and evolve products on the Observability Data Platform. You'll lead the design and delivery of shared platform capabilities that reduce developer friction, improve operational visibility, and enable engineering teams to move faster with confidence. This role combines deep distributed systems expertise with technical leadership across multiple teams, influencing platform strategy while remaining hands-on in the code. You'll have the opportunity to solve company-wide challenges spanning cost intelligence, operational tooling, platform health, and developer experience. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead strategic engineering initiatives that improve how product teams build, operate, and evolve services on the Observability Data Platform. Design and build scalable platform capabilities for cost intelligence, including cloud cost allocation, trend analysis, and optimization recommendations. Develop operational intelligence and self-service tooling that helps engineering teams understand platform health, troubleshoot incidents, and improve operational efficiency. Drive reusable platform services and developer workflows that increase engineering autonomy while reducing operational complexity across multiple products. Provide technical leadership across teams by influencing architecture, mentoring engineers, and raising engineering standards through hands-on technical contributions. Participate in the team's on-call rotation and continuously improve platform reliability, observability, and operational excellence. Who You Are: You have experience designing and building large-scale SaaS or cloud platforms with deep expertise i

javakubernetesai
View job →
P
Pinterest
📍 United StatesFull-timeRemoteFrom $177.2K/yr
1mo ago

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . We're seeking an exceptional Staff Software Engineer to join our Observability team at Pinterest. This role combines deep technical expertise in distributed systems and data engineering with a product-oriented mindset to build world-class observability solutions that empower our engineering organization. As a Staff Engineer on the Observability team, you'll be responsible for designing and building the infrastructure and tools that provide visibility into Pinterest's large-scale distributed systems, helping thousands of engineers understand, debug, and optimize their services. What you'll do: Define and execute the observability roadmap, treating it as a product. Understand engineering team needs and translate them into technical solutions with measurable impact Architect, build, and scale distributed observability infrastructure (me

REMOTEpythonjavaaws
View job →
A
3 days ago

Who We Are Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. Addepar unifies portfolio, market and client data in a total portfolio view and delivers AI-powered insights within investment and client workflows. More than 1,400 firms in nearly 60 countries use Addepar to manage and advise on nearly $9 trillion in assets. Its open platform integrates with nearly 650 software, data and consulting partners to power end-to-end investment operations across firms of all sizes and complexity. Addepar supports clients worldwide with offices in New York City, Salt Lake City, London, Edinburgh, Pune, Dubai, Geneva, Singapore and São Paulo. The Role We are currently seeking a Staff Software Engineer, Infrastructure to join the AI Platform team that powers seamless insights and interaction through natural language and data intelligence across our AI products. As a Staff Software Engineer, you’ll architect, build, and operate the backend and platform systems that power AI Platform. You’ll work across service design, distributed systems, cloud infrastructure, event-driven processing, observability, CI/CD, and production reliability, helping shape the technical direction of a platform that supports scalable, client-facing AI experiences. This role requires a strong software engineering foundation combined with deep infrastructure and systems thinking. We are looking for an engineer who can write high-quality production code, make sound architectural tradeoffs, and own platform capabilities end-to-end — not someone focused only on scripting, cloud configuration, or infrastructure tooling in isolation. You will collaborate closely with frontend, product, and AI/ML engineers to deliver reliable, secure, and scalable systems that align with Addepar’s standards of performance, resilience, and trust. Applicants must have legal authorization to work in the country where this role is based o

pythonjavasql
View job →
R
1mo ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Observability team's mission is to build and own Robinhood's full-stack observability platform — the foundation that keeps every product, service, and customer experience running reliably at scale. We design and operate the systems that give engineers deep visibility into how Robinhood's infrastructure behaves, ensuring that when something goes wrong, the right people know immediately and can act fast. Our work spans metrics, logs, distributed tracing, and alerting pipelines, and we partner closely with the Robinhood Command Center to ensure our observability systems meet or exceed 99.9% uptime. We believe in monitoring our own monitors — the observability infrastructure is a product, not just a tool. As a Staff Software Engineer on the Observability team, you will be the technical lead shaping the roadmap for how Robinhood observes itself at scale. You will own the observability control plane end-to-end, making architectural decisions that directly impact the reliability and operational health of Robinhood's entire product surface. You'll lead a team of six engineers, drive the strategy for cost-efficient telemetry ingestion, build in-house solutions where off-the-shelf

pythonvueaws
View job →
M
Mongodb
📍 IrelandFull-time
1mo ago

We are seeking a Staff engineer to design, build, and operate the internal and external Observability stack for the MongoDB platform. Tens of thousands of customers depend on our Observability stack to monitor their database clusters and to generate actionable alerts to safeguard critical workloads. This is an opportunity to join a team that is responsible for all Observability systems that support metrics, metric visualization, logs, traces, and alerts for MongoDB. We are looking for engineers with high standards, and experience in setting direction and technical leadership for large engineering teams in designing and operating complex distributed systems, with strict SLO on security, durability, availability and performance. As MongoDB Atlas and its supporting infrastructure continue to experience rapid growth, the demand for high-cardinality observability data for internal and external use cases means we need to continually innovate and scale our systems to the next level. For example, MongoDB Observability systems need to handle 10’s of billions of metrics time series, all whilst processing petabytes of logs, traces, and events. Our stack includes VictoriaMetrics, Splunk, Flink, WarpStream/Kafka, Java, Golang Fluentbit. In addition to owning critical components of our observability infrastructure, as a Staff engineer on the team, you’ll also work closely with other SWE, Product and SRE teams to promote and implement best practices in instrumenting and monitoring their services. This is a highly collaborative role, and you will get to own some of the most relied upon internal infrastructure at Mongo. Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to be a deeply technical leader on a collaborative team that applies low-level systems expertise to build the foundational infrastructure of a popular database, join us! Let’s build a faster, more reliable, and exceptionally observable database system together. W

javamongodbaws
View job →
CA
Careers at Tide
📍 LithuaniaFull-timeFrom €3.6K/yr
3 days ago

A BOUT TIDE At Tide, we help SMEs save time and money in the running of their businesses by not only offering business accounts and related banking services, but also a comprehensive set of highly usable and connected administrative solutions, from invoicing to accounting. Tide is transforming the small business banking market and now supports over 2 million members globally across the UK, India, Germany and France. Using advanced technology, all solutions are designed with SMEs in mind. With quick onboarding, low fees and innovative features, we thrive on making data driven decisions to serve our mission: to help SMEs save time and money so they can get back to doing what they love. Tide facts: Tide is available for UK, Indian, German and French SMEs Over 2 million members across UK and India Over $300 million raised in funding Over 2,800 Tideans globally Recognised with Great Place to Work certification three years in a row, and among India’s Top 50 Best Workplaces in Banking, Financial Services, and Insurance in 2026 We have offices in Central London, with a member support and technology centre in Sofia, Bulgaria, technology centres in Serbia, Romania, Lithuania and Hyderabad and offices in Gurugram, New Delhi, Berlin, Paris and Luxembourg ABOUT THE ROLE Tide is hiring a Senior Staff Software Engineer to lead the architecture of our agentic platform. You will shape the shared capabilities that allow AI systems to operate safely, reliably, and at scale across Tide. That includes context, orchestration, tool execution, trust controls, observability, and evaluation. You will participate in key build versus buy decisions, integrate external components where they accelerate us, and ensure we own the parts that matter most for Tide’s trust, data advantage, and long-term platform leverage. WHAT YOU WILL DO Define the architecture for Tide’s agentic platform Drive the design of shared services such as context APIs, tool layers, policy controls, and auditability Partner w

pythonjavaangular
View job →
🔔

Get new staff platform engineer observability jobs by email

Daily job updates · Unsubscribe anytime