Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Observability team's mission is to build and own Robinhood's full-stack observability platform — the foundation that keeps every product, service, and customer experience running reliably at scale. We design and operate the systems that give engineers deep visibility into how Robinhood's infrastructure behaves, ensuring that when something goes wrong, the right people know immediately and can act fast. Our work spans metrics, logs, distributed tracing, and alerting pipelines, and we partner closely with the Robinhood Command Center to ensure our observability systems meet or exceed 99.9% uptime. We believe in monitoring our own monitors — the observability infrastructure is a product, not just a tool. As a Staff Software Engineer on the Observability team, you will be the technical lead shaping the roadmap for how Robinhood observes itself at scale. You will own the observability control plane end-to-end, making architectural decisions that directly impact the reliability and operational health of Robinhood's entire product surface. You'll lead a team of six engineers, drive the strategy for cost-efficient telemetry ingestion, build in-house solutions where off-the-shelf
Salary not disclosed
Check market pay for comparable Lead Engineer roles before applying.
Role overview
Job description
Lead Engineer - Observability Platform — 49 Locations. Apply via Workday.
Hiring company
Cvshealth
Explore this employer's active roles, salary signals and company profile on Jobiba.
Keep exploring
Similar active roles
Fresh roles matched to this title and market.
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Overview We are looking for a strong software and platform engineer to join our Production Engineering team in Bangalore as an individual contributor in FA ProductionEng APJ. This role will help build and operate internal platforms that improve how we provision, observe, govern, and troubleshoot engineering infrastructure at scale. The fleet management use cases that give teams a single place to understand and operate the test infrastructure. If you enjoy building internal platforms that remove friction, improve visibility, and make engineering teams faster and more effective, this role is for you. Why This Role Is Unique This is not a typical application development role.You will work on internal platforms that directly shape how engineering teams consume and manage shared infrastructure. The role spans platform engineering, workflow automation, observability, API-driven services, and infrastructure lifecycle management. The right candidate will work on systems such as: Self-serviceability workflows and lease-based testbed governance. Developer Platform dashboards and APIs used for triage, visibility, and product trend observation. Testbed and workflow orchestration across fleet management domains. Impact This role is a high-leverage engineering investment. The work will improve how engineering teams provision testbeds, understand failures, operate shared infrastructure, and move faster with less friction. Better
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab
Become a part of our caring community Humana is seeking a self-driven and collaborative Lead Engineer to join our Interactive Voice Response (IVR) team. In this role, you will deliver innovative IVR solutions and develop robust omnichannel APIs for our enterprise platforms. You will have the opportunity to drive the success of a high-impact, customer-facing application within a Fortune 50 company, working closely with multiple teams throughout the software development lifecycle (SDLC). Lead Engineer –Omnichannel Humana is seeking a self-driven and collaborative Lead Engineer to join our Omnichannel team. In this role, you will design, develop, secure, and enhance enterprise APIs that support high-impact, member facing, applications across Humana's digital and voice channels. This role offers the opportunity to modernize and strengthen existing API capabilities while helping deliver resilient, scalable, and secure omnichannel solutions within a Fortune 50 organization. Key Responsibilities Design, develop, and maintain scalable Omnichannel APIs that support enterprise applications and customer-facing capabilities. Enhance the security, resiliency, performance, and reliability of existing APIs through modernization, improved architecture, observability, testing, and operational controls. Apply AI and AI-assisted engineering practices to accelerate development, improve quality, automate testing, enhance documentation, and identify opportunities for optimization. Partner with architecture, security, cloud, product, engineering, and operations teams to deliver secure, resilient, and enterprise-aligned API solutions. Collaborate with agile teams to plan, track, and deliver API enhancements, platform improvements, and cloud-based capabilities. Develop proofs of
Data Semantics is Datadog’s authority on semantic knowledge, providing shared infrastructure that powers both Datadog’s product experiences and AI capabilities. As Datadog continues its investment in OpenTelemetry-native observability, semantic interoperability, and AI-powered workflows, this team sits at the center of some of the company’s most strategic platform initiatives. As a Staff Engineer, you will serve as a technical leader for the team, balancing stewardship of critical production systems with the exploration of new platform capabilities that improve how telemetry is modeled, understood, and consumed across Datadog. You will work closely with engineering and product partners to define standards, drive technical direction, and deliver solutions that scale across Datadog’s observability platform. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the technical direction of the Data Semantics team while remaining deeply hands-on in design, implementation, and delivery. Build and scale semantic infrastructure that bridges OpenTelemetry, Datadog-native telemetry, cloud-provider telemetry, and customer-defined data models. Drive platform initiatives focused on schema evolution, telemetry standardization, data insights, and semantic interoperability across Datadog products. Partner with engineering and product teams across the platform to define standards, align stakeholders, and deliver high-leverage platform capabilities. Mentor engineers through design reviews, technical guidance, operational excellence, and long-term career development. Participate in on-call rotations and lead investigation and resolution efforts for complex production incidents affecting critical platform services. Who You Are: You have significant experience designing, opera
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. In this role you will: Develop interactive, data-rich user interfaces using React, TypeScript, and Vega, with a focus on integrating LLM-driven features (e.g., natural language querying, generative UI, and AI-assisted data storytelling). Lead the end-to-end delivery of substantial product features, ensuring AI outputs are presented with high reliability and low latency. Work closely with PMs, UX designers, and AI/ML engineers to bridge t
🔔 Get job alerts
New Lead Engineer - Observability Platform jobs in 49 Locations, straight to your inbox.
No spam · Unsubscribe anytime