We're building the platform that lets long-running autonomous agents operate safely inside NVIDIA's enterprise. These are not assistants on a developer's laptop. They are fleets of agents deployed in the cloud, running continuously at scale on shared accelerated compute. They take on real work across enterprise systems, so people get far more done than they could before. This role defines the constructs that agents are built from: the blueprints they start from, the tools, skills, and plugins that power them against enterprise data, the runtime safety harness that keeps them in bounds, and the connections into credential management, sandbox, memory, and observability. The team designs and ships these building blocks so that agent developers across the company can stand up a new agent, wire it in, and run it for days or weeks. Security and safe execution come out of the box, not something each team has to get right on its own. Today an agent runs inside a single harness. Claude, Codex, and open-source agent harnesses each work differently underneath, with their own execution model, tool interface, and telemetry shape. The platform smooths over those differences, so a single skill, safety policy, or trace works the same no matter which harness is running. We want to enable agents that act on a person's behalf, governed and secure, continuously evaluated and self-improving. These agents coordinate and hand work off to each other, with identity and policy following every hop. They route and tune themselves across harnesses from live eval signals, and get better from their own production telemetry instead of waiting on a human to retrain them. Have you run agents on a harness like Claude or Codex and hit the walls that show up when they run for real, for days, against live systems — and wanted them to learn from it on their own? We're building the platform that solves those problems once, for every team. What you'll b
Jobiba hiring network
Senior Staff Software Engineer Observability Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior staff software engineer observability jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: The Experience team is at the center of one of the most exciting transitions in software development history — the shift from human-driven to agent-driven product experiences. We own Pinecone's API, clients, authentication, revenue, and observability systems, and right now that means redesigning all of it for a world where AI agents are first-class users alongside humans. This is a wide-scope role. You'll own things end-to-end — from backend architecture to API design to SDK and web surfaces. You’ll be working closely with product, design, and other engineering teams to identify user needs and build the right thing, at the right abstraction level, at the right time. Along the way, you will be building high-leverage platform capabilities that accelerate Pinecone’s product development and user growth systems. We're looking for an engineer who sees this moment for what it is: a rare opportunity to shape how developers and agents interact with a category-defining product. You're not waiting to see how the industry figures out MCP, agentic workflows, and AI-native interfaces — you're already experimenting, already forming opinions, already building. You know that speed and leverage matter more than labor, and you've internalized AI-assisted development not as a productivity trick but as a fundamentally different way of working. Responsibilities: Pioneer our agent experience. Shape how AI agents interact with Pinecone — designing interfaces, protocols (MCP), and tooling that make Pinecone the easiest and most capable platform f
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: The Experience team is at the center of one of the most exciting transitions in software development history — the shift from human-driven to agent-driven product experiences. We own Pinecone's API, clients, authentication, revenue, and observability systems, and right now that means redesigning all of it for a world where AI agents are first-class users alongside humans. This is a wide-scope role. You'll own things end-to-end — from backend architecture to API design to SDK and web surfaces. You’ll be working closely with product, design, and other engineering teams to identify user needs and build the right thing, at the right abstraction level, at the right time. Along the way, you will be building high-leverage platform capabilities that accelerate Pinecone’s product development and user growth systems. We're looking for an engineer who sees this moment for what it is: a rare opportunity to shape how developers and agents interact with a category-defining product. You're not waiting to see how the industry figures out MCP, agentic workflows, and AI-native interfaces — you're already experimenting, already forming opinions, already building. You know that speed and leverage matter more than labor, and you've internalized AI-assisted development not as a productivity trick but as a fundamentally different way of working. Responsibilities: Pioneer our agent experience. Shape how AI agents interact with Pinecone — designing interfaces, protocols (MCP), and tooling that make Pinecone the easiest and most capable platform f
About Datadog: We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Opportunity: Datadog’s Senior Staff Engineers are technical leaders operating at the forefront of large-scale systems design, building the infrastructure that will support our next five years of growth and beyond. They do this in three major ways: As individual contributors, they bring world-class technical depth to build industry-leading systems in areas such as observability data platforms, distributed query engines, and real-time event streaming at global scale. As technical leaders, they apply broad architectural perspective and deep systems thinking to align design decisions across teams and domains. They work across complex, multi-team problem spaces to define long-term technical direction, drive large-scale initiatives forward, and ensure consistent execution. As engineering stewards, they play a key role in evolving our systems and engineering culture. They actively participate in Datadog’s senior technical community, bringing external insights and internal experience to elevate engineering standards and mentor the next generation of technical leaders. Examples of projects a Senior Staff Engineer may lead include designing and launching a new distributed data storage engine capable of handling hundreds of millions of records per second, building the real-time infrastructure behind a new observability product, or re-architecting a core service to support exponential growth in throughput and complexity. What You’ll Do: Be the technical owner of multiple critical systems or architecture areas, often spanning several t
A BOUT TIDE At Tide, we help SMEs save time and money in the running of their businesses by not only offering business accounts and related banking services, but also a comprehensive set of highly usable and connected administrative solutions, from invoicing to accounting. Tide is transforming the small business banking market and now supports over 2 million members globally across the UK, India, Germany and France. Using advanced technology, all solutions are designed with SMEs in mind. With quick onboarding, low fees and innovative features, we thrive on making data driven decisions to serve our mission: to help SMEs save time and money so they can get back to doing what they love. Tide facts: Tide is available for UK, Indian, German and French SMEs Over 2 million members across UK and India Over $300 million raised in funding Over 2,800 Tideans globally Recognised with Great Place to Work certification three years in a row, and among India’s Top 50 Best Workplaces in Banking, Financial Services, and Insurance in 2026 We have offices in Central London, with a member support and technology centre in Sofia, Bulgaria, technology centres in Serbia, Romania, Lithuania and Hyderabad and offices in Gurugram, New Delhi, Berlin, Paris and Luxembourg ABOUT THE ROLE Tide is hiring a Senior Staff Software Engineer to lead the architecture of our agentic platform. You will shape the shared capabilities that allow AI systems to operate safely, reliably, and at scale across Tide. That includes context, orchestration, tool execution, trust controls, observability, and evaluation. You will participate in key build versus buy decisions, integrate external components where they accelerate us, and ensure we own the parts that matter most for Tide’s trust, data advantage, and long-term platform leverage. WHAT YOU WILL DO Define the architecture for Tide’s agentic platform Drive the design of shared services such as context APIs, tool layers, policy controls, and auditability Partner w
A BOUT TIDE At Tide, we help SMEs save time and money in the running of their businesses by not only offering business accounts and related banking services, but also a comprehensive set of highly usable and connected administrative solutions, from invoicing to accounting. Tide is transforming the small business banking market and now supports over 2 million members globally across the UK, India, Germany and France. Using advanced technology, all solutions are designed with SMEs in mind. With quick onboarding, low fees and innovative features, we thrive on making data driven decisions to serve our mission: to help SMEs save time and money so they can get back to doing what they love. Tide facts: Tide is available for UK, Indian, German and French SMEs Over 2 million members across UK and India Over $300 million raised in funding Over 2,800 Tideans globally Recognised with Great Place to Work certification three years in a row, and among India’s Top 50 Best Workplaces in Banking, Financial Services, and Insurance in 2026 We have offices in Central London, with a member support and technology centre in Sofia, Bulgaria, technology centres in Serbia, Romania, Lithuania and Hyderabad and offices in Gurugram, New Delhi, Berlin, Paris and Luxembourg ABOUT THE ROLE Tide is hiring a Senior Staff Software Engineer to lead the architecture of our agentic platform. You will shape the shared capabilities that allow AI systems to operate safely, reliably, and at scale across Tide. That includes context, orchestration, tool execution, trust controls, observability, and evaluation. You will participate in key build versus buy decisions, integrate external components where they accelerate us, and ensure we own the parts that matter most for Tide’s trust, data advantage, and long-term platform leverage. WHAT YOU WILL DO Define the architecture for Tide’s agentic platform Drive the design of shared services such as context APIs, tool layers, policy controls, and auditability Partner w
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role The ML Infrastructure team is responsible for building & improving the core infrastructure for autonomy teams at Nuro. In this role, you will work closely with teams across Nuro, to design, build and deploy core infrastructure components in machine learning model life cycle, to push the autonomous future forward. You will have an opportunity to work across the full stack of machine learning solutions - from designing robust and scalable model & data pipelines to building to deploying the optimized models on Nuro’s fleet of self-driving robots! About the Work Design and develop ML workflow pipelines to train, optimize, validate, and deploy Nuro autonomy models. Develop and maintain continuous testing and monitoring systems for core ML infrastructure components. Develop observability to track ML model lifecycles from data generation to on-road validation. Maintain an in-house ML inference platform to serv
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Intro : Pinterest’s Service Configuration & Coordination team builds the platforms, APIs, and tooling that make dynamic configuration and service coordination safe, scalable, and reliable for every service across regions and clouds, serving as the central hub for cross-functional alignment and system orchestration. As a Senior Staff Software Engineer (IC17) , you’ll own the end-to-end technical strategy and long-term vision for our configuration and coordination infrastructure. You’ll define the future of these mission-critical systems, partnering closely with Traffic, Compute, Observability, and Cloud Architecture teams to deliver robust, high-performance solutions at scale. What you’ll do: Orchestrate the long-term technical vision for configuration ecosystems to ensure feature flags, ML settings, and experiments utilize unified pave
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Staff / Lead Software Engineer to architect, design, and scale our External Observability Platform . In this role, you will lead the technical strategy for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex cross-functional architecture initiatives, and skilled at mentoring senior engineers while holding your own with the brightest technical minds in the industry. Key Responsibilities Architect & Scale Distributed Infr
Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is experiencing rapid growth. You’ll play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features, while also balancing scalability, reusability, and performance. As a Staff Engineer on the Application Engineering team, you’ll serve as a senior technical leader - responsible for designing and delivering high-quality, scalable software systems that power Cohere Health’s core platform. You’ll act as a multiplier, elevating the technical bar for the team, mentoring engineers, and partnering with product, data, and clinical teams to deliver solutions that meet compliance, quality, and performance standards. This role is ideal for engineers who thrive on solving complex problems in healthcare, have deep expertise in building distributed systems, and want to influence architecture and engineering practices at scale. What you’ll do: Technical Leadership & Architecture Define and drive the architecture of large-scale, distributed application systems across the Cohere platform. Ensure solutions are secure, performant, maintainable, and compliant with NCQA, CMS, and payer requirements. Champion engineering best practices in CI/CD, testing, release management, and observability. Hands-On Engineering Write clean, maintainable, and well-tested code, primarily in modern frameworks (e.g., Python, TypeScript/React, Java/Kotlin). Lead the development of core features and APIs that directly impact providers, payers, and patients. Partner with DevOps and Data teams to ensure seamless integration, scalability, and operational readiness. Quality & Compliance Focus Embed automated testing, monitoring, and release safeguards into the development lifecycle. Proactively address compliance and audit-readiness requirements in application
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Staff Software Engineer, Core AI Infrastructure You'll join a high-performing team of engineers driving AI transformation at Coinbase as a Senior Software Engineer on the IT Operations team within ESTO. This team builds custom products through full-stack engineering and scales the infrastructure powering Coinbase's AI products, with direct exposure to senior leadership in a fast-paced, incubator-style environment. You'll own the reliability and automation of critical AI infrastructure, ensuring our systems are resilient, observable, and secure at scale. What you'll do: Own end-to-end delivery of AI products by building production-grade distributed systems, including serving infrastructure, data pipelines, and deployment orchestration across the full stack throughout the SDLC. Drive platform adoption by designing clean APIs, abstractions, and developer-facing tooling that enable product teams to integrate AI capabilities without bespoke infrastructure requests. Partner with engineering and product leadership across Platform and other product groups to align infrastructure requirements, resolve cross-team technical dependencies, and define shared platform contracts. Shape engineering standards and technical culture by establishing architectural patterns, mentoring engineers, and raising the bar on code quality, observability, and operational excellence. Build
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Senior Software Engineer, Access & Authorization As a Staff Software Engineer on the Access & Authorization team within the Platform group, you'll help shape the systems customers use to securely access Coinbase. This team builds and operates distributed services at the intersection of security, identity, developer infrastructure, and customer experience. You'll design foundational authentication and authorization capabilities that balance strong security with a simple experience, creating leverage across every Coinbase product and partner integration. What you'll do: Own end-to-end design, implementation, and operation of capabilities across authentication, OAuth, sessions, tokens, 2FA, account recovery, and edge authorization. Build secure, low-latency Go and gRPC services used across Coinbase's web, mobile, API, and partner experiences, with strong observability, failure testing, and measurable SLOs. Lead technical designs for complex initiatives such as signed authorization tokens, passkeys, biometric recovery, device-aware authentication, and mobile OAuth. Partner with product, security, fraud, and infrastructure teams to balance customer conversion, account protection, regulatory requirements, and engineering constraints. Simplify integration patterns and build self-service tooling that enables product teams to adopt access and authorization capabilitie
Get new senior staff software engineer observability jobs by email
Daily job updates · Unsubscribe anytime