At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - Container Platform (Menlo Park) About the Role We build the foundational container platform that runs Snowflake's production, AI/ML, and CI workloads across AWS, Azure, and GCP, including a rapidly growing AI/ML footprint. Hundreds of large Kubernetes clusters under management and growing. The work is to make that fleet reliable, automated, and invisible to the thousands of engineers building on top of it. This is a staff-level role on a senior, high-performing platform team. You'll own hard problems end to end, drive technical direction across teams, and build the automation and platform abstractions that make operating at this scale sustainable. There is significant unsolved work ahead: improving the developer experience for thousands of internal engineers and continuing to scale the platform to meet Snowflake's growth. What You'll Do Own the design and delivery of large, complex platform initiatives spanning cluster lifecycle management, multi-cloud automation, and internal developer tooling. Identify and drive cross-team technical improvements across the platform, from architecture through adoption. Make and defend architectural trade-offs grounded in reliability, scalability, and operational reality. Act as a technical anchor for the team, dev
Jobiba hiring network
Senior Staff Software Engineer Observability Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior staff software engineer observability jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Our Merchants organization focuses on onboarding and activation, catalog quality and health, merchant presence (including creator/UGC and trust signals), and ads upsell—helping merchants build healthy catalogs, reach the right audiences, and unlock clear paths to growth. In this role, you’ll lead high-impact engineering at the intersection of commerce platforms, catalog systems, and AI-native experiences. This includes applying ML, AI-assisted workflows, and GenAI where appropriate to improve metadata quality, merchant tooling, and operational efficiency. You’ll partner closely with Engineering Managers, Product Managers, Data Scientists, and other senior engineers to deliver systems with measurable business and user impact. What you’ll do: Own end-to-end technical delivery for cross-team initiatives—from problem framing and technical stra
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As an Intermediate Software Engineer in GitLab's Security section, you'll contribute to the data layer behind our vulnerability and dependency management capabilities. You'll design and develop well-scoped features within the systems that ingest, store, index, and query security report data, including integrations with key third-party systems. With a strong focus on PostgreSQL and scalable backend design, your work will support security features delivered by teams across GitLab. You'll work with clear context and strong support from the Senior Software Engineers and Staff Software Engineers around you. Yo
We are hiring an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build scalable security controls and services within MongoDB Atlas multi-cloud infrastructure. The team sits within the Site Reliability Engineering organization and works with other engineering teams to ensure that our infrastructure adheres to the highest security standards. This role can be based out of our New York City, Austin, Seattle or San Francisco offices, or work fully remotely on standard East Coast business hours. Responsibilities: Design and build core security primitives and services that protect MongoDB Atlas compute, networking, and identity across AWS, Azure, and GCP Build secure-by-default infrastructure using Linux security mechanisms (AppArmor, SELinux, seccomp, cgroups), Kubernetes, and eBPF to enforce runtime policies and gain deep visibility into systems behaviour Develop APIs, automation, and tooling that manage security posture at scale (CSPM, vulnerability management, workload identity) and provide monitoring, logging, and alerting pipelines that integrate with our tooling (Grafana, Splunk, Victoria Metrics.) Integrate security into our CI/CD and infrastructure-as-code workflows (Terraform) so that security controls are versioned, reviewed, and deployed just like any other code Lead complex projects end‑to‑end, from problem discovery and design docs to implementation, rollout, and long‑term ownership Collaborate with SRE, platform and product engineering teams to define secure architectures for new infrastructure and services Qualifications: You might be a great fit if you match some of the following: 5+ years of experience in Software Engineering, Site Reliability Engineering, or similar roles, preferably with relevant security work Proficiency with at least one programming language (Java, Golang, Rust, Python, or C/C++) and experience with infrastructure-as-code tools (Terraform) to automate security configurations
We are hiring an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build scalable security controls and services within MongoDB Atlas multi-cloud infrastructure. The team sits within the Site Reliability Engineering organization and works with other engineering teams to ensure that our infrastructure adheres to the highest security standards. This role can be based out of our Dublin office, or work fully remotely in Ireland. Responsibilities: Design and build core security primitives and services that protect MongoDB Atlas compute, networking, and identity across AWS, Azure, and GCP Build secure-by-default infrastructure using Linux security mechanisms (AppArmor, SELinux, seccomp, cgroups), Kubernetes, and eBPF to enforce runtime policies and gain deep visibility into systems behaviour Develop APIs, automation, and tooling that manage security posture at scale (CSPM, vulnerability management, workload identity) and provide monitoring, logging, and alerting pipelines that integrate with our tooling (Grafana, Splunk, Victoria Metrics.) Integrate security into our CI/CD and infrastructure-as-code workflows (Terraform) so that security controls are versioned, reviewed, and deployed just like any other code Lead complex projects end‑to‑end, from problem discovery and design docs to implementation, rollout, and long‑term ownership Collaborate with SRE, platform and product engineering teams to define secure architectures for new infrastructure and services Qualifications: You might be a great fit if you match some of the following: 5+ years of experience in Software Engineering, Site Reliability Engineering, or similar roles, preferably with relevant security work Proficiency with at least one programming language (Java, Golang, Rust, Python, or C/C++) and experience with infrastructure-as-code tools (Terraform) to automate security configurations and processes A deep understanding of Linux and networking concepts,
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. About the Role We are looking for Staff System Software Engineer in Test to join our team. In this role, you will be responsible for design, development, automation and reporting of Integration and system tests spanning across firmware and device drivers. This role requires you to have significant technical breadth and deep understanding of low-level system software specifically in server class systems. You will be part of a new team responsible for integration of different system software deliverables and development of system tests spanning all the components. You will contribute to shaping the test strategy , guide best practices and solve complex problems while maintaining a strong hands-on focus. You will partner with development and other QA teams to deliver high quality scalable and reliable solutions. About the Team Integration and system test team is responsible for verification and validation of integrated components across Board management controller (BMC), Firmware and Linux device driver. The team is also responsible for management and maintenance of common tools and pipel
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. POSITION SUMMARY CVS Health Digital is looking for hands-on, passionate people who want to join a high energy and growing team to make a difference in customers lives and who want to be on the forefront of digital innovation that aims to reinvent what a pharmacy and a health care company can be in the digital world. Currently, we are seeking a Staff Software Development Engineer, who will serve as a senior technical leader on a cross-functional product team responsible for designing, developing, and delivering scalable healthcare technology solutions. You will play a critical role in driving architecture decisions, mentoring engineers, and partnering with product, UX, and business stakeholders to build secure, cloud-native applications that enhance the digital healthcare experience. In this role, you will lead the development of distributed systems and microservices leveraging Java, Spring Boot, event-driven architectures, and modern cloud technologies. This Staff Engineer will also guide data migration and interoperability initiatives, supporting healthcare integrations that enable seamless and secure exchange of clinical data across platforms. The ideal candidate combines strong software engineering expertise with a passion for innovation, technical leadership, and delivering high-quality solutions in an Agile environment. Expectations for the Role Software
Scale’s rapidly growing International Public Sector team is focused on using AI to address critical challenges facing the public sector around the world. Our core work consists of: Creating custom AI applications that will impact millions of citizens Generating high-quality training data for custom LLMs Upskilling and advisory services to spread the impact of AI As a Full Stack Software Engineer (Forward Deployed), you’ll collaborate directly with public sector counterparts to quickly build full-stack, AI applications, to solve their most pressing challenges and achieve meaningful impact for citizens. At Scale, we’re not just building AI solutions—we’re enabling the public sector to transform their operations and better serve citizens through cutting-edge technology. If you’re ready to shape the future of AI in the public sector and be a founding member of our team, we’d love to hear from you. You will: Serve as the lead technical strategist for public sector engagements, converting ambiguous mission requirements into robust architectural roadmaps and guiding onsite implementation Architect the fundamental frameworks for production-grade AI applications, setting the gold standard for how interactive UIs, backend systems, and AI models are integrated at scale to deliver reliable outcomes. Guide the evolution of cloud infrastructure, ensuring security, global scalability, and long-term system integrity across all environments. Direct the development of core platforms and shared services, ensuring they solve cross-cutting needs for diverse global client use cases. Partner with cross-functional leadership to steer the technical roadmap, mentoring senior and junior staff and ensuring all products align with a cohesive, future-proof technical architecture. Bridge the gap between the field and the core platform by turning real-world client lessons into the reusable patterns that power the entire engineering team. Ideally you’d have: Masters or Phd in Computer Science or eq
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Security Software Engineer on the Production IAM team, you will set the technical direction for how identity and access work across Roblox's production infrastructure, from the mTLS-based identity that services use to authenticate to one another, to the privileged access controls that govern how engineers reach production. The team is accountable for Roblox's machine and workload identity platform, its centralized authorization engine, its production access management platform, production PKI and certificate lifecycle, and just-in-time privileged access for engineers. As an individual contributor in Production IAM, you will define multi-year strategy, drive alignment across Roblox Platform, mentor senior and staff engineers, and personally build the hardest parts of these systems. As AI agents become first-class actors in production, you will also help pioneer how they get identity, prove who they are, and receive safely-scoped access. You will Lead the architecture for production identity and access. Define and evolve the end-to-end design for machine, workload, human, and AI-agent identity across our hybrid on-prem and cloud fleet, making secure access invisible when
About the Team ChatGPT is a rapidly evolving system: new capabilities ship continuously, product surfaces change quickly, and usage patterns shift week-to-week. Supporting that pace requires infrastructure that can handle real production constraints—high concurrency, unpredictable traffic patterns, complex dependency graphs, and frequent change. The ChatGPT Infrastructure team builds and operates the platforms that enable fast iteration without compromising performance or reliability. We design shared systems, data paths, rollout mechanisms, and reliability guardrails that teams rely on to ship changes to ChatGPT at scale. We focus on high-leverage infrastructure: primitives and “golden paths” that incorporate operational lessons as defaults, so engineers don’t need to rediscover failure modes, latency pitfalls, or integration issues each time they build something new. About the Role We’re hiring Senior and Staff Engineers to design and build infrastructure systems that underlie ChatGPT and multiply the effectiveness of teams building user experiences. This is not a support-only role. It’s a platform-building role: you’ll define interfaces, develop core abstractions, and create tooling to make safe, fast iteration the norm. Your work will reduce friction, prevent regressions, improve performance, and ensure systems scale gracefully as the product grows. Where You Can Have Impact You might work on one or more of the following areas (without being restricted to any single area): Platform foundations & frameworks: Core libraries, service frameworks, and shared components that standardize system building, integration, and evolution. Scalability & performance primitives: Patterns and infrastructure that reduce tail latency, improve throughput, and keep costs predictable as demand increases. Reliability guardrails: Mechanisms that prevent outages by design—rate limiting, load shedding, dependency isolation, backpressure, safe fallbacks, and robust regression contr
Who We Are Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. Addepar unifies portfolio, market and client data in a total portfolio view and delivers AI-powered insights within investment and client workflows. More than 1,400 firms in nearly 60 countries use Addepar to manage and advise on nearly $9 trillion in assets. Its open platform integrates with nearly 650 software, data and consulting partners to power end-to-end investment operations across firms of all sizes and complexity. Addepar supports clients worldwide with offices in New York City, Salt Lake City, London, Edinburgh, Pune, Dubai, Geneva, Singapore and São Paulo. The Role We are seeking a Staff Full Stack Software Engineer to join the Advisor Experience team as our Technical Lead. Our team is focused on building tools for financial advisors to grow and sustain their business. We oversee bespoke products for advisors and develop advisor-focused capabilities throughout the Addepar platform. In this role, you will be the primary technical anchor for new capabilities including Secure Message Center — a compliant messaging experience built into Addepar's client portal that allows advisors and their clients to communicate directly within the platform. You will partner directly with Engineering Leadership and Product Management to build a modern, scalable architecture from the ground up. Beyond system design, you will act as a true engineering multiplier: setting technical standards, mentoring junior and mid-level engineers, and working alongside other senior engineers and AI specialists to deliver high-impact advisor tools. Applicants must be legally authorized to work in the United States for any employer without requiring current or future visa sponsorship (for example, employment-based visas such as H-1B, F-1/OPT, or similar), and must be authorized to begin work in the U.S. on their first day of employme
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Agent Gateway Team The Agent Gateway team owns the identity-aware infrastructure that connects enterprise AI agents to the tools, data, and services their organizations authorize. Every call from Claude, Agentforce, Codex, and internal/homegrown agents to a resource flows through us. We enforce authorization, isolate credentials, mint the right token per target, and produce the audit trails that security teams rely on. We are early in a rapidly evolving space. The standards for agent identity (MCP, OAuth token exchange, DCR) are being built under our feet. Our roadmap includes hardening the data plane for on-premises customer deployments, extending policy semantics beyond tool-level allowlists, adding native support for Agent-to-Agent brokered delegation, and scaling to tenants with thousands of virtual MCP servers. The Senior Software Engineer Opportunity Okta is looking for a Senior Software Engineer to help build the Agent Gateway. You will own features and components across the data and control planes, turning technical designs and product requirements into reliable production systems. Working alongside staff engineers, you will implement token exchange, request routing, credential resolution, and policy evaluation as agent identity specifications evolve. This is a hands-on software development role at the intersection of product, security, and infrastructure. You will ship services that handle agentic traffic reliably and performantly at scale. Wha
At Linear, we're building the product development system for teams and agents. AI is fundamentally changing how software gets built, and we’re shaping the tools this new era requires. Founded in 2019, Linear has become the platform of choice for more than 40,000 companies (including OpenAI, Coinbase, and Ramp) to plan, build, and ship their products. Today, our team is distributed across North America, Europe, and Australia, and we’re continuing to grow internationally. What unites us is relentless focus, fast execution, and a deep care for software craftsmanship. As a small team, we’re all generalists that work across the full stack (built in Typescript end-to-end). We’re looking for experienced engineers that thrive in an environment of autonomy and individual responsibility to help us build the future of product development. Location & work mode Linear is a remote-first company, with optional co-working offices in San Francisco, New York, and London. This role is open to candidates based in the US and Europe. You can work from anywhere within these regions. We value deep focus and async collaboration, with intentional moments to connect in person through team off-sites, optional co-working, and occasional travel. What you'll do Work closely with founders and design to implement new concepts and ideas Build AI-powered functionality into the core of Linear Update our realtime collaborative content editor used across all internal surfaces Build new user-facing features with beautiful and scalable UI components Obsessively improve application performance Refine our software development processes to keep the team operating at high velocity What we're looking for 5+ years of experience building customer-facing products at a high-quality software company Strong React and TypeScript fundamentals, with experience across the full stack (Browser technologies, Node, GraphQL, PostgreSQL) Track record of driving complex, end-to-end features (not just incremental improvemen
At Linear, we're building the product development system for teams and agents. AI is fundamentally changing how software gets built, and we’re shaping the tools this new era requires. Founded in 2019, Linear has become the platform of choice for more than 40,000 companies (including OpenAI, Coinbase, and Ramp) to plan, build, and ship their products. Today, our team is distributed across North America, Europe, and Australia, and we’re continuing to grow internationally. What unites us is relentless focus, fast execution, and a deep care for software craftsmanship. As a small team, we’re all generalists that work across the full stack (built in Typescript end-to-end). We’re looking for experienced engineers that thrive in an environment of autonomy and individual responsibility to help us build the future of product development. Location & work mode Linear is a remote-first company, with optional co-working offices in San Francisco, New York, and London. This role is open to candidates based in the US and Europe. You can work from anywhere within these regions. We value deep focus and async collaboration, with intentional moments to connect in person through team off-sites, optional co-working, and occasional travel. What you'll do Work closely with founders and design to implement new concepts and ideas Build AI-powered functionality into the core of Linear Update our realtime collaborative content editor used across all internal surfaces Build new user-facing features with beautiful and scalable UI components Obsessively improve application performance Refine our software development processes to keep the team operating at high velocity What we're looking for 5+ years of experience building customer-facing products at a high-quality software company Strong React and TypeScript fundamentals, with experience across the full stack (Browser technologies, Node, GraphQL, PostgreSQL) Track record of driving complex, end-to-end features (not just incremental improvemen
At Linear, we're building the product development system for teams and agents. AI is fundamentally changing how software gets built, and we’re shaping the tools this new era requires. Founded in 2019, Linear has become the platform of choice for more than 40,000 companies (including OpenAI, Coinbase, and Ramp) to plan, build, and ship their products. Today, our team is distributed across North America, Europe, and Australia, and we’re continuing to grow internationally. What unites us is relentless focus, fast execution, and a deep care for software craftsmanship. As a small team, we’re all generalists and constantly picking up new challenges. When it comes to code, we’re looking to work with experienced people who can pick a problem and solve it. We use TypeScript and build scalable systems so we can continuously make progress on a solid foundation. We don’t expect you to have a background in everything we use, but we do expect strong JavaScript fundamentals and a background working with React and TypeScript. Location & work mode Linear is a remote-first company, with optional co-working offices in San Francisco, New York, and London. This role is open to candidates based in North America and Europe. You can work from most timezones within these regions. We value deep focus and async collaboration, with intentional moments to connect in person through team off-sites, optional co-working, and occasional travel. What you'll do Build new user-facing features with everything from database models to GraphQL resolvers and UI components Optimize our data synchronization stack by applying better serialization protocols Add real-time collaborative editing to our content editor Improve performance by profiling and tweaking virtualized list rendering Add analytics, monitoring, and alerts to our service so that we can better respond to operational incidents Open-source any non-trivial innovations that come out of our work on the product Redefine best-in-class software deve
Get new senior staff software engineer observability jobs by email
Daily job updates · Unsubscribe anytime