A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We've been doing a lot of thinking as a team, and what we've landed on is that this person needs to be someone who takes ownership. Real ownership. Not just "I'll close the ticket", but "I'm going to make sure this never happens again" type of ownership. That's the kind of person we're looking for, and frankly, that's the kind of person who's going to thrive here at Palantir. As an IT Support Engineer, you take real pride of our internal ecosystem, from individual workstations to conference rooms to the systems employees rely on every day. You're the person people turn to when something isn't working, whether they're at their desk, travelling, or preparing for an important meeting. You're thorough when it comes to troubleshooting. You don't just close the ticket, you make sure the underlying issue is actually resolved. When something keeps coming back, you dig into why and strive to look for a permanent fix. You're proactive by nature and complacency isn't something you settle for, and it’s not something we settle for either. You're constantly looking for ways to eliminate friction, automate the mundane, and elevate the bar for TechOps. People feel comfortable coming to you with problems because you're approachable and follow through. You're familiar with the needs of executives and senior staff, and you check in proactively when you know something critical is on the horizon like a board meeting, an earnings call, an internal conference or a big reveal rather than taking a back seat. What sets you apart is how much you care about the growth of the people around you. When you work through a tough ticket with a colleague, you make it a teaching moment. When you
Jobiba hiring network
Senior Staff Software Engineer Observability Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior staff software engineer observability jobs. Use filters to narrow by work mode, employment type, experience and date posted.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We've been doing a lot of thinking as a team, and what we've landed on is that this person needs to be someone who takes ownership. Real ownership. Not just "I'll close the ticket", but "I'm going to make sure this never happens again" type of ownership. That's the kind of person we're looking for, and frankly, that's the kind of person who's going to thrive here at Palantir. As an IT Support Engineer, you take real pride of our internal ecosystem, from individual workstations to conference rooms to the systems employees rely on every day. You're the person people turn to when something isn't working, whether they're at their desk, traveling, or preparing for an important meeting. You're thorough when it comes to troubleshooting. You don't just close the ticket, you make sure the underlying issue is actually resolved. When something keeps coming back, you dig into why and strive to look for a permanent fix. You're proactive by nature and complacency isn't something you settle for, and it’s not something we settle for either. You're constantly looking for ways to eliminate friction, automate the mundane, and elevate the bar for TechOps. People feel comfortable coming to you with problems because you're approachable and follow through. You're familiar with the needs of executives and senior staff, and you check in proactively when you know something critical is on the horizon like a board meeting, an earnings call, an internal conference or a big reveal rather than taking a back seat. What sets you apart is how much you care about the growth of the people around you. When you work through a tough ticket with a colleague, you make it a teaching moment. Whe
MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of our Boston, New York City, Raleigh, Miami, Pittsburgh or remotely in the United States while physically based in an Eastern or Central time zone location. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure g
The Team MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of either our Dublin or Cork office or remotely in Ireland. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Build for reliability, making services and infrastructure avail
MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of our Toronto or Montreal office or remotely in the Canada while physically based in an Eastern or Central time zone location. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineerin
The Team This role can sit in our NYC HQ on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas platform. As a senior SRE, you will be expected to be able to design & build complex systems, operate with autonomy and act as owner for everything you do. The SRE Atlas team works alongside the various Atlas software engineering teams to provide expertise about running systems at scale, build new tooling and automation and perform essential maintenance of the Atlas fleet. This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions that have the ability to impact our customer’s most crucial workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This role requires engineers to have a customer-first mindset to ensure that everything we do results in a stronger product and a better experience for all Atlas customers. The ideal candidate should Have 5+ years of experience running critical systems at scale Value efficiency in processes and operations, and display a preference for automation over manual processes (“allergic to ops work”) Be familiar with a major cloud provider (AWS, Azure, or GCP) and possess the ability to build and operate systems in a multi-cloud environment A strong understanding of how to run a large scale Linux environment, including low level fundamentals Firm grasp of at least one modern programming language, beyond basic scripting (Go, Ruby, Python) Solid understanding of web and network protocols and standards (HTTP, TLS, DNS, etc) Special Requirements: Be a US Citizen Expectations Participate in the development of a reliable and resilient multi-cloud platform that hosts business critic
Senior: GBP 73,500 - 99,500 Staff: GBP 97,300 - 131,700 Subject to alignment to the responsibilities and duties of the role - we currently have multiple positions available at both Senior and Staff level. About the job Build the Linux distribution foundation that turns upstream software into trusted Graphcore platform releases. You will help create the Linux distribution that powers Graphcore AI systems. The team produces production-ready system images from proven upstream distributions. Your work will shape how releases are built, validated and prepared for deployment. You will strengthen the engineering path from upstream Linux software to dependable platform releases. You will build and improve automated pipelines, run established Linux test suites, and diagnose issues across build and validation flows. As the platform evolves, you will introduce controlled configuration and tuning changes with evidence-led validation. This is hands-on systems engineering with visible impact. You will help define reliable processes for a new team building a critical part of Graphcore’s platform. The team and culture You will join one of Graphcore’s newest engineering teams, helping shape its culture from the start. It is a small, co-located team where ownership matters and progress is visible. Work happens through close technical discussion, practical problem-solving and evidence-led decisions. Ideas are challenged openly, and the best path wins regardless of hierarchy. You will report to a leader who values technical credibility and invests in people’s growth. The team moves with pace, takes responsibility and changes direction when the evidence demands it. What we're looking for Strong practical experience working in Linux environments Experience building or maintaining automated CI/CD pipelines for reliable engineering workflows Proficiency in Python, Bash or similar
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. Intermediate Fullstack Engineer - Data Products An overview of this role Data Products integrates GitLab and third-party software development lifecycle data, and builds the dashboards, APIs, and data products that turn it into reliable, interoperable intelligence for customers and internal teams. You'll work across the stack, contributing to frontend experiences, backend services, APIs, and AI-enabled workflows. This is a hands-on product engineering role. You'll develop features, learn how we build and operate large-scale data systems, and work closely with Senior and Staff Engineers to deliver secure, reliable, and performant s
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Backend Engineer on the Chat Engine team, you'll build the engine behind GitLab Duo Chat, the conversational AI experience for GitLab. You'll work primarily in Python to build and maintain our agentic runtime: the Flow Registry, LangGraph flows, and the Duo Workflow Service. You'll also work in the GitLab Rails monolith, where Chat connects with the product. You'll own scoped parts of the system, ship small features and improvements with minimal guidance, and collaborate with the team on larger projects. You'll work alongside senior and staff engineers who will partner with you on design and support
Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: A senior engineering professional who independently applies advanced knowledge to complete complex assignments, who leads the design and development of complex software code, unit tests and integration tests for a subsystem. MAIN RESPONSIBILITIES • Leads, is accountable for, and serves as the technical subject matter expert for the engineering design and implementation for one or more software features, identifying process issues and recommending corrective measures. • Defines feature evolution, branching, integration and deployment strategy. • Defines structure of the source code files. • Ensures successful integration. • Implements hardware/interface simulation. • Analyzes user needs, product requirements, software requirements and provides input to stakeholders Education Level Major/Field of Study or Equivalent: Bachelors Degree (± 16 years) Experience/Background Minimum 6 years Minimum 6 - 12 Years of experience The base pay for this position is $99,300.00 – $198,700.00 In specific locations, the pay range may vary from the range posted. JOB FAMILY: Product Development DIVISION: ONCO Cancer Diagnostics LOCATION:<
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Staff Backend Engineer at GitLab, you will help shape a major investment in our Software Supply Chain Security offering. In this role, you'll serve as a senior technical leader for backend systems that help customers secure how software is built, verified, and delivered inside the GitLab platform. You'll work on foundational capabilities across package policy enforcement, build provenance, artifact signing, and malicious package detection, with a strong focus on enterprise-grade security and performance. You'll define architecture before systems are built, write clear technical proposals, and guide i
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a highly skilled and motivated Senior Design Verification Engineer to join our team. In this role, you will be responsible for the end-to-end verification of our IOMMU (Input/Output Memory Management Unit) IP. You will play a critical role in ensuring the functional correctness and performance of the design, taking ownership of the verification process right from the initial specification understanding down to final coverage closure. This role is hybrid, based out of Bangalore, India. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You have hands-on experience in ASIC or SoC verification using SystemVerilog and UVM. You enjoy debugging complex design and verification issues and working closely with cross-functional teams. You’re comfortable building verification environments from scratch and driving coverage closure. You have familiarity with standard bus protocols and modern verification tools. What We Need Experience owning block or subsystem-level verification from test planning to sign-off. Strong understanding of constrained-random verification, coverage analysis, and regression debugging. Familiarity with proto
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Staff DB SRE to build the runtime foundation for NVIDIA’s enterprise AI platforms — with a strong emphasis on database infrastructure at scale. This role blends large-scale database transformation with the building and development of GPU-accelerated platforms. You'll develop the software systems, automation frameworks, and high-performance database services that power NVIDIA’s AI workloads at scale. What you'll be doing: Design and operate highly available database clusters (MySQL, MSSQL, Oracle) with automated replication, failover, point-in-time recovery, and disaster-recovery strategies at enterprise scale. Drive database performance engineering — own query optimization, indexing strategies, connection pooling, lock-contention analysis, and storage-engine tuning for production systems handling millions of transactions. Build self-service database lifecycle automation — from one-click cluster provisioning and schema migrations to zero-downtime upgrades, blue-green deployments, and automated capacity scaling. Bridge relational and AI-native data infrastructure — extend traditional database exper
Get new senior staff software engineer observability jobs by email
Daily job updates · Unsubscribe anytime