Jobs in India

Senior Staff Software Engineer Observability in Bengaluru

356 active opportunities · Updated October 2026

Explore current senior staff software engineer observability jobs in Bengaluru. Filter by work mode, employment type, experience, department, date posted and distance.

Hiring demand

55/100

steady · 202 related jobs

Hiring trend

+24.4%

Job postings compared with the previous 30 days

Remote options

5.4%

Share of matching jobs listed as remote

G
📍 Bengaluru, India· Full-time
✓ High-confidence listingDemand 55/100
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. About the Role We are looking for Staff System Software Engineer in Test to join our team. In this role, you will be responsible for design, development, automation and reporting of Integration and system tests spanning across firmware and device drivers. This role requires you to have significant technical breadth and deep understanding of low-level system software specifically in server class systems. You will be part of a new team responsible for integration of different system software deliverables and development of system tests spanning all the components. You will contribute to shaping the test strategy , guide best practices and solve complex problems while maintaining a strong hands-on focus. You will partner with development and other QA teams to deliver high quality scalable and reliable solutions. About the Team Integration and system test team is responsible for verification and validation of integrated components across Board management controller (BMC), Firmware and Linux device driver. The team is also responsible for management and maintenance of common tools and pipel

PythonCI/CDGitLinux
G
📍 Bengaluru, India· Full-time
✓ Quality checkedCompany trend -82.6%

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. Intermediate Fullstack Engineer - Data Products An overview of this role Data Products integrates GitLab and third-party software development lifecycle data, and builds the dashboards, APIs, and data products that turn it into reliable, interoperable intelligence for customers and internal teams. You'll work across the stack, contributing to frontend experiences, backend services, APIs, and AI-enabled workflows. This is a hands-on product engineering role. You'll develop features, learn how we build and operate large-scale data systems, and work closely with Senior and Staff Engineers to deliver secure, reliable, and performant s

GitRestAIRuby
G
📍 Bengaluru, India· Full-time
✓ Quality checkedCompany trend -82.6%

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Staff Backend Engineer at GitLab, you will help shape a major investment in our Software Supply Chain Security offering. In this role, you'll serve as a senior technical leader for backend systems that help customers secure how software is built, verified, and delivered inside the GitLab platform. You'll work on foundational capabilities across package policy enforcement, build provenance, artifact signing, and malicious package detection, with a strong focus on enterprise-grade security and performance. You'll define architecture before systems are built, write clear technical proposals, and guide i

CI/CDGitRestAI
T
📍 Bengaluru, KARNATAKA, India
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a highly skilled and motivated Senior Design Verification Engineer to join our team. In this role, you will be responsible for the end-to-end verification of our IOMMU (Input/Output Memory Management Unit) IP. You will play a critical role in ensuring the functional correctness and performance of the design, taking ownership of the verification process right from the initial specification understanding down to final coverage closure. This role is hybrid, based out of Bangalore, India. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You have hands-on experience in ASIC or SoC verification using SystemVerilog and UVM. You enjoy debugging complex design and verification issues and working closely with cross-functional teams. You’re comfortable building verification environments from scratch and driving coverage closure. You have familiarity with standard bus protocols and modern verification tools. What We Need Experience owning block or subsystem-level verification from test planning to sign-off. Strong understanding of constrained-random verification, coverage analysis, and regression debugging. Familiarity with proto

G
📍 Bengaluru, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu

PythonLinuxAIC++
DC
📍 Bengaluru, KARNATAKA, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Role Overview Build reliable software services that power products, platforms, and business decisions. As a Senior Software Developer, you’ll design and deliver scalable applications, backend services, and integrations that perform well in production and evolve with changing business needs. You’ll apply strong software engineering practices across APIs, data-intensive applications, cloud services, AI-enabled solutions, and deployment pipelines. You’ll help shape technical solutions, improve system reliability, and contribute to a high-quality engineering culture. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design and develop scalable backend services and applications using Python or TypeScript. Lead the development of APIs, integrations, reusable software components, and AI-enabled features. Build reliable solutions for data ingestion, manipulation, service-to-service communication, and intelligent automation. Apply AI technologies and modern software engineering practices to improve product capabilities, developer productivity, and operational efficiency. Make sound technical decisions around architecture, performance, security, scalability, and maintainability. Deploy and operate applications using AWS services and CI/CD practices while improving testing, monitoring, documentation, and delivery standards. These are the essentials you’ll need to get an interview 5+ years of professional experience developing and delivering production software. Strong hands-on experience with Python; TypeScript or similar languages is also valuable. Proven experience building backend services, APIs, integrations, and service-oriented applications. Experience applying AI technologies, such as generative AI, machine learning services, intelligent automation, or AI-enabled application features. Strong understanding of software design principles, testing, debugging, performance optimization, and secure development. Experience working with cloud platfor

TypeScriptPythonAWSCI/CD
E
📍 Bengaluru, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As the Team Lead for Initiator & Protocol Engineering, you will spearhead the critical bridge between our industry-leading FlashArray and the Linux/VMWare ecosystems. You will drive the performance and reliability of our storage protocol stacks—spanning NVMe over Fabrics and Fibre Channel—ensuring Pure Storage remains the gold standard for enterprise connectivity. Collaborating closely with cross-functional hardware and software teams, you’ll mentor a high-caliber engineering squad to solve complex kernel-level challenges and influence the global Linux upstream community. WHAT YOU'LL DO Own the Protocol Lifecycle: Lead the development, maintenance, and optimization of Linux and VMWare initiator stacks (NVMeoF, FC-SCSI, iSCSI) and target drivers to ensure seamless, high-performance integration with Pure FlashArray. Drive System Resilience: Architect enhancements for Fibre Channel and NIC driver stacks that improve RAS (Reliability, Availability, and Serviceability), specifically focusing on multipathing logic and link health monitoring. Technical Leadership & Mentorship: Guide a team of senior and junior engineers through complex project deliveries, conducting deep-dive code reviews and setting the technical bar for C/C++ and Python development within the kernel space. Solve the Impossible: Act as the final escalation point for the most challenging system-level bugs found in the field or internal testing, u

PythonAWSLinuxRest
D
📍 Bengaluru, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. What You’ll Do: Architect the Future of AI Infrastructure: You will design, build, and own the end-to-end platform that supports the entire lifecycle of our ML models—from massive-scale distributed training to ultra-low-latency, highly-available inference. Optimize and Serve Cutting-Edge Models: You'll implement and scale sophisticated inference stacks for LLMs using frameworks like vLLM, TensorRT-LLM, or SGLang . You’ll solve complex challenges in throughput, latency, token streaming, and automated scaling to deliver a seamless user experience. Empower AI Innovation: You will act as a strategic partner to our AI Research and Data Science teams. You’ll create a seamless developer experience that accelerates their ability to experiment, fine-tune, and deploy groundbreaking models with velocity and confidence. Automate Everything: You'll develop robust CI/CD/CT (Continuous Training) pipelines using tools like Argo Workflows, ArgoCD, and GitHub Actions to automate model validation, deployment, and lifecycle management, ensuring our systems are both agile and rock-solid. What are we looking for Experience: 5+ years in infrastructure or software engineering, with at least 2+ years laser-focused on MLOps or ML infrastructu

PythonKubernetesCI/CDGit
G
📍 Bengaluru, India· Full-time
✓ Quality checkedCompany trend -82.6%

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As Director, Engineering, Platform Operations & Productivity, you'll own three functions that all require hands-on technical depth, not just people management. Platform Staff is a small, senior, AI-native team that moves to wherever the organization needs the most leverage, from standing up early scaffolding for an initiative, to taking on a high-impact customer request that doesn't fit any existing team's charter, to stepping directly into a production crisis until it's resolved. This role is for a technical engineering director who has personally built distributed systems, not only managed people wh

GitRestAIRust
O
📍 Bengaluru, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. Sr Staff Solutions Architect (Design Partner Program) The Challenge We're looking for a Solutions Architect to join our R&D Customer Engineering organisation and partner with strategic Design Partner customers. This is a highly customer-facing role that sits at the intersection of solution architecture, product strategy, and enterprise governance. You will work directly with customers to understand their business challenges, technical requirements, regulatory considerations, and strategic objectives. Acting as the bridge between customers, Product Management, Engineering, and Executive Leadership, you will help shape product direction, influence roadmap investments, and design scalable solutions that deliver measurable customer outcomes. Success in this role means helping customers operationalise trust, governance, privacy, risk management, consent, and responsible data use while accelerating adoption, improving implementation outcomes, and driving long-term platform value. Your Mission Partner directly with enterprise customers to understand business objectives, technical requirements, operational challenges, and

AWSGitAIGo
N
📍 Bengaluru, Bengaluru, India
✓ Quality checkedCompany trend -100%

NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services. What you’ll be doing: Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale. Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation. Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data. Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems. Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation. What we need to see: BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10&#43; years operating production infrastructure or platform services. Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges. <l

PythonDockerKubernetesLinux
N
📍 Bengaluru, Bengaluru, India
✓ Quality checkedCompany trend -100%

We are seeking a highly skilled and experienced Staff Network Site Reliability Engineer (SRE) to join our Enterprise Network Operations and SRE team. In this role, you will be pivotal in implementing our vision for a reliable and efficient network infrastructure. The ideal candidate is passionate about network operations and committed to enhancing the user experience. You'll have the opportunity to solve complex network challenges using hands-on debugging and by focusing on network automation, observability, documentation, and operational excellence. This is a critical position focused on ensuring user satisfaction and brilliance in network operations. What you'll be doing: Owning the operational aspect of the network infrastructure, ensuring its high availability and reliability, actively working on network incidents and service requests. Partnering with architecture and deployment teams to guarantee that new implementations are supportable and align with production standards. Advocating for and implementing automation to reduce toil and improve operational efficiency. Minimizing manual operational tasks to achieve and maintain Service Level Objectives (SLOs). Monitoring network performance, identifying areas for improvement, and collaborating with relevant teams to implement refinements. Proactively identifying and mitigating network risks to promote continuous improvement. Collaborating with domain experts across functions to resolve production issues swiftly and effectively, ensuring customer happiness. Conducting blameless postmortems and following through on Root Cause Analyses (RCAs). Discovering opportunities for operational improvements and teaming up with colleagues to devise solutions that enhance excellence and sustainability in network operations. Developing knowledge base articles for automa

PythonLinuxAnsible
O
📍 Bengaluru, India· Full-time
✓ Quality checkedCompany trend -74.1%

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta is searching for a SOX Senior IT Auditor to join its internal audit team and assist in the successful execution of Okta’s global SOX program. The ideal candidate is a self-motivated team player capable of overseeing IT general control and application control compliance walkthroughs and testing. Reporting to the SOX Program IT Manager, key operational responsibilities include: Perform SOX IT general control and application control walkthroughs and testing to determine whether internal controls over financial reporting are designed and operating effectively Actively follows and champions the SOX methodology with limited guidance Lead SOX IT auditors with confidence and help in their knowledge and development Able to pinpoint systemic causes of control breakdowns and the associated technical gap that generated or permitted the issue Review staff auditor work product and provide clear, actionable feedback to aid in their audit methodology understanding Identify opportunities, provide recommendations, and gain stakeholder agreement on root cause of issues and appropriate corrective actions Develop collaborative relationships with business and IT stakeholders Leverage technology in order to rationalize or automate control activities In addition to essential SOX responsibilities, individuals in this position are also responsible for assisting the Internal Audit team in risk-based operational audits Required Qualifications: BA/BS degree in accounting, finance

AWSRestMachine LearningAI
O
📍 Bengaluru, India
✓ Quality checkedCompany trend -74.1%

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Staff Backend Engineer We’re redefining Privileged Access Management (PAM) from the ground up, purpose-built for Cloud, SaaS, Databases, Containers, and any virtualized environment. Our mission is to simplify and secure workforce access with seamless, secure-by-default workflows that adapt dynamically to modern infrastructure. We eliminate standing privileges, enforce least privilege, and embed Zero Trust principles into every access workflow by default. About the Role We are seeking a Staff Backend Engineer to serve as the core technical anchor and senior Individual Contributor (IC) for our newly established engineering pod in India. At the P4 level, your primary sphere of influence will be at the team level —taking ownership of complex, ambiguous problems and defining how to solve them cleanly, securely, and efficiently. In this role, you will lead by example through hands-on architecture, high-velocity coding, and end-to-end execution. You will drive the implementation of secure database and network device connectors (routers, switches, firewalls) on top of our core Zero Standing Privileges (ZSP) platform. You will work closely with our local Technical Team Lead to elevate the pod’s engineering craft, acting as a technical multiplier for mid-level developers while ensuring tight architectural alignment with our global team. What You’ll Be Doing Execution & Technical Impact End-to-End Ownership: Consistently design, code, debug, test, moni

JavaPostgreSQLAWSDocker
🔔

Get new senior staff software engineer observability jobs in Bengaluru, India by email

Daily job updates · Unsubscribe anytime