Staff Software Engineer Bengaluru, Karnataka, India Opportunity Get Well is seeking a visionary and technically adept Staff Software Engineer to architect, design, develop, and optimize our cloud-native healthcare platform while driving the adoption of AI-First and AI-Augmented Software Engineering practices. This role is pivotal in shaping the future of software development at Get Well by combining deep technical expertise with modern AI-assisted engineering workflows. As we evolve toward an AI-First engineering organization, this leader will champion the use of Generative AI, AI development assistants, and Agentic AI to improve developer productivity, software quality, and engineering velocity. The ideal candidate brings deep expertise in software architecture, cloud-native application development, AI-enabled engineering, and distributed systems. This role provides technical leadership across multiple engineering teams, ensuring high standards for architecture, code quality, reliability, security, and AI adoption. This is a hands-on leadership role where strategic thinking meets deep engineering execution within a complex healthcare environment. This position reports to the Director, Product Development and requires close collaboration with software engineers, AI engineers, product managers, DevOps, QA, and compliance specialists. Key Responsibilities Technical Leadership Define and drive the architecture of scalable, distributed healthcare platforms. Champion AI-First Software Engineering practices across the development lifecycle. Lead the adoption of AI-Augmented Development , including spec-driven development, AI-assisted coding, code reviews, testing, and documentation. Establish engineering standards, Cursor/AI coding guidelines, reusable patterns, and governance for responsible AI usage. Provide hands-on leadership in architecture, coding, design reviews, debugging, and performance optimization. Mentor enginee
Jobs in India
Software Reliability Engineer Specialist in Bengaluru
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current software reliability engineer specialist jobs in Bengaluru. Filter by work mode, employment type, experience, department, date posted and distance.
Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debug compl
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role combines deep technical expertise with people leadership responsibilities, including team development, prioritisation, mentoring and delivery coordination across multiple projects and stakeholders. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debug complex issues, optimize workloads and continuously imp
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software, and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers, and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Validation leadership team, the Principal Execution and Quality Validation Engineer will be responsible for driving validation execution, automation, product quality, and release readiness across Graphcore silicon and platform technologies. The role combines deep expertise in hardware validation, test automation, and quality engineering with a strong focus on execution excellence. Working closely with architecture, design, verification, firmware, software, systems engineering, and validation teams, the successful candidate will develop scalable validation methodologies, improve test coverage and execution efficiency, and ensure products meet the highest standards of functionality, reliability, and performance before customer deployment. As a recognized technical leader within the validation organisation, this role will influence validation strategies, automation roadmaps, and quality practices across multiple projects and engineering disciplines. The Team The Validation Execution and Quality team sits within the Validation organisation and is responsible for improving validation effectiveness, te
Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu
Senior -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to deb
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software, and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers, and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Validation leadership team, the Senior Execution and Quality Validation Engineer will be responsible for executing validation plans, developing automation solutions, and improving product quality across Graphcore silicon and platform technologies. Working closely with architecture, design, verification, firmware, software, systems engineering, and validation teams, the successful candidate will contribute to scalable validation methodologies, improve test coverage and execution efficiency, and help ensure products meet high standards of functionality, reliability, and performance before customer deployment. The role requires strong technical skills, attention to detail, and a passion for improving validation quality through effective execution, automation, and continuous improvement. The Team The Validation Execution and Quality team sits within the Validation organisation and is responsible for improving validation effectiveness, test execution efficiency, product quality, and release readiness across Graphcore silicon and platform products. The team develops validation methodologies, automation
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. As a Staff Site Reliability Engineer (SRE) at GitLab, you’ll help keep all user-facing services and production systems reliable, scalable, and efficient. Our SREs combine a pragmatic operations mindset with strong software engineering practices to drive automation, reduce toil, and improve resilience across our platform. In the Environment Automation specialization, your focus is on operating and automating hundreds of GitLab environments—from initial provisioning to day-to-day maintenance tasks. Unlike other SRE roles, this position centers on automating the lifecycle of many tenant environments, ensuring they remain secur
Title: Staff Site Reliability Engineer, Product Area Focus Location: Noida / Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering team. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedite
JOB TITLE Site Reliability Engineer A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO You will play a highly critical operational role where you will apply a combination of software and systems engineering skills to develop and maintain a complex set of distributed, real-time systems that serve critical stakeholders in Point72’s Global Macro business. You will focus on optimizing the operations of existing systems and infrastructure in an efficient manner, through a strict adherence to automation and tooling Specifically, you will: Build out foundational technical components of an extensive SRE program across multiple complex systems, both new and existing • Collaborate with our development and quant teams to ensure that ongoing change is consistent with a pre-determined, measurable set of SLOs spanning multiple complex user interactions with our systems • Monitor system capacity and performance, identifying and addressing potential future bottlenecks and sources of instability before they become impactful to our stakeholders • Review and provide feedback on automation code developed by peers to maintain high standards of code quality and efficiency • Troubleshoot and resolve system issues, analyzing their impact on infrastructure and service operations • Participate in or lead design reviews with peers and stakeholders, evaluating and selecting the best technologies and automation strategies for our needs WHAT’S REQUIRED We are looking for highly motivated, proactive engineers
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng
Job Title Senior Software Technologist I - C++ Job Description Software Engineer II Your Role: • Design, develop, test, and maintain software components and applications using modern C++, C# in a Windows-based environment. • Participate in the full software development lifecycle including requirements analysis, design, implementation, testing, debugging, and maintenance. • Develop and maintain software applications using Visual Studio and associated C++, C# development tools. • Work with Windows operating system fundamentals including processes, services, registry, file system, User Account Control (UAC), and application configuration. • Create, enhance, and troubleshoot software modules while adhering to coding standards, design guidelines, and software development best practices. • Utilize GitHub for source control management including branching strategies, commits, pull requests, merges, rebasing, and code reviews. • Support and troubleshoot CI/CD pipeline issues using GitHub Actions and participate in continuous integration activities. • Manage software dependencies and package management using NuGet and Conan. • Configure and maintain build systems using CMake and Visual Studio project configurations. • Develop and execute unit tests using Google Test (GTest) to ensure software quality and reliability. • Perform debugging, root cause analysis, and defect resolution for software issues identified during development, testing, and field support activities. • Participate in peer code reviews and contribute to software quality, maintainability, and technical excellence. • Collaborate effectively with Software Verification, Product Management, DevOps, Architecture, and cross-functional teams to deliver high-quality software solutions. • Create and maintain
You’ll shape the future of a business‑critical platform as the technical lead across both product engineering and cloud infrastructure. You’ll modernize a mature .NET application running on AWS today, while steering its evolution toward a cloud‑native, React/Node.js, AI‑enabled architecture. If you enjoy owning architecture end‑to‑end, from backend and frontend through CI/CD, DevOps, and AWS infrastructure, this role gives you real influence at Staff Engineer level and the opportunity to set engineering standards that others follow. You’ll spend your time leading complex .NET and React features, designing scalable AWS infrastructure with Infrastructure as Code, and building automation that makes releases fast, safe, and repeatable. You’ll work on performance, reliability, and modernization in equal measure—fixing what’s slowing the platform down today and designing what it will look like in the next generation. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture and development of enterprise .NET services and APIs that power a business‑critical platform. Design and operate AWS infrastructure (using AWS CDK in TypeScript) to support secure, scalable, multi‑environment deployments. Build and optimize CI/CD pipelines (AWS CodePipeline, CodeBuild, Windows build agents) to make shipping .NET and React changes fast and reliable. Drive modernization initiatives across the stack, including clean architecture, refactoring legacy components, and reducing technical debt. Design and tune PostgreSQL and MSSQL database solutions for performance, scalability, and reliability. Mentor engineers and influence engineering practices across teams, raising the bar on cloud, DevOps, and software design. These are the essentials you’ll need to get an interview Significant experience (typically 8+ years) delivering and operating scalable enterprise software, owning both application code and cloud infrastructure. Deep hands‑on expertise with C
Role Overview Build reliable software services that power products, platforms, and business decisions. As a Senior Software Developer, you’ll design and deliver scalable applications, backend services, and integrations that perform well in production and evolve with changing business needs. You’ll apply strong software engineering practices across APIs, data-intensive applications, cloud services, AI-enabled solutions, and deployment pipelines. You’ll help shape technical solutions, improve system reliability, and contribute to a high-quality engineering culture. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design and develop scalable backend services and applications using Python or TypeScript. Lead the development of APIs, integrations, reusable software components, and AI-enabled features. Build reliable solutions for data ingestion, manipulation, service-to-service communication, and intelligent automation. Apply AI technologies and modern software engineering practices to improve product capabilities, developer productivity, and operational efficiency. Make sound technical decisions around architecture, performance, security, scalability, and maintainability. Deploy and operate applications using AWS services and CI/CD practices while improving testing, monitoring, documentation, and delivery standards. These are the essentials you’ll need to get an interview 5+ years of professional experience developing and delivering production software. Strong hands-on experience with Python; TypeScript or similar languages is also valuable. Proven experience building backend services, APIs, integrations, and service-oriented applications. Experience applying AI technologies, such as generative AI, machine learning services, intelligent automation, or AI-enabled application features. Strong understanding of software design principles, testing, debugging, performance optimization, and secure development. Experience working with cloud platfor
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are hiring a Staff Software Engineer for our Frontier Security AI team. Snowflake's Frontier Security AI teams develop production-grade LLM applications, intelligent agents, AI infrastructure, and evaluation systems for enterprise customers — products that must meet a high bar for quality, security, reliability, and efficiency while operating over sensitive data at large scale. In this role, you will lead the design and development of our Agentic Harness and agent evaluation platform, working across product, infrastructure, applied AI, security, and modeling teams to take new capabilities from prototype to dependable customer value. AS A STAFF SOFTWARE ENGINEER AT SNOWFLAKE, YOU WILL: Architect and build the Agentic Harness that executes complex, multi-step AI workflows across models, tools, data, and services. Design stable interfaces for tool execution, context construction, state management, memory, permissions, retries, fallbacks, and human review. Own agent quality end to end by building evaluation harnesses, representative datasets, automated graders, experiment pipelines, and release gates. Convert ambiguous reports such as "the agent feels worse" into measurable failure modes, reproducible tests, and durable fixes. Analyze production agent trajectories to identif
Other cities to consider
More places hiring for this role
Get new software reliability engineer specialist jobs in Bengaluru, India by email
Daily job updates · Unsubscribe anytime