Title: Senior Site Reliability Engineer - I, Product Area Focus Location: Noida (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering teams. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedit
Jobs in India
Senior Reliability Engineer in India
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current senior reliability engineer jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What you will be doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuous
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. As a Senior Site Reliability Engineer you will champion all things pertaining to reliability at Okta for Auth0. Working closely with the Product Engineers, Quality Engineers, Platform Engineers and Architecture teams, your primary focus will be on ensuring production systems remain operational at all times, while continually setting and achieving long-term performance, reliability and scalability goals in a platform with an exponential growth plan for the coming years. With Okta’s increased dedication to ensuring customer availability expectations are exceeded in every way, you will play a key role as we evolve our system architecture to meet the demands of enormous growth and support the hundreds of millions of users who rely on us to provide uninterrupted access to business-critical enterprise and consumer applications. Skills Exceptional communication skills, including technical writing in the English language Systematic problem-solving approach, coupled with a strong sense of ownership and drive Understanding of microservices, cloud infrastructure (AWS, Azure), databases (SQL, No-SQL, Key/Value), containers (docker, kubernetes), web technologies (web sockets, http) and networking (SSL, routing, VPN) Live and breathe SLIs, SLOs, error budgets and SLAs Strong belief in automating everything and reducing toil for yourself and teammates Loves to work as a team, but is able to work effectively in a remote environment where tasks may be self-driven Knowledge
We are seeking a highly skilled and experienced Staff Network Site Reliability Engineer (SRE) to join our Enterprise Network Operations and SRE team. In this role, you will be pivotal in implementing our vision for a reliable and efficient network infrastructure. The ideal candidate is passionate about network operations and committed to enhancing the user experience. You'll have the opportunity to solve complex network challenges using hands-on debugging and by focusing on network automation, observability, documentation, and operational excellence. This is a critical position focused on ensuring user satisfaction and brilliance in network operations. What you'll be doing: Owning the operational aspect of the network infrastructure, ensuring its high availability and reliability, actively working on network incidents and service requests. Partnering with architecture and deployment teams to guarantee that new implementations are supportable and align with production standards. Advocating for and implementing automation to reduce toil and improve operational efficiency. Minimizing manual operational tasks to achieve and maintain Service Level Objectives (SLOs). Monitoring network performance, identifying areas for improvement, and collaborating with relevant teams to implement refinements. Proactively identifying and mitigating network risks to promote continuous improvement. Collaborating with domain experts across functions to resolve production issues swiftly and effectively, ensuring customer happiness. Conducting blameless postmortems and following through on Root Cause Analyses (RCAs). Discovering opportunities for operational improvements and teaming up with colleagues to devise solutions that enhance excellence and sustainability in network operations. Developing knowledge base articles for automa
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . As a Software Dev Senior Engineer , you will own the reliability, scalability, and operational excellence of our Cloud-based services. You will define and enforce reliability standards, drive the adoption of SRE practices across engineering teams, and build the systems and tooling that keep our production infrastructure healthy. We follow a DevOps model: Development and Operations teams are integrated, and the SRE function acts as the reliability layer — setting Service Level Objectives, managing error budgets, and continuously reducing toil through engineering. Key Responsibilities: Define, publish, and continuously refine Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs ) for all critical services, partnering with product and engineering leadership. Own the error budget framework: track consumption, enforce error budget policies, and drive reliability investments when budgets are at risk. Lead the design and implementation of comprehensive observability platforms — metrics, structured logging, and distributed tracing — to ensure full visibility into pro
Who are we? FalconX is a pioneering team of operators, investors, and builders committed to revolutionizing institutional access to the crypto markets. Operating at the intersection of traditional finance and cutting-edge technology, FalconX addresses the industry's foremost challenges: Navigating the digital asset market can be complex and fragmented, with limited products and services that support trading strategies, structures, and liquidity found in conventional financial markets. As a comprehensive solution for all digital asset strategies from start to scale, FalconX operates as the connective tissue empowering clients with seamless navigation through the ever- evolving cryptocurrency landscape. Responsibilities Be part of a trading systems engineering team, dedicated to building out the core trading platforms. Work closely with cross functional teams to improve the system reliability, scalability and security. Engage in and improve the quality supporting the platform. Build and manage systems, infrastructure and applications through automation. Provide operational support to internal teams working on the platform. Work on improvements to bring in high efficiency, reduce latency, deploy systems faster. Practice sustainable incident response and blameless postmortems. Together with your engineering team, you will share an on-call rotation and be an escalation contact for service incidents. Implement and maintain rigorous security best practices across all infrastructure, with a focus on minimizing attack surface and ensuring data integrity. Monitor system health and performance with a keen eye for identifying and resolving issues before they affect trading activity. Manage user queries and service requests (often requiring in depth analysis of the technical and/or business logic of our systems). Proactive approach to problem analysis and resolution of production incidents. Manage Issue tracking and prioritisation of day to day production incidents. Manage platf
We are looking for a highly motivated AI/ML Software Engineer to join the Enterprise Agentic AI Platform team within IT. You will work closely with Business Analysts, and Engineering teams to design, develop, and deploy enterprise AI solutions that improve productivity and automate business workflows across Engineering, Operations, and Manufacturing. What you'll be doing: Design, develop, and deploy Agentic AI applications using Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and AI orchestration frameworks. Build scalable AI services and reusable components integrated with enterprise applications such as PLM, SAP, and other business systems. Collaborate with business and IT teams to translate business requirements into AI-driven solutions. Develop secure, scalable APIs and enterprise integrations to enable intelligent workflows and automation. Improve AI solution quality, performance, and reliability through prompt engineering, evaluation, and continuous optimization. Partner with cross-functional teams throughout the Software Development Lifecycle (SDLC), from solution design through deployment and production support. What we need to see: Bachelor's or Master's degree in Computer Science, Information Technology, AI/ML, or a related field. 6+ years of software engineering experience with strong proficiency in Python and backend application development. Hands-on experience with Generative AI, LLMs, RAG, AI agents, REST APIs, and cloud-native application development. Experience integrating enterprise applications and building scalable, production-ready software solutions. Strong analytical, problem-solving, communicatio
This role is responsible to maintain and enhance the operational excellence and reliability of all Control & Instrumentation systems associated with Boiler systems. This role is pivotal for conducting on-field maintenance activities including routine maintenance, timely overhauls, and calibration of equipment to prevent failures and ensure continuous operation. Source: Adani Group | Job ID: 54705
About the job Senior Backend Engineer |100% Remote We are searching for a seasoned Sr Backend Engineer who will be responsible for developing and maintaining our backend systems, ensuring their efficiency, scalability, and reliability. The ideal candidate has a strong background in PHP and Java, with optional experience in Golang and Node.js. You should be well-versed in working with databases such as MongoDB, Redis, and Postgres, and have a solid understanding of cloud technologies, specifically AWS. Responsibilities Design, develop, and maintain backend systems and APIs to support our application's functionality. Collaborate with cross-functional teams, including front-end developers, product managers, and designers, to deliver high-quality solutions. Write clean, scalable, and well-documented code that adheres to industry best practices and coding standards. Perform code reviews and provide constructive feedback to peers to ensure code quality and consistency. Optimize and improve the performance of existing backend systems. Troubleshoot and debug production issues, providing timely resolutions. Stay up-to-date with emerging technologies and industry trends, identifying opportunities for innovation and improvement. Collaborate with DevOps teams to ensure smooth deployment and operation of backend services in the AWS cloud environment. Requirements Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent work experience). 5-10 years of professional experience as a Backend Engineer. Strong proficiency in Java/Golang is a must. Experience with Node js is a plus. In-depth knowledge of database technologies, including MongoDB, Redis, and Postgres. Solid understanding of cloud computing platforms, particularly AWS. Familiarity with containerization technologies such as Docker and orchestration tools like Kubernetes. Proficiency in writing efficient and optimized SQL queries. Experience with version control systems, such as Git. Excellent pr
We are seeking a Senior Software Engineer with strong infrastructure expertise to design, build, and operate the next generation of our enterprise Observability, Automation, and AI-driven Reliability Platform. This role will build highly scalable distributed systems and platform services spanning Storage, Compute, Network, VMware, OpenShift, and bare-metal infrastructure. The engineer will help transform infrastructure operations from reactive monitoring and manual remediation to proactive, predictive, and AI-driven autonomous operations. What You Will Be Doing: Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at large scale. Develop reusable platform services, APIs, automation frameworks, and control planes that enable self-service, reduce operational toil, and automate infrastructure operations across multiple engineering teams. Build scalable telemetry and event-processing systems spanning metrics, logs, traces, events, topology, and alerts, with the performance and efficiency to process billions of infrastructure signals. Build intelligent and AI-native reliability capabilities, including agentic workflows for anomaly detection, forecasting, root-cause analysis, automated debugging, and closed-loop remediation. Drive technical architecture and engineering direction across Storage, Compute, Network, and Platform domains, solving complex and ambiguous problems that span multiple teams. Engineer for production at scale, with strong focus on software quality, scalability, security, performance, observability, maintainability, and operational readiness. Provide technical leadership and mentorship, influence engineerin
We are fueled by a moral imperative to advance mankind, and it all begins with our people, our product, and our purpose. Passion isn’t something we turn on and off; it’s woven into everything we do. If you thrive in high-challenge environments, are inspired by exceptional teammates, and are driven to grow beyond what you thought possible, MX is where you belong. Come build the future with us. Join an award-winning company that isn’t just shaping the financial industry, but transforming it in ways that create meaningful, lasting impact for millions of people. At MX, reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions, and customers feel every second of downtime. We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand. We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise, and leadership gets coverage and health as a program metric. This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought. Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, and data on PostgreSQL an
DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. About the job DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides an in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. In this role, you will Build core capabilities for our SaaS Platform across multiple clouds Drive development of functional enhancements for Data Discovery, Observability & Governance for both OSS and SaaS offering Lead efforts around non functional aspects like performance, scalability, reliability Lead and mentor junior engineers Work closely with PM, Customers and OSS community Requirements Over 8+ years of experience building and scaling backend systems, preferably in cloud-first or SaaS environments. Solve complex tech
DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You’ll Do: We are looking for a Senior Software Engineer – Platform Operations based in Pune, India, who will play a key role in ensuring the reliability, performance, and operational excellence of DeepIntent's platform and data ecosystem. This role requires a strong engineering mindset with the ability to troubleshoot complex technical issues, understand distributed data architectures, and collaborate across Engineering, Product, Analytics, and Customer-facing teams to deliver timely and effective solutions. As part of the Operations organization, you will work closely with Engineering to support production systems, improve operational processes, and drive platform stability. The ideal candidate is a self-motivated problem solver who is passionate about learning new technologies, improving system reliability, and delivering exceptional customer outcomes through engineering excellence. Serve as the engineering interface between Customer-facing teams, Analytics, Product, and Engineering organizations. Partner with Platform Support, Client Success, and other customer-facing teams to investigate and resolve complex platform-related issues. Analyze application, API, and data pipeline issues to identify root causes and drive timely resolution. Develop and standardize operational tools, and interfaces to support analytical and operational use cases. Monitor data pipeline executions, investigate failures, and implement corrective and pre
Role Overview Build reliable software services that power products, platforms, and business decisions. As a Senior Software Developer, you’ll design and deliver scalable applications, backend services, and integrations that perform well in production and evolve with changing business needs. You’ll apply strong software engineering practices across APIs, data-intensive applications, cloud services, AI-enabled solutions, and deployment pipelines. You’ll help shape technical solutions, improve system reliability, and contribute to a high-quality engineering culture. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design and develop scalable backend services and applications using Python or TypeScript. Lead the development of APIs, integrations, reusable software components, and AI-enabled features. Build reliable solutions for data ingestion, manipulation, service-to-service communication, and intelligent automation. Apply AI technologies and modern software engineering practices to improve product capabilities, developer productivity, and operational efficiency. Make sound technical decisions around architecture, performance, security, scalability, and maintainability. Deploy and operate applications using AWS services and CI/CD practices while improving testing, monitoring, documentation, and delivery standards. These are the essentials you’ll need to get an interview 5+ years of professional experience developing and delivering production software. Strong hands-on experience with Python; TypeScript or similar languages is also valuable. Proven experience building backend services, APIs, integrations, and service-oriented applications. Experience applying AI technologies, such as generative AI, machine learning services, intelligent automation, or AI-enabled application features. Strong understanding of software design principles, testing, debugging, performance optimization, and secure development. Experience working with cloud platfor
Other cities to consider
More places hiring for this role
Get new senior reliability engineer jobs in India by email
Daily job updates · Unsubscribe anytime