JOB TITLE Site Reliability Engineer A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO You will play a highly critical operational role where you will apply a combination of software and systems engineering skills to develop and maintain a complex set of distributed, real-time systems that serve critical stakeholders in Point72’s Global Macro business. You will focus on optimizing the operations of existing systems and infrastructure in an efficient manner, through a strict adherence to automation and tooling Specifically, you will: Build out foundational technical components of an extensive SRE program across multiple complex systems, both new and existing • Collaborate with our development and quant teams to ensure that ongoing change is consistent with a pre-determined, measurable set of SLOs spanning multiple complex user interactions with our systems • Monitor system capacity and performance, identifying and addressing potential future bottlenecks and sources of instability before they become impactful to our stakeholders • Review and provide feedback on automation code developed by peers to maintain high standards of code quality and efficiency • Troubleshoot and resolve system issues, analyzing their impact on infrastructure and service operations • Participate in or lead design reviews with peers and stakeholders, evaluating and selecting the best technologies and automation strategies for our needs WHAT’S REQUIRED We are looking for highly motivated, proactive engineers
Jobs in India
Reliability Engineer Iii in Bengaluru
114 active opportunities · Updated October 2026
Showing
15 jobs
Explore current reliability engineer iii jobs in Bengaluru. Filter by work mode, employment type, experience, department, date posted and distance.
JOB TITLE Data Reliability Engineer A CAREER WITH CUBIST Cubist Systematic Strategies, an affiliate of Point72, deploys systematic, computer-driven trading strategies across multiple liquid asset classes, including equities, futures, and foreign exchange. The core of our effort is rigorous research into a wide range of market anomalies, fueled by our unparalleled access to a wide range of publicly available data sources. What you’ll do Ensure smooth day-to-day implementation of a large research infrastructure and the timely delivery of comprehensive and error-free data to Cubist’s portfolio managers across the globe Serve as a frontline owner for mission-critical data ETL pipelines that power trading and investment decision-making, ensuring reliability, accuracy, and timeliness. Actively manage and resolve data incidents in a fast-paced trading environment, partnering closely with investment professionals, data scientists, and external data vendors. Design and build tooling, automation, and robust documentation to improve operational efficiency, scalability, and data quality across the platform. Play a hands-on role in daily data operations, including data validation, remediation, and enrichment, with opportunities to continuously improve and modernize workflows through engineering best practices. What’s REQUIRED Bachelor’s degree in computer science or a related field. Strong proficiency in SQL Server and Python programming, with experience in AWS and both Windows and Linux environments. Exceptional attention to detail with a strong appreciation for well-defined processes and systems. 3+ years of experience in a client-facing support or operations role. Excellent organizational, communication, and interpersonal skills. Commitment to the highest ethical standards About point72 Point72 is a leading global alternative investment firm led by Steven A. Cohen. Building on more than 30 years of investing experience, Poin
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta’s TDI Network Engineering team is responsible for the global corporate network, building and supporting a high-performing, reliable network at scale. As a member of this team, you will have a direct impact on network design, deployment, and reliability, enabling our employees to work effectively from any location globally. Your role ensures the overall security and integrity of our corporate network by leveraging network security best practices, innovative products, and rigorous security validation. Reporting to the Network Engineering Manager, this operations-focused role is distinct from core Network Engineering and Network Security, centering primarily on operational execution—including responding to alerts, maintaining service availability, and ensuring system health across our global enterprise network. You will drive the strategic reduction of systemic toil and technical debt across multiple teams, applying a systems-level perspective and leveraging deep expertise in Distributed Systems, Networking fundamentals, Infrastructure as Code, and observability to architect scalable platforms and lead technical efforts to ensure an "Always Secure. Always On." environment. You will own multi-quarter objectives and establish long-term strategies for network reliability. What you'll be doing : Design and Own the resilience, health and availability of our entire global corporate network domain, managing operational responsibilities such as responding to aler
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. As a Staff Site Reliability Engineer (SRE) at GitLab, you’ll help keep all user-facing services and production systems reliable, scalable, and efficient. Our SREs combine a pragmatic operations mindset with strong software engineering practices to drive automation, reduce toil, and improve resilience across our platform. In the Environment Automation specialization, your focus is on operating and automating hundreds of GitLab environments—from initial provisioning to day-to-day maintenance tasks. Unlike other SRE roles, this position centers on automating the lifecycle of many tenant environments, ensuring they remain secur
Title: Staff Site Reliability Engineer, Product Area Focus Location: Noida / Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering team. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedite
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. At Okta, our motto is "Always On" and nowhere do we embrace that more than in Technical Operations. We strive to build the most reliable and performant systems on the planet through the skillful use of automation. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. You will work on: Leading the architecture, design and rollout of the internal developer platform which would span CI/CD, tooling, infrastructure as code (IAC) integrations as well as modernization efforts Enhance existing automation for fleet management and build new features as per the needs of the infrastructure platform Spearhead initiatives and projects that enhance engineer productivity by identifying and mitigating bottlenecks in the development flow Mentoring, managing, and leading a team of SWEs & SREs with a broad range of expertise and experience Triaging and troubleshooting complex production issues to ensure reliability and performance. Working closely with our stakeholders across the organization to ensure our new capabilities are aligned to our competing constraints of reliability, security, and delivery velocity Partnering directly with recruiting and people ops to hire and retain the best talent
DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. About the Role We're seeking an experienced DevOps/ Site Reliability Engineering (SRE) Engineer to join DataHub and drive the reliability, scalability, and operational excellence of our platform offerings. In this role, you'll work on technical initiatives across DataHub Cloud and our emerging enterprise deployment solution, which provides customers with enhanced control and flexibility for running DataHub in their preferred environments. Key Responsibilities Enterprise Platform Development: Partner with product and engineering teams to influence the development of advanced deployment capabilities. Collaborate with cross-functional teams to help build systems for seamless installation, upgrade, and rollback processes across various environments. Influence the design and help implement comprehensive monitoring and health check systems for distributed deployments. Partner with engineering teams to help develop self-healing and automated remediation capabilities. Platform Reliability and Operations: Establish and maintain SLAs/SLOs for both cloud and enterprise offerings. Lead incident response and post-mortem processes to drive continuous improvement. Optimise system performance, capacity planning, and cost efficiency. Work closely with product, engineerin
Who are we? FalconX is a pioneering team of operators, investors, and builders committed to revolutionizing institutional access to the crypto markets. Operating at the intersection of traditional finance and cutting-edge technology, FalconX addresses the industry's foremost challenges: Navigating the digital asset market can be complex and fragmented, with limited products and services that support trading strategies, structures, and liquidity found in conventional financial markets. As a comprehensive solution for all digital asset strategies from start to scale, FalconX operates as the connective tissue empowering clients with seamless navigation through the ever- evolving cryptocurrency landscape. Responsibilities Be part of a trading systems engineering team, dedicated to building out the core trading platforms. Work closely with cross functional teams to improve the system reliability, scalability and security. Engage in and improve the quality supporting the platform. Build and manage systems, infrastructure and applications through automation. Provide operational support to internal teams working on the platform. Work on improvements to bring in high efficiency, reduce latency, deploy systems faster. Practice sustainable incident response and blameless postmortems. Together with your engineering team, you will share an on-call rotation and be an escalation contact for service incidents. Implement and maintain rigorous security best practices across all infrastructure, with a focus on minimizing attack surface and ensuring data integrity. Monitor system health and performance with a keen eye for identifying and resolving issues before they affect trading activity. Manage user queries and service requests (often requiring in depth analysis of the technical and/or business logic of our systems). Proactive approach to problem analysis and resolution of production incidents. Manage Issue tracking and prioritisation of day to day production incidents. Manage platf
Role Overview Build the software services that power products, platforms, and better business decisions. As a Software Engineer II, you’ll develop scalable backend applications, APIs, integrations, and AI-enabled features using Python and cloud technologies. You’ll contribute to solutions from design through production, helping improve reliability, performance, security, and developer productivity. This is an opportunity to solve meaningful engineering challenges, grow your technical ownership, and collaborate with experienced engineers across the development lifecycle. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design and build scalable backend services, REST APIs, integrations, and reusable software components using Python and, where relevant TypeScript. Develop data ingestion, transformation, service-to-service communication, and automation capabilities that support reliable product experiences. Contribute to AI-enabled features and use AI development tools responsibly to improve coding, testing, research, documentation, and delivery. Apply sound engineering practices across architecture, performance, security, testing, debugging, and maintainability. Deploy and operate services using AWS and CI/CD workflows, contributing to monitoring, troubleshooting, documentation, and continuous improvement. Partner with engineers and cross-functional colleagues through design discussions, code reviews, technical problem-solving, and knowledge sharing. These are the essentials you’ll need to get an interview 3–5 years of professional experience building and delivering production software in an agile environment. Strong hands-on experience with Python and backend development, including APIs, integrations, or service-oriented applications. Experience working with cloud platforms, preferably AWS, and familiarity with deployment or CI/CD practices. Working knowledge of software design principles, testing, debugging, performance optimization, an
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are hiring a Staff Software Engineer for our Frontier Security AI team. Snowflake's Frontier Security AI teams develop production-grade LLM applications, intelligent agents, AI infrastructure, and evaluation systems for enterprise customers — products that must meet a high bar for quality, security, reliability, and efficiency while operating over sensitive data at large scale. In this role, you will lead the design and development of our Agentic Harness and agent evaluation platform, working across product, infrastructure, applied AI, security, and modeling teams to take new capabilities from prototype to dependable customer value. AS A STAFF SOFTWARE ENGINEER AT SNOWFLAKE, YOU WILL: Architect and build the Agentic Harness that executes complex, multi-step AI workflows across models, tools, data, and services. Design stable interfaces for tool execution, context construction, state management, memory, permissions, retries, fallbacks, and human review. Own agent quality end to end by building evaluation harnesses, representative datasets, automated graders, experiment pipelines, and release gates. Convert ambiguous reports such as "the agent feels worse" into measurable failure modes, reproducible tests, and durable fixes. Analyze production agent trajectories to identif
Position Overview We are looking for a Software Engineer II to build and deliver scalable software solutions across our products. You will work on modern web applications and cloud-based services using Node.js, React, TypeScript, AWS, PostgreSQL, MSSQL, and Docker, while contributing to AI-enabled features and integrations. You will collaborate closely with other engineers, product managers, and cross-functional teams to develop reliable, maintainable, and production-ready solutions. This role provides an opportunity to work with modern AI technologies including Python, AWS Bedrock, MCP, RAG, and agentic AI workflows while developing strong expertise in cloud-native software engineering. What You'll Do Develop and maintain scalable backend services and APIs using Node.js, TypeScript, and JavaScript. Build responsive and maintainable frontend applications using React. Design and implement integrations with AWS services and contribute to cloud-native application development. Develop and maintain applications using PostgreSQL and MSSQL, including writing efficient queries and working with database schemas. Build, test, and deploy applications using Docker and modern CI/CD practices. Contribute to AI-enabled product features using Python, AWS Bedrock, RAG, MCP, and AI integration patterns. Work with the team to integrate LLM capabilities, APIs, tools, and data sources into production applications. Write clean, maintainable, and well-tested code following established engineering practices. Participate in code reviews, technical discussions, debugging, and production issue resolution. Develop unit and integration tests and contribute to improving application quality and reliability. Monitor application performance and troubleshoot issues across development and production environments. Collaborate with senior engineers and architects to implement technical solutions aligned with product and engineering requirements. Stay current with emerging technologies, particularly in
About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. About the role We are looking for a Quality Architect/Lead with hands-on experience building quality systems and has deep expertise in building and scaling test automation frameworks.The role requires an individual who applies systems thinking to solving complex problems. They should be able to understand the product from various perspectives and be able to effectively create testing programs that validate not just functionality but performance, reliability and user experience.DevRev is building a next generation AI native product that requires us to build novel test systems for the Agent AI platform. The role is mult-faceted and is going to continuously evolve with time. What you'll do Test Case Design and Documentation Actively use AI and intelligent agents to accelerate test generation, test maintenance, and coverage expansion. Leverage LLMs to convert requirements, user stories, and production incidents into high-quality automated test cases. Reduce reliance on manual test case creation by introducing AI-assisted automation workflows, with human review and ownership. Apply AI to optimize test selection, prioritization, and execution based on risk, code changes, and historical failures. Use AI to assist in identi
DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. About the job DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides an in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. In this role, you will Build core capabilities for our SaaS Platform across multiple clouds Drive development of functional enhancements for Data Discovery, Observability & Governance for both OSS and SaaS offering Lead efforts around non functional aspects like performance, scalability, reliability Lead and mentor junior engineers Work closely with PM, Customers and OSS community Requirements Over 8+ years of experience building and scaling backend systems, preferably in cloud-first or SaaS environments. Solve complex tech
Toradex is a global company strongly focused on engineering & technology. We’re powered by a diverse & uniquely gifted workforce. We pursue the best people to propel our innovative vision of embedded computing and IoT. If you’re interested in being a driving force at an agile technology company, engineering clever computing solutions & helping other companies bring their products to life, we should talk. Description We are looking for a DevOps Engineer to strengthen our cloud operations and engineering practices, with a focus on reliable website delivery, secure AWS foundations, and fast but controlled delivery of new services. The position combines AWS operations, infrastructure as code, CI/CD, automation, and pragmatic software engineering. The person should be confident working with services for edge delivery, compute, storage, databases, DNS, security, and observability without relying on manual console changes as the default operating model. The role also supports on-premises to cloud migration, global service optimization, and practical responses to increasing AI-driven traffic. We value candidates who can use modern AI-assisted development effectively to spin up proof-of-concept projects quickly, while still applying disciplined Git, review, security, and deployment practices. About you You enjoy building stable, secure, and maintainable infrastructure that supports business-critical services. You can work independently and take ownership of cloud environments, deployments, and operational improvements. You are comfortable balancing speed, reliability, cost, and security when making technical decisions. You communicate clearly with technical and non-technical stakeholders and explain trade-offs in a practical way. You document your work well and create clear runbooks and support material for future maintenance. You are methodical when troubleshooting incidents and stay calm when systems are under pressure. You are curious about modern traffic patt
Other cities to consider
More places hiring for this role
Get new reliability engineer iii jobs in Bengaluru, India by email
Daily job updates · Unsubscribe anytime