DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. About the Role We're seeking an experienced DevOps/ Site Reliability Engineering (SRE) Engineer to join DataHub and drive the reliability, scalability, and operational excellence of our platform offerings. In this role, you'll work on technical initiatives across DataHub Cloud and our emerging enterprise deployment solution, which provides customers with enhanced control and flexibility for running DataHub in their preferred environments. Key Responsibilities Enterprise Platform Development: Partner with product and engineering teams to influence the development of advanced deployment capabilities. Collaborate with cross-functional teams to help build systems for seamless installation, upgrade, and rollback processes across various environments. Influence the design and help implement comprehensive monitoring and health check systems for distributed deployments. Partner with engineering teams to help develop self-healing and automated remediation capabilities. Platform Reliability and Operations: Establish and maintain SLAs/SLOs for both cloud and enterprise offerings. Lead incident response and post-mortem processes to drive continuous improvement. Optimise system performance, capacity planning, and cost efficiency. Work closely with product, engineerin
Jobs in India
Engineering Manager Platform Reliability in Bengaluru
392 active opportunities · Updated October 2026
Showing
15 jobs
Explore current engineering manager platform reliability jobs in Bengaluru. Filter by work mode, employment type, experience, department, date posted and distance.
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the FlashArray team to build high-performance, resilient storage software that powers mission-critical applications worldwide. As a core member of an agile engineering group, you will design and deliver zero-downtime algorithms that directly shape our enterprise storage and public cloud offerings. In this role, you will partner closely with global systems engineers and product teams to translate complex technical challenges into scalable, real-world solutions. Your work will directly impact how thousands of global enterprises manage, protect, and scale their data seamlessly. WHAT YOU’LL DO Design & Deliver Resilient Systems: Architect and implement high-performance algorithms for enterprise storage products, ensuring platform reliability, end-to-end delivery from concept to release, and six-nines availability. Expand Cloud Architecture: Extend core platform capabilities into public cloud environments (such as AWS), driving performance and agility for both traditional IT and cloud-native applications. Drive Problem Solving & Quality: Analyze and resolve complex systems software challenges, optimizing storage internals and contributing to continuous continuous platform upgrades. Collaborate & Mentor: Partner with multidisciplinary engineering peers to review code, refine architecture, and maintain high engineering standards across distributed software projects. WHAT YOU BRING Systems Programming Exp
DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. About the job DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides an in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. In this role, you will Build core capabilities for our SaaS Platform across multiple clouds Drive development of functional enhancements for Data Discovery, Observability & Governance for both OSS and SaaS offering Lead efforts around non functional aspects like performance, scalability, reliability Lead and mentor junior engineers Work closely with PM, Customers and OSS community Requirements Over 8+ years of experience building and scaling backend systems, preferably in cloud-first or SaaS environments. Solve complex tech
Toradex is a global company strongly focused on engineering & technology. We’re powered by a diverse & uniquely gifted workforce. We pursue the best people to propel our innovative vision of embedded computing and IoT. If you’re interested in being a driving force at an agile technology company, engineering clever computing solutions & helping other companies bring their products to life, we should talk. Description We are looking for an Integration Platform Specialist to manage company-wide integrations using Workato. The role focuses on building reliable workflows, supporting API-based connections, and working with teams to automate business processes. The position requires strong knowledge of REST APIs, GraphQL, Postman, and the Workato platform. Experience with Workato MCP functionality and coding skills in Python, Ruby, or similar languages are preferred. About You You enjoy solving complex system and process problems with practical, scalable solutions. You can work independently and take ownership of integrations from discovery through deployment and support. You communicate clearly with both technical and non-technical stakeholders. You document your work well and create clear support material for future maintenance. You are curious, hands-on, and willing to investigate issues until you find the root cause. You care about reliability, data quality, security, and a good internal user experience. You are comfortable working across teams and balancing business priorities with technical constraints. Key Responsibilities Own the integration platform roadmap and day-to-day operation, ensuring business-critical automations are reliable, observable, and maintainable. Partner with business and application owners to turn process gaps into pragmatic integration designs and delivery plans. Build Workato recipes, custom connectors, and reusable patterns that reduce manual work and improve data flow between systems. Maintain and modernize existing integrations,
JOB TITLE Data Reliability Engineer A CAREER WITH CUBIST Cubist Systematic Strategies, an affiliate of Point72, deploys systematic, computer-driven trading strategies across multiple liquid asset classes, including equities, futures, and foreign exchange. The core of our effort is rigorous research into a wide range of market anomalies, fueled by our unparalleled access to a wide range of publicly available data sources. What you’ll do Ensure smooth day-to-day implementation of a large research infrastructure and the timely delivery of comprehensive and error-free data to Cubist’s portfolio managers across the globe Serve as a frontline owner for mission-critical data ETL pipelines that power trading and investment decision-making, ensuring reliability, accuracy, and timeliness. Actively manage and resolve data incidents in a fast-paced trading environment, partnering closely with investment professionals, data scientists, and external data vendors. Design and build tooling, automation, and robust documentation to improve operational efficiency, scalability, and data quality across the platform. Play a hands-on role in daily data operations, including data validation, remediation, and enrichment, with opportunities to continuously improve and modernize workflows through engineering best practices. What’s REQUIRED Bachelor’s degree in computer science or a related field. Strong proficiency in SQL Server and Python programming, with experience in AWS and both Windows and Linux environments. Exceptional attention to detail with a strong appreciation for well-defined processes and systems. 3+ years of experience in a client-facing support or operations role. Excellent organizational, communication, and interpersonal skills. Commitment to the highest ethical standards About point72 Point72 is a leading global alternative investment firm led by Steven A. Cohen. Building on more than 30 years of investing experience, Poin
Who are we? FalconX is a pioneering team of operators, investors, and builders committed to revolutionizing institutional access to the crypto markets. Operating at the intersection of traditional finance and cutting-edge technology, FalconX addresses the industry's foremost challenges: Navigating the digital asset market can be complex and fragmented, with limited products and services that support trading strategies, structures, and liquidity found in conventional financial markets. As a comprehensive solution for all digital asset strategies from start to scale, FalconX operates as the connective tissue empowering clients with seamless navigation through the ever- evolving cryptocurrency landscape. Responsibilities Be part of a trading systems engineering team, dedicated to building out the core trading platforms. Work closely with cross functional teams to improve the system reliability, scalability and security. Engage in and improve the quality supporting the platform. Build and manage systems, infrastructure and applications through automation. Provide operational support to internal teams working on the platform. Work on improvements to bring in high efficiency, reduce latency, deploy systems faster. Practice sustainable incident response and blameless postmortems. Together with your engineering team, you will share an on-call rotation and be an escalation contact for service incidents. Implement and maintain rigorous security best practices across all infrastructure, with a focus on minimizing attack surface and ensuring data integrity. Monitor system health and performance with a keen eye for identifying and resolving issues before they affect trading activity. Manage user queries and service requests (often requiring in depth analysis of the technical and/or business logic of our systems). Proactive approach to problem analysis and resolution of production incidents. Manage Issue tracking and prioritisation of day to day production incidents. Manage platf
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Portworx team to build and deliver our highest-quality product suite. In this role, you will write clean, scalable code with a strong focus on quality, reliability, and user-centric design. You will directly contribute to building a new SaaS platform that delivers a secure, consistent, and best-in-class experience for customers purchasing and managing Portworx offerings. As a core developer, you will take ownership of designing and implementing critical features across the entire Portworx portfolio. WHAT YOU’LL DO Design & Scale SaaS Microservices: Develop, test, and integrate high-performance microservices and features into the Portworx product suite, ensuring high availability in distributed systems. Drive End-to-End Delivery: Lead software lifecycle activities including architectural design, code reviews, unit/functional testing, documentation, and continuous integration and deployment (CI/CD). Partner Across Teams: Collaborate with product managers, cross-functional engineering peers, and early-adopter customers to transform requirements into production-ready software. Own Product Quality & Iteration: Take full ownership of feature stability by proactively incorporating customer feedback and rapidly resolving issues identified during testing and deployment. Innovate & Experiment: Research emerging technologies and cloud infrastructure tools to push performance boundaries and continuously i
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Portworx team to build and deliver our highest-quality product suite. In this role, you will write clean, scalable code with a strong focus on quality, reliability, and user-centric design. You will directly contribute to building a new SaaS platform that delivers a secure, consistent, and best-in-class experience for customers purchasing and managing Portworx offerings. As a core developer, you will take ownership of designing and implementing critical features across the entire Portworx portfolio. WHAT YOU’LL DO Design & Scale SaaS Microservices: Develop, test, and integrate high-performance microservices and features into the Portworx product suite, ensuring high availability in distributed systems. Drive End-to-End Delivery: Lead software lifecycle activities including architectural design, code reviews, unit/functional testing, documentation, and continuous integration and deployment (CI/CD). Partner Across Teams: Collaborate with product managers, cross-functional engineering peers, and early-adopter customers to transform requirements into production-ready software. Own Product Quality & Iteration: Take full ownership of feature stability by proactively incorporating customer feedback and rapidly resolving issues identified during testing and deployment. Innovate & Experiment: Research emerging technologies and cloud infrastructure tools to push performance boundaries and continuously i
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. You Will: Data Architecture and Design: Designing and overseeing the architecture of scalable and reliable data platforms, including data pipelines, storage solutions, and processing systems Data Modelling and Management:Developing and implementing data models, ensuring data quality, and establishing data governance policies Data Pipeline Development: Building and optimising data pipelines for ingesting, processing, and transforming large datasets from various sources Performance Optimisation: Identifying and resolving performance bottlenecks in data pipelines and systems, ensuring efficient data retrieval and processing Technology Evaluation and Innovation: Staying abreast of emerging data technologies and exploring opportunities for innovation to improve the organisation’s data infrastructure Troubleshooting and Problem Solving: Diagnosing and resolving complex data-related issues, ensuring the stability and reliability of the data platform Data Security and Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data Perform other duties as assigned You Have: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field. 10+ years of experience in data engineering or a similar role. Enterprise SaaS software solutions with high availability and scalability Solution handling large scale structured and unstructured data from varied data sources Experience in building and maintaining data platform systems such as distributed compute,
About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge We're looking for a Staff Software Engineer to lead technical direction across a major feature area or system domain at OneTrust. Staff Engineers here own outcomes, not just designs; they decide how ambiguous, cross-cutting problems get solved when no existing playbook applies, and their judgment carries weight across teams they don't formally manage. Your Mission Technical Leadership & Architecture Lead architecture and design for systems with significant scope and blast radius, ensuring decisions hold up under real growth, compliance, and reliability constraints; not just initial requirements. Paying attention to application performance Exercise judgment on where AI-assisted tooling accelerates delivery and where deeper human design thinking is required Cross-Team Collaboration Partner with Product, UX, and other engineering teams early, shaping problems before solutions are locked in. Build working relationships and technical credibility beyond your immediate team. Quality & Standards Set engineering practices for code revie
Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge We're looking for a Staff Software Engineer to lead technical direction across a major feature area or system domain at OneTrust. Staff Engineers here own outcomes, not just designs; they decide how ambiguous, cross-cutting problems get solved when no existing playbook applies, and their judgment carries weight across teams they don't formally manage. Your Mission Technical Leadership & Architecture Lead architecture and design for systems with significant scope and blast radius, ensuring decisions hold up under real growth, compliance, and reliability constraints; not just initial requirements. Paying attention to application performance Exercise judgment on where AI-assisted tooling accelerates delivery and where deeper human design thinking is required Cross-Team Collaboration Partner with Product, UX, and other engineering teams early, shaping problems before solutions are locked in. Build working relationships and technical credibility beyond your immediate team. Quality & Standards Set engineering practices for code revie
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . We are seeking an experienced Principal Software Engineer to lead the architecture, development, and evolution of a large-scale malware analysis and cybersecurity platform. This role will own the end-to-end technical architecture across Python-based analysis pipelines, message-driven job orchestration systems, Windows VM sandbox environments, and web-based interfaces. The ideal candidate brings deep expertise in software engineering, malware analysis, and distributed systems, with a proven ability to take ownership of complex legacy platforms and drive modernization initiatives. You will lead the design and implementation of advanced detection capabilities, including behavioral analysis, malware sandboxing, machine learning-based classification, and support for new file types, while ensuring platform scalability, reliability, and performance. This position requires strong hands-on development skills in Python with Linux-based production environments, messaging architectures such as RabbitMQ, virtualization technologies, and cloud-native deployment practices. As the technical leader for the platform,
Staff Software Engineer Bengaluru, Karnataka, India Opportunity Get Well is seeking a visionary and technically adept Staff Software Engineer to architect, design, develop, and optimize our cloud-native healthcare platform while driving the adoption of AI-First and AI-Augmented Software Engineering practices. This role is pivotal in shaping the future of software development at Get Well by combining deep technical expertise with modern AI-assisted engineering workflows. As we evolve toward an AI-First engineering organization, this leader will champion the use of Generative AI, AI development assistants, and Agentic AI to improve developer productivity, software quality, and engineering velocity. The ideal candidate brings deep expertise in software architecture, cloud-native application development, AI-enabled engineering, and distributed systems. This role provides technical leadership across multiple engineering teams, ensuring high standards for architecture, code quality, reliability, security, and AI adoption. This is a hands-on leadership role where strategic thinking meets deep engineering execution within a complex healthcare environment. This position reports to the Director, Product Development and requires close collaboration with software engineers, AI engineers, product managers, DevOps, QA, and compliance specialists. Key Responsibilities Technical Leadership Define and drive the architecture of scalable, distributed healthcare platforms. Champion AI-First Software Engineering practices across the development lifecycle. Lead the adoption of AI-Augmented Development , including spec-driven development, AI-assisted coding, code reviews, testing, and documentation. Establish engineering standards, Cursor/AI coding guidelines, reusable patterns, and governance for responsible AI usage. Provide hands-on leadership in architecture, coding, design reviews, debugging, and performance optimization. Mentor enginee
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About the Role Okta is the identity standard. The Okta Identity Cloud is an independent and neutral platform that securely connects the right people to the right technologies at the right time. We help organizations secure and manage their extended enterprise while transforming their customers’ experiences. With thousands of global customers, 7,000+ app integrations, and over 200 million registered users, we are only getting started. As a member of the Developer Productivity Engineering team, you will tackle high-impact challenges across development environments, AI enablement for engineering, scalability, and stability. Grounded in Okta’s core value— Always Secure. Always On. —your work directly powers developer velocity, system reliability, and software quality at scale. You will act as a force multiplier for our engineering teams by identifying workflow bottlenecks, pioneering AI integrations, establishing best practices for code organization, and maintaining performant development environments. What You’ll Do Design & Automation: Build and ship automated solutions that allow developers to deliver features rapidly without compromising quality, stability, or security standards. Environment Performance & Tuning: Analyze local development workflows, build tools to track operational metrics, and profile/tune development environments (including codebase modifications). AI & Tooling Enablement: Leverage AI technologies across the development stack
Other cities to consider
More places hiring for this role
Get new engineering manager platform reliability jobs in Bengaluru, India by email
Daily job updates · Unsubscribe anytime