Jobs in India

Reliability Engineer in India

376 active opportunities · Updated October 2026

Explore current reliability engineer jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.

D
📍 Bengaluru, KARNATAKA, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. About the job DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides an in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. In this role, you will Build core capabilities for our SaaS Platform across multiple clouds Drive development of functional enhancements for Data Discovery, Observability & Governance for both OSS and SaaS offering Lead efforts around non functional aspects like performance, scalability, reliability Lead and mentor junior engineers Work closely with PM, Customers and OSS community Requirements Over 8+ years of experience building and scaling backend systems, preferably in cloud-first or SaaS environments. Solve complex tech

JavaCI/CDRestAI
SL
📍 Noida, Uttar Pradesh, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Staff Software Engineer - Testing & Automation Exceptional software engineering is challenging. Amplifying it to ensure that multiple teams can concurrently create and manage a vast, intricate product escalates the complexity. As a Staff Engineer within the Verification Platform team at Sumo Logic, you will drive the implementation and optimization for our verification platform as well as the modernization of our CI/CD pipelines. Your mission is to develop and sustain automated tooling for all testing, verification, and functional requirements, leveraging AI reasoning and machine learning models to predict and prevent delivery issues, while integrating advanced security validation and non-functional requirements into our delivery lifecycle. You will contribute significantly to establishing automated delivery pipelines, empowering autonomous teams to create independently deployable services, and progressing Sumo Logic’s internal Platform-as-a-Service. This role sits at the intersection of Platform Engineering, Quality Engineering, DevSecOps, and Developer Productivity, helping teams deliver secure, reliable, and independently deployable services at scale. Responsibilities Strategy & Leadership: Drive technical direction and design for a modern Quality Engineering platform, driving the adoption of AI reasoning for enhanced automation of all testing, verification, and functional requirements. Pipeline Modernization: Lead the modernization of CI/CD pipelines to include automated security validation, compliance checks, and other critical non-functional requirements, with a focus on integrating AI/ML for intelligent pipeline optimization and risk prediction. Framework Ownership: Own the delivery pipeline and release automation framework for all Sumo services, ensuring improvements in developer productivity, deployment frequency, and release reliability. Cross-Team Collaboration: Educate and collaborate with teams during design and development phases to ensur

PythonJavaReactAWS
T
📍 Bengaluru, KARNATAKA, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Toradex is a global company strongly focused on engineering & technology. We’re powered by a diverse & uniquely gifted workforce. We pursue the best people to propel our innovative vision of embedded computing and IoT. If you’re interested in being a driving force at an agile technology company, engineering clever computing solutions & helping other companies bring their products to life, we should talk. Description We are looking for a DevOps Engineer to strengthen our cloud operations and engineering practices, with a focus on reliable website delivery, secure AWS foundations, and fast but controlled delivery of new services. The position combines AWS operations, infrastructure as code, CI/CD, automation, and pragmatic software engineering. The person should be confident working with services for edge delivery, compute, storage, databases, DNS, security, and observability without relying on manual console changes as the default operating model. The role also supports on-premises to cloud migration, global service optimization, and practical responses to increasing AI-driven traffic. We value candidates who can use modern AI-assisted development effectively to spin up proof-of-concept projects quickly, while still applying disciplined Git, review, security, and deployment practices. About you You enjoy building stable, secure, and maintainable infrastructure that supports business-critical services. You can work independently and take ownership of cloud environments, deployments, and operational improvements. You are comfortable balancing speed, reliability, cost, and security when making technical decisions. You communicate clearly with technical and non-technical stakeholders and explain trade-offs in a practical way. You document your work well and create clear runbooks and support material for future maintenance. You are methodical when troubleshooting incidents and stay calm when systems are under pressure. You are curious about modern traffic patt

JavaScriptTypeScriptPythonJava
D
📍 Pune, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You’ll Do: We are looking for a Senior Software Engineer – Platform Operations based in Pune, India, who will play a key role in ensuring the reliability, performance, and operational excellence of DeepIntent's platform and data ecosystem. This role requires a strong engineering mindset with the ability to troubleshoot complex technical issues, understand distributed data architectures, and collaborate across Engineering, Product, Analytics, and Customer-facing teams to deliver timely and effective solutions. As part of the Operations organization, you will work closely with Engineering to support production systems, improve operational processes, and drive platform stability. The ideal candidate is a self-motivated problem solver who is passionate about learning new technologies, improving system reliability, and delivering exceptional customer outcomes through engineering excellence. Serve as the engineering interface between Customer-facing teams, Analytics, Product, and Engineering organizations. Partner with Platform Support, Client Success, and other customer-facing teams to investigate and resolve complex platform-related issues. Analyze application, API, and data pipeline issues to identify root causes and drive timely resolution. Develop and standardize operational tools, and interfaces to support analytical and operational use cases. Monitor data pipeline executions, investigate failures, and implement corrective and pre

PythonJavaSQLAWS
D
📍 Pune, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You’ll Do: We are looking for a Software Engineer to help build and scale our core backend systems and data infrastructure. In this role, you will work hands-on to develop the foundational data pipelines, storage solutions, and robust architectures that drive our healthcare advertising solutions and support our core products, reporting APIs, and analytics initiatives. This is an excellent opportunity for a growth-oriented engineer to work with massive datasets, modern cloud technologies, and cross-functional teams to deliver high-performance, fault-tolerant solutions. Build & Operate: Develop, test, and maintain highly reliable, scalable, and cost-optimized distributed systems and data architectures. Enable Self-Service Data: Create automated ingestion, storage, and transformation pipelines that make it simple for downstream users to access and utilize new datasets. Empower Machine Learning: Design and operate data pipelines specifically tailored to support the complex workflows of our Data Scientists and Machine Learning Engineers. Drive Operational Excellence: Help implement and champion DataOps and DevOps practices across the team to ensure system reliability and smooth deployments. Contribute to Best Practices: Play an active role in establishing and refining formal data practices, architectures, and engineering standards for the organization. Cross-Functional Collaboration: Partner effectively with business stakeholders,

PythonJavaSQLAWS
P
📍 Bengaluru, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

A Career with Point72's Technology Team As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. What you'll do Lead the design, development, and operation of scalable, enterprise-grade AI/ML architectures and systems with a strong emphasis on reliability, availability, and performance. Lead and mentor a team of engineers, driving technical direction, code quality, and iterative delivery of large-scale solutions. Partner closely with data scientists, engineers, product teams, and compliance to integrate AI/ML solutions into existing and new products. Own the end-to-end lifecycle of GenAI services, including LLM inference, model serving, and proxy/gateway layers that support multiple downstream applications. Define and uphold engineering best practices around observability, scalability, security, and cost efficiency for AI/ML platforms. Evaluate tools, technologies, and processes to ensure the highest quality and performance of AI/ML systems. Stay abreast of the latest advancements in AI/ML technologies and methodologies, and translate them into pragmatic solutions for the business. Ensure compliance with industry standards and best practices in AI/ML. What's required Bachelor's or Master's degree in Computer Science, Engineering, or a related field. 10+ years of experience in software/AI/ML engineering, with a proven track record of successful delivery of complex, production-grade systems. Demonstrated experience building large-scale enterprise-grade services with high reliability, availability, and observability (SLO/SLA-driven en

PythonJavaAWSAzure
P
📍 Bengaluru, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

JOB TITLE Observability Engineer A CAREER WITH POINT72'S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. WHAT YOU’LL DO Design observability capabilities that give engineering teams clear insight into application health, platform performance, and issues affecting users Build scalable collection pipelines for metrics, logs, and traces across cloud-based and on-premises environments Develop actionable alerting standards that reduce noise, shorten incident response, and highlight the most important signals Partner with application and infrastructure teams to define service health indicators and improve operational readiness before production launches Automate monitoring configuration, dashboard deployment, and reliability checks to support consistent observability across the technology environment Analyze production incidents to identify telemetry gaps and improve detection, diagnosis, and recovery Create dashboards and reporting views that help teams understand trends, capacity risks, and reliability outcomes Establish practical observability

PythonAgileAIHR
DC
📍 Bengaluru, KARNATAKA, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

You’ll shape the future of a business‑critical platform as the technical lead across both product engineering and cloud infrastructure. You’ll modernize a mature .NET application running on AWS today, while steering its evolution toward a cloud‑native, React/Node.js, AI‑enabled architecture. If you enjoy owning architecture end‑to‑end, from backend and frontend through CI/CD, DevOps, and AWS infrastructure, this role gives you real influence at Staff Engineer level and the opportunity to set engineering standards that others follow. You’ll spend your time leading complex .NET and React features, designing scalable AWS infrastructure with Infrastructure as Code, and building automation that makes releases fast, safe, and repeatable. You’ll work on performance, reliability, and modernization in equal measure—fixing what’s slowing the platform down today and designing what it will look like in the next generation. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture and development of enterprise .NET services and APIs that power a business‑critical platform. Design and operate AWS infrastructure (using AWS CDK in TypeScript) to support secure, scalable, multi‑environment deployments. Build and optimize CI/CD pipelines (AWS CodePipeline, CodeBuild, Windows build agents) to make shipping .NET and React changes fast and reliable. Drive modernization initiatives across the stack, including clean architecture, refactoring legacy components, and reducing technical debt. Design and tune PostgreSQL and MSSQL database solutions for performance, scalability, and reliability. Mentor engineers and influence engineering practices across teams, raising the bar on cloud, DevOps, and software design. These are the essentials you’ll need to get an interview Significant experience (typically 8+ years) delivering and operating scalable enterprise software, owning both application code and cloud infrastructure. Deep hands‑on expertise with C

TypeScriptReactNode.jsSQL
DC
📍 Bengaluru, KARNATAKA, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Role Overview Build reliable software services that power products, platforms, and business decisions. As a Senior Software Developer, you’ll design and deliver scalable applications, backend services, and integrations that perform well in production and evolve with changing business needs. You’ll apply strong software engineering practices across APIs, data-intensive applications, cloud services, AI-enabled solutions, and deployment pipelines. You’ll help shape technical solutions, improve system reliability, and contribute to a high-quality engineering culture. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design and develop scalable backend services and applications using Python or TypeScript. Lead the development of APIs, integrations, reusable software components, and AI-enabled features. Build reliable solutions for data ingestion, manipulation, service-to-service communication, and intelligent automation. Apply AI technologies and modern software engineering practices to improve product capabilities, developer productivity, and operational efficiency. Make sound technical decisions around architecture, performance, security, scalability, and maintainability. Deploy and operate applications using AWS services and CI/CD practices while improving testing, monitoring, documentation, and delivery standards. These are the essentials you’ll need to get an interview 5+ years of professional experience developing and delivering production software. Strong hands-on experience with Python; TypeScript or similar languages is also valuable. Proven experience building backend services, APIs, integrations, and service-oriented applications. Experience applying AI technologies, such as generative AI, machine learning services, intelligent automation, or AI-enabled application features. Strong understanding of software design principles, testing, debugging, performance optimization, and secure development. Experience working with cloud platfor

TypeScriptPythonAWSCI/CD
CH
📍 Hyderabad, TELANGANA, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Opportunity Overview: This is a unique opportunity to join a software engineering team that is growing quickly. You will build impactful healthcare technology on a modern stack utilizing your full stack software engineering background. Last but not least: People who succeed here are empathetic teammates who are candid, kind, caring, and embody our core values and principles . We believe that diverse, inclusive teams make the most impactful work. Cohere is deeply invested in ensuring that we have a supportive, growth-oriented environment that works for everyone. What you’ll do: Lead technical design and architecture for complex features across our business intelligence platform Take full ownership of end-to-end feature releases, platform enhancements, and system optimization initiatives Drive technical decision-making by conducting thorough analysis, evaluating trade-offs, and making data-driven architectural choices Architect and implement highly scalable cloud-based data services, APIs, ETL pipelines, and user interfaces Build mission-critical backend systems supporting analytics, reporting, and complex data query capabilities at scale Design and develop intuitive UI platforms and interactive dashboards for data visualization and business intelligence Champion engineering excellence by establishing and enforcing best practices, design patterns, and coding standards Optimize system performance, identify bottlenecks, and implement solutions for scalability and reliability Define and maintain comprehensive test strategies, ensuring high code quality and test coverage Actively participate in ensuring Cohere maintains a disciplined approach to healthcare security, compliance, and data governance Mentor and coach junior and mid-level engineers, fostering their technical growth and career development Collaborate cross-functionally with product, data science, analytics, and business stakeholders to deliver robust technical solutions ISMS roles and respon

JavaScriptTypeScriptPythonJava
CH
📍 Hyderabad, TELANGANA, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Opportunity Overview: We are seeking a Senior Data Engineer to contribute to the design and delivery of our cloud-native healthcare data platform. You will implement scalable data solutions built on AWS, Apache Iceberg, Lake Formation, Glue Catalog, Athena, dbt, and modern orchestration frameworks. This role combines strong hands-on engineering with collaboration across platform, analytics, and business teams. What You'll Do Data Engineering Delivery Deliver complex data engineering projects in collaboration with cross-functional teams Drive technical execution from design through production deployment Implement scalable data patterns and reusable frameworks Design and implement batch and near-real-time pipelines Build reusable ingestion, transformation, validation, and publishing frameworks Support modernization of legacy workloads Contribute to Apache Iceberg implementation and optimization Apply standards for schema evolution, partitioning, compaction, and metadata management Ensure efficient storage and query performance Implement data quality frameworks and validation layers Support observability and monitoring practices Contribute to operational excellence and reliability improvements Participate in architecture and design discussions Conduct and participate in code reviews Mentor junior engineers and share best practices ISMS roles and responsibilities Good knowledge of Information security Oversee specific business processes within the ISMS. Responsible to manage the ISMS documentation, conduct risk assessments, and implement risk treatment plans. Risk Owners are responsible for identifying, assessing, and managing risks within their areas of responsibility. They are also responsible for implementing risk treatment plans. Conduct the BCP and other test related to information security continuity along with CISO Responsible for monitoring and reporting on the performance of the ISMS. Responsible for implementation of security policies and procedures and report

PythonSQLAWSAI
CH
📍 Hyderabad, TELANGANA, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Opportunity Overview: We are seeking a Lead Data Engineer to drive the design and delivery of our cloud-native healthcare data platform. You will lead the implementation of scalable data solutions built on AWS, Apache Iceberg, Lake Formation, Glue Catalog, Athena, dbt, and modern orchestration frameworks. This role combines deep hands-on engineering with technical leadership and collaboration across platform, analytics, and business teams. What you’ll do: Lead Data Engineering Initiatives Lead delivery of complex data engineering projects across multiple teams Drive technical execution from design through production deployment Establish scalable implementation patterns Build and Optimize Data Platforms Design and implement batch and near-real-time pipelines Build reusable ingestion, transformation, validation, and publishing frameworks Support modernization of legacy workloads Lakehouse Engineering Lead Apache Iceberg implementation and optimization Define standards for schema evolution, partitioning, compaction, and metadata management Ensure efficient storage and query performance Data Quality and Reliability Implement data quality frameworks Drive observability and monitoring practices Improve operational excellence and reliability Technical Leadership Review architecture and design proposals Conduct code reviews and engineering reviews Mentor engineers and establish best practices ISMS roles and responsibilities: Good knowledge of Information practices. Assist the manager in all the information security activities implementation and maintenance process. Ensuring the team and imparted with Competence related to Information security Responsible for implementation of security policies and procedures and report any issues to the Information Security Manager. Required Qualifications: 8–12 years of Data Engineering experience. Experience leading enterprise-scale data initia

PythonSQLAWSAI
S
📍 Bengaluru, KARNATAKA, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . We are seeking an experienced Principal Software Engineer to lead the architecture, development, and evolution of a large-scale malware analysis and cybersecurity platform. This role will own the end-to-end technical architecture across Python-based analysis pipelines, message-driven job orchestration systems, Windows VM sandbox environments, and web-based interfaces. The ideal candidate brings deep expertise in software engineering, malware analysis, and distributed systems, with a proven ability to take ownership of complex legacy platforms and drive modernization initiatives. You will lead the design and implementation of advanced detection capabilities, including behavioral analysis, malware sandboxing, machine learning-based classification, and support for new file types, while ensuring platform scalability, reliability, and performance. This position requires strong hands-on development skills in Python with Linux-based production environments, messaging architectures such as RabbitMQ, virtualization technologies, and cloud-native deployment practices. As the technical leader for the platform,

PythonSQLMySQLMongoDB
I
📍 Bengaluru, KARNATAKA, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Instawork is on a mission to create meaningful economic opportunities for skilled hourly professionals in communities around the globe. Our AI-powered labor marketplace helps local businesses scale, and enables global technology companies to push the frontiers of robotics and AI. Backed by world-class investors like Benchmark, Spark Capital, Craft Ventures, Greylock, Y Combinator, and others, we’re looking for exceptional talent to reimagine the way the world works. About IRL (Instawork Robotics Labs) Researchers at UC Berkeley have identified a “100,000-year data gap”—the gulf between what trained AI language models and what physical robots actually have to learn from. Closing that gap is the defining infrastructure challenge of the physical AI era. IRL is Instawork’s answer to it. We deploy skilled workers into real commercial and residential environments—kitchens, warehouses, hotel floors, and homes—to capture the high-fidelity task data that the world’s leading robotics labs use to train their foundation models. About the Role Instawork Robotics creates the highest-quality, highest-diversity datasets for robotics and physical AI. We work with leading robotics builders and research labs to solve one of the most important challenges in robotics: closing the data gap. As a Platform Engineer, you will build the products, services, tools, and infrastructure that power IRL’s data collection, processing, and quality workflows. This is a hybrid product and infrastructure role: you will design platform capabilities for application engineers while also owning the reliability, scalability, cost efficiency, and operation of the systems behind them. You will help shape the platform roadmap by understanding the needs of application engineers, prioritizing high-impact problems, and delivering simple, reliable, and scalable solutions. Who You Are - 5+ years of experience building and operating production software platforms. - Experience designing platform products, backend serv

PythonAWSAzureKubernetes
AI
📍 India· Full-time· Remote
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About AlphaSense: The world’s most sophisticated companies rely on AlphaSense to remove uncertainty from decision-making. With market intelligence and search built on proven AI, AlphaSense delivers insights that matter from content you can trust. Our universe of public and private content includes equity research, company filings, event transcripts, expert calls, news, trade journals, and clients’ own research content. The acquisition of Tegus by AlphaSense in 2024 advances our shared mission to empower professionals to make smarter decisions through AI-driven market intelligence. Together, AlphaSense and Tegus will accelerate growth, innovation, and content expansion, with complementary product and content capabilities that enable users to unearth even more comprehensive insights from thousands of content sets. Our platform is trusted by over 6,000 enterprise customers, including a majority of the S&P 500. Founded in 2011, AlphaSense is headquartered in New York City with more than 2,000 employees across the globe and offices in the U.S., U.K., Finland, India, Singapore, Canada, and Ireland. Come join us! The Role: We're looking for a Senior Engineering Manager to lead the teams building shared product platforms at AlphaSense. The domain covers two connected areas: the API-first platforms that power multiple product surfaces, integrations, and agentic workflows, and the usage platform that meters and reports AI credit consumption across our product features. Both areas are built as products, not as internal plumbing. That means well-designed APIs, clear contracts, strong reliability guarantees, and a roadmap driven by real product needs. Our consumption platform is business-critical: it underpins how AlphaSense packages, prices, and reports on AI capabilities, so accurate, timely, and auditable metering is a hard requirement rather than an implementation detail. This is a people-management role. The immediate team is largely based in EMEA/US, so glo

AWSRestGraphqlAI
🔔

Get new reliability engineer jobs in India by email

Daily job updates · Unsubscribe anytime