The Shift Engineer - C&I monitors C&I operations, manages maintenance, responds to emergencies, ensures safety, collaborates with stakeholders, and drives continuous improvement to optimize C&I system performance and reliability. Source: Adani Group | Job ID: 51314
Jobs in India
Reliability Engineer in India
376 active opportunities · Updated October 2026
Showing
15 jobs
Explore current reliability engineer jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.
The Shift Engineer - Electrical performs shift-based maintenance, monitors equipment, and executes preventive maintenance schedules. They provide technical support, conduct safety inspections, and maintain detailed documentation to ensure operational continuity and system reliability. Source: Adani Group | Job ID: 56189
The Shift Engineer - Electrical performs shift-based maintenance, monitors equipment, and executes preventive maintenance schedules. They provide technical support, conduct safety inspections, and maintain detailed documentation to ensure operational continuity and system reliability. Source: Adani Group | Job ID: 56191
Graphcore Senior Principal AI SoC Validation (Bring-up lead) Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Bengaluru which will play a central role in Graphcore's work building the future of AI computing. We are developing the next generation of AI compute, a large-scale system-on-chip (SoC) designed to power future high-performance AI systems. As the SoC Validation Lead, you will be responsible for enabling pre-production software to run reliably on new silicon quickly and efficiently, before showing that the silicon meets the highest standards of quality, reliability and functionality, ready for production deployment. You will lead a team delivering post-silicon validation across the full AI SoC, working across silicon, firmware, and platform levels. The role requires a deep technical understanding, strong hands-on debug experience, and the ability to collaborate effectively with hardware, software, and systems engineering teams. Key responsibilities Define and lead post-silicon validation strategy Develop and refine the overall post-silicon validation approach for our AI SoCs, ensuring reliable and timely delivery of validated silicon, architectural correctness, feature robustness, and at-scale system reliability. Drive cross-domain debug and issue resolution Lead investigation and resolution of complex issues spanning silicon, firmware, operating systems, and platform interactions. Ensure that fixes are effective and sustainable. Promote collaboration and shared understanding Work closely with
Executive - Building Management Systems (BMS) is responsible for overseeing the engineering, maintenance, and operational management of all BMS systems within the airport. This includes ensuring the proper functioning, reliability, and performance of systems such as HVAC, lighting, fire alarm systems, access control, and other integrated mechanical, electrical, and civil systems that are controlled through BMS. The role ensures preventive and corrective maintenance, and manage troubleshooting for BMS-related issues across airport facilities. Source: Adani Group | Job ID: 46328
Couchbase, the operational data platform for AI, empowers businesses to succeed by bringing data to life in new ways. Major market-leading companies rely on Couchbase for mission critical operational, analytical, mobile and AI workloads. Built to replace legacy infrastructure and fragmented data services, Couchbase empowers enterprises with a unified platform architected for performance, flexibility and global scale. With Couchbase, organizations bring their data to life, launching game‑changing customer experiences, exploring the limitless potential of AI, and seamlessly extending applications from the cloud to the edge and beyond. Couchbase’s AI‑ready technology and enterprise partnership model eliminate complexity and reduce total cost of ownership, enabling teams to stay agile, innovative and secure. Couchbase believes data should never slow you down, but act as the foundation for your next breakthrough. Discover why Couchbase is trusted to help the world’s biggest players scale, move fast and stay resilient, no matter what’s next on their roadmap. Visit couchbase.com and follow us on LinkedIn and X. Want to be part of our story? Apply today! AI Platform Engineering Location: Bangalore (Hybrid - in office at least 3 days/week) About the Role We are seeking an experienced and visionary technology leader to lead the development and scaling of our Operational AI platform capabilities. This role will own the strategy, architecture, delivery, and operational excellence of Couchbase AI Cloud Platform capabilities. This is a strategic leadership role at the intersection of distributed systems, cloud-native platforms, and AI. You will lead a large, multi-layered engineering organization responsible for delivering The Operational Data Platform for AI, while partnering closely with Product, Design, SRE, and Go-To-Market teams. Your leadership will directly influence company growth, customer adoption, platform reliability, and Couchbase’s competitive pos
Principal Product Manager - Agentic Investigation & Reliability Experiences Sumo Logic is hiring a Principal Product Manager to lead how engineers and operators investigate incidents, understand reliability risk, and act on their operational and security telemetry. The observability category was built around collecting telemetry and giving people tools to navigate it: dashboards, queries, monitors, traces, and alerts. Customer expectations are now shifting. Teams don't just want more dashboards; they want help getting from a signal to a resolution, understanding what's broken, why, what's impacted, and what to do next. As AI agents move into production operations, this role owns how Sumo Logic brings intelligent, agent-assisted investigation and reliability workflows to customers, grounded in evidence, context, and enterprise governance. This is a senior, high-ownership role. It requires genuine observability domain background. You should have lived in this space and understand how monitoring, troubleshooting, and reliability actually work, combined with the ambition to define a new category of experience on top of it. What You Will Own The current data experiences. Log Search, Live Tail, query and query optimization, Metrics Search, Tracing, Dashboards, and the data-experience UI. This is a live, revenue-generating product with real customers, and keeping it strong is part of the job. You own its health, roadmap, and competitiveness today while steering it toward an AI-native future, focusing new investment where it strengthens investigation, speed, and value for both new and power users. The reliability and alerting surface. Monitors, Alerts, SLOs, Scheduled Searches, and the reliability workflows around them. You will own alerting accuracy, noise reduction, and operational health signals both as capabilities customers depend on today and as the foundation for more automated, agent-assisted detection and investigation. The agentic investigation experience. You
Principal Product Manager - Agentic Investigation & Reliability Experiences Sumo Logic is hiring a Principal Product Manager to lead how engineers and operators investigate incidents, understand reliability risk, and act on their operational and security telemetry. The observability category was built around collecting telemetry and giving people tools to navigate it: dashboards, queries, monitors, traces, and alerts. Customer expectations are now shifting. Teams don't just want more dashboards; they want help getting from a signal to a resolution, understanding what's broken, why, what's impacted, and what to do next. As AI agents move into production operations, this role owns how Sumo Logic brings intelligent, agent-assisted investigation and reliability workflows to customers, grounded in evidence, context, and enterprise governance. This is a senior, high-ownership role. It requires genuine observability domain background. You should have lived in this space and understand how monitoring, troubleshooting, and reliability actually work, combined with the ambition to define a new category of experience on top of it. What You Will Own The current data experiences. Log Search, Live Tail, query and query optimization, Metrics Search, Tracing, Dashboards, and the data-experience UI. This is a live, revenue-generating product with real customers, and keeping it strong is part of the job. You own its health, roadmap, and competitiveness today while steering it toward an AI-native future, focusing new investment where it strengthens investigation, speed, and value for both new and power users. The reliability and alerting surface. Monitors, Alerts, SLOs, Scheduled Searches, and the reliability workflows around them. You will own alerting accuracy, noise reduction, and operational health signals both as capabilities customers depend on today and as the foundation for more automated, agent-assisted detection and investigation. The agentic investigation experience. You
Toradex is a global company strongly focused on engineering & technology. We’re powered by a diverse & uniquely gifted workforce. We pursue the best people to propel our innovative vision of embedded computing and IoT. If you’re interested in being a driving force at an agile technology company, engineering clever computing solutions & helping other companies bring their products to life, we should talk. Description We are looking for an Integration Platform Specialist to manage company-wide integrations using Workato. The role focuses on building reliable workflows, supporting API-based connections, and working with teams to automate business processes. The position requires strong knowledge of REST APIs, GraphQL, Postman, and the Workato platform. Experience with Workato MCP functionality and coding skills in Python, Ruby, or similar languages are preferred. About You You enjoy solving complex system and process problems with practical, scalable solutions. You can work independently and take ownership of integrations from discovery through deployment and support. You communicate clearly with both technical and non-technical stakeholders. You document your work well and create clear support material for future maintenance. You are curious, hands-on, and willing to investigate issues until you find the root cause. You care about reliability, data quality, security, and a good internal user experience. You are comfortable working across teams and balancing business priorities with technical constraints. Key Responsibilities Own the integration platform roadmap and day-to-day operation, ensuring business-critical automations are reliable, observable, and maintainable. Partner with business and application owners to turn process gaps into pragmatic integration designs and delivery plans. Build Workato recipes, custom connectors, and reusable patterns that reduce manual work and improve data flow between systems. Maintain and modernize existing integrations,
Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. This position is based at our office in Chennai, India. Appian was built on a culture of in-person collaboration, which we believe is a key driver of our mission to be the best. You will be the product manager working closely with the team whose mission is to strengthen and optimize site infrastructure by delivering essential upgrades, resource efficiency, and scalable solutions. You will be responsible for the direction and roadmap of a component of the Appian Cloud data plane that ensures reliable, high-performance operations for all Appian Cloud customer sites. This role is specifically focused on the cloud-native persistence and messaging layer. You will oversee the backend sub-systems—including technologies like S3 and Redis—that power the platform's internal data plane and core services. This component of the software is not directly user-facing but has strong implications on the scalability and reliability requirements our customers expect. What you will be doing: Prioritize and Define: Work on an agile team to prioritize, define, and ensure the success of infrastructure and managed services features for a high-level strategic roadmap. Stakeholder Collaboration: Prioritize what we should build and when by collaborating with stakeholders on product vision and strategy, while taking customer feedback into account. Technical Discussions: Define how infrastructure features will work through close collaboration with engineers in design sessions an
Role Purpose: At Jumio, the Software Engineer II (QA) will focus on ensuring the quality and performance of highly scalable web and backend applications. In this role, you will design and implement automated tests for web (Playwright/Selenium) and API-based solutions, leveraging your knowledge of Java or JavaScript. Collaborating closely with development and product teams, you will ensure Jumio's products meet the highest standards of quality and reliability. You’ll have an opportunity to learn and grow in a fast-paced environment while exploring new tools and methodologies. A problem-solving mindset, willingness to innovate, and attention to detail will make you successful in this role. T-Shaped Engineering Expectation: As part of Jumio’s engineering culture, you will adopt a T-shaped engineering approach. In addition to developing expertise in test automation and quality engineering, you will collaborate across the development lifecycle, including understanding software architecture, contributing to design discussions, and ensuring robust and scalable test solutions. Role Value: This role is critical to ensuring the reliability, scalability, and security of Jumio’s products. By building and maintaining automated testing frameworks, you will enable faster releases and higher confidence in the quality of our software. Example Responsibilities: Develop, maintain, and execute automated test scripts for web applications using Playwright, Selenium, or similar automation frameworks. Create and execute API test suites using tools such as Postman, REST Assured, or equivalent testing frameworks. Design and execute functional, regression, integration, and exploratory test cases based on business and technical requirements. Validate application functionality, backend services, APIs, and data flows across different environments. Identify, document, track, and verify defects, working closely with developers to ensure timely resolution. Execute automated test suites as part of C
Are you ready to do your life’s work at the heart of the autonomous revolution? NVIDIA’s SWQA organization is seeking a world-class Software QA Test and Tool Developer to join our Automotive Platform team, where the code you validate ensures the safety of millions on the road. In this role, you won't just be testing software; you will be architecting the security and reliability of the next generation of intelligent vehicles. We are looking for engineers who are as comfortable navigating low-level product architecture as they are deep-diving into complex product use cases with passion for quality. This is a high-impact, hands-on position focused on our industry-leading automotive products, offering a rare opportunity to influence the core of our tech stack. You will also build the tools and frameworks that define performance standards for systems running on Linux and QNX. What you’ll be doing: Design, execute, and automate comprehensive test cases and test scenarios to validate our automotive platforms using various test methodologies to identify and track actionable defects and track them to closure. Participate in deep-dive reviews of product requirements and technical designs, providing critical feedback to ensure features are built for testability and security from day one. Partner closely with project management, hardware teams, and software developers to provide rigorous technical analysis of bugs and publish data-driven statistical reports for global team members. Architect and maintain a distributed test automation framework capable of managing high-concurrency workloads across an extensive automation farm of hundreds of concurrent systems. Develop sophisticated test libraries and automation solutions to accelerate development cycles and expand automated test coverage for re
NVIDIA is looking for Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer! What you'll be doing: Primary responsibilities will include building AI/HPC infrastructure for new and existing customers. Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting. Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement. Maintain services once they are live by measuring and monitoring availability, latency, and overall system health. Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements. What we need to see: BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields. At least 5+ years of professional experience in networking fundamentals, Ethernet or InfiniBand World. Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS, etc. Possess solid working knowl
Job Title Senior Software Technologist I - C++ Job Description Software Engineer II Your Role: • Design, develop, test, and maintain software components and applications using modern C++, C# in a Windows-based environment. • Participate in the full software development lifecycle including requirements analysis, design, implementation, testing, debugging, and maintenance. • Develop and maintain software applications using Visual Studio and associated C++, C# development tools. • Work with Windows operating system fundamentals including processes, services, registry, file system, User Account Control (UAC), and application configuration. • Create, enhance, and troubleshoot software modules while adhering to coding standards, design guidelines, and software development best practices. • Utilize GitHub for source control management including branching strategies, commits, pull requests, merges, rebasing, and code reviews. • Support and troubleshoot CI/CD pipeline issues using GitHub Actions and participate in continuous integration activities. • Manage software dependencies and package management using NuGet and Conan. • Configure and maintain build systems using CMake and Visual Studio project configurations. • Develop and execute unit tests using Google Test (GTest) to ensure software quality and reliability. • Perform debugging, root cause analysis, and defect resolution for software issues identified during development, testing, and field support activities. • Participate in peer code reviews and contribute to software quality, maintainability, and technical excellence. • Collaborate effectively with Software Verification, Product Management, DevOps, Architecture, and cross-functional teams to deliver high-quality software solutions. • Create and maintain
About Us: Sauce Labs is the world’s largest full-lifecycle, test automation platform, and the company behind Selenium. Trusted by 80% of the world’s top ten largest financial institutions and over 300,000 enterprise users, Sauce Labs provides the only AI platform capable of turning business intent into autonomous testing and quality assurance. With a proprietary dataset of 8.7 billion test runs, Sauce Labs empowers the Fortune 2000 to bridge the gap between AI-driven code generation and enterprise-grade software quality. Learn more at saucelabs.com . The Role: We are seeking an innovative and experienced AI Architect to join our engineering leadership team. This is a strategic role that will be instrumental in designing and building the next generation of AI-powered features for our continuous testing platform. You will be responsible for architecting scalable and robust AI solutions that transform how our customers gain insights from their test data and production environments, and how they create tests. Responsibilities: Define AI Architecture: Lead the design and architecture of cutting-edge AI/ML solutions for new product offerings, ensuring scalability, performance, quality and reliability within a cloud-native environment. AI-Powered Insights (Test & Production): Architect AI systems to derive actionable insights from vast quantities of test run logs and analytics data. This includes identifying patterns, anomalies, and performance trends. Production Error Reporting Integration: Design AI solutions that integrate with our existing error reporting product to analyze production issues for mobile and web applications, providing deeper understanding and predictive capabilities. Unified Data Intelligence: Develop architectures for combining insights from both test runs and production data, creating a holistic view of application quality and user experience. Automated Failure Analysis & Remediation: Architect AI models and systems t
Other cities to consider
More places hiring for this role
Get new reliability engineer jobs in India by email
Daily job updates · Unsubscribe anytime