Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu
Jobs in India
Ai Systems Engineer in India
3,705 active opportunities · Updated October 2026
Showing
15 jobs
Explore current ai systems engineer jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Debug Validation Engineer will drive post-silicon debug and validation activities for next-generation AI compute silicon and systems. The role is responsible for leading teams focused on identifying, reproducing, analysing and resolving complex silicon, firmware and system-level issues during bring-up, characterization and product readiness. This position combines deep technical debugging expertise with strong cross-functional collaboration across multiple engineering disciplines. The role will work closely with architecture, RTL, firmware, software and systems teams to improve debug methodologies, accelerate issue resolution and strengthen validation coverage. The role will work closely with architecture, RTL, firmware, software, systems and platform teams to improve debug methodologies, accelerate issue resolution and strengthen validation coverage. The Team The Post-Silicon Debug and Validation team sits within the Architecture and Validation organisation and is responsible for bring-up, debug and validation of Graphcore silicon and systems.
Senior -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to deb
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is now at one of the most consequential inflection points in its history. Enterprises are rapidly shifting from human-driven API workflows to multi-agent systems — where AI agents autonomously discover, call, and collaborate with APIs, tools, CLIs, and each other. Postman already owns the API design, governance, and lifecycle layer for the enterprise; the next frontier is owning the runtime layer for how agents actually interact with all of it . Where Postman today is the API platform for human-first integration , we are building Postman into the agent interface fabric for AI-first integration . That transformation starts with Fabric Gateway — a brand new product, being built from scratch, that will serve as the policy-driven control plane governing how agents connect to APIs, services, and other agents at scale. This is a rare 0-to-1 opportunity inside a company with 40+ million developers already in the ecosystem. About the Team The Fabric Gateway team is building the API gateway for the AI era — one of Postman's most ambitious greenfield infrastructure products. We're a small, high-ownership team wo
About Us: Paytm is India's leading mobile payments and financial services distribution company. Pioneer of the mobile QR payments revolution in India, Paytm builds technologies that help small businesses with payments and commerce. Paytm’s mission is to serve half a billion Indians and bring them to the mainstream economy with the help of technology. Job Summary: Build systems for collection & transformation of complex data sets for use in production systems Collaborate with engineers on building & maintaining back-end services Implement data schema and data management improvements for scale and performance Provide insights into key performance indicators for the product and customer usage Serve as team's authority on data infrastructure, privacy controls and data security Collaborate with appropriate stakeholders to understand user requirements Support efforts for continuous improvement, metrics and test automation Maintain operations of live service as issues arise on a rotational, on-call basis Verify whether data architecture meets security and compliance requirements and expectations .Should be able to fast learn and quickly adapt at rapid pace. java/scala, SQL, Minimum Qualifications: Bachelor's degree in computer science, computer engineering or a related field, or equivalent Experience 6+ years of progressive experience demonstrating strong architecture, programming and engineering skills. Firm grasp of data structures, algorithms with fluency in programming languages like Java, Python, Scala. Strong SQL language and should be able to write complex queries. Strong Airflow like orchestration tools. Demonstrated ability to lead, partner, and collaborate cross functionally across many engineering organizations Experience with streaming technologies such as Apache Spark, Kafka, Flink. Backend experience including Apache Cassandra, MongoDB and relational databases such as Oracle, PostgreSQL AWS/GCP Solid hands on with 4+ years of experience. St
DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. About the Role We're seeking an experienced DevOps/ Site Reliability Engineering (SRE) Engineer to join DataHub and drive the reliability, scalability, and operational excellence of our platform offerings. In this role, you'll work on technical initiatives across DataHub Cloud and our emerging enterprise deployment solution, which provides customers with enhanced control and flexibility for running DataHub in their preferred environments. Key Responsibilities Enterprise Platform Development: Partner with product and engineering teams to influence the development of advanced deployment capabilities. Collaborate with cross-functional teams to help build systems for seamless installation, upgrade, and rollback processes across various environments. Influence the design and help implement comprehensive monitoring and health check systems for distributed deployments. Partner with engineering teams to help develop self-healing and automated remediation capabilities. Platform Reliability and Operations: Establish and maintain SLAs/SLOs for both cloud and enterprise offerings. Lead incident response and post-mortem processes to drive continuous improvement. Optimise system performance, capacity planning, and cost efficiency. Work closely with product, engineerin
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary The Workplace Coordinator is responsible for the day-to-day operation and maintenance of workplace infrastructure, ensuring that all building systems operate efficiently, safely, and reliably. This role oversees HVAC systems, including chillers and AHUs, Building Management Systems (BMS), electrical panels, and coordinates preventive and corrective maintenance activities to support uninterrupted business operations. Key Responsibilities Technical Operations · Monitor and maintain HVAC systems including Chillers, AHUs, FCUs, and ventilation systems. · Operate and monitor Building Management System (BMS) for alarms, trends, and equipment performance. · Inspect electrical LT panels, UPS systems, DG synchronization (if applicable), and power distribution systems. · Monitor critical utilities including temperature, humidity, pressure, and energy consumption. · Ensure uninterrupted operation of critical infrastructure and respond promptly to system failures. Preventive & Corrective Maintenance · Plan and execute preventive maintenance schedules for HVAC and electrical systems. · Coordinate breakdown maintenance with
Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team... Here at GoDaddy, the ML Engineering (MLE) team exists as the backbone of our machine learning infrastructure, enabling ML scientists and product teams across Domains to ship models to production reliably, efficiently, and at scale. This team owns the full lifecycle of ML systems — from CI/CD pipelines and model serving infrastructure to GPU workload orchestration and observability. Through disciplined engineering practices, thoughtful system design, and close collaboration with ML scientists, data engineers, and product teams, we deliver the platform that powers domain search, pricing, recommendations, and emerging AI experiences for millions of customers worldwide. We are currently looking for an experienced, highly motivated Senior Engineering Manager to lead our ML Engineering team based in India. This is an established team with existing engineers — we expect the candidate to ramp up quickly on our ML infrastructure stack, build strong relationships with the team, and partner with both India-based teams and US-based teams to drive execution and grow the team further. This individual will join us on our journey to build and scale ML infrastructure that serves real-time predictions at low latency, automates model deployment and promotion, and provides the observability and reliability guarantees that production ML systems demand. Become part of a team that bridges the gap between ML research and production engineering — shipping systems that directly impact GoDaddy's core revenue. What you'll get to do... Lead a team o
About the Role Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. You’ll work across two areas: Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack. Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve. We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure. This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation , solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure. responsible for delivering the software but also for operating and supporting it in production. Why this Role You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective. You’ll have the opportunity to build foundational systems from the ground up, work on infrastructur
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Our mission is to enable effective financial decisions through reliable data, increased efficiency, and automation. We support Marketing, Sales, Seller, Accounting, Tax, Finance and Strategy (F&S), Finance Operations (FinOps), and Treasury functions across automation, data insights, and process improvements. You may work on a wide variety of critical business areas including Seller Systems — Responsible for building the systems and tooling that make sellers and internal revenue teams at Stripe dramatically more productive and effective. We partner with Sales, Finance, Legal, and Product to deliver a single "plane of glass" selling experience that spans deal creation and modeling, negotiation and approvals, contracting, onboarding, and activation. The Seller Systems team composes first‑party, custom Stripe components with best‑in‑class third‑party business systems to deliver configurable, auditable, and globally scalable workflows. Engineers on Seller Systems build services, APIs, integrations, data pipelines, and internal UIs that power seller productivity, reduce time‑to‑activation, improve deal velocity, and enable AI‑driven assistive workflows. Use cases include deal modeling and pricing engines, approval and orchestration platforms, CLM/CPQ integrations, onboarding automation, seller analytics, and AI‑assisted seller tooling. Finance Engineering — Responsible for building the robust and scalable infrastructure that powers
A Career with Point72’s Technology Team As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. What you’ll do Build software applications and deliver software enhancements and projects supporting fund accounting and trade processing technology. Work closely with business stakeholders to develop software solutions using test-driven and agile software development methodologies. Be responsible for system upgrades and features supporting resiliency and capacity improvements, automation and controls, and integration with internal and external vendors and services. Driving architecture of core platforms and accelerating modernizing leveraging AI tools. Work with DevOps teams to manage and resolve operational issues and leverage CI/CD platforms while following DevOps practices within the team and projects. Continuously improve the platforms using the latest technologies and software development ideas. Participate in initiatives to transition select applications to cloud platforms, enhancing stability, scalability and performance of the existing platform. Develop close wo
Want to be a bswifter? At bswift we’ve been transforming benefits administration since 1996, making it simpler, smarter, and more human. Our state-of-the-art, cloud-based technology and services empower employees to understand, manage, and love their benefits. From downtown Chicago, and remotely across the country, we serve thousands of companies and millions of people nationwide, reducing administrative burdens and freeing HR teams to focus on creating thriving, people-first workplaces. We’re looking for motivated and goal-driven individuals who share our passion for delivering excellence and creating solutions that make a difference. The reward is a fun, flexible and creative environment with ample opportunity for professional and personal growth. If you love the bswift values of pursue excellence, embrace accountability, deliver superior service, and be a great place to work, we want to hear from you! About the Role We are looking for an AI Engineer II to design, build, and deploy cutting-edge generative AI applications , including agentic workflows and intelligent chat experiences , that transform how employees interact with their benefits. This role goes beyond individual contribution—you will own complex AI features end-to-end , influence architecture decisions around LLMs and agent systems , and mentor junior engineers . You will play a key role in bringing secure, scalable, and high-performance AI solutions to production using AWS cloud technologies. You will collaborate closely with software engineers, architects, product managers, and cross-functional teams to deliver seamless and impactful user experiences Key Responsibilities 1. Solution Design & Development Lead the design and development of scalable, secure, and high-performance applications . Architect robust backend systems and integrate them with frontend applications and third-party services. Drive technical decisions and participate in design reviews alongside senio
We're hiring an AI Support Engineer to work directly with the founder and build the systems that power customer support at Bolna. This isn't a traditional support role — you'll use AI to make support scale, and you'll partner closely with the business team on the customer conversations that matter most. What you'll do - Work directly with the founder to design and continuously improve how customer support runs at Bolna - Pull and collate data from Intercom to spot patterns, recurring issues, and gaps in how customers are being helped - Build AI-powered workflows that triage, answer, and resolve customer support queries with less manual effort - Design the systems and processes behind a streamlined, scalable support flow — from triage to escalation to resolution - Step in directly on critical customer support situations alongside the business team when it matters - Turn recurring support themes into feedback for product and engineering What we're looking for - 1–3 years of experience in a support, ops, or technical customer-facing role — ideally somewhere that rewarded building your own tools and process, not just following a playbook - Hands-on comfort with AI tools/workflows (prompting, automations, agent builders) — you don't need to be an ML engineer, but you should be someone who reaches for AI to solve a workflow problem - Experience with Intercom or a similar support/helpdesk tool - Sharp, structured communicator — equally comfortable writing to customers and to the founder - Comfortable with ambiguity — this role is being built as you build it Nice to have - Experience setting up support automations, chatbots, or AI agents in a real product company - Familiarity with SQL or basic scripting to pull/analyze support data - Startup experience, especially in a 0-to-1 function
AI Engineer - Enterprise Search Overview We are looking for an experienced Enterprise Search Lead to build and optimize our multi-tenant enterprise search solution. This role focuses on creating scalable systems that integrate seamlessly with customer environments, leveraging cutting-edge AI and ML technologies. You will design and manage the enterprise knowledge graph, implement personalized search experiences, and drive AI- powered innovations to enhance search relevance and ITSM workflows. What You Will Do: Build and structure our enterprise knowledge graph to organize content, people, and activity into meaningful relationships for better search relevance. Develop and refine personalized ranking models that adapt to user behavior and improve search results over time. Design ways to adapt AI language models to each customerʼs data for enhanced accuracy and context. Explore innovative methods to combine LLMs with search engines for answering complex queries Write clean, robust, and maintainable code that integrates smoothly with multi-tenant systems. Collaborate with cross-functional teams to align search capabilities with ITSM workflows. Mentor junior engineers or learn from experienced ones to grow as a technical leader. What You Should Have 3-5 years of experience working on enterprise search products with AI/ML integration. Expertise in multi-tenant systems and securely integrating with external customer systems. Hands-on experience with tools like Elasticsearch, Solr, or similar search platforms. Strong coding skills in Python, Java, or equivalent languages. A passion for solving complex problems with AI and delivering intuitive user experiences. Important notice for candidates: Job scams are on the rise. Please keep these guidelines in mind when applying for any open roles at Atomicwork. Only apply through official Atomicwork channels. We do not use third-party agencies or individuals who ask for payments in exchange for interviews or offer letter
About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
Other cities to consider
More places hiring for this role
Get new ai systems engineer jobs in India by email
Daily job updates · Unsubscribe anytime