About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
Jobs in India
Ai Infrastructure Engineer in Bengaluru
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current ai infrastructure engineer jobs in Bengaluru. Filter by work mode, employment type, experience, department, date posted and distance.
About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. What You’ll Do: Architect the Future of AI Infrastructure: You will design, build, and own the end-to-end platform that supports the entire lifecycle of our ML models—from massive-scale distributed training to ultra-low-latency, highly-available inference. Optimize and Serve Cutting-Edge Models: You'll implement and scale sophisticated inference stacks for LLMs using frameworks like vLLM, TensorRT-LLM, or SGLang . You’ll solve complex challenges in throughput, latency, token streaming, and automated scaling to deliver a seamless user experience. Empower AI Innovation: You will act as a strategic partner to our AI Research and Data Science teams. You’ll create a seamless developer experience that accelerates their ability to experiment, fine-tune, and deploy groundbreaking models with velocity and confidence. Automate Everything: You'll develop robust CI/CD/CT (Continuous Training) pipelines using tools like Argo Workflows, ArgoCD, and GitHub Actions to automate model validation, deployment, and lifecycle management, ensuring our systems are both agile and rock-solid. What are we looking for Experience: 5+ years in infrastructure or software engineering, with at least 2+ years laser-focused on MLOps or ML infrastructu
A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO As Database Support Engineer, you’ll support various critical database platforms across Development, QA, UAT, and Production environments. The role partners closely with application teams, application support, and database engineers and operates within a Follow‑the‑Sun model to ensure availability, performance, and reliability of database services. Key responsibilities include: • Provide operational support for enterprise database platforms in both on-prem private cloud and public cloud • Monitor database health, capacity, performance, and availability, and respond to alerts, diagnose issues, and perform timely remediation • Perform routine maintenance activities (patching, upgrades, housekeeping etc) • Troubleshoot database‑related incidents and collaborate on root cause analysis • Work closely with application owners, application support teams, and DB Engineers • Provide guidance on database best practices and operational standards • Participate in cross‑team problem resolution and continuous improvement initiatives • Contribute to design, implementation and testing of automation and self service capabilities of DB platforms • Drive continuous improvement, identifying opportunities to reduce toil and increase platform efficiency. • Participate in a Follow‑the‑Sun operating model, including shift‑based coverage and handoffs WHAT’S REQUIRED • Bachelor’s degr
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Exa team and lead the charge in redefining enterprise storage by unifying block, file, and object protocols across hybrid-cloud environments. You will combine deep technical expertise in distributed systems with hands-on people leadership to guide architectural decisions and mentor high-impact engineers. This is a unique opportunity to build new engineering teams from the ground up and drive industry-leading innovation alongside Product and Architecture partners. Your work will directly impact how customers consume, scale, and operate mission-critical storage infrastructure. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management, validation, and support teams to align technical roadmaps, establish architectural standards, and drive enterprise
Couchbase, the operational data platform for AI, empowers businesses to succeed by bringing data to life in new ways. Major market-leading companies rely on Couchbase for mission critical operational, analytical, mobile and AI workloads. Built to replace legacy infrastructure and fragmented data services, Couchbase empowers enterprises with a unified platform architected for performance, flexibility and global scale. With Couchbase, organizations bring their data to life, launching game‑changing customer experiences, exploring the limitless potential of AI, and seamlessly extending applications from the cloud to the edge and beyond. Couchbase’s AI‑ready technology and enterprise partnership model eliminate complexity and reduce total cost of ownership, enabling teams to stay agile, innovative and secure. Couchbase believes data should never slow you down, but act as the foundation for your next breakthrough. Discover why Couchbase is trusted to help the world’s biggest players scale, move fast and stay resilient, no matter what’s next on their roadmap. Visit couchbase.com and follow us on LinkedIn and X. Want to be part of our story? Apply today! AI Platform Engineering Location: Bangalore (Hybrid - in office at least 3 days/week) About the Role We are seeking an experienced and visionary technology leader to lead the development and scaling of our Operational AI platform capabilities. This role will own the strategy, architecture, delivery, and operational excellence of Couchbase AI Cloud Platform capabilities. This is a strategic leadership role at the intersection of distributed systems, cloud-native platforms, and AI. You will lead a large, multi-layered engineering organization responsible for delivering The Operational Data Platform for AI, while partnering closely with Product, Design, SRE, and Go-To-Market teams. Your leadership will directly influence company growth, customer adoption, platform reliability, and Couchbase’s competitive pos
WHO ARE WE? We are a bunch of super enthusiastic, passionate, and highly driven people, working to achieve a common goal! We believe that work and the workplace should be joyful and always buzzing with energy! CloudSEK , one of India’s most trusted Cyber security product companies, is on a mission to build the world’s fastest and most reliable AI technology that identifies and resolves digital threats in real-time. The central proposition is leveraging Artificial Intelligence and Machine Learning to create a quick and reliable analysis and alert system that provides rapid detection across multiple internet sources, precise threat analysis, and prompt resolution with minimal human intervention. Founded in 2015, headquartered at Singapore, we are proud to say that we’ve grown at a frenetic pace and have been able to achieve some accolades along the way, including: CloudSEK’s Product Suite: CloudSEK XVigil constantly maps a customer’s digital assets, identifies threats and enriches them with cyber intelligence, and then provides workflows to manage and remediate all identified threats including takedown support. A powerful Attack Surface Monitoring tool that gives visibility and intelligence on customers’ attack surfaces. CloudSEK's BeVigil uses a combination of Mobile, Web, Network and Encryption Scanners to map and protect known and unknown assets. CloudSEK’s Contextual AI SVigil identifies software supply chain risks by monitoring Software, Cloud Services, and third-party dependencies. CloudSEK’s AIVigil is an AI-native Attack Surface Monitoring platform that continuously discovers, monitors, and secures exposed AI infrastructure, MCP servers, leaked AI credentials, vector databases, agentic workflows, and shadow AI across the internet. Key Milestones: 2016 : Launched our first product. 2018 : Secured Pre-series A funding. 2019 : Expanded operations to India, Southeast Asia, and the Americas. 2020 : Won the NASSCOM-DSCI Excellence Award for Security Product Company
WHO ARE WE? We are a bunch of super enthusiastic, passionate, and highly driven people, working to achieve a common goal! We believe that work and the workplace should be joyful and always buzzing with energy! CloudSEK , one of India’s most trusted Cyber security product companies, is on a mission to build the world’s fastest and most reliable AI technology that identifies and resolves digital threats in real-time. The central proposition is leveraging Artificial Intelligence and Machine Learning to create a quick and reliable analysis and alert system that provides rapid detection across multiple internet sources, precise threat analysis, and prompt resolution with minimal human intervention. Founded in 2015, headquartered at Singapore, we are proud to say that we’ve grown at a frenetic pace and have been able to achieve some accolades along the way, including: CloudSEK’s Product Suite: CloudSEK XVigil constantly maps a customer’s digital assets, identifies threats and enriches them with cyber intelligence, and then provides workflows to manage and remediate all identified threats including takedown support. A powerful Attack Surface Monitoring tool that gives visibility and intelligence on customers’ attack surfaces. CloudSEK's BeVigil uses a combination of Mobile, Web, Network and Encryption Scanners to map and protect known and unknown assets. CloudSEK’s Contextual AI SVigil identifies software supply chain risks by monitoring Software, Cloud Services, and third-party dependencies. CloudSEK’s AIVigil is an AI-native Attack Surface Monitoring platform that continuously discovers, monitors, and secures exposed AI infrastructure, MCP servers, leaked AI credentials, vector databases, agentic workflows, and shadow AI across the internet. Key Milestones: 2016 : Launched our first product. 2018 : Secured Pre-series A funding. 2019 : Expanded operations to India, Southeast Asia, and the Americas. 2020 : Won the NASSCOM-DSCI Excellence Award for Security Product Company
WHO ARE WE? We are a bunch of super enthusiastic, passionate, and highly driven people, working to achieve a common goal! We believe that work and the workplace should be joyful and always buzzing with energy! CloudSEK , one of India’s most trusted Cyber security product companies, is on a mission to build the world’s fastest and most reliable AI technology that identifies and resolves digital threats in real-time. The central proposition is leveraging Artificial Intelligence and Machine Learning to create a quick and reliable analysis and alert system that provides rapid detection across multiple internet sources, precise threat analysis, and prompt resolution with minimal human intervention. Founded in 2015, headquartered at Singapore, we are proud to say that we’ve grown at a frenetic pace and have been able to achieve some accolades along the way, including: CloudSEK’s Product Suite: CloudSEK XVigil constantly maps a customer’s digital assets, identifies threats and enriches them with cyber intelligence, and then provides workflows to manage and remediate all identified threats including takedown support. A powerful Attack Surface Monitoring tool that gives visibility and intelligence on customers’ attack surfaces. CloudSEK's BeVigil uses a combination of Mobile, Web, Network and Encryption Scanners to map and protect known and unknown assets. CloudSEK’s Contextual AI SVigil identifies software supply chain risks by monitoring Software, Cloud Services, and third-party dependencies. CloudSEK’s AIVigil is an AI-native Attack Surface Monitoring platform that continuously discovers, monitors, and secures exposed AI infrastructure, MCP servers, leaked AI credentials, vector databases, agentic workflows, and shadow AI across the internet. Key Milestones: 2016 : Launched our first product. 2018 : Secured Pre-series A funding. 2019 : Expanded operations to India, Southeast Asia, and the Americas. 2020 : Won the NASSCOM-DSCI Excellence Award for Security Product Company
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... We’re in an unbelievably exciting area of tech, fundamentally reshaping the cloud-native, modern virtualization, and AI infrastructure landscape. Here, you’ll lead with innovative thinking, grow alongside us, and work with some of the smartest minds in the industry. We’re looking for engineers passionate about system testing, distributed systems, Kubernetes, storage, modern virtualization, AI workloads, and automation. You’ll work on complex, real-world scenarios involving HA, resiliency, disaster recovery, scalability, and failure testing across large-scale Kubernetes environments. As enterprises modernize their infrastructure, containers, virtual machines, and AI/ML workloads are increasingly converging on Kubernetes. From traditional enterprise applications and VMs to GPU-accelerated AI training, inference, and data-intensive workloads, Kubernetes is rapidly becoming the common platform for running the next generation of applications. This role gives you the opportunity to test and influence how Portworx delivers enterprise-grade storage, data protection, availability, and resiliency across these workloads—including containerized applications, KubeVirt/OpenShift Virtualization VMs, and demanding AI/ML workloads running on Kubernetes. WHAT YOU WILL DO: Own system-level quality for Portworx Enterprise across Kubernetes, storage, and modern virtualization environments. Develop comprehens
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... We’re in an unbelievably exciting area of tech, fundamentally reshaping the cloud-native, modern virtualization, and AI infrastructure landscape. Here, you’ll lead with innovative thinking, grow alongside us, and work with some of the smartest minds in the industry. We’re looking for engineers passionate about system testing, distributed systems, Kubernetes, storage, modern virtualization, AI workloads, and automation. You’ll work on complex, real-world scenarios involving HA, resiliency, disaster recovery, scalability, and failure testing across large-scale Kubernetes environments. As enterprises modernize their infrastructure, containers, virtual machines, and AI/ML workloads are increasingly converging on Kubernetes. From traditional enterprise applications and VMs to GPU-accelerated AI training, inference, and data-intensive workloads, Kubernetes is rapidly becoming the common platform for running the next generation of applications. This role gives you the opportunity to test and influence how Portworx delivers enterprise-grade storage, data protection, availability, and resiliency across these workloads—including containerized applications, KubeVirt/OpenShift Virtualization VMs, and demanding AI/ML workloads running on Kubernetes. WHAT YOU WILL DO: Own system-level quality for Portworx Enterprise across Kubernetes, storage, and modern virtualization environments. Develop comprehens
WHO ARE WE? We are a bunch of super enthusiastic, passionate, and highly driven people, working to achieve a common goal! We believe that work and the workplace should be joyful and always buzzing with energy! CloudSEK , one of India’s most trusted Cyber security product companies, is on a mission to build the world’s fastest and most reliable AI technology that identifies and resolves digital threats in real-time. The central proposition is leveraging Artificial Intelligence and Machine Learning to create a quick and reliable analysis and alert system that provides rapid detection across multiple internet sources, precise threat analysis, and prompt resolution with minimal human intervention. Founded in 2015, headquartered at Singapore, we are proud to say that we’ve grown at a frenetic pace and have been able to achieve some accolades along the way, including: CloudSEK’s Product Suite: CloudSEK XVigil constantly maps a customer’s digital assets, identifies threats and enriches them with cyber intelligence, and then provides workflows to manage and remediate all identified threats including takedown support. A powerful Attack Surface Monitoring tool that gives visibility and intelligence on customers’ attack surfaces. CloudSEK's BeVigil uses a combination of Mobile, Web, Network and Encryption Scanners to map and protect known and unknown assets. CloudSEK’s Contextual AI SVigil identifies software supply chain risks by monitoring Software, Cloud Services, and third-party dependencies. CloudSEK’s AIVigil is an AI-native Attack Surface Monitoring platform that continuously discovers, monitors, and secures exposed AI infrastructure, MCP servers, leaked AI credentials, vector databases, agentic workflows, and shadow AI across the internet. Key Milestones: 2016 : Launched our first product. 2018 : Secured Pre-series A funding. 2019 : Expanded operations to India, Southeast Asia, and the Americas. 2020 : Won the NASSCOM-DSCI Excellence Award for Security Product Company
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the high-growth FlashBlade team as a foundational engineering leader building our next-generation, scale-out all-flash file and object storage platform. In this role, you will architect, design, and deliver high-performance distributed systems optimized for modern, data-intensive workloads like AI, log analytics, and cluster computing. You will own critical technological roadmaps end-to-end, directly influencing our core architecture while mentoring top-tier talent. Partnering closely with Product Management, System Validation, and Customer Support, your mission is to drive scalable innovation that redefines enterprise data storage and delivers six-nines reliability to our global customers. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management,
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the high-growth FlashBlade team as a foundational engineering leader building our next-generation, scale-out all-flash file and object storage platform. In this role, you will architect, design, and deliver high-performance distributed systems optimized for modern, data-intensive workloads like AI, log analytics, and cluster computing. You will own critical technological roadmaps end-to-end, directly influencing our core architecture while mentoring top-tier talent. Partnering closely with Product Management, System Validation, and Customer Support, your mission is to drive scalable innovation that redefines enterprise data storage and delivers six-nines reliability to our global customers. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management,
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the high-growth FlashBlade team as a foundational engineering leader building our next-generation, scale-out all-flash file and object storage platform. In this role, you will architect, design, and deliver high-performance distributed systems optimized for modern, data-intensive workloads like AI, log analytics, and cluster computing. You will own critical technological roadmaps end-to-end, directly influencing our core architecture while mentoring top-tier talent. Partnering closely with Product Management, System Validation, and Customer Support, your mission is to drive scalable innovation that redefines enterprise data storage and delivers six-nines reliability to our global customers. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management,
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the high-growth FlashBlade team as a foundational engineering leader building our next-generation, scale-out all-flash file and object storage platform. In this role, you will architect, design, and deliver high-performance distributed systems optimized for modern, data-intensive workloads like AI, log analytics, and cluster computing. You will own critical technological roadmaps end-to-end, directly influencing our core architecture while mentoring top-tier talent. Partnering closely with Product Management, System Validation, and Customer Support, your mission is to drive scalable innovation that redefines enterprise data storage and delivers six-nines reliability to our global customers. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management,
Other cities to consider
More places hiring for this role
Get new ai infrastructure engineer jobs in Bengaluru, India by email
Daily job updates · Unsubscribe anytime