DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You'll Do: Deploy, configure, and maintain Kubernetes clusters for our microservices architecture. Utilize Git and Helm for version control and deployment management. Implement and manage monitoring solutions using Prometheus and Grafana. Work on continuous integration and continuous deployment (CI/CD) pipelines. Containerize applications using Docker and manage orchestration. Manage and optimize AWS services, including but not limited to EC2, S3, RDS, and AWS CDN. Maintain and optimize MySQL databases, Airflow, and Redis instances. Write automation scripts in Bash or Python for system administration tasks. Perform Linux administration tasks and troubleshoot system issues. Utilize Ansible and Terraform for configuration management and infrastructure as code. Demonstrate knowledge of networking and load-balancing principles. Collaborate with development teams to ensure applications meet reliability and performance standards. Who you are: Bachelor’s degree in engineering (CS / IT) or equivalent degree from a well-known Institute / University. 2+ years of experience in a Site Reliability Engineer role or similar. Proven experience with Kubernetes, Git, Helm, Prometheus, Grafana, CI/CD, Docker, and microservices architecture. Strong knowledge of AWS services, MySQL, Airflow, Redis, AWS CDN. Proficient in scripting languages such as Bash or Python. Hands-on experience with Linux administration. Familiarity with Ansible and Terraform fo
Jobs in India
Cluster Chef in India
92 active opportunities · Updated October 2026
Showing
15 jobs
Explore current cluster chef jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity AI is a top strategic priority for New Relic, and the Bengaluru design team is at the center of it. The team works across three closely related product clusters: autonomous incident response (SRE Agent, Autopilot, and Intelligent RCA), the intelligence and platform layer that powers them (Ground Truth, Agentic Platform, and New Relic AI), and AIOps for event correlation and incident management. These products share a common design challenge: users need to trust systems that act autonomously, and building that trust through good design is genuinely hard work. This is an on-site role in Bengaluru. Your designers are there, and many of your engineering and product partners are too. Being present — in standups, reviews, and the quick conversations before a decision gets made — is part of how you'll build the relationships that make design effective. You'll also collaborate with design, product, and engineering partners in the US and Spain, so operating across time zones and communicating well in writing are part of the job. You'll manage a small team of designers and own the quality of the work across these products. The design questions here don't have established answers — how do you make an autonomous system legible? How do you build user trust in AI-generated root cause analysis? How do you design a handoff from machine decision to human judgment? If you're already working in AI product design, or actively building toward it, and you want to lead a team
Key Responsibilities : Primary responsibilities :- Installation and configuration of MySQL instances on single or multiple ports. Hands-on experience of working with MysQL 5.7 and MySQL 8. Clear understanding of MysQL Replication process flows , threads , setting up multi node clusters and basic troubleshooting. Understanding of at least one of the backup and recovery methods for MySQL . Strong fundamentals of SQL and able to understand and tune complex SQL queries when needed. Strong fundamentals on the linux system side and monitoring tools like top , iostats , sar etc. At Least couple of years of production hands on experience on medium to big sized MySQL databases. Setting up and maintaining users and privileges management system and troubleshooting relevant access issues. Some exposure to external tools like Percona , ProxySQL , HAP etc. Understand the transaction flows and ACID compliance. Basic understanding of networking concepts . Performing on-call support and should be able to provide the first level support . Excellent verbal and written communication skills. Strong shell scripting skills . Good to have Python . Secondary responsibilities. :- Able to configure and setup NOSQL databases like Mongodb and Cassandra. Ability to learn new technologies along with a team and a positive outlook to understand problems from the business point of view. Qualifications: Proficiency in database management systems such as , MySQL or NoSQL databases. SQL programming and database design skills. Knowledge of database performance tuning and optimization techniques. Familiarity with database security best practices. Scripting and automation skills (Good to have- Python). Good problem-solving and analytical skills. Excellent communication and teamwork skills.
Key Responsibilities : Primary responsibilities :- Installation and configuration of MySQL instances on single or multiple ports. Hands-on experience of working with MysQL 5.7 and MySQL 8. Clear understanding of MysQL Replication process flows , threads , setting up multi node clusters and basic troubleshooting. Understanding of at least one of the backup and recovery methods for MySQL . Strong fundamentals of SQL and able to understand and tune complex SQL queries when needed. Strong fundamentals on the linux system side and monitoring tools like top , iostats , sar etc. At Least couple of years of production hands on experience on medium to big sized MySQL databases. Setting up and maintaining users and privileges management system and troubleshooting relevant access issues. Some exposure to external tools like Percona , ProxySQL , HAP etc. Understand the transaction flows and ACID compliance. Basic understanding of networking concepts . Performing on-call support and should be able to provide the first level support . Excellent verbal and written communication skills. Strong shell scripting skills . Good to have Python . Secondary responsibilities. :- Able to configure and setup NOSQL databases like Mongodb and Cassandra. Ability to learn new technologies along with a team and a positive outlook to understand problems from the business point of view. Qualifications: Bachelor's degree in Computer Science, Information Technology, or a related field (or equivalent experience). Proficiency in database management systems such as , MySQL or NoSQL databases. SQL programming and database design skills. Knowledge of database performance tuning and optimization techniques. Familiarity with database security best practices. Scripting and automation skills (Good to have- Python). Good problem-solving and analytical skills. Excellent communication and teamwork ski
About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
NVIDIA is looking for Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer! What you'll be doing: Primary responsibilities will include building AI/HPC infrastructure for new and existing customers. Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting. Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement. Maintain services once they are live by measuring and monitoring availability, latency, and overall system health. Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements. What we need to see: BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields. At least 5+ years of professional experience in networking fundamentals, Ethernet or InfiniBand World. Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS, etc. Possess solid working knowl
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities, such as reducing costs and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Staff Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimising costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimised for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimise AWS cloud infrastructure, including EKS, ECS, S3, VPCS, RDS, IAM, and more. I
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities, such as reducing costs and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Staff Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimising costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimised for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimise AWS cloud infrastructure, including EKS, ECS, S3, VPCS, RDS, IAM, and more. I
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE You will join the Portworx team in Everpure, which is responsible for delivering the highest quality Portworx Enterprise products. You will be contributing clean & robust code, be customer oriented and put quality first.A WHAT YOU'LL DO Designing and developing cloud native microservices and integrating new features to Portworx products Bringing a focus on design, development, unit/functional testing, code reviews, documentation, continuous integration and continuous deployment Debug product and performance issues in large scale clusters using AI tooling Collaborating with peers and stake-holders to take solutions from initial design to production Take full ownership of design and development activity by adapting to customer feedback and handling issues found in unit testing, system testing and customer deployments Experimenting with new technologies in order to push the state-of-the-art and innovate new solutions. We are primarily an in-office environment and therefore, you will be expected to work from the Bangalore office in compliance with Everpure's policies, unless you are on PTO, or work travel, or other approved leave. WHAT YOU BRING BS in Computer Science 8+ years of experience in Designing, Development and Testing of Enterprise products (Golang preferred). Good understanding of Microservice Architectures and Cloud Native platforms Designing and owning micro services to operate and scale in a di
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE You will join the Portworx team in Everpure, which is responsible for delivering the highest quality Portworx Enterprise products. You will be contributing clean & robust code, be customer oriented and put quality first.A WHAT YOU'LL DO Designing and developing cloud native microservices and integrating new features to Portworx products Bringing a focus on design, development, unit/functional testing, code reviews, documentation, continuous integration and continuous deployment Debug product and performance issues in large scale clusters using AI tooling Collaborating with peers and stake-holders to take solutions from initial design to production Take full ownership of design and development activity by adapting to customer feedback and handling issues found in unit testing, system testing and customer deployments Experimenting with new technologies in order to push the state-of-the-art and innovate new solutions. We are primarily an in-office environment and therefore, you will be expected to work from the Bangalore office in compliance with Everpure's policies, unless you are on PTO, or work travel, or other approved leave. WHAT YOU BRING BS in Computer Science 8+ years of experience in Designing, Development and Testing of Enterprise products (Golang preferred). Good understanding of Microservice Architectures and Cloud Native platforms Designing and owning micro services to operate and
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE You will join the Portworx team in Everpure, which is responsible for delivering the highest quality Portworx Enterprise products. You will be contributing clean & robust code, be customer oriented and put quality first.A WHAT YOU'LL DO Designing and developing cloud native microservices and integrating new features to Portworx products Bringing a focus on design, development, unit/functional testing, code reviews, documentation, continuous integration and continuous deployment Debug product and performance issues in large scale clusters using AI tooling Collaborating with peers and stake-holders to take solutions from initial design to production Take full ownership of design and development activity by adapting to customer feedback and handling issues found in unit testing, system testing and customer deployments Experimenting with new technologies in order to push the state-of-the-art and innovate new solutions. We are primarily an in-office environment and therefore, you will be expected to work from the Bangalore office in compliance with Everpure's policies, unless you are on PTO, or work travel, or other approved leave. WHAT YOU BRING BS in Computer Science 7+ years of experience in Designing, Development and Testing of Enterprise products (Golang preferred). Good understanding of Microservice Architectures and Cloud Native platforms Designing and owning micro services to operate and scale in a di
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Pure Solutions team as a Senior MLOps Solutions Engineer to architect and build high-scale, enterprise-grade AI/ML solutions. You will be instrumental in integrating Pure Storage platforms with the evolving open-source MLOps ecosystem (Kubeflow, MLflow, Ray) to operationalize the complete machine learning lifecycle. This role requires a creative technologist with deep Python expertise to drive innovation and enable our customers and partners to achieve production AI success. WHAT YOU'LL DO Design and Automate MLOps Pipelines: Lead the development of end-to-end MLOps workflows using CI/CD tools (Git/Jenkins) and orchestration platforms (MLflow/Kubeflow), specifically integrating Pure Storage's FlashBlade, FlashArray, and Portworx as the high-performance data plane for data ingestion, training, and inference. Build High-Performance AI/ML Reference Architectures: Create validated, repeatable deployment models using Infrastructure as Code (e.g., Ansible, Terraform) for AI/ML environments spanning bare metal, virtual machines, and GPU-accelerated Kubernetes clusters, ensuring optimal performance for distributed training. Optimize and Operationalize GPU Inference: Architect and implement solutions for high-throughput, low-latency model serving, utilizing technologies like NVIDIA Triton Inference Server and advanced optimization techniques (quantization, model sharding like DeepSpeed/Megatron-LM, and dynamic bat
Roles and Responsibilities Installation and configuration of NoSQL instances on single or multiple ports. ? Hands on experience of production on medium to big sized NoSQL databases Setting up and maintaining users and privileges management systems and Troubleshooting relevant access issues. Understand the transaction flowsand ACID compliance. Performing on-call support and should be able to provide the first level support . Configure and setup NOSQL databases like mongodb and Cassandra. Automation of repetitive tasks. Qualifications & Experience 3-6 years of Hands-on experience of working with NoSQL DBA . Some exposure to external tools like Percona , ProxySQL , HAP etc. Understanding of networking concepts . verbal and written communication skills. Experience in tools like shell , python . perl etc for automation. fundamentals on the linux system side and monitoring tools like top , iostats , sar etc. Clear understanding of NoSQL Replication process flows , threads , setting up multi node clusters and basic troubleshooting. Understanding of at least one of the backup and recovery methods for MySQL, fundamentals of SQL. Understand and tune complex SQL queries when needed.
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As the Engineering Manager, GitLab Delivery - Operate , you’ll guide a globally distributed team focused on making it easier for customers to deploy, upgrade, and run GitLab reliably in their own infrastructure. You’ll help shape the systems and tooling that support environments ranging from single-node virtual machines to large Kubernetes clusters, with a focus on reliability , operational simplicity , upgrade velocity , and zero-downtime capabilities across GitLab.com , GitLab Dedicated , and self-managed deployments. In this role, you’ll partner closely with a Product Manager and work across Infrastruc
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. We are seeking a versatile marketing leader to manage our marketing engagement across India. As our first local marketing person in the region, you will design and execute a comprehensive GTM strategy that accounts for the diversity of the Indian landscape. You will be responsible for identifying high-growth clusters, navigating complex regulatory environments, and building a unified voice that resonates across multiple borders. Your work spans social, influencers, partner and event marketing. You will be a fulcrum of all the region's outbound activity. You Will: Holistic Regional Strategy: Design a scalable marketing framework that can be adapted for Tier 1 markets (high volume) and Tier 2/3 markets (high growth). Agile GTM Execution: Lead product launches across the region, prioritizing markets based on data, barrier to entry, and product-market fit. Cross-Border Community & Influencer Growth: Leverage the region’s high social media penetration to build a loyal community through influencer partnerships and localized social channels (WhatsApp, Instagram, TikTok). Policy & Comms Collaboration: Work closely with legal and policy teams to navigate the varied regulatory lan
Other cities to consider
More places hiring for this role
Get new cluster chef jobs in India by email
Daily job updates · Unsubscribe anytime