Jobs in India

Ansible in India

41 active opportunities · Updated October 2026

Explore current ansible jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.

N
📍 Bengaluru, India
✓ Quality checkedCompany trend -100%

We are seeking a highly skilled and experienced Staff Network Site Reliability Engineer (SRE) to join our Enterprise Network Operations and SRE team. In this role, you will be pivotal in implementing our vision for a reliable and efficient network infrastructure. The ideal candidate is passionate about network operations and committed to enhancing the user experience. You'll have the opportunity to solve complex network challenges using hands-on debugging and by focusing on network automation, observability, documentation, and operational excellence. This is a critical position focused on ensuring user satisfaction and brilliance in network operations. What you'll be doing: Owning the operational aspect of the network infrastructure, ensuring its high availability and reliability, actively working on network incidents and service requests. Partnering with architecture and deployment teams to guarantee that new implementations are supportable and align with production standards. Advocating for and implementing automation to reduce toil and improve operational efficiency. Minimizing manual operational tasks to achieve and maintain Service Level Objectives (SLOs). Monitoring network performance, identifying areas for improvement, and collaborating with relevant teams to implement refinements. Proactively identifying and mitigating network risks to promote continuous improvement. Collaborating with domain experts across functions to resolve production issues swiftly and effectively, ensuring customer happiness. Conducting blameless postmortems and following through on Root Cause Analyses (RCAs). Discovering opportunities for operational improvements and teaming up with colleagues to devise solutions that enhance excellence and sustainability in network operations. Developing knowledge base articles for automa

PythonLinuxAnsible
G
📍 Pune, Maharashtra, India
✓ High-confidence listingCompany trend +33.3%
Quick readStrong listing-quality and freshness signals

Location Details: Pune, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a hybrid position. You’ll divide your time between working remotely from your home and an office, so you should live within commuting distance. Hybrid teams may work in-office as much as a few times a week or as little as once a month or quarter, as decided by leadership. The hiring manager can share more about what hybrid work might look like for this team. Join our Team Our team builds and operates the foundational infrastructure platforms that power GoDaddy's engineering organization. We own critical services including secrets management, software distribution, host security controls, and live patching for thousands of Linux systems running on OpenStack. This role sits at the intersection of Linux engineering, platform engineering, reliability engineering, and security. You will help define how core infrastructure services are designed, operated, automated, and scaled across the enterprise! What you'll get to do... Design, build, and operate highly available, scalable, and secure infrastructure platforms supporting large-scale Linux environments, with a focus on reliability, resiliency, and operational efficiency Lead the architecture, implementation, and operation of infrastructure services, including OpenStack, enterprise secrets management, package management, software promotion pipelines, and platform lifecycle management Develop and maintain automation solutions using infrastructure-as-code, Ansible, Python, Go, and self-service capabilities to improve efficiency and reduce operational overhead Build and improve observability and reliability practices through monitoring, logging, alerting, dashboards, managing incidents, analyzing underlying causes, disaster recovery, and service health reporting

PythonGitLinuxAI
NS
📍 Gurugram, Haryana, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

NK Securities Research is a leading financial firm that leverages cutting edge technology and sophisticated algorithms to trade the financial markets. Founded in 2011, we have gained invaluable experience in the field of High Frequency Trading across different asset classes. Key Responsibilities: As a Software Developer - Platform, you will play a vital role in building and enhancing tools that empower our Quant, Infrastructure, Compliance, and Operations teams. In addition to web application development, you will be responsible for automating critical infrastructure processes, optimizing tasks, and delivering scalable solutions for our trading ecosystem. Application Development: Develop and maintain in-house software tailored for business and trading requirements. Enhance internal tools and applications to improve the user experience for different teams. Ensure robust and reliable trade monitoring systems through continuous innovation. Automation Development: Design and develop frameworks to automate infrastructure provisioning, configuration, and deployment using latest industry best practices Automate exchange-specific tasks such as connectivity management, order book monitoring, and trade execution workflows. Task Optimization: Identify and optimize repetitive tasks through scripting and configuration management. Ensure scalable and adaptable solutions to support multiple exchanges and regions. Infrastructure as Code (IaC): Use Ansible to codify infrastructure configurations, ensuring consistency and repeatability. Manage playbooks for server setups, network configurations, and middleware deployment. System Monitoring and Maintenance: Develop tools for system health checks, performance monitoring, and logging. Automate response mechanisms for critical alerts and incidents. Collaboration: Work closely with infrastructure, network, and trading teams to gather requirements and deliver robust solutions. Coordinate with exchange connectivity teams to ensure complianc

JavaScriptPythonJavaReact
TA
📍 India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine

PythonKubernetesCI/CDAI
TA
📍 Bengaluru, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw

PythonKubernetesLinuxAI
GR
📍 Gurugram, Haryana, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Role: VPN Engineer Location: Gurgaon Graviton is a privately funded quantitative trading firm striving for excellence in financial markets' research. We are seeking a Network Engineer for our team in Gurgaon. Graviton trades across a multitude of asset classes and trading venues using a gamut of concepts and techniques ranging from time series analysis, filtering, classification, stochastic models, pattern recognition to statistical inference analysing terabytes of data to come up with ideas to identify pricing anomalies in financial markets. Responsibilities Manage and support corporate network infrastructure across multiple locations, including routers, firewalls, switches, wireless access points, VPN gateways, Internet links, and LAN/WAN connectivity. Configure and troubleshoot VPN technologies such as IPsec, SSL VPN, site-to-site VPN, remote-access VPN, WireGuard, OpenVPN, and FortiClient/FortiGate VPN, including issues related to authentication, tunnels, routing, DNS, packet loss, performance, split tunnelling, and firewall policies. Manage secure connectivity between offices, data centres, cloud environments, and remote users. Configure and maintain office LAN infrastructure, including VLANs, trunk/access ports, inter-VLAN routing, DHCP, DNS, NAT, ACLs, static routing, BGP, and OSPF where required. Manage multiple ISP connections, including primary and backup Internet links, automatic failover, and monitoring of utilization, latency, jitter, packet loss, and link availability. Coordinate with ISPs and telecom providers for new circuits, link failures, bandwidth upgrades, routing issues, packet-loss investigations, and service escalations. Manage firewall policies, NAT rules, VPN policies, network objects, and routing, while regularly reviewing and removing unnecessary access. Implement network segmentation across user, server, management, guest, and other business networks, while maintaining secure administrative access to network equip

PythonLinuxRestAI
GR
📍 Gurugram, Haryana, India· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Role: Network Engineer Location: Gurgaon Graviton is a privately funded quantitative trading firm striving for excellence in financial markets' research. We are seeking a Network Engineer for our team in Gurgaon. Graviton trades across a multitude of asset classes and trading venues using a gamut of concepts and techniques ranging from time series analysis, filtering, classification, stochastic models, pattern recognition to statistical inference analysing terabytes of data to come up with ideas to identify pricing anomalies in financial markets. Key Responsibilities Design, deploy, operate, and troubleshoot low-latency network infrastructure used by trading firms. Manage connectivity to global stock exchanges, brokers, market-data providers, and ISPs. Build and maintain colocation infrastructure including routers, switches, Layer-1 devices (added advantage), structured cabling and cross connects. Configure and support Cisco Nexus, Arista and similar platform devices. Design and troubleshoot Layer 2 and Layer 3 networks including: VLANs, VRFs, BGP, OSPF, Static routing, PIM, IGMP, Multicast, SSM, ACLs and QoS. Troubleshoot packet loss, multicast issues, duplicate packets, IGMP/PIM and multicast/BGP routing. Monitor and optimize latency, jitter, packet loss, interface errors, congestion, and network performance. Work with ultra-low-latency technologies including: Cut-through switching, Layer-1 switches, FPGA-based network devices, Kernel-bypass networking, ExaNIC/Solarflare NICs, Hardware timestamping. Configure and troubleshoot PTP and clock synchronization infrastructure. Perform server and network equipment installation in exchange and third-party data centres. Manage rack layout, patching, cable optimization, optics, DACs, cross-connects, and inventory. Coordinate network changes with exchanges, telecom providers, brokers, vendors, and data-centre teams. Plan and execute production changes during approved maintenance windows. Perform pre-change validation, c

PythonAIGoAnsible
N
📍 Mumbai, India
✓ Quality checkedCompany trend -100%

NVIDIA is looking for Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer! What you'll be doing: Primary responsibilities will include building AI/HPC infrastructure for new and existing customers. Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting. Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement. Maintain services once they are live by measuring and monitoring availability, latency, and overall system health. Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements. What we need to see: BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields. At least 5+ years of professional experience in networking fundamentals, Ethernet or InfiniBand World. Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS, etc. Possess solid working knowl

PythonLinuxAIAnsible
N
📍 Bengaluru, India
✓ Quality checkedCompany trend -100%

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What you will be doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuous

PythonDockerKubernetesAI
N
📍 Bengaluru, India
✓ Quality checkedCompany trend -100%

We are seeking a Senior Software Engineer with strong infrastructure expertise to design, build, and operate the next generation of our enterprise Observability, Automation, and AI-driven Reliability Platform. This role will build highly scalable distributed systems and platform services spanning Storage, Compute, Network, VMware, OpenShift, and bare-metal infrastructure. The engineer will help transform infrastructure operations from reactive monitoring and manual remediation to proactive, predictive, and AI-driven autonomous operations. What You Will Be Doing: Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at large scale. Develop reusable platform services, APIs, automation frameworks, and control planes that enable self-service, reduce operational toil, and automate infrastructure operations across multiple engineering teams. Build scalable telemetry and event-processing systems spanning metrics, logs, traces, events, topology, and alerts, with the performance and efficiency to process billions of infrastructure signals. Build intelligent and AI-native reliability capabilities, including agentic workflows for anomaly detection, forecasting, root-cause analysis, automated debugging, and closed-loop remediation. Drive technical architecture and engineering direction across Storage, Compute, Network, and Platform domains, solving complex and ambiguous problems that span multiple teams. Engineer for production at scale, with strong focus on software quality, scalability, security, performance, observability, maintainability, and operational readiness. Provide technical leadership and mentorship, influence engineerin

PythonKubernetesAITerraform
N
📍 Bengaluru, India
✓ Quality checkedCompany trend -100%

NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services. What you’ll be doing: Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale. Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation. Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data. Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems. Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation. What we need to see: BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10&#43; years operating production infrastructure or platform services. Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges. <l

PythonDockerKubernetesLinux
JT
📍 India
✓ Quality checkedCompany trend -100%

- Proven experience deploying and managing Kubernetes clusters for AI/ML workloads. Experience of at scale deployments with Azure Kubernetes. Experience level - 5 Years or more Positions - 2 Proven experience deploying and managing Kubernetes clusters for AI/ML workloads. - Experience of at scale deployments with Azure Kubernetes Service, RedHat OpenShift, Microk8s and Helm Charts. - Expertise with infrastructure and resource management and virtualization tools such as VMWare/EXSi, KVM, Ansible, Redfish. - Strong understanding of Run:AI platform, including job scheduling, quota management, and GPU virtualization. - Knowledge of NVIDIA AI Enterprise components including, NIM, NeMO, TAO, Triton and Nucleus Servers - Familiarity with DGX systems, Jetson, and NVIDIA’s AI Factory components. - Proficiency in Python, C++, and optionally .NET/C# for enterprise integration.

PythonAzureKubernetesAI
O
📍 India· Full-time
✓ Quality checkedCompany trend -68.5%

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities, such as reducing costs and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Staff Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimising costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimised for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimise AWS cloud infrastructure, including EKS, ECS, S3, VPCS, RDS, IAM, and more. I

PythonAWSDockerKubernetes
O
📍 India· Full-time
✓ Quality checkedCompany trend -68.5%

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities, such as reducing costs and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Staff Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimising costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimised for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimise AWS cloud infrastructure, including EKS, ECS, S3, VPCS, RDS, IAM, and more. I

PythonAWSDockerKubernetes
D
📍 Pune, India
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You'll Do: Deploy, configure, and maintain Kubernetes clusters for our microservices architecture. Utilize Git and Helm for version control and deployment management. Implement and manage monitoring solutions using Prometheus and Grafana. Work on continuous integration and continuous deployment (CI/CD) pipelines. Containerize applications using Docker and manage orchestration. Manage and optimize AWS services, including but not limited to EC2, S3, RDS, and AWS CDN. Maintain and optimize MySQL databases, Airflow, and Redis instances. Write automation scripts in Bash or Python for system administration tasks. Perform Linux administration tasks and troubleshoot system issues. Utilize Ansible and Terraform for configuration management and infrastructure as code. Demonstrate knowledge of networking and load-balancing principles. Collaborate with development teams to ensure applications meet reliability and performance standards. Who you are: Bachelor’s degree in engineering (CS / IT) or equivalent degree from a well-known Institute / University. 2+ years of experience in a Site Reliability Engineer role or similar. Proven experience with Kubernetes, Git, Helm, Prometheus, Grafana, CI/CD, Docker, and microservices architecture. Strong knowledge of AWS services, MySQL, Airflow, Redis, AWS CDN. Proficient in scripting languages such as Bash or Python. Hands-on experience with Linux administration. Familiarity with Ansible and Terraform fo

PythonMySQLRedisAWS
🔔

Get new ansible jobs in India by email

Daily job updates · Unsubscribe anytime