Description: Graviton is a privately funded quantitative trading firm striving for excellence in financial markets' research. We are seeking a Quantitative Researcher for our team in Gurgaon. This team trades across a multitude of asset classes and trading venues using a gamut of concepts and techniques ranging from time series analysis, filtering, classification, stochastic models, pattern recognition to statistical inference analysing terabytes of data to come up with ideas to identify pricing anomalies in financial markets. Responsibilities: Develop new or improve existing trading models using in-house platforms Use advanced mathematical techniques to model and predict market movements Analyse large financial datasets to identify trading opportunities Provide real time analytical support to experienced traders Requirements: Possess a degree in a highly analytical field, such as Engineering, Mathematics, Computer Science from IITs schools Quantitative bent of mind A working knowledge of Linux/Unix Programming experience, preferably in C++ or C No prior knowledge of financial markets is needed but must have a strong interest in learning about financial markets. Have a strong work ethic Hard Working Benefits: Our open and collaborative work culture gives you the freedom to innovate and experiment. Our cubicle free offices, non-hierarchical work culture and insistence to hire the very best creates a melting pot for great ideas and technological innovations. Everyone on the team is approachable, there is nothing better than working with friends! Our perks have you covered. Competitive compensation Annual international team outing Fully covered commuting expenses Best-in-class health insurance Delightful catered breakfasts and lunches A well-stocked kitchen 4 week annual leaves along with market holidays Gym and sports club memberships Regular social events and clubs After work parties
Jobs in India
Inference Technical Lead in India
186 active opportunities · Updated October 2026
Showing
15 jobs
Explore current inference technical lead jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.
Role: Network Engineer Location: Gurgaon Graviton is a privately funded quantitative trading firm striving for excellence in financial markets' research. We are seeking a Network Engineer for our team in Gurgaon. Graviton trades across a multitude of asset classes and trading venues using a gamut of concepts and techniques ranging from time series analysis, filtering, classification, stochastic models, pattern recognition to statistical inference analysing terabytes of data to come up with ideas to identify pricing anomalies in financial markets. Key Responsibilities Design, deploy, operate, and troubleshoot low-latency network infrastructure used by trading firms. Manage connectivity to global stock exchanges, brokers, market-data providers, and ISPs. Build and maintain colocation infrastructure including routers, switches, Layer-1 devices (added advantage), structured cabling and cross connects. Configure and support Cisco Nexus, Arista and similar platform devices. Design and troubleshoot Layer 2 and Layer 3 networks including: VLANs, VRFs, BGP, OSPF, Static routing, PIM, IGMP, Multicast, SSM, ACLs and QoS. Troubleshoot packet loss, multicast issues, duplicate packets, IGMP/PIM and multicast/BGP routing. Monitor and optimize latency, jitter, packet loss, interface errors, congestion, and network performance. Work with ultra-low-latency technologies including: Cut-through switching, Layer-1 switches, FPGA-based network devices, Kernel-bypass networking, ExaNIC/Solarflare NICs, Hardware timestamping. Configure and troubleshoot PTP and clock synchronization infrastructure. Perform server and network equipment installation in exchange and third-party data centres. Manage rack layout, patching, cable optimization, optics, DACs, cross-connects, and inventory. Coordinate network changes with exchanges, telecom providers, brokers, vendors, and data-centre teams. Plan and execute production changes during approved maintenance windows. Perform pre-change validation, c
Research Engineer, Applied AI Location: Bangalore (or throughout India remote-friendly with travel) About EnCharge AI: EnCharge AI is building the next generation AI platform. Our novel in-memory-computing architecture delivers a 10x step-function improvement in compute energy efficiency and performance for AI inference workloads. As the demands of artificial intelligence move beyond today's models, we believe fundamental underlying infrastructure must evolve. We are an experienced team of AI researchers, silicon & systems engineers, and architects backed by leading investors, poised to become the essential platform for the next wave of AI innovation. The Opportunity: Modern AI workloads—from large language models to diffusion-based generators to multimodal systems—represent some of the most compute-intensive frontiers in AI, and some of the most promising applications for our hardware’s energy efficiency advantages. We’re building a vertically integrated AI stack that will showcase the transformative potential of our silicon while delivering real value to customers today. We are seeking a Research Engineer to push the boundaries of AI model capability, quality, and efficiency. You’ll build fine-tuning and post training pipelines, develop rigorous benchmarking frameworks, and work at the intersection of ML research and hardware-aware optimization—ensuring our models run beautifully on our silicon. This is a role for someone who thrives at the boundary between research and engineering. You’ll read papers, implement techniques, and ship production-quality code—all in service of making AI inference faster, cheaper, and better. Key Responsibilities: Algorithmic Acceleration: Research and implement state-of-the-art techniques to accelerate AI inference—quantization, sparsity,
EnCharge AI is a leader in advanced AI hardware and software systems for edge-to-cloud computing. EnCharge’s robust and scalable next-generation in-memory computing technology provides orders-of-magnitude higher compute efficiency and density compared to today’s best-in-class solutions. The high-performance architecture is coupled with seamless software integration and will enable the immense potential of AI to be accessible in power, energy, and space constrained applications. EnCharge AI launched in 2022 and is led by veteran technologists with backgrounds in semiconductor design and AI systems. Senior Emulation Engineer Location: India - Remote Job Description: At EnCharge AI, we are building the next generation of AI compute silicon — purpose-built for high-performance, low-power, and scalable AI inference. As an Emulation Engineer, you will play a critical role in validating complex AI accelerator architectures on emulation platforms before tape-out. This position is ideal for someone passionate about bridging the gap between hardware and software in fast-paced, deep tech environments. Responsibilities: • Set up and maintain Siemens Veloce emulation and prototyping platforms • Adapt SoC designs for Emulation and Prototyping • Develop and debug emulation testbenches and system-level environments • Support pre-silicon validation, power/performance analysis, and early software bring-up. Participate in silicon bring-up and validation. • Collaborate with design and verification teams to isolate design issues and accelerate debug. • Optimize performance of the emulation workloads and reduce turnaround time. • Work with firmware/software teams to enable use of emulators for OS and driver testing. Required Background: • BS/MS/Ph.D. in EE, CS, or related field with 7+ years of SoC design experience. • Experience with emulation platforms (Veloce, Palladium, or ZeBu) and FPGA-based prototyping systems (proFPGA, HAPS, or Protium) • Experience with emula
AI Software Engineer, Agent Harness Location: Bengaluru, Karnataka (or throughout India remote-friendly with travel) About EnCharge AI EnCharge AI is building the next generation AI platform. Our novel in-memory-computing architecture delivers a 10x step-function improvement in compute energy efficiency and performance for AI inference workloads. As the demands of artificial intelligence move beyond today's models, we believe fundamental underlying infrastructure must evolve. We are an experienced team of AI researchers, silicon & systems engineers, and architects backed by leading investors, poised to become the essential platform for the next wave of AI innovation. The Opportunity We serve open-weight models and our own bespoke checkpoints on EnCharge hardware. The models change often, and the harness around them needs to keep up. You own this layer that runs agents against files, tools, documents with permissions, memory, unattended execution, and real outputs. It will be assembled from a combination of open-source and bespoke code. Key Responsibilities Own the harness architecture end to end — agent loop, safe execution, context management, knowledge base, memory, permissions, orchestration, outputs, interfaces, observability — one component per layer, with clear interfaces so layers can be swapped. Build the pieces with no open-source equivalent e.g. session semantics, enforced permissions, memory in a human-editable file, orchestrator, and outputs. Keep pace with the models: adapters, prompt formats, tool-call schemas, stop conditions, benchmarking and evaluation. Make tool use reliable across models of uneven tool-calling quality — validation, repair, retries, fallbacks. Develop agents, tools, and MCP servers for internal and customer use cases, and review them for security before they ship. Build the evaluation harness: task suites, regression runs on every model or harness change, cost and latency per task alongside quality. Define the interfaces:
About Us Paytm is India's payment Super App offering consumers and merchants the most comprehensive payment services. As the pioneer of the mobile QR payments revolution in India, today, Paytm stands as India’s largest payment company by Users, Merchants, Payment Transactions, and Revenue. Paytm’s mission is to drive financial inclusion in India and bring half a billion Indians into the mainstream economy through technology-led financial Services. Paytm enables commerce for small merchants and distributes various financial services offerings to its consumers and merchants in partnership with financial institutions. Paytm has been a pioneer in the merchant space by introducing innovative solutions like QR codes to accept payments and Sound-box to reconcile payments via voice alerts. We are also distributing loans to these partners via our ‘Paytm for Business’ App. About The Team Paytm Intelligence is building next-generation AI platforms across two key areas: AI inference infrastructure and agentic AI solutions. Our focus is on enabling enterprises to move from AI experimentation to production-grade deployment with measurable business outcomes. In the agentic AI space, we are building solutions such as Outreach Manager, CLM agents, and domain- specific AI agents that automate complex customer journeys across sales, service, and operations. These systems go beyond traditional chatbots, orchestrating workflows, decisions, and actions across enterprise systems. About the Role We are hiring an AI Agentic Solutions Architect to drive enterprise adoption of AI-powered workflow automation. This role sits at the intersection of business process transformation, AI systems, and enterprise sales. You will work directly with enterprise customers to understand their workflows, identify automation opportunities, and design agentic solutions that deliver measurable outcomes such as cost reduction, conversion uplift, and improved customer experience. Key Responsibilities Customer
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Job Description/ Responsibilities: Designing, developing and maintaining stable and reliable AI/ML Ops platforms / pipelines Minimum experience of 4-6 Years required in AI ML Ops Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production. Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow. Security & Compliance: Implementing data security measures, ensuring compliance with data governance
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... We’re in an unbelievably exciting area of tech, fundamentally reshaping the cloud-native, modern virtualization, and AI infrastructure landscape. Here, you’ll lead with innovative thinking, grow alongside us, and work with some of the smartest minds in the industry. We’re looking for engineers passionate about system testing, distributed systems, Kubernetes, storage, modern virtualization, AI workloads, and automation. You’ll work on complex, real-world scenarios involving HA, resiliency, disaster recovery, scalability, and failure testing across large-scale Kubernetes environments. As enterprises modernize their infrastructure, containers, virtual machines, and AI/ML workloads are increasingly converging on Kubernetes. From traditional enterprise applications and VMs to GPU-accelerated AI training, inference, and data-intensive workloads, Kubernetes is rapidly becoming the common platform for running the next generation of applications. This role gives you the opportunity to test and influence how Portworx delivers enterprise-grade storage, data protection, availability, and resiliency across these workloads—including containerized applications, KubeVirt/OpenShift Virtualization VMs, and demanding AI/ML workloads running on Kubernetes. WHAT YOU WILL DO: Own system-level quality for Portworx Enterprise across Kubernetes, storage, and modern virtualization environments. Develop comprehens
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... We’re in an unbelievably exciting area of tech, fundamentally reshaping the cloud-native, modern virtualization, and AI infrastructure landscape. Here, you’ll lead with innovative thinking, grow alongside us, and work with some of the smartest minds in the industry. We’re looking for engineers passionate about system testing, distributed systems, Kubernetes, storage, modern virtualization, AI workloads, and automation. You’ll work on complex, real-world scenarios involving HA, resiliency, disaster recovery, scalability, and failure testing across large-scale Kubernetes environments. As enterprises modernize their infrastructure, containers, virtual machines, and AI/ML workloads are increasingly converging on Kubernetes. From traditional enterprise applications and VMs to GPU-accelerated AI training, inference, and data-intensive workloads, Kubernetes is rapidly becoming the common platform for running the next generation of applications. This role gives you the opportunity to test and influence how Portworx delivers enterprise-grade storage, data protection, availability, and resiliency across these workloads—including containerized applications, KubeVirt/OpenShift Virtualization VMs, and demanding AI/ML workloads running on Kubernetes. WHAT YOU WILL DO: Own system-level quality for Portworx Enterprise across Kubernetes, storage, and modern virtualization environments. Develop comprehens
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Pure Solutions team as a Senior MLOps Solutions Engineer to architect and build high-scale, enterprise-grade AI/ML solutions. You will be instrumental in integrating Pure Storage platforms with the evolving open-source MLOps ecosystem (Kubeflow, MLflow, Ray) to operationalize the complete machine learning lifecycle. This role requires a creative technologist with deep Python expertise to drive innovation and enable our customers and partners to achieve production AI success. WHAT YOU'LL DO Design and Automate MLOps Pipelines: Lead the development of end-to-end MLOps workflows using CI/CD tools (Git/Jenkins) and orchestration platforms (MLflow/Kubeflow), specifically integrating Pure Storage's FlashBlade, FlashArray, and Portworx as the high-performance data plane for data ingestion, training, and inference. Build High-Performance AI/ML Reference Architectures: Create validated, repeatable deployment models using Infrastructure as Code (e.g., Ansible, Terraform) for AI/ML environments spanning bare metal, virtual machines, and GPU-accelerated Kubernetes clusters, ensuring optimal performance for distributed training. Optimize and Operationalize GPU Inference: Architect and implement solutions for high-throughput, low-latency model serving, utilizing technologies like NVIDIA Triton Inference Server and advanced optimization techniques (quantization, model sharding like DeepSpeed/Megatron-LM, and dynamic bat
About the Role Adobe is seeking a Machine Learning Engineer to join the Adobe Genuine Engineering team. This group protects Adobe's ecosystem from fraud, abuse, and misuse using intelligent systems worldwide. In this position, you will build and develop machine learning models from scratch, including custom transformer-based frameworks, to identify fraudulent actions, stop account sharing, and protect the experience of hundreds of millions of users. You will manage the entire model lifecycle: raw behavioral data and feature engineering, architecture development, large-scale GPU training, deployment, and monitoring. The team is actively building in-house behavioral foundation models that learn identity-preserving representations from long sequences of user activity. This is a role for an engineer who wants to own deep learning systems end-to-end — not consume pre-built ones. Key Responsibilities Build and train deep learning models from scratch, including custom transformer and attention-based architectures for long behavioral event sequences. Own the full training stack: event tokenization, temporal and positional embeddings, self-supervised pretraining (e.g., masked modeling, contrastive learning), and downstream fine-tuning. Train large models efficiently on GPU infrastructure using mixed-precision training, gradient accumulation/checkpointing, efficient attention, and distributed strategies (DDP, FSDP, or equivalent). Build and optimize feature pipelines on Databricks and Spark, transforming raw behavioral events into high-quality model inputs. Translate prototypes into production ML systems — scalable, reliable, and observable — and drive inference performance through architectural and serving-side optimization. Contribute to MLOps practices: experiment tracking, model versioning, CI/CD, automated retraining, and
About the Role Adobe is seeking a Machine Learning Engineer to join the Adobe Genuine Engineering team. This group protects Adobe's ecosystem from fraud, abuse, and misuse using intelligent systems worldwide. In this position, you will build and develop machine learning models from scratch, including custom transformer-based frameworks, to identify fraudulent actions, stop account sharing, and protect the experience of hundreds of millions of users. You will manage the entire model lifecycle: raw behavioral data and feature engineering, architecture development, large-scale GPU training, deployment, and monitoring. The team is actively building in-house behavioral foundation models that learn identity-preserving representations from long sequences of user activity. This is a role for an engineer who wants to own deep learning systems end-to-end — not consume pre-built ones. Key Responsibilities Build and train deep learning models from scratch, including custom transformer and attention-based architectures for long behavioral event sequences. Own the full training stack: event tokenization, temporal and positional embeddings, self-supervised pretraining (e.g., masked modeling, contrastive learning), and downstream fine-tuning. Train large models efficiently on GPU infrastructure using mixed-precision training, gradient accumulation/checkpointing, efficient attention, and distributed strategies (DDP, FSDP, or equivalent). Build and optimize feature pipelines on Databricks and Spark, transforming raw behavioral events into high-quality model inputs. Translate prototypes into production ML systems — scalable, reliable, and observable — and drive inference performance through architectural and serving-side optimization. Contribute to MLOps practices: experiment tracking, model versioning, CI/CD, automated retraining, and
About the Role Adobe is seeking a Machine Learning Engineer to join the Adobe Genuine Engineering team. This group protects Adobe's ecosystem from fraud, abuse, and misuse using intelligent systems worldwide. In this position, you will build and develop machine learning models from scratch, including custom transformer-based frameworks, to identify fraudulent actions, stop account sharing, and protect the experience of hundreds of millions of users. You will manage the entire model lifecycle: raw behavioral data and feature engineering, architecture development, large-scale GPU training, deployment, and monitoring. The team is actively building in-house behavioral foundation models that learn identity-preserving representations from long sequences of user activity. This is a role for an engineer who wants to own deep learning systems end-to-end — not consume pre-built ones. Key Responsibilities Build and train deep learning models from scratch, including custom transformer and attention-based architectures for long behavioral event sequences. Own the full training stack: event tokenization, temporal and positional embeddings, self-supervised pretraining (e.g., masked modeling, contrastive learning), and downstream fine-tuning. Train large models efficiently on GPU infrastructure using mixed-precision training, gradient accumulation/checkpointing, efficient attention, and distributed strategies (DDP, FSDP, or equivalent). Build and optimize feature pipelines on Databricks and Spark, transforming raw behavioral events into high-quality model inputs. Translate prototypes into production ML systems — scalable, reliable, and observable — and drive inference performance through architectural and serving-side optimization. Contribute to MLOps practices: experiment tracking, model versioning, CI/CD, automated retraining, and
At Bolna, we’re building tools that change the way teams leverage Voice AI. We’re looking for a Founding Machine Learning Engineer to own the end-to-end lifecycle of building, evaluating, deploying, and improving models that power millions of production conversations. This is a high-impact, high ownership role where you won’t just work on Bolna’s ML stack—you’ll help build the foundation it scales on. Our team includes IIT alumni with experience at Bain, Atlassian, Uber, Zomato, and LinkedIn, and is backed by leading investors. Responsibilities: Build the data engine - Design pipelines to source and clean conversational voice data across Indian languages, accents, and telephony conditions. Fine-tune models that ship - Fine tune and train models to improve accuracy, speed, and reliability across different use-cases. Define what "good" means - Build evaluation datasets and benchmarks for transcription accuracy, voice naturalness, interruption handling, latency, and end-to-end conversation quality. Set up human-in-the-loop pipelines to capture subjective quality at scale. Ship to production - Work with the engineering team to deploy models into a latency-sensitive, high-volume system. Monitor performance in the wild, debug regressions, and iterate fast. Required Skills: 3+ years of hands-on ML experience with deep practical real-world experience in training models. Strong Python and PyTorch fundamentals with exposure in distributed training, and modern fine-tuning techniques (LoRA, QLoRA, DPO, RLHF, etc.). Training data as a first-class problem. Experience designing data pipelines from collection, cleaning, labeling, deduplication, augmentation and treating data quality as a core engineering discipline. Rigorous about evaluation. You know that "looks good in a demo" is not a benchmark. You build the evals before you trust the model. Speech model experience is a plus with real-time / streaming inference experience where you would have contributed to latency optimization
Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time, others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team At GoDaddy, we believe better workforce decisions start with trusted data. We're growing our People Analytics capabilities and looking for a People Business Intelligence Analyst to help us build clear, reliable reporting that supports our People team and business leaders. Are you someone who enjoys finding patterns in complex datasets and transforming information into meaningful stories? Do you like solving challenges that directly influence how organisations hire, develop, and support employees? You'll own the data mapping work that connects source systems (HRIS, ATS, payroll, engagement surveys, etc.) into our Redshift warehouse, and build the dashboards, reports, and AI solutions that people leaders, HRBPs, and executives actually use to make decisions! Working with HR systems, analytics platforms, and AI-powered tools, this role helps us understand workforce trends, improve reporting, and maintain trusted data across the employee lifecycle. We are committed to creating an inclusive hiring process and will provide reasonable accommodations for candidates who need additional support. We'd love to hear from you! What you'll get to do... Map and document data flows from HR source systems into Redshift, including field definitions, lineage, and business logic (e.g., how "headcount," "attrition," or "time-to-fill" are calculated). Maintain a data dictionary and metric glossary to ensure definitions remain consistent across teams and dashboards Build and maintain dashboards in QuickSight and Power BI covering key people metrics inclu
Other cities to consider
More places hiring for this role
Get new inference technical lead jobs in India by email
Daily job updates · Unsubscribe anytime