We are investing in agentic AI and need a Senior AI Engineer to lead the design and delivery of these systems. This is a foundational hire: you will own both the agent-facing workstreams — pipelines, orchestration, conversational interfaces — and the underlying context layer that makes them reliable, including memory management, knowledge graph integration, and retrieval infrastructure. You will work closely with data engineers, project leads, and client stakeholders, and play a key role in shaping how Lynx builds and ships AI solutions at scale. What This Involves: Lead the architecture and delivery of agentic AI systems end-to-end: agents, orchestration, tool use, and multi-step reasoning workflows. Own the context layer: design and implement memory architectures (episodic, semantic, working memory) and integrate GraphRAG and knowledge graph retrieval into agentic pipelines. Build robust RAG systems — including vector retrieval, graph traversal, and hybrid search — and ensure retrieval quality through evaluation frameworks. Translate client requirements into technical designs, presenting approaches and trade-offs to both technical and non-technical stakeholders. Define standards and reusable patterns for agentic AI development that other engineers at Lynx can build on. Set up observability, evaluation, and monitoring pipelines to ensure AI systems perform correctly in production. Requirements: 5–8 years of software or ML engineering experience, with at least 2–3 years building LLM-based or agentic AI systems in production. Deep hands-on experience with agentic frameworks (LangChain, LlamaIndex, AutoGen, CrewAI, or similar) and LLM APIs (OpenAI, Anthropic, etc.). Strong understanding of agent design patterns: ReAct, planning loops, tool use, multi-agent coordination, and memory architectures. Practical experience with GraphRAG or knowledge graph-based retrieval (e.g., Neo4j, Microsoft GraphRAG) and vector databases (Pinecone, Weaviate, Qdrant, etc.). Proficiency in
Jobiba hiring network
Software Engineer Ml Infrastructure Platform Salary India Jobs
6,326 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current software engineer ml infrastructure platform salary india jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary As a research engineer at Graphcore, you will contribute to the advancement of AI research, investigating new ideas that push the limits on important AI/ML problems. Specialised hardware has been the key driver of the progress of AI over the last decade, and we believe that hardware-aware AI algorithms and AI-aware hardware developments will continue to be critical to advancing this exciting field. We are therefore looking for individuals who combine strong machine learning experience with practical engineering skills to deliver impactful AI research. We are seeking AI researchers with strong software engineering experience, particularly in lower-level programming and performance optimisation for hardware efficiency. Our research spans a broad range of topics, including efficient training and inference, world models, life sciences, reinforcement learning, and beyond. You will work closely with researchers to generate ideas and translate them into scalable implementations, contributing to publications and projects that help to steer the future of AI hardware. The Team Graphcore Research participates in both fundamental and applied research, to characterise the computational requirements of machine intelligence a
We are expanding our agentic AI capability and are looking for an AI Engineer to join the team. You will work alongside senior engineers to build and maintain AI systems — contributing to agentic pipelines, retrieval infrastructure, and the integrations that tie these systems together. This is a hands-on implementation role with real ownership of components. You will grow your skills in a fast-moving AI practice, working on production systems that directly affect client outcomes. What This Involves: Build and maintain agentic pipelines and workflows under the guidance of senior engineers: tool use, orchestration, and multi-step reasoning. Implement and tune RAG pipelines — including embedding, chunking strategies, vector retrieval, and retrieval evaluation. Contribute to memory and context layer components: integrating vector databases, supporting knowledge graph pipelines, and helping maintain state management across agentic systems. Write clean, well-tested Python code and participate in code reviews. Debug and improve existing AI systems based on evaluation results and production feedback. Collaborate with data engineers and domain experts to integrate AI components with upstream data sources and downstream applications. Document implementations clearly and contribute to shared internal tooling. Requirements: 2–4 years of software or ML engineering experience, with at least 1 year working with LLMs or AI systems in a professional setting. Working knowledge of LLM APIs (OpenAI, Anthropic, or similar) and at least one agentic or RAG framework (LangChain, LlamaIndex, or equivalent). Solid Python skills and comfort with software engineering basics: version control, testing, REST APIs. Familiarity with vector databases or embedding-based search. Curiosity about agentic AI — you follow developments in the space and are eager to apply new techniques. Excellent communication and collaboration skills — comfortable working across cross-functional and client-facing te
About Anyscale At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure. As part of this role, you will Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices We'd love to hear from you if you have Familiarity with running ML inference at large scale with high throughput and low latency Familiarity with deep learning and deep learning frameworks (e.g. PyTorch) Solid understanding of distributed systems, ML inference challenges Bonus points
About the Team The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. Our mission is to push the frontier of code generation and agentic reasoning, and deploy these capabilities in real-world products such as ChatGPT and the API, as well as in next-generation tools specifically designed for agentic coding. We operate across research, engineering, product, and infrastructure—owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. About the Role As a Performance & Systems Engineer on the Codex team, you will be responsible for whole-system optimization across a complex, evolving stack. Codex spans LLM inference, cloud orchestration, agentic work management, and multiple product surfaces. Your job will be to identify and land high-leverage changes—across infrastructure, modeling, and product layers—that make Codex agents significantly faster and cheaper to serve. We’re looking for generalists who thrive in ambiguity and love chasing performance bottlenecks to ground. This is a high-ownership role where your work will directly improve the experience of millions of users. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Hunt down and address inefficiencies across the Codex system stack, from agent behavior to LLM inference to container orchestration, and beyond. Build tooling to measure, profile, and optimize system performance at scale. Collaborate with researchers and engineers to land high-ROI changes that improve latency and cost. You might thrive in this role if you: Have experience operating across both ML systems and cloud infrastructure. Enjoy diving into messy, ambiguous problems and emerging with clear wins. Think holistically about performance, balancing spee
Senior Cloud Security Engineer At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale's security needs are growing as we operate more production and cloud infrastructure for larger and more demanding customers. We're looking for a Senior Cloud Security Engineer to own the security of that infrastructure. This is a hands-on, high-ownership role: you will own how our production and cloud environments are hardened, isolated, and monitored. You will set and drive the direction for infrastructure and production security, reporting to the Head of Security and partnering closely with the wider engineering organization. This role is based in the San Francisco, Bay Area. In your first year, success looks like hardened and well-segmented production environments, strong runtime security coverage across our container footprint, and a clear, defensible story for how we secure the infrastructure our customers rely on. What You'll Do Own the security posture of Anyscale's production and cloud infrastructure across AWS and Azure, including hardening, network segmentation, and tenant isolation. Own runtime security coverage across our Kubernetes environments, from deployment through detection of anomalous activity. Partner with engineering on secure infrastructure architectur
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. As a Data Engineer, you’ll be responsible for laying the foundation for a best-in-class analytics function. You’ll partner closely with our engineering team and business stakeholders to ensure that our analytics stack and processes meet the business needs today with an eye towards the future. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll do as a Senior Data Engineer at Vanta: Design and deploy data infrastructure needed to drive data-driven decision-making solutions Design and implement complex data orchestration models, modeling metadata, scaling reporting tools for data science and ML products users Be the company’s expert on data administration, data management and scalable data systems Write highly tuned, scalable SQL queries running over large-scale, heterogeneous data warehouses Work with the Product and Enterprise Engineering system teams to structure source systems for reporting consumption across the enterprise Help maintain CDC pipelines to power customer reporting Help develop front end applications to expose analytical data sets enterprise wide How to be successful in this role: Have at least four years of experience working with data and two years of experience in Software Engineering or a related field. Have experience with common analytics tooling (e.g. Stitch/Fivetran, Snowflake/BigQuery/Redshift, dbt, Airflow, Dagster). Have good working knowledge of AWS data infra systems and Terraform. Bring a system-oriented and software engineering mindset to the Data Engineering practice. We’re looking to build frameworks that manage data, and minimize bespoke queries Deep kn
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX (Hybrid) About the role You'll design and build the core infrastructure that powers AI inference across Cloudflare's global network — real-time voice, frontier open LLMs, and customer-deployed models running on a heterogeneous fleet of GPUs and next-generation accelerators in hundreds of cities worldwide. Working alongside AI/ML engineers, hardware partners, and Cloudflare product teams, you'll solve hard prob
About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are looking for a Principal Architect to define and drive the architectural vision of the software stack for the Graphcore ML accelerator. In this role, you will shape the architecture of our software ecosystem and maintain a deep understanding of the product’s hardware and software components, their interfaces, and how they interact. You are an excellent communicator, and you proactively convey the software architecture. You bring a pragmatic, trade-off-aware approach to decision-making, fully recognising the impact of architectural choices on product direction and engineering outcomes. The Team The software architecture team is responsible for defining, maintaining and communicating the overarching architecture of our software stack, from firmware to ML frameworks. The team works within the wider software organisation, partnering closely with engineering teams who deliver against this architectural vision. Responsibilities and Duties Define & document the software architecture of the software stack. Work across different software
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Design and implement novel research ideas, ship state of the art models to production, and maintain deep connections to academia and the government. We have one of the highest ratio of compute to engineers in the world. We do not delineate strongly between engineering and research. Everyone will contribute to writing production code and conducting research depending on individual interest and organizational needs. We have all the compute, data, and talent available for you to do your best work. As a Member of Technical Staff - Sovereign AI, you will: Design, build and scale agentic AI systems for serving mission critical use cases. Research, implement, and experiment with ideas on our supercompute and data infrastructure. Learn from and work with the best researchers in the field. Execute across the full AI stack and ship products to serve public interest. You may be a good fit if you have: Canadian citizenship and eligibility for security clearance ( required for this role ). Extremely strong software engineering skills. Proficiency in Python and related ML frameworks. Experience training, evaluating, and using (
The ML Observability team builds cutting-edge tools to monitor, explain, and improve AI systems in production, particularly those leveraging Large Language Models (LLMs) and generative AI. We provide robust, scalable observability for AI workloads, including drift detection and model evaluation, and behavior tracing, enabling customers to ship AI with confidence. As a Staff Engineer, you’ll lead the development of new features and foundational capabilities within Datadog’s LLM Observability product. You will shape product direction, drive experimentation, and apply your deep understanding of both AI systems and software engineering to solve open-ended problems in the fast-moving AI landscape. Your work will directly impact how our customers monitor, troubleshoot, and optimize LLM-based applications in production. Join us in building the foundational tools that make AI systems observable, understandable, and reliable in the real world. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Drive design and implementation of LLM observability features. Ideate, prototype, and scale new product features to provide insights and drive improvements for generative AI systems Work cross-functionally with other eng teams, product, UX, and applied science to iterate fast and find product-market fit Develop and extend tools for tracing, evaluating, and debugging LLMs Influence architecture decisions and mentor engineers to build resilient, high-performance systems Stay close to customer pain points and use those insights to guide product and engineering priorities Stay current with industry trends and advancements in machine learning and observability, driving innovation within the team Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering or r
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Who We Are: Shape the future of Roblox’s virtual economy. The Economy ML team is building the machine learning backbone that powers Roblox’s Marketplace, Developer Monetization, and Payments ecosystems. From intelligent pricing and personalized storefronts to dynamic layout optimization and avatar understanding, we’re reimagining how the Roblox economy drives user engagement, monetization, and creator success at scale. As a Principal Software Engineer (Data Systems) , you will architect, build and deploy high-scale, reliable real-time and batch data systems for personalization, search and recommendation across various product surfaces in Marketplace, Developer Monetization and Payments. You will be involved in key data projects from architecting event taxonomies and logging interfaces to real-time feature serving across multiple search and recommendation surfaces. What You’ll Do Act as data engineering lead for Economy ML, setting standards for batch vs streaming feature pipelines, table design, observability, and documentation used across the Economy group. Work as a hands-on contributor on our data systems to power content recommendation, search and personalization across Economy product
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Mapping team at Lyft is tasked with building a digital representation of the physical world - a map. We collect and serve the freshest and most accurate mapping data possible, along with algorithms, models, platform services, and map-based user experiences that power Lyft’s current and future transportation offerings. Mapping represents a huge opportunity for Lyft’s business, but also a big challenge. We build and scale systems that deal with large data storage, real-time data processing, machine / deep learning pipelines, routing and ETA models, driver and passenger location tracking, and more. We built beautiful and magical user experiences on top of all those services, and compete with companies that have been in the mapping business for decades. To strengthen our efforts, we are hiring a Senior ML Engineer who will work end-to-end on creating and improving new capabilities to detect changes in the environment and reflect them in our Lyft map using a wide variety of input sources from the Lyft fleet. For this we are looking for someone who values software engineering best practices, loves the algorithmic and geospatial side of the challenge and is data-driven from start to end. Our technology stack ranges from basic machine learning models to large language models and running them at scale on millions of images. You will work with incredibly passionate and talented colleagues from machine learning, data science, and engineering on projects that delight our passengers and drivers – powered by an up to date map. Responsibilities: Partner with Engineers, Data Scientists, Product Managers, and Business Partners to apply machine learning for business and user impact Perform data analysis and build proof-of-concept to explore and propose ML solutions to both new and existing proble
About Snorkel At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data. We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler! In September 2026 we raised a $350 million Series E at a $3.5 billion valuation , and we are scaling our engineering and research teams to meet demand. The role Frontier AI data is expensive to make and hard to measure. Every task we deliver is tested against the strongest models, often through many long-running agent rollouts. Your job is to make that process faster, cheaper, and more rigorous with ML and AI You will be one of the early members of ML & Research Engineering at Snorkel. You will study how frontier-grade data is generated and evaluated, form hypotheses, validate them against real production data, and ship the winners at scale. You will shape the discipline's direction, its standards, and the team that grows around it. What you'll work on Efficient agentic evals. Cut the cost of long-horizon agent evaluation with adaptive sampling, statistically grounded early stopping, model cascades, caching, and cheap-first gating. AI model routing. Route every eval and judge call to the cheapest model that clears the quality bar, with fallback, monitoring, and cost attribution. Fine-tuned small models. Fine-tune and serve open-weight models (LoRA and other
We’re looking for a Staff Software Engineer with deep experience in GenAI/ML to join Datadog’s Application Performance Monitoring (APM) team. APM is a product which provides deep visibility into applications, enabling users to identify performance bottlenecks, troubleshoot issues, and optimize services. With distributed tracing, profiling, out-of-the-box dashboards, and seamless correlation with other telemetry data, Datadog APM provides some of the deepest and most structured visibility into the health and performance of applications. This context sets us up for an opportunity to be the world leaders in agentic investigations and incident troubleshooting. You’ll act as a technical leader within the APM group, focused on agentic workflows. You’ll lead efforts to design, train, evaluate, and deploy GenAI/ML models at scale. We’re looking for a product-minded ML engineer with strong technical expertise, excellent communication skills, and a track record of driving impactful initiatives end to end. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Act as a technical leader within the APM organization, driving GenAI/machine learning projects from concept to production. Build and benchmark GenAI/ML models using state-of-the-art techniques. Collaborate with cross-functional teams to build automated investigation and triaging tools. Influence product direction by bringing a strong product mindset to your work, always advocating for the end user. Guide teams through ambiguity, scaling challenges, and evolving requirements with clear technical direction. Actively mentor engineers and influence engineering culture through leadership in design reviews, technical talks, and working groups. Who You Are: You have a BS/MS/PhD in a scientific field or equiva
Get new software engineer ml infrastructure platform salary india jobs by email
Daily job updates · Unsubscribe anytime