We’re looking for a Staff Software Engineer with deep experience in GenAI/ML to join Datadog’s Application Performance Monitoring (APM) team. APM is a product which provides deep visibility into applications, enabling users to identify performance bottlenecks, troubleshoot issues, and optimize services. With distributed tracing, profiling, out-of-the-box dashboards, and seamless correlation with other telemetry data, Datadog APM provides some of the deepest and most structured visibility into the health and performance of applications. This context sets us up for an opportunity to be the world leaders in agentic investigations and incident troubleshooting. You’ll act as a technical leader within the APM group, focused on agentic workflows. You’ll lead efforts to design, train, evaluate, and deploy GenAI/ML models at scale. We’re looking for a product-minded ML engineer with strong technical expertise, excellent communication skills, and a track record of driving impactful initiatives end to end. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Act as a technical leader within the APM organization, driving GenAI/machine learning projects from concept to production. Build and benchmark GenAI/ML models using state-of-the-art techniques. Collaborate with cross-functional teams to build automated investigation and triaging tools. Influence product direction by bringing a strong product mindset to your work, always advocating for the end user. Guide teams through ambiguity, scaling challenges, and evolving requirements with clear technical direction. Actively mentor engineers and influence engineering culture through leadership in design reviews, technical talks, and working groups. Who You Are: You have a BS/MS/PhD in a scientific field or equiva
Jobiba hiring network
Lead Software Engineer Inference Performance Optimization Salary India Jobs
6,753 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current lead software engineer inference performance optimization salary india jobs. Use filters to narrow by work mode, employment type, experience and date posted.
You’ll shape the future of a business‑critical platform as the technical lead across both product engineering and cloud infrastructure. You’ll modernize a mature .NET application running on AWS today, while steering its evolution toward a cloud‑native, React/Node.js, AI‑enabled architecture. If you enjoy owning architecture end‑to‑end, from backend and frontend through CI/CD, DevOps, and AWS infrastructure, this role gives you real influence at Staff Engineer level and the opportunity to set engineering standards that others follow. You’ll spend your time leading complex .NET and React features, designing scalable AWS infrastructure with Infrastructure as Code, and building automation that makes releases fast, safe, and repeatable. You’ll work on performance, reliability, and modernization in equal measure—fixing what’s slowing the platform down today and designing what it will look like in the next generation. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture and development of enterprise .NET services and APIs that power a business‑critical platform. Design and operate AWS infrastructure (using AWS CDK in TypeScript) to support secure, scalable, multi‑environment deployments. Build and optimize CI/CD pipelines (AWS CodePipeline, CodeBuild, Windows build agents) to make shipping .NET and React changes fast and reliable. Drive modernization initiatives across the stack, including clean architecture, refactoring legacy components, and reducing technical debt. Design and tune PostgreSQL and MSSQL database solutions for performance, scalability, and reliability. Mentor engineers and influence engineering practices across teams, raising the bar on cloud, DevOps, and software design. These are the essentials you’ll need to get an interview Significant experience (typically 8+ years) delivering and operating scalable enterprise software, owning both application code and cloud infrastructure. Deep hands‑on expertise with C
Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is experiencing rapid growth. You’ll play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features, while also balancing scalability, reusability, and performance. As a Staff Engineer on the Application Engineering team, you’ll serve as a senior technical leader - responsible for designing and delivering high-quality, scalable software systems that power Cohere Health’s core platform. You’ll act as a multiplier, elevating the technical bar for the team, mentoring engineers, and partnering with product, data, and clinical teams to deliver solutions that meet compliance, quality, and performance standards. This role is ideal for engineers who thrive on solving complex problems in healthcare, have deep expertise in building distributed systems, and want to influence architecture and engineering practices at scale. What you’ll do: Technical Leadership & Architecture Define and drive the architecture of large-scale, distributed application systems across the Cohere platform. Ensure solutions are secure, performant, maintainable, and compliant with NCQA, CMS, and payer requirements. Champion engineering best practices in CI/CD, testing, release management, and observability. Hands-On Engineering Write clean, maintainable, and well-tested code, primarily in modern frameworks (e.g., Python, TypeScript/React, Java/Kotlin). Lead the development of core features and APIs that directly impact providers, payers, and patients. Partner with DevOps and Data teams to ensure seamless integration, scalability, and operational readiness. Quality & Compliance Focus Embed automated testing, monitoring, and release safeguards into the development lifecycle. Proactively address compliance and audit-readiness requirements in application
We’re looking for a Senior Staff Software Engineer with deep experience in GenAI/ML to join Datadog’s Application Performance Monitoring (APM) team. APM is a product which provides deep visibility into applications, enabling users to identify performance bottlenecks, troubleshoot issues, and optimize services. With distributed tracing, profiling, out-of-the-box dashboards, and seamless correlation with other telemetry data, Datadog APM provides some of the deepest and most structured visibility into the health and performance of applications. This context sets us up for an opportunity to be the world leaders in agentic investigations and incident troubleshooting. You’ll act as a technical leader within the APM group, focused on agentic workflows. You’ll lead efforts to design, train, evaluate, and deploy GenAI/ML models at scale. We’re looking for a product-minded ML engineer with strong technical expertise, excellent communication skills, and a track record of driving impactful initiatives end to end. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Serve as the technical owner for GenAI initiatives within APM, leading design, development, and deployment of ML/AI-powered features across multiple teams. Guide long-term strategy and technical direction for GenAI workflows across APM and related products. Build and benchmark GenAI/ML models using state-of-the-art techniques. Contribute to Datadog’s broader senior engineering community through thought leadership and collaboration on company-wide initiatives. Collaborate with cross-functional teams to build automated investigation and triaging tools. Influence product direction by bringing a strong product mindset to your work, always advocating for the end user. Guide teams through ambiguity, sc
The ML Observability team builds cutting-edge tools to monitor, explain, and improve AI systems in production, particularly those leveraging Large Language Models (LLMs) and generative AI. We provide robust, scalable observability for AI workloads, including drift detection and model evaluation, and behavior tracing, enabling customers to ship AI with confidence. As a Staff Engineer, you’ll lead the development of new features and foundational capabilities within Datadog’s LLM Observability product. You will shape product direction, drive experimentation, and apply your deep understanding of both AI systems and software engineering to solve open-ended problems in the fast-moving AI landscape. Your work will directly impact how our customers monitor, troubleshoot, and optimize LLM-based applications in production. Join us in building the foundational tools that make AI systems observable, understandable, and reliable in the real world. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Drive design and implementation of LLM observability features. Ideate, prototype, and scale new product features to provide insights and drive improvements for generative AI systems Work cross-functionally with other eng teams, product, UX, and applied science to iterate fast and find product-market fit Develop and extend tools for tracing, evaluating, and debugging LLMs Influence architecture decisions and mentor engineers to build resilient, high-performance systems Stay close to customer pain points and use those insights to guide product and engineering priorities Stay current with industry trends and advancements in machine learning and observability, driving innovation within the team Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering or r
Join the Atlas Search team to design and develop the next generation of Semantic and Vector Search infrastructure. Atlas Search is a growing cloud service that allows users to execute complex search and vector search queries using the MongoDB Query Language. Our users can focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team is building a cloud-based distributed system responsible for the core components of search including data ingestion, performance, query language, query execution, for both relevance-based search and vector search. Our product is being adopted quickly and there are many interesting projects. This is a technical role where you will be responsible for the infrastructure and features enabling our at-scale cloud service powering vector and semantic search. We are looking to speak to candidates who are based in the San Francisco Bay Area for our hybrid working model. What You’ll Do Lead complex projects across the MongoDB ecosystem, for instance, development of a new Search deployment framework within the MongoDB managed cloud Set project level strategy, architect features, and lead projects to successful execution Identify, design, and implement features enhancing our reliability, performance, security and efficiency Perform code reviews with peers and make recommendations on how to improve our software development processes Influence and grow team members through active mentoring and leading by example What We Look For 5+ years experience in data management/search systems, ideally with a strong distributed systems and infrastructure background Experienced in the development and maintenance of concurrent, stateful services Eager to shape the technological direction of a complex system and have the ability to lead initiatives through collaboration with others Experienced in writing features, debugging and optimizing multithreaded applications written in Java Familiarity with LLM
Role Overview You’ll be the Principal Software Engineer driving the next generation of a large-scale enterprise SaaS platform. In this role, you combine deep hands-on engineering with high-impact technical leadership, shaping how cloud-native and AI-enabled products are designed and built. You’ll design and deliver secure, scalable, serverless systems on AWS using TypeScript and Node.js, modernize critical platform components, and set the technical direction for multiple teams. You’ll also lead how AI capabilities are integrated across the product ecosystem, ensuring they are transparent, observable, and compliant. If you enjoy system-level thinking, complex distributed architectures, and mentoring senior engineers while still staying close to the code, this role gives you company-wide impact and the opportunity to define the long-term technical vision. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture and delivery of secure, scalable, serverless applications on AWS using TypeScript/Node.js. Define and evolve the platform architecture, driving modernization, performance, resilience, and maintainability. Design and operate distributed, event-driven systems using services like Lambda, DynamoDB, Aurora, S3, and EventBridge. Shape and implement AI-enabled solutions, embedding governance, observability, and responsible AI practices into the platform. Own Infrastructure as Code (e.g., Terraform, AWS CDK, CloudFormation) to reliably provision and manage cloud infrastructure. Mentor senior engineers, influence technical decisions across teams, and clearly communicate complex concepts to diverse stakeholders. These are the essentials you’ll need to get an interview Extensive experience (typically 12+ years) building secure, production-grade software systems. Proven track record architecting and delivering cloud-native, serverless applications on AWS. Strong expertise in Node.js, TypeScript, REST API design, and at leas
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview Join a team focused on transforming how Mastercard's payment systems are built, scaled, and operated. As a Senior Software Engineer, you will lead the design and development of cloud-ready applications, microservices, and distributed systems that support large-scale payment processing platforms while helping advance modernization, automation, and engineering excellence across the organization. In this role, you will contribute to software architecture decisions, drive technical design discussions, and partner with engineers to deliver scalable, resilient, and maintainable software solutions. You'll have the opportunity to solve complex technical challenges, mentor other engineers, and influence how software is designed, developed, tested, and supported across critical technology platforms. What You Will Do •Design software solutions and contribute to software architecture decisions that support scalability, maintainability, and operational excellence. •Translate complex product requirements into technical designs and implementation plans. •Lead development of modular, extensible, high-performance applications. •Design and implement comprehensive unit, functional, and integration testing strategies. •Analyze, optimize, and improve application performance, scal
Role Summary: Datadog is seeking a Staff Software Engineer to help shape the future of our Bring Your Own Cloud (BYOC) Logs offering by unifying observability pipelines with log management software that customers deploy and manage in their own infrastructure. This role will focus on building and scaling systems that process, route, and store high-volume observability data within customer-managed infrastructure. You will operate as a hands-on technical leader, driving architecture, cross-team delivery, and product direction across a complex and evolving space. This is a high-impact opportunity to influence product strategy, mentor engineers, and solve deeply technical challenges at scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Make customer-controlled deployments feel like a managed Datadog product: deployment, upgrades, configuration, observability, diagnostics, reliability, and secure operation across diverse customer cloud environments Build and scale high-throughput systems for log processing, routing, and transformation across distributed environments Lead cross-team initiatives, aligning engineers, product managers, and stakeholders to deliver complex, multi-team projects Design and implement software that runs reliably that customers deploy and operate within their own cloud infrastructure. Improve system performance, scalability, and cost efficiency through thoughtful trade-off analysis and capacity planning Contribute hands-on to critical code paths, debugging, and deployment challenges in customer environments Who You Are: You have significant experience building software that is installed, deployed, and operated in customer environments rather than only as a fully managed SaaS service. You have strong expertise in distributed systems,
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Lyft Business organization builds the platform that lets organizations move the people they depend on employees, clients, and customers. From the foundational mechanics of a company funding and provisioning a ride to the tailored workflows that get an employee to a client meeting or home after a late shift, we sweat the small stuff to make Lyft the transportation solution businesses build their programs on. As a Staff Software Engineer on the Lyft Business Team, you will act as a critical technical leader, taking holistic ownership of complex systems that span the rider experience, the administrator experience, and the enterprise integrations behind them — defining strategic roadmaps, driving cross-functional alignment, and driving engineering excellence to improve how organizations move their people. Responsibilities: Shape the long-term architecture for systems, taking accountability for both short-term functionality and long-term health Translate high-level business goals into actionable engineering projects. Own the technical roadmap from conception to delivery, managing cross-team dependencies and mitigating risks Champion improvements in system, observability, performance, and tech debt reduction, extending your influence beyond your immediate team Establish best practices for deployment, alerting, and on-call health. Take holistic ownership of the platform's stability, tracking down issues root causes and building preventive safeguards for your immediate scope and beyond. Drive alignment across product, design, and operations. Proactively resolve bottlenecks and make decisive trade-offs to protect system architecture from competing organizational priorities Mentor and level up the engineers around you. Delegate stretch opportunities, lead cross-team reviews, and foster a culture of s
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Global Support & Partnerships team is a critical and unique part of helping Lyft succeed in our purpose to Serve and Connect. We support Lyft's international growth by developing integration platforms that seamlessly link global customers to an exceptional customer care experience. We go beyond simply supporting customers; we build the foundational infrastructure that turns every support interaction into a moment of genuine connection. Whether supporting our established North American community or welcoming new users worldwide, we ensure a unified, reliable, and efficient experience for riders, drivers, applicants and support agents alike. We are looking for an experienced technical leader who can support our global ambition by designing, owning, and scaling the Global Support Platform. Responsibilities: Set the technical vision and strategy for a rapidly evolving product area, making high-judgment calls on architecture and technology direction Drive the adoption of AI and machine learning solutions to optimize customer support workflows and enhance operational efficiency across the platform. Lead a team of talented engineers who ship code and tackle hard engineering problems, maintaining a high bar for technical excellence Translate high-level business goals into actionable engineering projects. Own the technical roadmap from conception to delivery, managing cross-team dependencies and mitigating risks. Drive the responsible adoption of AI development tools across engineering teams - modeling effective use, establishing best practices, and mentoring engineers to improve productivity without compromising code quality or security. Champion improvements in system, observability, performance, and tech debt reduction, extending your influence beyond your immediate team. Establish best practices f
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. About the role The App Runtime team is building a platform that lets every Snowflake user go from an idea to a live, deployed web application (Node.js first) in minutes. We own the end-to-end experience of building and running apps, from Cortex AI-assisted development to deployment in a secure, scalable infrastructure. Apps built on the platform inherit Snowflake's security, governance, lineage, and access controls by default. See Deploy Faster with Snowflake Apps . As a Staff/Principal Engineer, you will shape the product and its architecture. It’s a high leverage role - you will lead a team to scale and harden an early-stage, public-preview product. Your decisions will have a long-lasting impact on the future of Snowflake as an app platform. This role requires a unique combination of deep hands-on expertise in scalability, performance and security, great product instincts, customer obsession and the organizational influence to drive cross-team programs. Responsibilities Own the roadmap and technical decisions to evolve the public-preview platform into a production-grade one, adding features like horizontal scaling, cost efficiency via suspend/resume, support for stacks beyond Node.js, and stronger access and data-governance controls Stay deeply hands-on by authoring specs
Want to be a bswifter? At bswift we’ve been transforming benefits administration since 1996, making it simpler, smarter, and more human. Our state-of-the-art, cloud-based technology and services empower employees to understand, manage, and love their benefits. From downtown Chicago, and remotely across the country, we serve thousands of companies and millions of people nationwide, reducing administrative burdens and freeing HR teams to focus on creating thriving, people-first workplaces. We’re looking for motivated and goal-driven individuals who share our passion for delivering excellence and creating solutions that make a difference. The reward is a fun, flexible and creative environment with ample opportunity for professional and personal growth. If you love the bswift values of pursue excellence, embrace accountability, deliver superior service, and be a great place to work, we want to hear from you! About the Role We are looking for an AI Engineer II to design, build, and deploy cutting-edge generative AI applications , including agentic workflows and intelligent chat experiences , that transform how employees interact with their benefits. This role goes beyond individual contribution—you will own complex AI features end-to-end , influence architecture decisions around LLMs and agent systems , and mentor junior engineers . You will play a key role in bringing secure, scalable, and high-performance AI solutions to production using AWS cloud technologies. You will collaborate closely with software engineers, architects, product managers, and cross-functional teams to deliver seamless and impactful user experiences Key Responsibilities 1. Solution Design & Development Lead the design and development of scalable, secure, and high-performance applications . Architect robust backend systems and integrate them with frontend applications and third-party services. Drive technical decisions and participate in design reviews alongside senio
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Cortex CoWork team is defining the future of AI for enterprise data. Our mission is to transform how the world’s largest enterprises interact with their data through flagship products like Snowflake (CoWork) Intelligence . As a Principal AI Engineer , you will be a technical North Star for our AI initiatives. You won't just execute on a roadmap; you will help define it. You will tackle the most complex, "frontier" problems in agentic reasoning, NL-to-SQL, and enterprise-scale RAG, ensuring our AI products are not only innovative but fundamentally reliable and scalable for the Fortune 500. What you will do in this role: Technical Strategy & Architecture: Define the long-term technical vision for Snowflake Intelligence. Lead the architectural design of multi-agent systems, complex tool-use frameworks, and self-correcting NL-to-SQL engines. Drive Industry-Leading Reliability: Move beyond simple evals to build world-class, automated "hill-climbing" infrastructure. You will establish the methodology for how Snowflake measures and guarantees LLM performance across diverse customer schemas. Cross-Functional Influence: Partner with Product and Engineering leadership to align AI capabilities with business goals. You will bridge the gap between Research (modeling) and Product
As a Staff Engineer on Datadog's Compute – Disruption and Workload Placement team, you'll help define how our Kubernetes fleet scales to meet the demands of rapidly growing AI and cloud-native workloads. You'll work on the systems that ensure engineering teams have the right compute capacity, in the right region, at the right time across AWS, Google Cloud, and Azure. This is a highly technical, high-impact role where you'll shape the future of capacity orchestration, influence platform architecture, and solve infrastructure challenges that directly support Datadog's continued growth. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead the technical direction of capacity management and workload placement for Datadog's Kubernetes platform spanning 100,000+ virtual machines across multiple cloud providers. Design and build systems that optimize how engineering workloads are scheduled and deployed across regions while balancing capacity constraints, reliability, and performance. Partner across infrastructure teams to evolve multi-region and multi-cloud capacity orchestration as Datadog continues to scale. Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency. Use data and capacity signals to influence infrastructure decisions, forecast growth, and improve workload placement strategies. Who You Are: You have significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems. You are an experienced software engineer with strong programming skills, ideally in Go or a comparable systems programming language. You have hands-on experience with at least one major cloud provider (AWS, Google Cloud, or Azure) and understand distributed cloud infrastructure. Yo
Get new lead software engineer inference performance optimization salary india jobs by email
Daily job updates · Unsubscribe anytime