Jobiba hiring network

Lead Software Engineer Inference Performance Optimization Salary India Jobs

6,753 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead software engineer inference performance optimization salary india jobs. Use filters to narrow by work mode, employment type, experience and date posted.

D
Datadog
📍 Madrid• Full-time
1mo ago

As a Staff Engineer on Datadog's Compute – Disruption and Workload Placement team, you'll help define how our Kubernetes fleet scales to meet the demands of rapidly growing AI and cloud-native workloads. You'll work on the systems that ensure engineering teams have the right compute capacity, in the right region, at the right time across AWS, Google Cloud, and Azure. This is a highly technical, high-impact role where you'll shape the future of capacity orchestration, influence platform architecture, and solve infrastructure challenges that directly support Datadog's continued growth. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead the technical direction of capacity management and workload placement for Datadog's Kubernetes platform spanning 100,000+ virtual machines across multiple cloud providers. Design and build systems that optimize how engineering workloads are scheduled and deployed across regions while balancing capacity constraints, reliability, and performance. Partner across infrastructure teams to evolve multi-region and multi-cloud capacity orchestration as Datadog continues to scale. Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency. Use data and capacity signals to influence infrastructure decisions, forecast growth, and improve workload placement strategies. Who You Are: You have significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems. You are an experienced software engineer with strong programming skills, ideally in Go or a comparable systems programming language. You have hands-on experience with at least one major cloud provider (AWS, Google Cloud, or Azure) and understand distributed cloud infrastructure. Yo

awsazurekubernetes
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role OpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency of our systems. In this role, you'll apply deep technical expertise to optimize infrastructure and application-level performance across mission-critical products like ChatGPT and our developer API. You’ll work cross-functionally with teams building core services, training models, and developing real-time user experiences to push our latency, throughput, and cost-efficiency to the next level. We are looking for engineers who thrive in ambiguous environments, value deep systems understanding, and are motivated by delivering measurable impact. This is a highly technical, individual contributor role focused on root-cause analysis, profiling, instrumentation, and architecture-level performance improvements across our stack. In this role, you will: Analyze and optimize performance across application, middleware, runtime, and infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond. Develop tooling and metrics that provide deep observability into system performance. Collaborate closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements. Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale. Lead investigations into high-impact performance regressions or scalability issues in production. Drive performance testing strategies and help define SLAs/SLOs around latency and throughput for critical systems. You might thrive in this role if you: Have 7+ years of experience in software engineering with a strong tr

pythonawsrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve people's lives. About the Role We are seeking a lead thermal simulation engineer to help accelerate the design and development of next-generation robotic systems through modeling, simulation, and analysis. You will work closely with mechanical, electrical, controls, and robotics engineers to evaluate designs before hardware is built, identify risks early, and guide critical architecture decisions. This role spans structural and thermal analysis and design across robotic subsystems including actuators, mechanisms, structures, electronics, and integrated systems. You will develop simulation workflows that improve engineering velocity, increase confidence in design decisions, and help us build more capable, reliable, and manufacturable robotic platforms. This role is based in San Francisco, CA. This role will be expected to be in office 4 days per week and offer relocation assistance to new employees. In this role, you will: Perform thermal simulations to assess heat generation, cooling strategies, thermal interfaces, and system-level thermal performance Partner with mechanical, electrical, and controls engineers to influence design decisions early in development Build simulation models to evaluate robotic actuators, transmissions, mechanisms, structures, soft goods, and integrated assemblies Correlate simulation results with physical testing and develop methodologies to improve model accuracy Support architecture trade studies by evaluating design concepts before hardware is built Develop simulation workflows, standards, and best practices that scale across the robotics o

awsgitrest
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
21 days ago

Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is experiencing rapid growth. You’ll play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features while also balancing scalability, reusability, and performance. As a Staff Engineer on the Application Engineering team, you’ll serve as a senior technical leader - responsible for designing and delivering high-quality, scalable software systems that power Cohere Health’s core platform. You’ll act as a multiplier, elevating the technical bar for the team, mentoring engineers, and partnering with product, data, design , clinical and payment teams to deliver solutions that meet compliance, quality, and performance standards. This role is ideal for engineers who thrive on solving complex problems in healthcare, have deep expertise in building distributed systems including data solutions, and want to influence architectural decisions for security and scale, drive cross-collaborations for alignment, establish technical standards for consistency and evolve both application and data engineering best practices at scale. What you’ll do: Technical Leadership & Architecture Define and drive the architecture of large-scale, distributed application systems across the Cohere platform. Ensure solutions are secure, performant, maintainable, and compliant with NCQA, CMS, and payer requirements. Champion platform engineering best practices in CI/CD, testing, release management, and observability. Hands-On Engineering Write clean, maintainable, and well-tested code, primarily in modern frameworks (e.g., Python, TypeScript/React, Java/Kotlin). Lead the development of core features and APIs that directly impact providers, payers, and patients. Partner with DevOps and Data teams to ensure seamless integration, scalability, and operati

typescriptpythonjava
View job →
G
Gitlab
📍 Poland• Full-time• Remote
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. About the role As a Staff Backend Engineer, you will provide technical leadership across your team and adjacent teams, solving the highest-scope and most complex problems in your area. You will lead large, cross-cutting backend initiatives, drive our modular architecture strategy, and define the standards that let teams move faster without compromising quality, security, reliability, or operability. This is a technical leadership role that combines deep backend expertise, systems judgment, product judgment, and influence across organizational boundaries. You will work with Product, Frontend, Infrastructure, Security, Data, Engine

REMOTEsqlpostgresqlkubernetes
View job →
G
Gitlab
📍 India• Full-time• Remote
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. About the role As a Staff Backend Engineer, you will provide technical leadership across your team and adjacent teams, solving the highest-scope and most complex problems in your area. You will lead large, cross-cutting backend initiatives, drive our modular architecture strategy, and define the standards that let teams move faster without compromising quality, security, reliability, or operability. This is a technical leadership role that combines deep backend expertise, systems judgment, product judgment, and influence across organizational boundaries. You will work with Product, Frontend, Infrastructure, Security, Data, Engine

REMOTEsqlpostgresqlkubernetes
View job →

Job Details: Job Description: The Role and Impact As a Systems and Solutions Engineer, you will drive the design, development, and integration of systems that combine software, firmware, board, and silicon/SoC components to meet specific customer needs. In this role, you will play a key part in defining, implementing, and optimizing solutions to ensure high performance, reliability, and quality across the system lifecycle. Your work will directly impact the seamless functionality and user experience of cutting-edge technologies, enhancing Intel's position in delivering innovative systems to global customers. Business Group You will be joining the Silicon and Platform Engineering Group (SPE), an organization committed to advancing Intel's mission of delivering world-class silicon and platform solutions. The group focuses on developing integrated systems that align with customer needs and support Intel's broader goals of leadership in technology innovation. SPE collaborates across diverse domains to ensure Intel platforms meet performance, reliability, and scalability requirements. Key Responsibilities - Design and develop software, firmware, and hardware solutions that integrate seamlessly across system components. - Lead the definition and implementation of system architecture, translating business opportunities into technical requirements and use cases. - Evaluate technical risk and optimize systems for ease of use, reliability, security, availability, and sustainability. - Drive technical solutions to address customer challenges, deploying systems and conducting benchmarks to validate performance. - Collaborate with cross-functional teams to influence next-generation requirements and solutions, guiding research and academic collaborations as needed. - Conduct lab experiments to simulate real-life environments, analyze prototype performance, and refine system

recruitment
View job →

NVIDIA is looking for a hands-on Solutions Architect Manager to lead a team of GPU, networking & software solution architects and engineers. Do you want to build and lead a group that designs, debugs, and deploys new AI hardware and software technologies into production in customer data centers? As part of the NVIDIA SA organization, you will drive people and technical leadership for end-to-end solutions deployments at some of NVIDIA's most strategic technology customers, while directly contributing to designs and deep-dive debugging and shaping our product roadmap with customer feedback. What you will be doing: Recruit & manage a team of solutions architects, system/network and software engineers focused on large-scale GPU and AI networking deployments. Set priorities, allocate resources, mentor, and ensure high-quality customer delivery across multiple concurrent projects - while remaining directly involved in key technical reviews, design decisions, and critical debug efforts. Provide deep subject-matter expertise in advanced GPU and network systems and serve as the senior technical point of contact for strategic customers. Personally lead and guide complex compute/network configuration and performance debugging, working side-by-side with your team to deliver performant, reliable clusters. Guide your team as they lead network / compute / software architecture discussions, and support server, network, and cluster bring-up, including on-site data center work where needed. Systematically collect and synthesize customer-specific requirements across your portfolio. Partner with GPU/Network Systems Engineering, Product Management, and Sales to influence roadmap priorities and packaging of reference designs and solutions. Demonstrate SME in advanced GPU & network systems and be a trusted technical advisor to NVIDIA's strategic customers. Bring customer-sp

DC
Diligent Corporation
📍 Vancouver• Full-time• From C$250K/yr
21 days ago

Overview We are seeking a hands-on Director of AI Software Engineering to lead and scale AI engineering efforts supporting multiple business units across Governance, Risk, and Compliance (GRC). This role sits at the intersection of product delivery, platform evolution, and applied AI—driving real-world impact across core workflows. This is not a pure management role. We are looking for a builder who leads from the front, someone who has recently written production code, shipped systems end-to-end, and can operate comfortably in ambiguity while aligning teams and stakeholders. What You’ll Do Lead AI Engineering Across GRC Own delivery of AI-powered capabilities embedded directly into business unit workflows (e.g., risk analysis, compliance automation, reporting, due diligence) Partner with product, data, and platform teams to translate business problems into scalable AI systems Stay Hands-On Contribute to architecture, code reviews, and critical path implementation Prototype and validate new approaches (LLMs, agents, retrieval systems, classification pipelines, etc.) Set engineering standards for performance, reliability, and cost efficiency Build and Scale Teams Lead and mentor a high-performing team of AI/ML and software engineers Drive hiring, coaching, and career development Establish a culture of ownership, speed, and technical excellence Drive Execution Deliver production-grade systems—not experiments Balance speed with rigor (security, privacy, compliance) Operate across multiple concurrent initiatives with clear prioritization Communicate and Influence Act as a bridge between engineering and business stakeholders Clearly articulate trade-offs, risks, and outcomes to senior leadership Align cross-functional teams around shared goals and timelines What We’re Looking For Proven Builder 10+ years in software engineering, with recent hands-on coding experience Demonstrated track record of shipping production systems at scale Experience with modern

pythonjavaaws
View job →

The mission of OrgStore is to provide an easy-to-use, fully-managed platform to store and search data. With Postgres as our cornerstone technology, we focus on transactional (OLTP) data use-cases while providing a custom data plane and a flexible administrative layer. Our users are the thousands of engineers representing hundreds of teams producing products at Datadog. As a platform infrastructure team, our goal is to accelerate product teams by helping them rapidly stand up new features and make step change improvements to their application's performance. Data storage is ubiquitous and the service we are providing is critical for Datadog as the business scales and expands its product offerings. As the Engineering Manager for the OrgStore Blueprint Team, you will lead a high-performing team of software engineers, guiding them in building and maintaining these mission-critical systems. Collaboration is key in this role, needing to work closely with other engineering, product, and support teams across Datadog. Your strategic vision will not only guide your team's day-to-day operations but also influence the long-term roadmap of support within Datadog. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Guide and mentor a diverse team of 4-8 software engineers, fostering their career growth while ensuring high team performance. Drive the technical roadmap in collaboration with your team, product managers, and support teams ensuring it aligns with company objectives and directly contributes to improving customer satisfaction. Engage strategically with complex technical problems and work with your team to produce well-defined and actionable plans. Be a core contributor and reviewer of code, lead design decisions, and participate i

aigorust
View job →
E
Everpure
📍 Bengaluru• Full-time
21 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... In this role as an Engineering Manager, you will lead a team of engineers located in Bengaluru, India. You will focus on driving and shaping the direction of our observability software and enabling product engineers to deliver high-quality, reliable software to our customers. You will provide technical leadership and direction, mentor engineers on your team, and collaborate with product, engineering, and cross-functional stakeholders to deliver successful outcomes. The successful candidate must understand the dynamics of global R&D, possess deep knowledge of local culture, and have the ability to champion Pure values and leadership attributes. This role requires the ability to lead and influence multiple stakeholders across cross-functional teams and drive alignment across complex, distributed engineering environments. The team will help build and evolve observability capabilities that provide actionable insights into the health, performance, capacity, and reliability of Pure's products and infrastructure. WHAT YOU'LL NEED TO BRING TO THIS ROLE... 12+ years of combined experience as a software developer and manager 3+ years of technical management experience while staying hands-on 7+ years of hands-on software development experience Strong exposure to one or more of the following areas: distributed systems, systems programming, observability/telemetry, data platforms, or solving prob

awsrestai
View job →

Observability Pipelines (OP) is Datadog's on-premise, vendor-agnostic telemetry pipeline product. As an Engineering Manager on the team, you'll own people management and engineering execution for one of OP's core missions, spanning areas like Integrations (ingesting from and routing to the many source and destination systems customers rely on), streaming insights, cost control, or pipeline capabilities, reliability and scalability. You'll partner directly with Product to help shape the roadmap, and work closely with your peer EMs and senior ICs to define how OP operates and grows. This is an opportunity to build your management craft while having real influence over the technical direction of a fast-growing product area. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Own people management and engineering execution Establish a strong operating rhythm for the team Drive high standards for on-call rotations and incident response Partner with Product on the roadmap, balancing product priorities with technical realities Lead, coach, and grow the careers of engineers on your team Who You Are: Experienced managing engineers directly, comfortable owning a team’s operating rhythm end-to-end, from planning through execution and stakeholder communication to incident and on-call ownership Have a technical background in distributed systems and data infrastructure Have experience with on-premises or customer-installed software concepts A product-minded partner to have on the team — you enjoy working with Product on strategy Experience with high-performance or Rust-based data pipeline systems Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications o

aigorust
View job →
M
1mo ago

At Playlist, life's richest moments happen when people step away from screens to move, connect, explore, and play. We're building the definitive platform for intentional living, connecting people with inspiring experiences in fitness, wellness, and beyond. With popular brands like Mindbody and ClassPass, Playlist empowers businesses and individuals, making it effortless for aspirations to become actions. Join us in reshaping technology's role to foster meaningful, real-world connections. Mindbody equips wellness entrepreneurs with technology to support thriving businesses and create exceptional experiences. Innovation and curiosity drive our culture, connecting businesses and individuals through cutting-edge solutions. Join us if you're passionate about enhancing wellness through technology. The Role You'll Play: Build and maintain performant backend systems and applications that drive real-world experiences Partner with Product, Design, and QA to bring features to life from ideation through deployment, always iterating with the end-user in mind. Champion engineering best practices—automated testing, peer reviews, observability, and elegant design Lead and influence architecture decisions that prioritize scalability and simplicity <span data-contrast=&quo

javascripttypescriptpython
View job →
B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab

kubernetesmachine learningai
View job →

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. At Baseten, we are building the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale, the network is the computer. We are looking for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building block in our infrastructure and unlocking the next generation of distributed inference optimizations. THE OPPORTUNITY Networking and compute are no longer separate disciplines; they are converging. The massive throughput of H100, B200, and NVL72 architectures enables and demands a new approach where communication is co-optimized alongside computation. We are entering an era where the network is an active accelerator, leveraging smart hardware offloads and direct interconnects to ensure that data movement operates at wire-speed. In this role, you will go beyond network configuration to architect the software fabric that unifies thousands of GPUs into a cohesive operating system. While you will leverage the best of the open-source ecosystem, you won't be limited by it. Where off-the-shelf solutions stop, you will build from scratch, engineering the primitives required to co-optimize communication and compute for Disaggregated Serving, Wide Expert Parallelism (WideEP), and lightening cold starts. WHAT YOU'LL DO Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack,

pythonkubernetesmachine learning
View job →
🔔

Get new lead software engineer inference performance optimization salary india jobs by email

Daily job updates · Unsubscribe anytime