Biohub is the first large-scale initiative bringing frontier AI models, massive compute, and frontier experimental capabilities under one roof. We're building a general-purpose system to accelerate scientific discovery, integrating frontier AI models, biological foundation models, and lab capabilities, with the ultimate goal of curing disease. Our technology powers scientists around the world, translating AI capabilities into tools that accelerate research everywhere. The Team The AI Cluster Production Engineering team is part of the AI Compute Platform organization at Biohub, a non-profit research lab committed to open science and open-source AI. We own the design, operation, and reliability of large-scale multi-GPU AI clusters that power frontier AI biology research: protein language models, genomic foundation models, and scientific reasoning systems built to be shared, not monetized. Our clusters run Slurm on Kubernetes infrastructure and support everything from day-to-day AI researcher workflows to multi-node hero training runs at thousands of GPUs. The team works at the intersection of AI tooling, distributed systems, HPC, and frontier AI, debugging deep AI infrastructure problems and building AI systems critical to the entire AI organization. The Opportunity CZ Biohub's mission is to cure or prevent all human disease. Achieving that requires training frontier-scale AI biology models, and that demands reliable, high-performance compute infrastructure. This is production engineering work at a frontier AI lab, with the twist that the mission is biology and the science is open. You'll keep GPU clusters running at high utilization, debug the toughest distributed systems failures, and build the operational foundations for scaling to multi-thousand GPU hero runs. The technical problems are genuinely hard (e.g., multi-node distributed training, InfiniBand fabrics, large-scale storage, Slurm at scale) inside an organization where the work is aimed at helping peop
Jobiba hiring network
Staff Technical Program Manager Site Reliability Engineering Salary India Jobs
3,415 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current staff technical program manager site reliability engineering salary india jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale is seeking a Staff Software Engineer to lead the technical vision for our Infrastructure team. As a Staff Engineer, you will be responsible for the architectural evolution of our control plane and data plane, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. You will act as a force multiplier, setting the standards for Kubernetes-based cloud-native infrastructure while mentoring engineers and driving cross-functional alignment across the Ray open-source community and our proprietary product teams. Key Responsibilities Architectural Leadership: Define and drive the multi-year technical roadmap for services that orchestrate Ray clusters across diverse cloud and on-premises environments. Systemic Optimization: Lead the design and optimization of high-performance control plane components specifically tailored for large-scale, heterogeneous AI/ML workloads. Platform Reliability: Establish the organization-wide standards for the reliability, scalability, and observability of Anyscale-managed infrastructure. Strategic Integration: Direct the long-term strategy for accelerator integration (GPUs, TPUs) and container management to ens
At Secureframe , we are not just a company; we are at the forefront of revolutionizing cybersecurity compliance. Recognized as one of the industry's most innovative and trusted providers, Secureframe has consistently received accolades for our advanced technology solutions and commitment to excellence. With a robust portfolio of products that safeguard thousands of businesses worldwide, we have been featured in major publications such as Forbes’ next billion dollar startups , TechCrunch, and The Wall Street Journal for our transformative impact on the way companies achieve and maintain compliance standards. As we continue to grow, our mission remains clear: to provide seamless, secure solutions that enable businesses to focus on what they do best. Joining Secureframe means becoming part of a team dedicated to professional excellence and continuous learning in an environment that values creativity and forward-thinking. Secureframe is backed by top VCs including Kleiner Perkins, Accomplice, Gradient Ventures (Google’s AI Fund), BoxGroup, Village Global, and many more. Join Secureframe's Engineering team and help shape the future of compliance automation. Secureframe highly values having employees working in-office to foster a collaborative work environment and company culture. For office-based employees (employees who live within a defined radius of a Secureframe office), Secureframe considers working in the office, approximately 30% of the time under current policy, to be an essential function of the employee's role. Benefits Medical, dental, and vision benefits for you and your dependent(s) Flexible PTO 401(k) Paid family leave Ground floor opportunity as an early member of the team What you'll do ... Lead technical scoping, design, and implementation of new end-to-end functionality Perform detailed code reviews and provide technical mentorship to engineers Lead architecture and architectural discussions for core parts of the Secureframe application Own multi-team c
About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We’re entering a world where AI agents don’t just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents’ ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova’s throughput, correctness, and operational rigor grows dramatically. We’re looking for a Staff Software Engineer who wants to go deep on both the engine internals and the infrastructure underneath it. You’ll work across the full stack of a modern OLAP system — query planning and execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — while driving meaningful improvements to performance, cost-efficiency, and reliability at scale. You’ll influence technical direction through your work, your design reviews, and your mentorship of other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of customers. What You’ll Do Build and evolve core query engine infrastructure Work across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Design and implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Design for high-throughput automated quer
About The Role & Team We’re looking for a Fullstack Engineer to join the Core Analytics team, the group behind our flagship analytics product loved by thousands of product teams worldwide. You’ll work with our Product Manager and Designer to define, build, and ship end-to-end features that elevate the Analytics experience. You’ll own your impact from shaping product direction to delivering performant, reliable, and delightful user experiences. Our engineers are customer-focused and solve complex data problems to deliver a better user experience. We deliver value quickly and iteratively. If you’re passionate about building exceptional data-powered experiences that help businesses understand their customers, we’d love to meet you. The team’s mission is to deliver a fast, intuitive, and reliable Analytics platform that empowers customers to confidently uncover insights about their products while driving infrastructure initiatives that strengthen and scale the platform. As a Staff Engineer, you will: Work with Product Management and Design to generate and turn novel ideas for solving customer problems into engineering solutions Drive business impact through owning the end-to-end delivery of projects Have a strong critical thinking and problem-solving mindset with attention to detail Get involved in performance optimization and scaling efforts Drive business impact through leading the highest-leverage projects Have interest or experience in technical leadership of an engineering team Mentor and contribute to the success of other engineers on the team You'll be a great addition to the team if you: 7+ years of experience building and improving robust, scalable backends and developing interactive, user-facing applications or websites. Understand the whole stack and flow of user-facing web applications. Backend experience with application backends, microservices, and supporting high-throughput ingestion systems Strong critical thinking and problem-solving mindset with at
ABOUT THE TEAM The AI Foundations Team at Mural is pioneering how generative AI transforms visual collaboration and decision-making. We’re a remote-first group of engineers, designers, and product thinkers focused on helping teams work together more effectively. Our goal isn’t to replace human creativity. It’s to amplify it, building AI that enhances how people align, communicate, and make decisions visually. YOUR MISSION You will design and build the core AI systems and platforms that enable Mural’s next wave of agentic, AI-driven collaboration experiences. Rather than building isolated AI features, you’ll work on the core backend systems that power Mural’s agent platform, including agent orchestration, durable execution, contextual memory, tool integration, observability, and evaluation. Your work will enable intelligent agents to reason over product context, act on behalf of users, and operate reliably and safely at scale. Our stack at Mural includes Azure OpenAI, React, Node, MongoDB. WHAT YOU'LL DO Build the core backend systems that power Mural’s agent platform, including orchestration, durable execution, tool execution, memory, observability, and evaluation infrastructure Design scalable services and APIs that allow AI agents to retrieve context, coordinate multi-step workflows, interact with Mural data, and act reliably on behalf of users Develop the agent memory layer, including systems for conversation context, product context, retrieval, summarization, compaction, and long-term context management Create infrastructure to monitor, debug, and improve agent behavior through traces, metrics, feedback loops, and offline evaluation Translate complex, open-ended product needs into clear backend architectures, service boundaries, data models, and implementation plans that align technical capabilities with user value Help define the technical direction for agentic AI at Mural, contributing to long-term architecture and strategy Champion engineering excellence, men
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a highly skilled Staff AI Engineer – AI Product to join our ClickUp Engineering team. In this role, you will drive the development of intelligent, user-facing features that leverage the latest advancements in AI and large language models (LLMs). You will work closely with product, design, and engineering teams to deliver seamless, impactful AI-powered experiences that delight our users and differentiate ClickUp in the productivity space. This is a product-focused engineering role requiring deep expertise in AI, LLMs, and building scalable, production-ready applications. Key Responsibilities: Lead the design, development, and deployment of AI-powered features and products that directly impact ClickUp users. Collaborate with product managers, designers, and engineers to identify opportunities for AI-driven innovation and translate user needs into technical solutions. Integrate and orchestrate multiple LLMs and AI models to deliver robust, context-aware, and personalized user experiences. Prototype, test, and iterate on new AI features, leveraging user feedback and data to drive continuous improvement. Ensure the reliability, scalability, and performance of AI-powered features in production environments. Stay at the forefront of AI research and product trends, incorporating the latest advancements into ClickUp’s product roadmap. Address AI privacy, security, and compliance challenges, ensuring responsible and ethical use of AI in user-facing applications. Mentor and guide other engineers in best practices for building AI-powered products. Qualifications: Proven experience bui
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role This is a zero-to-one Product Engineering role on Replit's Money team. You'll bring our financial partnerships alive from scoping to integrating and operating them so that the builders on Replit, and the Apps and Agents they ship can transact reliably anywhere in the world. Money at Replit isn’t traditional payments. Builders are publishing apps that earn revenue. Agents are spending and being paid for work. New protocols for agentic commerce now exist to power the future of commerce: Shared Payment Tokens, the Agentic Commerce Protocol (ACP), the Universal Commerce Protocol (UCP), the Machine Payments Protocol (MPP). Replit is one of the platforms that will define what they look like in practice and lower the barrier to entry. To make any of that real, we need someone who can sit between Replit engineering and our financial partners including billing platforms, payment processors, agentic-commerce protocol partners, tax and compliance vendors to turn signed contracts into live, reliable integrations. You'll be the engineer partners ask for, and the engineer the rest of Replit relies on when a new monetization surface needs to ship. This role is a fit if you like writing code and you also like being in the room when a partnership is being scoped, because you know that the design decisions made in that room are the ones that bind the integration for years. You will Take financial partnerships from zero-to-one: scope the integration with the partner, pressure-test data and protocol specs, design the system, build it, ship it, and operate it. Own the technical relationship with Replit's financial partners across the full lifecycle: billing platforms, payment processors, agentic-commerce protocol partners, t
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We’re looking for someone to oversee internal AI use for non-coding. In this role, you’ll partner with our go-to-market, operations, and technical teams to build and deploy AI functionality to increase efficiency and outcomes. Responsibilities Identify and validate high-impact AI opportunities for internal teams, like Sales, GTM, Finance, HR, etc and build solutions to increase efficiency and outcomes. Train teams to use these tools, monitor usage, and improve as opportunities present themselves. This role will work closely with our CEO Qualifications : 5+ years of product management or engineering experience, including >1 year building AI tools Experience working with GTM teams is a plus [nice to have] Founder experience Our mission at Plaid is to unlock financial freedom for everyone. To support that mission, we seek to build a diverse team of driven individuals who care deeply about making the financial ecosystem more equitable. We recognize that strong qualifications can come from both prior work experiences and lived experiences. We encourage you to apply to a role even if your experience doesn't fully match the job description. We are always looking for team members that will bring somethin
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Making data driven decisions is key to Plaid's culture. To support that, we need to scale our data systems while maintaining correct and complete data. We provide tooling and guidance to teams across engineering, product, and business and help them explore our data quickly and safely to get the data insights they need, which ultimately helps Plaid serve our customers more effectively. Engineers on Data Infrastructure are domain experts in Data Warehouse, Data Lakehouse, Spark, Workflow Orchestration, and Streaming technologies. We scale our existing data pipelines in a performant and cost efficient way while creating the necessary abstractions to make developing on top of this platform extremely simple for other engineers at Plaid. Responsibilities Contribute towards the long-term technical roadmap for data-driven and machine learning iteration at Plaid Leading key data infrastructure projects such as improving ML development golden paths, implementing offline streaming solutions for data freshness, building net new ETL pipeline infrastructure, and evolving data warehouse or data lakehouse capabilities. Working with stakeholders in other teams and functions to define technical roadmaps for key backe
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. The Data Governance team makes sure Plaid handles consumer and customer data responsibly — and can prove it. Our mission is to enforce Plaid's privacy commitments and regulatory obligations in the systems themselves rather than in policy documents: we build the platform and controls that govern how data flows through Plaid — where it lives, who can use it, for what purpose, and for how long. That includes verifiable deletion of consumer data on request, enforcement of data-use restrictions so downstream systems can only use data in permitted ways, and the cataloging and classification that let Plaid know what data it holds and how sensitive it is. We operate at the scale of Plaid's entire data footprint, and correctness and auditability matter to us as much as throughput. As a Staff Software Engineer on Data Governance, you will set the technical direction for how Plaid enforces data governance at scale. You'll lead the design of distributed backend systems that reliably delete, restrict, and track data across dozens of services, making architectural decisions whose blast radius spans the whole company. You'll drive multi-quarter initiatives from ambiguous privacy and regulatory requirements through
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are seeking an experienced and technically influential Senior Software Development Engineer to join our Cloud Tooling and Pipelines team. This pivotal team is responsible for the design, development, and maintenance of our core Continuous Delivery (CD) platform (leveraging Spinnaker and custom tooling), Infrastructure as Code (IaC) execution engines (primarily Terraform), and a suite of supporting microservices. These systems are critical for enabling and managing our extensive resource footprint across AWS ECS and EKS . As a Senior Software Development Engineer, you will be a key contributor, driving the implementation of scalable, reliable, and secure software solutions that automate infrastructure provisioning and application deployments. Your deep expertise in software engineering principles and cloud-native development will be essential in building and enhancing our critical tooling for infrastructure provisioning, vulnerability management, and IaC deployments. You will also play a vital role in mentoring other engineers and influencing the team's technical roadmap. If you have a strong passion for building robust software systems that empower operational efficiency at scale, we encourage you to apply. Key Responsibilities Design and Develop Core Platform Components: Lead the design and development of scalable and reliable microservices and tools that form the backbone of Okta's Continuous Delivery (CD) platform (including components for Spinnaker,
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Device Identity and Access Organization The Device Identity and Access organization is at the forefront of Okta’s Zero Trust vision. As a foundational pillar within Okta Research and Development (ORD), our mission is to transform the device itself into a secure, trusted, and effortless identity factor. We are the teams responsible for ensuring users can seamlessly interact with their work from any endpoint, anywhere in the world. Our organization is comprised of engineers who thrive at the intersection of deep client-side platform engineering and massive-scale distributed systems. The work we do secures millions of enterprise endpoints globally, prevents modern identity attacks, and fundamentally changes how people work by making world-class security completely invisible to the end user. The Opportunity We are seeking a highly impactful and influential Staff Software Engineer to join our engineering team. The ideal candidate will leverage their deep expertise in distributed systems, particularly Java, to architect, build, and scale the critical server-side software and services at the heart of our security and identity platform. This is a high-visibility, hands-on opportunity to define the architectural vision, pioneer new capabilities, and drive the technical strategy to solve complex, company-wide challenges and shape the future of Okta's device identity ecosystem. You will act as a key technical leader—a player-coach and force multiplier—by setting the t
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Team : Have you ever considered what powers the intelligent features behind seamless product experiences? The GenAI team is at the forefront of enabling AI-powered security and intelligent innovation across our organization. From crafting AI powered security services, to intuitive generative AI-powered chat experiences that provide instant product support, and developing the best developer experience around authentication for generative AI and AI agents, our team is instrumental in bringing the transformative power of AI to life. We collaborate closely with the Machine Learning team and various product teams to ensure the seamless and secure delivery of AI-enhanced features that provide real value to our users. The Opportunity : As a Staff Machine Learning Engineer on the Generative AI team, you will help shape, architect, and accelerate our Generative AI strategy by contributing across the stack of model development, infrastructure, and platform services. You’ll drive design and implementation of production-ready AI/ML systems at scale: ranging from LLM-powered features to reusable components that other teams across Okta can build on. You will have the opportunity to: Architect, design, and deploy robust Machine Learning & GenAI systems, ensuring seamless integration with diverse platform services and establishing scalable LLMOps pipelines in production. Drive technical decision making while striving to hit the right balance between factors s
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Streaming Foundations team builds services and operates data pipeline infrastructure to support event streaming, messaging, and analytics use cases. We are looking for a Software Engineer who is passionate about distributed systems, platform engineering, and solving data-intensive problems at scale. In this high-impact role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Auth0 to scale for years to come. What you’ll be doing Help set the technical direction for the team and influence the engineering roadmap for the Platform’s streaming capabilities Design and lead the implementation of our most complex and critical systems for data-intensive use cases. Research and champion new technologies and architectural patterns to solve strategic challenges and scale the platform. Lead and influence cross-functional initiatives, ensuring technical alignment and successful execution across multiple teams. Improve the operational posture of our systems by designing for observability, reliability, and scalability, and by mentoring others in operational best practices. Coach and mentor senior engineers and act as a technical leader across the engineering organization. Collaborate with different stakeholders like product teams whenever needed. What you’ll bring to our teams 7+ years of software development experience in a fast-paced, agile environment Experience working with Golang or Java is preferred H
Get new staff technical program manager site reliability engineering salary india jobs by email
Daily job updates · Unsubscribe anytime