Jobiba hiring network

Engineering Manager Platform Reliability Salary India Jobs

8,358 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current engineering manager platform reliability salary india jobs. Use filters to narrow by work mode, employment type, experience and date posted.

N
1mo ago

NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as &#34;the AI computing company.&#34; We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent Senior Engineering Manager to lead a large firmware engineering organization delivering end-to-end manageability firmware for NVIDIA's next generation Data Center Compute Systems. This role owns HGX product line and OpenBMC-based management firmware and MCU firmware components in data center platforms, including architecture, execution, quality, reliability, telemetry, and customer readiness. We are seeking an experienced senior leader with strong technical depth, broad system perspective, and a proven ability to lead large teams through complex product cycles. This role is onsite in Santa Clara, CA, USA. If you're creative and autonomous, we want to hear from you! What you'll be doing: Lead a large firmware engineering organization delivering OpenBMC based firmware and MCU firmware for next-generation Data Center Compute Systems. Own HGX platform as a lead for Firmware and System software readiness working across the organization. Define and drive the long-term firmware roadmap, balancing architectural innovation with product execution and delivery milestones. Drive architecture strategy across BMC, MCU, platform software, manageability, health management, and data center firmware interfaces. <spa

pythongitlinux
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $243.9K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . The Core Infrastructure team within Coinbase's Platform product group builds the foundational systems that keep Coinbase online, secure, and scalable, owning the compute and networking platforms that power every product and service across the company. As the Group Product Manager for Core Infrastructure & Reliability, you'll own the product vision and multi-year strategy for Coinbase's cloud infrastructure, driving the design, operation, and scaling of the systems that underpin hundreds of billions of dollars in annual transaction volume. You'll partner deeply with Engineering, SRE, Security, and Finance to ensure Coinbase's infrastructure is reliable, cost-efficient, and resilient across multiple cloud environments and regions. What you’ll do: Own the product strategy and roadmap for Core Infrastructure, spanning compute, networking, multi-region and multi-cloud architecture, and platform reliability. Strengthen infrastructure reliability and resilience programs, defining platform-level SLOs, capacity planning, failover capabilities, and incident reduction targets to meet the uptime demands of a global financial platform. Lead evaluation and adoption of cloud infrastructure technologies (Kubernetes, service mesh, distributed storage, observability, infrastructure-as-code), making build-vs-buy decisions that balance cost, speed, and long-term scalability. A

REMOTEawskubernetesai
View job →
M
Mongodb
📍 Gurugram• Full-time
1mo ago

We are seeking an Engineering Manager to join our growing Gurugram Product & Technology team to provide technical direction, direct architecture, and implement core parts of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As an Engineering Manager on this new team, you will be responsible for leading and growing an engineering team, taking on challenging, high-visibility projects that improve and enhance the performance, scalability, and reliability of the distributed systems infrastructure for this new product. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Position Expectations Provide technical leadership and mentorship to a team of engineers, fostering a culture of innovation, quality, and continuous improvement Drive the architectural vision for the platform, ensuring it runs equally well on public clouds, private cloud environments and on-premise Work with product managers, program managers, design & analytics teams and other teams to define, prioritize and deliver new features that delight our users and drive platform improvements Take responsibility for the planning and execution of major features, raise delivery risks Own the monitoring, operations, and maintenance of the systems your team develops Enable the team to operate efficiently by removing technical obstacles, coordinating with other teams on dependencies, and prioritizing the team's overall well-being Contribute to planning for organizational growth, including allocation of engineering resources, participate in hiring and assignment of projects Qualifications 8+ years of experience of building distributed systems, and/or foundatio

pythonjavamongodb
View job →
S
1mo ago

Supabase is the open-source Postgres development platform that 7M+ developers and thousands of enterprises depend on every day. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the role We’re hiring a Product Manager to own the platform primitives every Supabase product runs on, including internal features like compute , disks , networking, and the API gateway, and external features like read replicas , custom domains , PrivateLink , and Bring Your Own Cloud. What you'll be responsible for: Talk to customers across the full spectrum. Indie developers running a single nano project, fast-growing startups whose costs are dominated by compute and disk, enterprises walking through a network-architecture review, and partners building on top of Supabase. Find the real blockers and bring them back to the roadmap. Own the problem statement and requirements behind every platform bet. Capture the customer evidence behind each decision, name the cost, capacity, and reliability constraints, and give the team a target it can hit. Decide what gets built, what gets deferred, and what gets cut. Every quarter you're choosing between enterprise unlocks blocking deals, reliability and cost wins for the long tail of projects, and net-new capabilities that change what Supabase can run. Set the priorities and defend them. Define how each launch is measured before it ships. Set the metric, agree on the threshold, and track it after launch. Know whether a feature moved enterprise deal velocity, project economics, or platform reliability. Use that to sharpen the next call. Keep engineering, design, and leadership aligned. The platform touches every other Supabase product, every region, and every customer tier. Write the roadmap, surface dependencies before they become blockers, and keep decisions moving. You might be a good fit if you: Have 7+ years of produ

aigorust
View job →

Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team... Here at GoDaddy, the ML Engineering (MLE) team exists as the backbone of our machine learning infrastructure, enabling ML scientists and product teams across Domains to ship models to production reliably, efficiently, and at scale. This team owns the full lifecycle of ML systems — from CI/CD pipelines and model serving infrastructure to GPU workload orchestration and observability. Through disciplined engineering practices, thoughtful system design, and close collaboration with ML scientists, data engineers, and product teams, we deliver the platform that powers domain search, pricing, recommendations, and emerging AI experiences for millions of customers worldwide. We are currently looking for an experienced, highly motivated Senior Engineering Manager to lead our ML Engineering team based in India. This is an established team with existing engineers — we expect the candidate to ramp up quickly on our ML infrastructure stack, build strong relationships with the team, and partner with both India-based teams and US-based teams to drive execution and grow the team further. This individual will join us on our journey to build and scale ML infrastructure that serves real-time predictions at low latency, automates model deployment and promotion, and provides the observability and reliability guarantees that production ML systems demand. Become part of a team that bridges the gap between ML research and production engineering — shipping systems that directly impact GoDaddy's core revenue. What you'll get to do... Lead a team o

typescriptpythonaws
View job →
G
Godaddy
📍 India• Full-time
1mo ago

Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings Join Our Team Ready to build the future of eCommerce? Join GoDaddy's global eCommerce platform team and help power the critical systems behind catalog, pricing, ordering, provisioning, and renewals for millions of customers worldwide. Reporting to a Director of Product Management, you'll help enhance every stage of the customer journey while driving innovative agentic solutions that improve experiences, unlock growth opportunities, and redefine what's possible in digital commerce. If that sounds exciting, we'd love to meet you What you'll get to do... You'll lead product initiatives across GoDaddy's eCommerce platform, shaping the roadmap and driving execution for capabilities spanning catalog, pricing, ordering, provisioning, and renewals. We'll count on you to guide features from discovery through launch, making sure requirements are well-defined and that what we ship meets both customer and business goals You'll partner closely with our engineering teams and other cross-functional collaborators to define requirements, prioritize the backlog, and deliver solutions that balance customer needs, business goals, technical dependencies, and platform investments. We rely on data, experimentation, and customer insight to guide our decisions, so you'll use these tools to validate opportunities, define success metrics, and continuously refine our approach based on what we learn Because our team spans the US, Romania, India, and other partner organizations, you'll build alignment across time zones and cultures, and you'll help support platform reliability

gitmicroservicesai
View job →

We are looking for a strong technical leader to join the Private Action Runner team, part of the larger Action Platform group and help shape one of the core execution layers behind Datadog’s action-taking and remediation capabilities. Private Action Runner (PAR) enables Datadog products and AI agents to securely run actions inside customer infrastructure with controls for authentication, permissions, auditing and safe execution. The role will be hands-on, covering architecture, implementation, reliability and collaboration with teams integrating PAR across Datadog. It also offers leadership exposure through leading a team of 3 engineers, with the expectation that the role will quickly transition into a formal Engineering Manager 1 position as the team grows. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: (Describe role responsibilities here) Solve a scaling bottleneck in a critical service Deploy a new feature to production, progressively rolling it out with feature flags Investigate and fix a production issue from a service your team owns Design a way to scale up a service for more traffic With your team, plan the most important projects to work on next Mentor and lead a small team of 3 engineers Who You Are: (Describe role qualifications here) You have significant experience in one or more languages You value code simplicity and performance You can design architecture to solve problems at high scale You have a BS/MS/PhD in a scientific field or equivalent experience You want to work in a fast, high-growth startup environment that respects its engineers and customers You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how You have demonstrated ability to use AI

aigorust
View job →
DU
DoorDash USA
📍 San Francisco• Full-time
16 days ago

About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus

pythonjavasql
View job →
DU
16 days ago

About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the technical direction for our in-house Spark deployment and shape the architecture that will run DoorDash's data, analytics, and ML compute for the next five years and beyond. You will own the deep, cross-cutting problems that span the runtime, the shuffle service, the scheduler, and the overall service reliability — making the architectural calls that compound across the platform's lifetime. You will partner with the Engineering Manager on technical roadmap, hiring, and team shape, and act as the senior technical voice in cross-team partnerships with Data Engineering, ML Platform, and product engineering teams that depend on the platform. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Set the multi-year technical direction for an in-house Spark-on-Kubernetes platform — runtime, shuffle, scheduler, reliability — and make the architectural calls that compound for years. Own the deepest distributed-systems problems on the team: shuffle architecture, multi-tenant scheduling, runtime performance, and the failure modes that only show up at scale. Partner with the Engineering Manager on technical roadmap, hiring, inte

pythonjavasql
View job →

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Staff Engineer on the GitLab Delivery - Upgrades team, you’ll guide the technical direction for GitLab’s self-managed deployment strategy so customers can deploy, upgrade, and run GitLab reliably in their own infrastructure with minimal disruption. You’ll serve as a technical anchor for the team, working closely with your engineering manager, product manager, and partners across Site Reliability Engineering, Release, Security, and Development to shape cloud-native, operator-driven deployment patterns that reduce operational complexity and upgrade friction. In your first year, you’ll help define the a

sqlpostgresqlkubernetes
View job →
D
Discord
📍 San Francisco Bay Area• Full-time• $180K – $220K/yr
1mo ago

Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. The Realtime Infrastructure team is responsible for building and maintaining some of Discord’s highest scale and most critical services. Those systems are at the core of our text chat infrastructure and facilitate the dispatching of every update to our users sessions. This role will have a significant impact on Discord’s overall reliability and performance. It will also help our product teams build new features on top of our infrastructure. This team is small but critical, and its work has a direct impact on Discord's success and ability to scale. This role reports to the Senior Engineering Manager of Realtime Infrastructure. What You'll Be Doing Build and operate large-scale, reliable and performant distributed systems. Collaborate with product teams to create new features. Ensure Discord “just works”. Write code but also manage our infrastructure. Work with a talented team of engineers who have built one of the largest communication platforms in the world. What you should have 2+ years of experience writing and designing backend systems. Experience solving complex distributed system problems. Experience operating and maintaining critical tier 0 services. Knowledge of monitoring and alerting best practices. Familiar with open source software, and not afraid to dig into the source code of a library to find the answer you’re looking for. Bonus Points Experience with Elixir or Rust. Experience working with systems deployed in a cloud environment (GCP, AWS, etc.) Knowledge of devops tools like Salt,Terraform or k8s. You have built or contributed to open source projects. You are a Discord power user and hav

awsgcprest
View job →
DU
DoorDash USA
📍 San Francisco• Full-time
16 days ago

About the Team The Code Quality team sits within the Developer Platform organization and owns the systems that keep DoorDash's codebase healthy and secure as it scales: static analysis, quality gates, test frameworks, regression infrastructure, and tooling. Our job is to make sure the signals engineers rely on before shipping — test results, coverage, performance feedback etc — are fast and trustworthy. The decisions we make about tooling and standards directly shape how confidently and quickly engineering teams at DoorDash can ship to production. About the Role We're looking for Software Engineers to help build and maintain the systems that validate code quality across DoorDash's engineering org, treating our tooling as a critical product for the engineers who rely on it every day: static analysis and quality gates, test frameworks and regression infrastructure. You’ll design the tooling and automation that will help derive trustworthy quality signals, integrate them into the development lifecycle, and make it easy for engineers to execute reliable, repeatable workflows. You will collaborate across the engineering org, partnering directly with the teams who use what you build to understand the accuracy, reliability and performance of their functionality. You will report into the Engineering Manager on our Code Quality team in our Developer Platform organization. You must be located in either San Francisco, CA, Sunnyvale, CA, Los Angeles, CA, Seattle, WA, or New York, NY. You're excited about this opportunity because you will… Build and maintain quality tooling — static analysis, quality gates, coverage reporting, test frameworks, regression infrastructure — and integrate it directly into our developer workflows and CI/CD pipelines Define and derive quality signals - flakiness, pass rate, coverage, performance, scale readiness etc - Build tooling that improves everyday engineering workflows, including local development, CI/CD, debugging, and rollou

awsci/cdgit
View job →
D
Discord
📍 San Francisco Bay Area• Full-time• $196K – $220K/yr
1mo ago

Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. We're looking for a Senior Software Engineer to join the Growth team at Discord. Our team owns how new users discover, understand, and get started with Discord — from the first moment someone encounters us on the web through the onboarding experience that turns them into engaged members. You'll work across the stack to build the systems that acquire and activate users at scale, whether that means improving how Discord shows up across the web or designing the in-product experiences that help new users find their footing. This is a high-impact role where you'll directly influence how Discord grows. This person will report to the Senior Engineering Manager for Growth. What you'll be doing: Build and optimize user-facing experiences that improve how people discover, understand, and get started with Discord Develop scalable backend systems that serve dynamic, high-performance content at scale Work across the full stack — frontend (TypeScript, React) and backend (Python) — wherever the problem leads Design and run A/B experiments to measure the impact of your work on user acquisition and activation Partner with Product, Design, and Data Science to identify high-impact growth opportunities and iterate quickly Maintain performance and reliability standards for user-facing properties Write code that's readable, reviewable, and built with the next engineer in mind What you should have: At least 5 years of professional software engineering experience Strong expertise in TypeScript and React with a track record of delivering high-performance frontend experiences Practical experience with Python and backend framewor

typescriptpythonreact
View job →
D
Discord
📍 San Francisco Bay Area• Full-time• $196K – $220K/yr
1mo ago

Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. We're looking for a Senior Software Engineer to join the Growth team at Discord. Our team owns how new users discover, understand, and get started with Discord — from the first moment someone encounters us on the web through the onboarding experience that turns them into engaged members. You'll work across the stack to build the systems that acquire and activate users at scale, whether that means improving how Discord shows up across the web or designing the in-product experiences that help new users find their footing. This is a high-impact role where you'll directly influence how Discord grows. This person will report to the Senior Engineering Manager for Growth. What you'll be doing: Build and optimize user-facing experiences that improve how people discover, understand, and get started with Discord Develop scalable backend systems that serve dynamic, high-performance content at scale Work across the full stack — frontend (TypeScript, React) and backend (Python) — wherever the problem leads Design and run A/B experiments to measure the impact of your work on user acquisition and activation Partner with Product, Design, and Data Science to identify high-impact growth opportunities and iterate quickly Maintain performance and reliability standards for user-facing properties Write code that's readable, reviewable, and built with the next engineer in mind What you should have: At least 5 years of professional software engineering experience Strong expertise in TypeScript and React with a track record of delivering high-performance frontend experiences Practical experience with Python and backend framewor

typescriptpythonreact
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The SRE Leadership Team The SRE Leadership Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great infrastructure is invisible—it just works. Our team champions a culture of continuous learning, data-driven decision-making, and blameless incident response. We work at the intersection of product engineering, architecture, and operations to ensure Auth0 remains the trusted authentication platform for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with a focus on scalability, resilience, and empowering engineers to grow as technical leaders. What You'll Be Doing Lead the SRE team's technical direction , translating organizational vision into actionable roadmaps while driving complex, cross-functional initiatives across product and platform teams Operate at scale through hands-on participation in 24/7 on-call rotations (follow-the-sun weekdays, shared weekends), directly troubleshooting and remediating incidents on critical systems Build infrastructure resilience , designing and implementing monitoring, alerting, and automation improvements that reduce toil and elevate operational efficiency Champion reliability best practices , establishing policies and cultural standards that embed observability, resilience, and software engineering rigor into all engineering efforts Mentor and develop SRE talent , elevati

pythonawsazure
View job →
🔔

Get new engineering manager platform reliability salary india jobs by email

Daily job updates · Unsubscribe anytime