About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
Jobiba hiring network
Software Engineer Distributed Systems Salary India Jobs
6,428 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current software engineer distributed systems salary india jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Role Overview You’ll be the Principal Software Engineer driving the next generation of a large-scale enterprise SaaS platform. In this role, you combine deep hands-on engineering with high-impact technical leadership, shaping how cloud-native and AI-enabled products are designed and built. You’ll design and deliver secure, scalable, serverless systems on AWS using TypeScript and Node.js, modernize critical platform components, and set the technical direction for multiple teams. You’ll also lead how AI capabilities are integrated across the product ecosystem, ensuring they are transparent, observable, and compliant. If you enjoy system-level thinking, complex distributed architectures, and mentoring senior engineers while still staying close to the code, this role gives you company-wide impact and the opportunity to define the long-term technical vision. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture and delivery of secure, scalable, serverless applications on AWS using TypeScript/Node.js. Define and evolve the platform architecture, driving modernization, performance, resilience, and maintainability. Design and operate distributed, event-driven systems using services like Lambda, DynamoDB, Aurora, S3, and EventBridge. Shape and implement AI-enabled solutions, embedding governance, observability, and responsible AI practices into the platform. Own Infrastructure as Code (e.g., Terraform, AWS CDK, CloudFormation) to reliably provision and manage cloud infrastructure. Mentor senior engineers, influence technical decisions across teams, and clearly communicate complex concepts to diverse stakeholders. These are the essentials you’ll need to get an interview Extensive experience (typically 12+ years) building secure, production-grade software systems. Proven track record architecting and delivering cloud-native, serverless applications on AWS. Strong expertise in Node.js, TypeScript, REST API design, and at leas
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? The Data Infrastructure team at Cohere is responsible for the storage and data movement layer underlying every model training run. We're building the unified storage layer that feeds our training workloads. It needs to serve petabytes of training data and model checkpoints fast enough to keep thousands of GPUs busy across several training clusters. In this role, you’d have an opportunity to build this system from the ground up. You’d be a key contributor, working on a problem few teams have had to solve at this scale. In this role, you will: Design, build, and operate the distributed storage system that feeds model training and evaluation. Run this system multiple on Kubernetes clusters at petabyte scale. Work with researchers and training-infra teams on how jobs actually read and write data, and turn that into throughput, latency, and durability requirements Work through the networking, I/O, and consistency problems of moving large datasets and checkpoints across regions and backends, with GPU idle time and time-to-insight as the measures of success You may be a good fit if you have: Strong storage fundamentals,
The MongoDB Atlas team is a diverse group of contributors working together to help our users manage MongoDB at global scale. We are responsible for MongoDB Atlas: our database as a service offering and fastest growing product which allows users to deploy fault-tolerant, globally distributed MongoDB clusters in just minutes. We're seeking a Senior Engineer to join the Atlas Identity and Access Management (IAM) team. IAM is a platform and a product team. We serve internal engineers by providing them a secure and durable suite of services, and we serve external customers by providing them user facing features and products. We are the owners of Atlas’ authentication (OAuth, SSO, Federated Identity) and authorization (RBAC, ABAC) systems, along with many others. The IAM team’s mission is to enable customers to securely build their applications with Atlas through our best in class user experience. We are looking to speak to candidates who are based in New York City, NY for our hybrid working model. Role Responsibilities Design, architect, build, and deliver core pieces of IAM Lead projects from specification to delivery Mentor and grow other team members Improve our codebase, best practices, and design principles Define your top priorities and focuses, communicate them, and execute against them Lead and contribute to complex technical projects and initiatives Candidate Profile 5+ years experience of software engineering, primarily focused on backend systems Proficient in a modern compiled programming language (Java, Go, C#, C++, etc.) Willingness to learn JavaScript and/or TypeScript along with modern frontend technologies (React, Redux, etc.); prior experience a plus Excellent communication skills, both written and verbal Desire to collaborate with colleagues and mentor fellow engineers Is curious, collaborative, empathetic, and intellectually honest Has a passion for problem solving and learning new things in the domains of computer science and software engineering Expe
About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! This role operates on a hybrid schedule requiring two days of in-office collaboration per week. The position must be based in the San Francisco Bay Area. About the Role We're hiring a Software Engineer II within our Fulfillment organization — the backend systems that get the right job to the right Tasker and see it through to completion. You'll join Fulfillment Lifecycle, the team that decides how jobs are matched to Taskers for our partner and marketplace business, increasingly using unstructured data and experimentation to make matching smarter and fulfillment more reliable. The team is part of a company-wide platform modernization effort, breaking a legacy monolith into well-bounded, API-first services. We're hiring for a strong backend engineer who thrives on complex, data-intensive problems, is comfortable with ambiguity, and takes pride in well-tested, observable, production-ready code. What You'll Work On B
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. The Infrastructure product organization develops New Relic infrastructure instrumentation agents, next generation data processing and management services, vulnerability management, and security testing capabilities for on-prem and cloud customers. We work with data at a scale using a diverse tech stack (Go, Java, JavaScript, React GraphQL, Kubernetes, many public cloud web services, and more). As a senior backend engineer, you will help us build and extend next generation solutions such as a control plane for customers to manage their data pipelines at scale. New Relic is looking for engineers who are interested in building a brand-new observability experience. This high-impact engineering position is a phenomenal opportunity to own and build a set of next generation services and capabilities for the company. We are searching for a motivated engineer who is ready for a career-defining role in their next opportunity. We look forward to talking with you! What you'll do ● Design, Build, maintain, and scale back-end services and their support tools. ● Participate in architectural definitions with a high degr
About Anyscale At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure. As part of this role, you will Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices We'd love to hear from you if you have Familiarity with running ML inference at large scale with high throughput and low latency Familiarity with deep learning and deep learning frameworks (e.g. PyTorch) Solid understanding of distributed systems, ML inference challenges Bonus points
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We're building the data infrastructure behind some of the most demanding AI training workloads in the world, and we want sharp, curious people to help us do it. In this role, you'll build and maintain the high-performance data layer our Modeling teams rely on for training and evaluation jobs. As a Software Engineer, Data Infrastructure, you will: Work directly on petabyte-scale storage infrastructure, and the networking and performance challenges that come with it. Collaborate daily with researchers and engineers who are some of the best in the world at what they do. You may be a good fit if you have: 4+ years of experience working on data storage infrastructure Strong command of Python Kubernetes experience, especially on the storage side (Persistent Volumes, CSI drivers, etc.) The ability to transform unstructured data into performant datasets across diverse storage backends including S3, GCS, and POSIX Experience with distributed data processing frameworks such as Apache Beam, Spark, or Flink [Nice-to-have] Familiarity with modern analytics tooling such as BigQuery, Airflow, or dbt Genuine excitement about AI.
About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale with trillions of data points per day, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams for tens of thousands of companies globally. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team The Datadog Security Libraries team owns the customer-side integrations behind our run-time security products App & API Protection , Workload Protection , and Code Security . Our libraries let customers automatically manage application security risk with continuous, real-time monitoring of vulnerabilities and threats against their web applications, serverless applications, and APIs, in production. Automatically integrated with Application Performance Monitoring (APM) distributed tracing and code-level context, our software empowers development, operations, and security teams to build and run secure applications. As a polyglot team we ship and maintain the security capabilities of Datadog's tracing libraries across .NET , Java , Go , Node.js , Python , Ruby , and PHP , on top of a shared C++ core and a set of HTTP proxy integrations (primarily Envoy, NGINX, and HAProxy). Our code runs inside thousands of production applications around the world. Recent work spans exploit prevention (RASP) and WAF detections, API Security, code security (IAST and SCA), and AI-assisted ("agentic") onboarding, always measured by real product outcomes and operational telemetry. The Opportunity We're looking for a senior, polyglot engineer to contribute across several of our security libraries, with .NET or Java expertise. You'll design and build security integrations and detection features, take them from prototype to production-hardened, and own them operationally as they instrument thousands of applications. As a se
About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale with trillions of data points per day, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams for tens of thousands of companies globally. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team The Datadog Security Libraries team owns the customer-side integrations behind our run-time security products App & API Protection , Workload Protection , and Code Security . Our libraries let customers automatically manage application security risk with continuous, real-time monitoring of vulnerabilities and threats against their web applications, serverless applications, and APIs, in production. Automatically integrated with Application Performance Monitoring (APM) distributed tracing and code-level context, our software empowers development, operations, and security teams to build and run secure applications. As a polyglot team we ship and maintain the security capabilities of Datadog's tracing libraries across .NET , Java , Go , Node.js , Python , Ruby , and PHP , on top of a shared C++ core and a set of HTTP proxy integrations (primarily Envoy, NGINX, and HAProxy). Our code runs inside thousands of production applications around the world. Recent work spans exploit prevention (RASP) and WAF detections, API Security, code security (IAST and SCA), and AI-assisted ("agentic") onboarding, always measured by real product outcomes and operational telemetry. The Opportunity We're looking for a senior, polyglot engineer to contribute across several of our security libraries, with .NET or Java expertise. You'll design and build security integrations and detection features, take them from prototype to production-hardened, and own them operationally as they instrument thousands of applications. As a se
MongoDB’s mission is to empower innovators to create, transform, and disrupt industries by unleashing the power of software and data. We enable organizations of all sizes to easily build, scale, and run modern applications by helping them modernize legacy workloads, embrace innovation, and unleash AI. Our industry-leading developer data platform, MongoDB Atlas, is the only globally distributed, multi-cloud database and is available in more than 115 regions across AWS, Google Cloud, and Microsoft Azure. Atlas allows customers to build and run applications anywhere—on premises, or across cloud providers. With offices worldwide and over 175,000 new developers signing up to use MongoDB every month, it’s no wonder that leading organizations, like Samsung and Toyota, trust MongoDB to build next-generation, AI-powered applications. MongoDB is seeking a Software Engineer 3 to join the Atlas Clusters Organization. The organization is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. This includes developing software to interface with the three major cloud providers (AWS, Azure, and GCP) in order to bring security, durability, availability, and performance to all deployments of MongoDB. The Atlas Clusters Security team creates a first-in-class cloud database security experience for our wide range of sophisticated customers. Our team develops and maintains systems for cluster networking, data encryption, database authentication, and more–enabling countless mission critical applications across the world. We are looking to speak to candidates who are based in New York for our hybrid working model. What you’ll do Build and design new features for MongoDB Atlas Contribute to and lead complex technical projects Work closely with product and design teams, considering the user’s perspective while building technical solutions
The MongoDB Customer Observability Team is a diverse group of contributors working together to help our users manage MongoDB at global scale. The team is responsible for MongoDB Atlas: our database-as-a-service offering and fastest-growing product, which allows users to deploy globally distributed MongoDB clusters in just minutes. We're seeking a Senior Software Engineer to join our team to tackle exciting challenges within the Observability space. You'll contribute to developing tools and platforms that help our users understand the health and performance of their MongoDB deployments. This includes collecting metrics, monitoring slow queries, and offering actionable insights such as index and schema suggestions that improve the speed, efficiency, and overall reliability of their databases. This role provides a unique opportunity to drive engineering excellence across both dimensions of observability, leveraging technologies being developed within the Customer Observability group and contributing directly to the success of our customers and our product teams. If you're passionate about large-scale systems, digging deep into telemetry data, and building tools that make a real impact both internally and externally, we’d love to have you on board! We are looking to speak to candidates who are based in Dublin for our hybrid working model. We're looking for someone who Has at least 5 years of experience as a backend or full stack engineer Enjoys collaboration and being part of a team Is approachable, curious, and intellectually honest Is a backend engineer with a willingness to take on frontend tasks or a full-stack developer with a bias towards backend Has written backend systems in a compiled language (Java, C#, Go, etc.) Has experience with the design and architecture of a modern, scalable web application Enjoys chasing down difficult problems in a distributed environment and on an database diagnostic/operation level Always strives to expand their knowledg
The MongoDB Query Execution Team is hiring software engineers who want to join us in developing a high performing, reliable and modular distributed query system. Our engineers work on implementing and maintaining execution algorithms, building new query language features, tuning database performance, and more to power our customers' critical workloads. This role can be based out of our Dublin office or remotely in Ireland. Relocation can be supported. Position Expectations Understand and improve current functionality of the MongoDB query engine Contribute high quality C++ code and give and solicit feedback in code reviews Identify, design, implement, test, and support new features related to query performance and robustness, query language enhancements, diagnostics for query performance problems, and integration with other products and tools Work constructively with peers to deliver excellent technical solutions Candidate Profile 5+ years of experience in systems programming Experience in databases and/or data management systems is a huge plus, but not a requirement Hands-on experience building industrial-strength software Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Experience with large code bases, preferably in C++, C, Rust or a similar compiled language B.Sc in Computer Science or similar field, or equivalent practical experience Interest in the theory and practice of database query engines. Hands-on experience or M.Sc./Ph.D in the domain is a plus Success Measures In three months you’ll have contributed to the development of a project slated for the next major version, as well as fixed a few bugs in a minor version of our latest stable release series In six months, you’ll have taken on code review responsibilities and are independently delivering complex functionality and squashing bugs independently In twelve months, you’re leading the development of a new major feature and are h
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! The Build Systems team within Figma’s Developer Experience organization owns Figma’s build and CI infrastructure, enabling engineers to ship changes to production quickly and safely. We build and operate core platforms across our polyglot monorepo, including build systems, artifact repositories, merge queues, test frameworks, and CI pipelines. We’re looking for an experienced technical leader to help shape these platforms, uplevel the team, and deliver high-impact platforms that accelerate engineering velocity. The ideal candidate has deep experience with large monolithic codebases, builds durable and scalable systems, and is motivated by solving high-leverage problems that amplify productivity across the engineering organization. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: Drive technical roadmap and strategy for the Build Systems team Partner with cross-functional teams and leadership to identify developer pain points and design elegant, scalable solutions Lead complex, multi-quarter initiatives reducing build/test times and improving CI reliability, all while balancing technical excellence with pragmatic delivery Design, build, and maintain modern developer tools including scalable build systems, distributed CI platforms, and test frameworks that serve thousands of engineers Architect and implement large-scale infrastructure on AWS that powers our entire build pipeline to ensure reliability, performance, and cost
Ignite your curiosity. Solve the unsolvable. At Leidos, we do more than write code—we decode the unknown. Our San Diego-based research and engineering team takes on some of the nation’s toughest defense challenges using advanced signal processing, ocean remote sensing, and high-performance computing. We’re seeking a Software Engineer / Computer Scientist who enjoys solving complex problems and pushing the limits of performance. In this role, you’ll work alongside a multidisciplinary team of scientists and engineers with expertise in hydrodynamics, physics, acoustics, and signal processing to build impactful software that turns massive, complex data sets into meaningful insight. If you are motivated by innovation, energized by collaboration, and excited to see your work support real-world missions, this could be the right opportunity for you. What You’ll Do Collaborate with scientists and engineers to design, develop, and optimize advanced algorithms for next-generation radar, optical, and infrared sensor systems. Build scalable, high-performance backend systems for scientific computing in distributed environments. Integrate, refactor, and improve scientific codebases to increase efficiency and scalability. Translate and optimize existing code for GPU/CUDA acceleration and parallel or distributed execution. Test, document, maintain, and enhance complex software in Linux/Unix environments. Contribute in a collaborative environment that values technical excellence, creativity, and continuous growth. Required Qualifications Bachelor’s degree in Computer Science, Applied Mathematics, Physics, or a related field with 4+ years of backend software development experience, or a Master’s degree with 2+ years of experience. Equivalent experience may be considered in place of a degree. U.S. citizenship and the ability to obtain a Top Secret clearance; active Top Secret clea
Get new software engineer distributed systems salary india jobs by email
Daily job updates · Unsubscribe anytime