At Bolna, we’re building tools that change how businesses leverage voice AI. We’re looking for a Software Engineer to build reliable, scalable systems that power millions of production conversations across languages, industries, and telephony environments. This is a high-impact, high-ownership role where you’ll work on core platform problems across distributed systems, real-time communication, developer infrastructure, and customer-facing products. Our team includes IIT alumni with experience at Bain, Atlassian, Uber, Zomato, and LinkedIn, and is backed by leading investors. Responsibilities Build systems that operate at scale: Design and build backend services that support high-volume, real-time voice AI conversations with strong reliability, performance, and fault tolerance. Own features end to end: Take problems from product requirements and technical design through implementation, testing, deployment, monitoring, and iteration. Improve platform reliability: Build systems that are observable, resilient, and easy to debug. Identify bottlenecks, reduce failure rates, and improve system availability. Work on real-time infrastructure: Solve problems across telephony, streaming audio, webhooks, queues, scheduling, concurrency, and low-latency communication. Build for developers and customers: Improve APIs, SDKs, integrations, dashboards, and internal tools that make the Bolna platform easier to use and operate. Raise the engineering bar: Contribute to technical design reviews, code quality, testing standards, documentation, incident response, and engineering best practices. Required Skills Strong engineering fundamentals: Solid understanding of data structures, algorithms, databases, networking, operating systems, and distributed systems. Backend development experience: 2+ years of experience building and operating production backend systems using Python, Go, Java, Node.js, or a similar language. Production ownership: Experience shipping software to production and own
Jobiba hiring network
Software Engineer Distributed Systems Salary India Jobs
6,428 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current software engineer distributed systems salary india jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team At OpenAI, we’re building safe and beneficial artificial general intelligence. We deploy our models through ChatGPT, our APIs, and other cutting-edge products. Behind the scenes, making these systems fast, reliable, and cost-efficient requires world-class infrastructure. The Caching Infrastructure team is responsible for building a caching layer that powers many critical use cases at OpenAI. We aim to provide a high-availability, multi-tenant cache platform that scales automatically with workload, minimizes tail latency, and supports a diverse range of use cases. We’re looking for an experienced engineer to help design and scale this critical infrastructure. The ideal candidate has deep experience in distributed caching systems (e.g., Redis, Memcached), networking fundamentals, and Kubernetes-based service orchestration. In This Role, You Will: Design, build, and operate OpenAI’s multi-tenant caching platform used across inference, identity, quota, and product experiences. Define the long-term vision and roadmap for caching as a core infra capability, balancing performance, durability, and cost. Collaborate with other infra teams (e.g., networking, observability, databases) and product teams to ensure our caching platform meets their needs. You Might Thrive In This Role If You: Have 5+ years of experience building and scaling distributed systems, with a strong focus on caching, load balancing, or storage systems. Have deep expertise with Redis, Memcached, or similar solutions, including clustering, durability configurations, client-side connection patterns, and performance tuning. Have production experience with Kubernetes, service meshes (e.g., Envoy), and autoscaling systems. Think rigorously about latency, reliability, throughput, and cost in designing platform capabilities. Thrive in a fast-paced environment and enjoy balancing pragmatic engineering with long-term technical excellence. About OpenAI OpenAI is an AI research and deployment company d
About the Role The Engineering Acceleration Delivery / Continuous Deployment team builds and operates the systems that safely ship OpenAI’s infrastructure and product code to production. We own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across OpenAI to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale infrastructure orchestration. In This Role, You Will Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces, and alerts to detect regressions and trigger safe rollbacks. Build systems that support agent-assisted or autonomous deployment workflows using modern AI tooling. Technologies commonly used in this environment include: Kubernetes for large-scale container orchestration and runtime infrastructure Python and FastAPI for internal services Terraform for infrastructure as code GitOps-based deployment workflows (e.g., ArgoCD, Flux, or similar systems) Buildkite for CI orchestration You may be a strong fit if you: Have worked with Kubernetes-based deployment systems at scale Have experience building or operating continuous deployment platforms Are familiar with GitOps tooling such as
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We're hiring Software Engineers for our Data Platform team to build and evolve Snowflake's real-time stream processing and data transformation platform. If you've built or researched high-throughput streaming systems or scalable transformation engines - in industry or graduate work - we want to talk. AS A SOFTWARE ENGINEER, DATA PLATFORM AT SNOWFLAKE, YOU WILL: Design and implement low-latency stream processing and in-flight data transformation systems at global scale Own correctness and performance of streaming execution (watermarks, exactly-once delivery, out-of-order handling) Build transformation primitives and operators that run reliably at cloud-scale throughput Contribute to architectural decisions for our next-generation streaming and transformation platform Write production-quality systems code and collaborate across product and engineering teams OUR IDEAL CANDIDATE WILL HAVE: BS/MS/PhD in Computer Science or related field — graduate research in streaming, distributed systems, or query/transformation engines is a strong differentiator Deep distributed systems fundamentals: fault tolerance, consistency, state management Hands-on or experience with stream processing or transformation systems (Flink, Kafka Streams, Spark Structured Streaming, or similar) Proficiency i
Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform providing APIs for knowledge retrieval, inference, evaluation, and more. We are seeking a strong Senior Full-Stack Engineer to help us build, scale, and refine our rapidly growing product. The ideal candidate is deeply grounded in software engineering best practices and experienced in developing and scaling modern web applications end-to-end. You will work across the stack—from React/TypeScript frontends to Python-based backends—while integrating with LLMs and machine learning systems. You will solve complex challenges in scalability, reliability, and product experience while owning significant product areas in a fast-paced environment. What You’ll Do Own major full-stack product areas , driving features from design through production deployment. Build modern frontend experiences using React and TypeScript, ensuring performance, usability, and responsiveness. Develop reliable backend services in Python, working with distributed systems, data pipelines, and ML/LLM components. Integrate with LLMs, vector databases, and AI infrastructure to power intelligent product experiences. Deliver experiments and new features quickly , maintaining high quality and tight feedback loops with customers. Collaborate across product, ML, and infrastructure teams to shape the direction of Scale GP. Adapt quickly —learning new technologies, frameworks, and tools as needed across the stack. Ideal Experience 5+ years of full-time engineering experience , post-graduation. Strong experience developing full-stack applications using React, TypeScript, and Python . Experience scaling or shipping products at high-growth startups . Familiarity with LLMs, vector databases, embeddings, or other modern AI tooling (tinkering or production experience welcome). Proficiency with SQL and modern API development. Experience with Kubernetes , containerization, and microservice architectures. Experience working with at leas
About the Team The Storage teams build and operate online stateful systems and abstractions that are reliable, efficient, secure and easy to use for DoorDash Engineering. The teams are responsible for understanding Product Engineering’s evolving needs and developing platform and infrastructure capabilities to serve them. The team currently supports CockroachDB, Cassandra, Kafka and Redis as well as data abstraction services to reduce the complexity of interacting with storage systems for Product Engineers. About the Role The Storage team is building and operating a high-performance, scalable, and reliable data abstraction layer that optimizes both efficiency and reliability. Our goal is to create a platform that manages itself and fades into the background—empowering engineers to focus on delivering product experiences our customers love. This role is available across two teams within Storage, each solving unique and high-impact challenges: One team is building the orchestration layer for DoorDash’s storage platform—unifying lifecycle management, operations, and self-serve APIs for databases and streaming systems, turning complex, stateful infrastructure into reliable, developer-friendly services used across the company. One team builds and operates the distributed data platform powering DoorDash's largest stateful workloads -- including Cassandra, which backs critical product surfaces across DoorDash, Wolt, and Roo. You'll design high-throughput data abstractions, smart clients, and platform services that make distributed data reliable and easy to work with at multi-petabyte, multi-million-QPS scale, with opportunities to go deep on distributed systems internals and contribute to the open-source Cassandra ecosystem. If you're passionate about distributed systems, developer experience, and building foundational infrastructure at scale, we'd love to hear from you. You must be located in San Francisco, Sunnyvale, Seattle, or the New York Metro Area for this hybrid pos
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G
Asana’s rapid growth brings new challenges in keeping our systems fast, reliable, and resilient. As our product evolves, we’re making a major investment in reliability – and building a brand new SRE team in Warsaw is a key part of that strategy. This is your chance to help shape it from day one. This isn’t a traditional “ops” role – we’re looking for strong software engineers who are passionate about building reliable, distributed systems. You’ll work closely with a small SRE team in San Francisco, infrastructure engineers in Reykjavik, and an established infrastructure team in Warsaw. Warsaw will be a significant hub for our future infrastructure engineering and operations. As one of the first engineers here, you’ll have a real say in how we build reliable infrastructure, manage incidents, and support the rest of the company. This role is based in our Warsaw office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do, and your recruiter can share more about the in-office requirements. We offer a Contract of Employment (UoP) for our employees in Poland. What you’ll achieve Influence the future of Asana’s SRE practice, especially as we grow the Warsaw team. Lead reliability-focused projects across our stack – from infrastructure to tooling to incident response. Define and implement Asana’s incident management process – we’re investing here, and you’ll help shape how it works. Build internal platforms and frameworks that help other teams improve the reliability of their services. Be part of (and help shape) a sustainable on-call rotation – shared across teams in Warsaw, San Francisco, and Reykjavik. On average, we handle ~1 page per day, but it’s not constant, and we care about keeping things sane. Work with our stack: AWS, Kubernetes (EKS), Datadog, MySQL (RDS), ElasticSearch (OpenSearch), Redis
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G
About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability Team Focus Areas Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely. Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads. Databases: Building high performance, distributed database systems that power all of OpenAI's product stack. In this
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we're building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Anyscale is looking for a Software Engineer to join the ML Developer Experience (MLDevX) team. MLDevX owns the experience layer of the Anyscale platform: the interfaces through which users and coding agents discover, configure, run, observe, debug, and productionize AI workloads. Every user journey crosses this layer through the CLI, SDKs, APIs, UI, Workspaces, MCP, or the workflows and integrations built on top of them. Together, these form the user’s primary interface into Anyscale, turning distributed computing from a systems problem back into a coding problem. We build the common contracts, tools, control-plane services, and architecture that power these surfaces. You will work across the stack from developer tooling to cloud infrastructure and the Ray runtime. Manage long-running operations and make failures across jobs, tasks, actors, nodes, and GPUs easier to diagnose. The systems you build must scale with the platform, remain predictable through failures, and be intuitive for developers, programmable for applications, and operable by coding agents. This is a high impact individual-contributor role with end-to-end ownership. You will work directly with users and field teams to identify high
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: As a Site Reliability Engineer, you will play a crucial role in ensuring the smooth operation of all user-facing services and other Anyscale production systems. Anyscale values diversity and inclusion, and we encourage applications from individuals of all backgrounds. This includes processes for provisioning, negotiating prices, managing costs, seeing opportunities for teams to reduce wastage by finding applications across the company. You will apply sound engineering principles, operational discipline, and mature automation to our environments and the Anyscale codebase as we scale. As part of this role, you will: Develop a unified perspective on how cloud components are utilized across the company, taking into account diverse needs and requirements. Ensure that deployment methodologies align with the company's reliability goals. Build systems that promote understanding of production environments, facilitating quick identification of issues through robust observability infrastructure for metrics, logging, and tracing. Create monitoring and alerting systems at different levels, enabling teams to easily contribute and enhance the overall monitoring capabilities. Establish testing infrastructure to s
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About Ray Data Team: Ray Data is Python-native data processing engine that is a one stop shop for all AI data processing needs. Ray Data provides performant, first-class integration with cutting edge AI frameworks using both multi-modal and structured data. The Ray Data team currently develops and maintains Ray Data . We are a team of engineers passionate about building a Data processing engine which is a one-stop shop for all of your ML/AI needs. We are looking for exceptional engineers to build, optimize, and scale Ray for modern and increasingly complex AI workloads. As part of this role, you will: Improve the performance of Ray Data and multi-modal batch inference use cases. Ensure efficient scaling across different stages of the Data pipeline in a heterogeneous environment. Building data loading solutions for production training workloads. Focus on stability and fault tolerance at high scale Working with customers and new age AI native companies in scaling their AI workloads. We'd love to hear from you if have: At least 3-4 years of relevant work experience Solid background in building scalable and fault-tolerant distributed systems Experience with data processing, database internals. Passionate about large
About Anyscale: At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we're building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Anyscale is looking for a Software Engineer to join the Platform and Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the team, we build the scalable, secure, and robust backbone that enables this vision, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads. We are seeking a talented Software Engineer with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale's cloud platform. You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impa
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About Ray Data Team: Ray Data is Python-native data processing engine that is a one stop shop for all AI data processing needs. Ray Data provides performant, first-class integration with cutting edge AI frameworks using both multi-modal and structured data. The Ray Data team currently develops and maintains Ray Data . We are a team of engineers passionate about building a Data processing engine which is a one-stop shop for all of your ML/AI needs. We are looking for exceptional engineers to build, optimize, and scale Ray for modern and increasingly complex AI workloads. As part of this role, you will: Improve the performance of Ray Data and multi-modal batch inference use cases. Ensure efficient scaling across different stages of the Data pipeline in a heterogeneous environment. Building data loading solutions for production training workloads. Focus on stability and fault tolerance at high scale Working with customers and new age AI native companies in scaling their AI workloads. We'd love to hear from you if have: At least 3-4 years of relevant work experience Solid background in building scalable and fault-tolerant distributed systems Experience with data processing, database internals. Passionate about large
Get new software engineer distributed systems salary india jobs by email
Daily job updates · Unsubscribe anytime