About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
Jobs in India
Ai Infrastructure System Engineer Bangalore in India
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current ai infrastructure system engineer bangalore jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.
Job Title Automation Engineer - C# Job Description Automation Engineer - C# As an Automation engineer, you will ensure that the complete and integrated MR systems meet the requirements as defined in the System Requirements Specifications and that all features are implemented and verified correctly. To improve test efficiency and coverage, selected verification tests and regression test suites are automated using C# .NET. The Test Automation Engineer plays a key role in the development, execution, and maintenance of automated test suites, working in close collaboration with verification engineers and the test automation team located in Bangalore, India. Your Role: 5+ years of proven experience in software development and/or test automation within a complex, high‑tech environment. Effectively communicates with stakeholders, escalates or removes impediments, supports risk management, and drives continuous improvement. Keeps technical knowledge up to date and translates emerging trends (e.g., Model‑Based Testing, AI‑driven testing) into practical applications within a high‑tech environment. Coaches and mentors team members on test automation practices, tools, and processes. Strong expertise in software development, testing, and debugging, with a quality‑first mindset. Expertise in C#. Good understanding of modern test automation trends, frameworks, and best practices. Experience in setting up, evolving, and maintaining test automation infrastructure. Strong quality drive, with attention to robustness, reliability, and maintainability of test solutions. Experience working in global, multicultural teams, collaborating across sites and disciplines. Demonstrates a continuous improvement mindset and leads by example. Strong communication and documen
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. About the Role As a Site Reliability Engineer at Ema, you will own the stability, availability, and operational health of our agentic AI platform across customer environments. You'll work closely with Engineering and DevOps to provision infrastructure, drive deployment excellence, and keep production running at the quality bar our enterprise customers expect — 99.9%+ uptime, proactive incident response, and continuous improvement. What You'll Do Infrastructure & Deployment Design and provision cloud infrastructure (GCP, Azure, AWS) tailored to customer environments, with security, scalability, and compliance built in Execute on-call SaaS deployments with minimal downtime; automate and optimize deployment workflows end-to-end Production Stability & Observability Monitor logs, alerts, and metrics to maintain SLA commitments and catch issues before they escalate Diagnose and resolve production incidents with speed and rigor; drive root cause analysis and permanent fixes Collaborate with DevOps to enhance monitoring dashboards and alerting frameworks; deliver clear system health reporting to internal and customer stakeholders Documentation & Knowledge Management Maintain de
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity We are looking for a Staff Engineer to join our Observability team. This role is ideal for a highly technical engineer who thrives on uncovering the truth behind complex system behaviors, diagnosing difficult production challenges, and driving platform-wide improvements. As a Staff Engineer, you will operate as a force multiplier across engineering teams, helping Postman build world-class observability capabilities while improving reliability, performance, and developer productivity. You will partner closely with Infrastructure, Platform, Product, Security, and Data teams to identify systemic issues, establish operational excellence, and ensure engineering teams have the visibility they need to operate at scale. What You'll Do Drive the technical vision and architecture for Postman's observability platform. Design and build scalable solutions for metrics, logging, tracing, alerting, and operational analytics. Investigate complex production issues, identify root causes, and drive long-term corrective actions. Partner with engineering teams to improve service reliability, availability, performance, and operational mat
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Infrastructure Engineer Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity As a Senior Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with 5–8 years of relevant experience who can leverage open-source and custom tooling to orchestrate and maintain large-scale Kubernetes environments. You will play a "Captain" role—leading critical deliverables and mentoring junior engineers while maintaining the reliability of our global fleet. What You'll Do Architectural Leadership: Drive the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management. Platform Orchestration: Lead complex, large-scale infrastructure shifts. Operational Excellence: Take ownership of incident response, author comprehensive retrospectives, and implement systemic hardening to prevent recurrence using advanced overcommit strategies. This Role Requires Experience: 5–8 years in a DevOps, Site Reliability, or Infrastructure Engineering role. Kubernetes Mastery: Deep internals knowledge of Kubernetes and hands-on experience writing custom operators. Tooling Proficiency: Strong experience building production-grade tools and services, specifically for infrastructure automation. Operations-Heavy Mindset: A proven track record of Day 1/Day 2 operations for a large-scale Kubernetes fleet, handling high-severity incidents, and improving SLA compliance through auto
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies — from the world's largest enterprises to the most ambitious startups — use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the Organization The Core Infrastructure organization operates the foundational systems that power Stripe globally — including databases (MongoDB, PostgreSQL), high availability and disaster recovery (HADR), AWS cloud infrastructure, Linux servers, container orchestration, mesh networking, service discovery, and network edge infrastructure. Within Core Infra, the Regional Enablement Platform (REP) team helps Stripe launch and operate new regions without learning about broken dependencies from users. REP builds the regionalization, validation, deploy-safety, and operator tooling needed to answer practical launch-readiness questions: can critical payment paths run from the new region, which services still depend on a remote control plane, what breaks under packet loss or failover, and what must be fixed before deploys, launches, traffic shifts, or failovers proceed. The team uses traffic replay, synthetics, failover drills, dependency analysis, CI/CD gates, and incident data to turn those findings into platform fixes, service-owner asks, and reusable readiness checks across networking, HADR, and service teams. This role is based in Bangalore and serves as a senior technical anchor for Core Infrastructure in India, with direct cross-region influence across AMER, EU, and APAC. What you'll do As a Staff Engineer on REP, you will play a key leadership role in enabling Stripe's infrastructure to power all of our products, globally and at scale. You will
Research Engineer, Applied AI Location: Bangalore (or throughout India remote-friendly with travel) About EnCharge AI: EnCharge AI is building the next generation AI platform. Our novel in-memory-computing architecture delivers a 10x step-function improvement in compute energy efficiency and performance for AI inference workloads. As the demands of artificial intelligence move beyond today's models, we believe fundamental underlying infrastructure must evolve. We are an experienced team of AI researchers, silicon & systems engineers, and architects backed by leading investors, poised to become the essential platform for the next wave of AI innovation. The Opportunity: Modern AI workloads—from large language models to diffusion-based generators to multimodal systems—represent some of the most compute-intensive frontiers in AI, and some of the most promising applications for our hardware’s energy efficiency advantages. We’re building a vertically integrated AI stack that will showcase the transformative potential of our silicon while delivering real value to customers today. We are seeking a Research Engineer to push the boundaries of AI model capability, quality, and efficiency. You’ll build fine-tuning and post training pipelines, develop rigorous benchmarking frameworks, and work at the intersection of ML research and hardware-aware optimization—ensuring our models run beautifully on our silicon. This is a role for someone who thrives at the boundary between research and engineering. You’ll read papers, implement techniques, and ship production-quality code—all in service of making AI inference faster, cheaper, and better. Key Responsibilities: Algorithmic Acceleration: Research and implement state-of-the-art techniques to accelerate AI inference—quantization, sparsity,
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is now at one of the most consequential inflection points in its history. Enterprises are rapidly shifting from human-driven API workflows to multi-agent systems — where AI agents autonomously discover, call, and collaborate with APIs, tools, CLIs, and each other. Postman already owns the API design, governance, and lifecycle layer for the enterprise; the next frontier is owning the runtime layer for how agents actually interact with all of it . Where Postman today is the API platform for human-first integration , we are building Postman into the agent interface fabric for AI-first integration . That transformation starts with Fabric Gateway — a brand new product, being built from scratch, that will serve as the policy-driven control plane governing how agents connect to APIs, services, and other agents at scale. This is a rare 0-to-1 opportunity inside a company with 40+ million developers already in the ecosystem. About the Team The Fabric Gateway team is building the API gateway for the AI era — one of Postman's most ambitious greenfield infrastructure products. We're a small, high-ownership team wo
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Platform Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning/sharding strat
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta is The World’s Identity Company. We enable everyone to safely use any technology anywhere, on any device or application. Our Workforce and Customer Identity Clouds provide secure, flexible access, authentication, and automation, transforming how people navigate the digital world and placing Identity at the core of business security and growth. At Okta, we are committed to celebrating a variety of perspectives and experiences. We are looking for lifelong learners and individuals whose unique experiences will make our organization better. The Technology Data and Intelligence (TDI) team at Okta is focused on boosting internal efficiency through the implementation of secure, scalable, and innovative systems. About the Role The team is seeking a talented Senior Software Engineer to become part of our TDI team in Bangalore, who can effectively balance technical excellence with a disciplined approach to the software development lifecycle. In this role, you will be responsible for designing and developing customizations, extensions, configurations, and integrations required to meet the company’s strategic business objectives. Candidates will work collaboratively with business stakeholders, business analysts, and engineers on different infrastructure layers, from proposal development to deployment and support. Therefore, a commitment to collaborative problem-solving and delivering high-quality solutions is essential. In addition, your product owner
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know Okta Team Okta is The World’s Identity Company. We enable everyone to safely use any technology anywhere, on any device or application. Our Workforce and Customer Identity Clouds provide secure, flexible access, authentication, and automation, transforming how people navigate the digital world and placing Identity at the core of business security and growth. At Okta, we are committed to celebrating a variety of perspectives and experiences. We are looking for lifelong learners and individuals whose unique experiences will make our organization better. The Technology Data and Intelligence (TDI) team at Okta is focused on boosting internal efficiency through the implementation of secure, scalable, and innovative systems. About the Role The team is seeking a talented Senior Software Engineer to become part of our TDI team in Bangalore, who can effectively balance technical excellence with a disciplined approach to the software development lifecycle. In this role, you will be responsible for designing and developing customizations, extensions, configurations, and integrations required to meet the company’s strategic business objectives. Candidates will work collaboratively with business stakeholders, business analysts, and engineers on different infrastructure layers, from proposal development to deployment and support. Therefore, a commitment to collaborative problem-solving and delivering high-quality solutions is essential. In a
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Overview We are looking for a strong software and platform engineer to join our Production Engineering team in Bangalore as an individual contributor in FA ProductionEng APJ. This role will help build and operate internal platforms that improve how we provision, observe, govern, and troubleshoot engineering infrastructure at scale. The fleet management use cases that give teams a single place to understand and operate the test infrastructure. If you enjoy building internal platforms that remove friction, improve visibility, and make engineering teams faster and more effective, this role is for you. Why This Role Is Unique This is not a typical application development role.You will work on internal platforms that directly shape how engineering teams consume and manage shared infrastructure. The role spans platform engineering, workflow automation, observability, API-driven services, and infrastructure lifecycle management. The right candidate will work on systems such as: Self-serviceability workflows and lease-based testbed governance. Developer Platform dashboards and APIs used for triage, visibility, and product trend observation. Testbed and workflow orchestration across fleet management domains. Impact This role is a high-leverage engineering investment. The work will improve how engineering teams provision testbeds, understand failures, operate shared infrastructure, and move faster with less friction. Better
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Search Team at Postman is responsible for enabling users to quickly find and get started with the APIs that they are looking for. Postman is growing at a rapid pace, and this manifests into an ever-increasing volume of data that users create and consume, within their teams and in the Public API Network. We focus on improving discovery and ease of consumption over this data. We are looking for a Senior Engineer with 6+ years of experience deep backend expertise on search and ETL systems and a strong product mindset, to lead core initiatives on our search platform. In this role, you'll work at the intersection of infrastructure, relevance, and developer experience—designing systems that power search across the platform. You’ll bring a bias for action, a strong backend foundation, and the curiosity to explore beyond traditional boundaries, including areas like high performance web services, high volume data pipelines, machine learning, and relevance tuning. What You'll Do Own end to end architecture and roadmap of search platform consisting of distributed indexing pipelines, storage infra and high performance web serve
TEGNA Inc. helps people thrive in their local communities by providing the trusted local news and services that matter most. With 64 television stations in 51 U.S. markets, TEGNA reaches more than 100 million people monthly across web, mobile apps, streaming, and linear television, while also maintaining a strong global presence in India with offices in Bangalore and Chennai that support technology, product, and business operations initiatives. Together, we are building a sustainable future for local news. Hybrid Cloud Engineer Position Overview TEGNA is looking for a Hybrid Cloud Engineer to join the dynamic engineering team. The TEGNA IT Infrastructure team represents a wide range of technologies and competencies. Working in tandem with peers, other teams, departments, business units and outside vendors, the team assists in the evaluation, implementation, and support of technology solutions toward business objectives. The analyst in this position will contribute to the design, implementation, and support of on premise and cloud components to core business applications and will participate in migrations to multiple public cloud platforms. As part of a team, the candidate will build, maintain, support technical systems and applications for business at data centres, cloud, and television stations. This requires comprehensive knowledge of server, storage, cloud, and networking components required to maintain the business. What You’ll Do Design, implement and maintain server and storage components for on-premises and public cloud. Working with technical and non-technical internal customers, provide in-depth consultation and communication for technology solutions. Lead the design, building, testing, documentation and implementation of systems management practices associated with AWS
Other cities to consider
More places hiring for this role
Get new ai infrastructure system engineer bangalore jobs in India by email
Daily job updates · Unsubscribe anytime