Jobiba hiring network

Ai Infrastructure System Engineer Bangalore Jobs

15 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current ai infrastructure system engineer bangalore jobs. Use filters to narrow by work mode, employment type, experience and date posted.

TA
Together AI
📍 BengaluruFull-time
3 days ago

About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw

pythonkuberneteslinux
View job →
P
1mo ago

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity At Postman, we are revolutionizing the way developers build, trace, and automate API workflows with Postman Flows , a powerful visual programming tool designed to simplify the development and sharing of API-powered applications. With an intuitive drag-and-drop interface, Postman Flows enables teams to collaborate and showcase their APIs regardless of technical expertise. We are looking for a Software and Systems Engineer to help scale and maintain the Flows runtime system. This system runs mission-critical automations in the cloud, with a focus on low latency, high throughput, and high availability. You’ll play a vital role in developing, deploying, and operating our backend services and infrastructure in a Kubernetes-based cloud environment. We’re looking for an experienced engineer who is excited not only about hands-on building as we ship and iterate on a weekly basis to get our product ready for GA, but who can also serve as a role model and mentor to other engineers. This role involves making key technical decisions and improvements to the system, as well as effectively making impact through influence wit

node.jsawsazure
View job →
P
Postman
📍 New YorkFull-time
1mo ago

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. About Fern, a Postman Company Fern helps software companies build a world-class API experience. Our customers include industry leaders like Nvidia, Square, and Twilio, as well as fast-growing AI companies like ElevenLabs and OpenRouter. In the next year, the majority of API integrations will be implemented by AI agents. Agents don’t read marketing pages. They need strongly typed schemas, structured endpoints, deterministic contracts, and documentation that can be fed into a context window. Our team is small and talent-dense. We’re a team of builders from Google, Palantir, Amazon, and Uber, working together in New York City. About the Role As a Staff Software Engineer at Fern, you’ll build APIs, scale AI infrastructure, and design developer experiences that reach millions of people. Set technical direction: Drive system design decisions around data modeling, performance, correctness, and reliability. Identify scaling bottlenecks before they become problems. Build for durability and scale: Ensure our systems are performant, observable, secure, and resilient under real-world production load across infrastructur

typescriptawsgit
View job →

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and performance of Postman’s AI-powered API and agentic systems in production. This role focuses on monitoring, availability, incident response, and automation to support AI services and tools trusted by millions of developers globally. What You’ll Do Develop and manage reliability metrics (SLOs) for AI-driven API services and agentic AI platform features Implement comprehensive observability and monitoring systems for real-time performance and fault detection Design and drive automated failover, recovery, and incident response strategies for high-availability AI infrastructure Optimize resource utilization, particularly GPU/accelerator efficiency, ensuring cost-effective AI system operation Collaborate closely with engineering, platform, and product teams to align reliability efforts with broader organizational goals Lead efforts to build internal tooling and automation focused on AI system stability and operational excellence Drive continuo

aigorust
View job →
S
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. When was the last occasion you had the opportunity to contribute to a company that is shaping an industry and empowering individuals to translate ideas into tangible impact with speed? Smartsheet's core mission is to empower everyone to enhance their work processes. Our business model is founded on identifying exceptional talent and providing them with the autonomy to develop our acclaimed Software as a Service (SaaS) offering. With a user base exceeding 10 million, our platform is utilized across various industries, including construction, retail, and software development, presenting us with unique technical challenges. Smartsheet is seeking a Senior Business DevOps Engineer to join our Corporate Systems Development team in Bangalore. This role will focus on building and scaling our CI/CD pipelines, infrastructure automation, monitoring frameworks, and deployment processes supporting mission-critical integrations across Finance, People, Sales, Legal, IT, and Engineering systems. You’ll work across a variety of systems and platforms (AWS, GitLab, DataDog, Terraform, Boomi, UiPath) to streamline deployment of backend integrations and automation solutions. If you thrive on optimizing developer velocity, ensuring system reliability, and automating everything from build to deploy, this role is for you. The position reports to the Senior Manager, Systems Development and collaborates closely with global developers, architects, and application administrators to ensure our platform foundations are secure, efficient, and scalable

javascripttypescriptpython
View job →
P
28 days ago

Job Title Automation Engineer - C# Job Description Automation Engineer - C# As an Automation engineer, you will ensure that the complete and integrated MR systems meet the requirements as defined in the System Requirements Specifications and that all features are implemented and verified correctly. To improve test efficiency and coverage, selected verification tests and regression test suites are automated using C# .NET. The Test Automation Engineer plays a key role in the development, execution, and maintenance of automated test suites, working in close collaboration with verification engineers and the test automation team located in Bangalore, India. Your Role: 5+ years of proven experience in software development and/or test automation within a complex, high‑tech environment. Effectively communicates with stakeholders, escalates or removes impediments, supports risk management, and drives continuous improvement. Keeps technical knowledge up to date and translates emerging trends (e.g., Model‑Based Testing, AI‑driven testing) into practical applications within a high‑tech environment. Coaches and mentors team members on test automation practices, tools, and processes. Strong expertise in software development, testing, and debugging, with a quality‑first mindset. Expertise in C#. Good understanding of modern test automation trends, frameworks, and best practices. Experience in setting up, evolving, and maintaining test automation infrastructure. Strong quality drive, with attention to robustness, reliability, and maintainability of test solutions. Experience working in global, multicultural teams, collaborating across sites and disciplines. Demonstrates a continuous improvement mindset and leads by example. Strong communication and documen

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. About the Role As a Site Reliability Engineer at Ema, you will own the stability, availability, and operational health of our agentic AI platform across customer environments. You'll work closely with Engineering and DevOps to provision infrastructure, drive deployment excellence, and keep production running at the quality bar our enterprise customers expect — 99.9%+ uptime, proactive incident response, and continuous improvement. What You'll Do Infrastructure & Deployment Design and provision cloud infrastructure (GCP, Azure, AWS) tailored to customer environments, with security, scalability, and compliance built in Execute on-call SaaS deployments with minimal downtime; automate and optimize deployment workflows end-to-end Production Stability & Observability Monitor logs, alerts, and metrics to maintain SLA commitments and catch issues before they escalate Diagnose and resolve production incidents with speed and rigor; drive root cause analysis and permanent fixes Collaborate with DevOps to enhance monitoring dashboards and alerting frameworks; deliver clear system health reporting to internal and customer stakeholders Documentation & Knowledge Management Maintain de

awsazuregcp
View job →

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity We are looking for a Staff Engineer to join our Observability team. This role is ideal for a highly technical engineer who thrives on uncovering the truth behind complex system behaviors, diagnosing difficult production challenges, and driving platform-wide improvements. As a Staff Engineer, you will operate as a force multiplier across engineering teams, helping Postman build world-class observability capabilities while improving reliability, performance, and developer productivity. You will partner closely with Infrastructure, Platform, Product, Security, and Data teams to identify systemic issues, establish operational excellence, and ensure engineering teams have the visibility they need to operate at scale. What You'll Do Drive the technical vision and architecture for Postman's observability platform. Design and build scalable solutions for metrics, logging, tracing, alerting, and operational analytics. Investigate complex production issues, identify root causes, and drive long-term corrective actions. Partner with engineering teams to improve service reliability, availability, performance, and operational mat

pythonjavanode.js
View job →
P
Postman
📍 New YorkFull-time
1mo ago

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. About Fern, a Postman Company Fern helps software companies build a world-class API experience. Our customers include industry leaders like Nvidia, Square, and Twilio, as well as fast-growing AI companies like ElevenLabs and OpenRouter. In the next year, the majority of API integrations will be implemented by AI agents. Agents don’t read marketing pages. They need strongly typed schemas, structured endpoints, deterministic contracts, and documentation that can be fed into a context window. Our team is small and talent-dense. We’re a team of builders from Google, Palantir, Amazon, and Uber, working together in New York City. About the Role As a Senior Software Engineer at Fern, you’ll build APIs, scale AI infrastructure, and design developer experiences that reach millions of people. Scale infrastructure to keep up with growth: Work on AI systems deployed on Vercel, AWS, and Turbopuffer that must be performant, reliable, and secure under real-world load. Stay close to customers: We work directly with API teams at companies like Nvidia, Twilio, and Square. You’ll debug real production edge cases, shape APIs ba

typescriptawsgit
View job →
S
Smartsheet
📍 Bangalore, INDIAFull-timeHybrid
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Automation is the key to creating highly reliable and secure large-scale software systems. Are you someone who engineers solutions to problems rather than simply fixing the same thing over and over again? Can you protect Smartsheet against attackers? We are looking for a Senior DevSecOps Engineer to join our global Security Operations team. In this critical role, you will be a leader in maturing our security and reliability posture by treating both as software engineering challenges. You will engineer and operate a highly reliable, scalable, and defensible production environment, directly impacting our ability to deliver a world-class service to our customers 24/7. This is a unique opportunity to blend deep expertise in Site Reliability Engineering (SRE) and modern Security Operations, working at the intersection of infrastructure, automation, and security to build a platform that is resilient and secure by design. You Will: Engineer Secure and Resilient Infrastructure: Design, build, maintain, and improve secure, scalable, and highly available infrastructure in our multi-cloud environment (primarily AWS) using Infrastructure as Code (IaC) principles with tools like Terraform, Kubernetes, and Helm. Automate Proactive Security: Engineer an

pythonawskubernetes
View job →

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Infrastructure Engineer Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning

pythonsqlaws
View job →

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity As a Senior Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with 5–8 years of relevant experience who can leverage open-source and custom tooling to orchestrate and maintain large-scale Kubernetes environments. You will play a "Captain" role—leading critical deliverables and mentoring junior engineers while maintaining the reliability of our global fleet. What You'll Do Architectural Leadership: Drive the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management. Platform Orchestration: Lead complex, large-scale infrastructure shifts. Operational Excellence: Take ownership of incident response, author comprehensive retrospectives, and implement systemic hardening to prevent recurrence using advanced overcommit strategies. This Role Requires Experience: 5–8 years in a DevOps, Site Reliability, or Infrastructure Engineering role. Kubernetes Mastery: Deep internals knowledge of Kubernetes and hands-on experience writing custom operators. Tooling Proficiency: Strong experience building production-grade tools and services, specifically for infrastructure automation. Operations-Heavy Mindset: A proven track record of Day 1/Day 2 operations for a large-scale Kubernetes fleet, handling high-severity incidents, and improving SLA compliance through auto

awsazurekubernetes
View job →
S
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies — from the world's largest enterprises to the most ambitious startups — use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the Organization The Core Infrastructure organization operates the foundational systems that power Stripe globally — including databases (MongoDB, PostgreSQL), high availability and disaster recovery (HADR), AWS cloud infrastructure, Linux servers, container orchestration, mesh networking, service discovery, and network edge infrastructure. Within Core Infra, the Regional Enablement Platform (REP) team helps Stripe launch and operate new regions without learning about broken dependencies from users. REP builds the regionalization, validation, deploy-safety, and operator tooling needed to answer practical launch-readiness questions: can critical payment paths run from the new region, which services still depend on a remote control plane, what breaks under packet loss or failover, and what must be fixed before deploys, launches, traffic shifts, or failovers proceed. The team uses traffic replay, synthetics, failover drills, dependency analysis, CI/CD gates, and incident data to turn those findings into platform fixes, service-owner asks, and reusable readiness checks across networking, HADR, and service teams. This role is based in Bangalore and serves as a senior technical anchor for Core Infrastructure in India, with direct cross-region influence across AMER, EU, and APAC. What you'll do As a Staff Engineer on REP, you will play a key leadership role in enabling Stripe's infrastructure to power all of our products, globally and at scale. You will

sqlpostgresqlmongodb
View job →
EA
4 days ago

Research Engineer, Applied AI Location: Bangalore (or throughout India remote-friendly with travel) About EnCharge AI: EnCharge AI is building the next generation AI platform. Our novel in-memory-computing architecture delivers a 10x step-function improvement in compute energy efficiency and performance for AI inference workloads. As the demands of artificial intelligence move beyond today's models, we believe fundamental underlying infrastructure must evolve. We are an experienced team of AI researchers, silicon & systems engineers, and architects backed by leading investors, poised to become the essential platform for the next wave of AI innovation. The Opportunity: Modern AI workloads—from large language models to diffusion-based generators to multimodal systems—represent some of the most compute-intensive frontiers in AI, and some of the most promising applications for our hardware’s energy efficiency advantages. We’re building a vertically integrated AI stack that will showcase the transformative potential of our silicon while delivering real value to customers today. We are seeking a Research Engineer to push the boundaries of AI model capability, quality, and efficiency. You’ll build fine-tuning and post training pipelines, develop rigorous benchmarking frameworks, and work at the intersection of ML research and hardware-aware optimization—ensuring our models run beautifully on our silicon. This is a role for someone who thrives at the boundary between research and engineering. You’ll read papers, implement techniques, and ship production-quality code—all in service of making AI inference faster, cheaper, and better. Key Responsibilities: Algorithmic Acceleration: Research and implement state-of-the-art techniques to accelerate AI inference—quantization, sparsity,

pythonaigo
View job →

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet is hiring a Senior Machine Learning Operations Engineer to architect our machine learning production lifecycle. Your mission is to maintain and deploy ML models to a scalable, reliable, and secure production environment. You will design and maintain the infrastructure, automation, and monitoring systems that ensure our AI products are high-performing and cost-effective. You will report to our Director, Analytics Engineering & Data Governance and work from our Bangalore, India office. You Will: Model and Pipeline Automation Automate the deployment and retraining of ML models, from training through to production inference, by building and managing complete CI/CD/CT (Continuous Training) pipelines, adhering to MLOps best practices. Build, fine-tune, or use pre-trained LLMs, deep learning models or traditional machine learning models. Evaluate and recommend AI or ML solutions for the product using any combination of vendor solutions and/or custom-built models. Governance & Compliance Implement model versioning, lineage tracking, and auditing to ensure compliance with security and ethical standards. Performance Monitoring Continuously monitor the health and performance of production machine learning models, proactively identifying and correcting model drift, staleness, and performance degradation. Incorporate user feedback for iterative improvements and manage necessary model retraining cycles. Cross-Functional Collaboration Act as the "glue" between Data Scientists (who build models

pythonawsazure
View job →
🔔

Get new ai infrastructure system engineer bangalore jobs by email

Daily job updates · Unsubscribe anytime