Datadog's Software Engineers with Systems depth leverage their experience with systems and tooling to build software that ensures Datadog remains reliable, performant, and secure. For this track, their Software Engineering experience may resemble the Distributed Systems track, but is typically applied in combination with their systems experience to build and run internal platforms and tools that our products are built on. These people typically have deep experience building and managing large cloud infrastructure deployments, or leading reliability efforts for orgs similar to ours, or building release machinery to allow hundreds or thousands of devs to do their jobs without stepping on each others' toes. The systems and tooling where they may have experience depth may include (but not limited to): bazel, build tooling, cassandra, CDN, chef, configuration management, container orchestration, consul, docker, elasticsearch envoy, haproxy, kafka, kubernetes, load balancing, network architecture, postgres, redis, release management, RPC frameworks, service discovery, spinnaker, terraform, zookeeper. Bonus: You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. #LI-KM5 Datadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan. Th
Jobs in United States
Ai Infrastructure System Engineer Bangalore in New York
648 active opportunities · Updated October 2026
Showing
15 jobs
Explore current ai infrastructure system engineer bangalore jobs in New York. Filter by work mode, employment type, experience, department, date posted and distance.
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for a Detection & Response Engineer to build the systems that help us identify, investigate, and respond to threats across our platform. This is an engineering role focused on automation. You'll build detections, investigation tooling, and response capabilities that scale with our infrastructure, using AI where it meaningfully improves signal, investigation speed, and operational effectiveness. You'll work closely with infrastructure, platform, and security engineers to ensure every incident makes the platform more resilient. What You'll Work On: Detection Engineering Design and build high-fidelity detections for attacks, abuse, and anomalous behavior across our infrastructure and production systems Continuously improve detections based on telemetry, threat intelligence, and lessons learned from incidents Improve visibility across cloud infrastruc
From $131K/yr
Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We’re looking for an Infrastructure Security Engineer to design and secure the core systems that power our platform. This role focuses on building security directly into our infrastructure—from container isolation and orchestration to identity and secrets management in a multi-tenant, cloud-native environment. You’ll work closely with engineering teams to define secure primitives and ensure our platform is resilient, scalable, and trustworthy by design. This is a hands-on, deeply technical role focused on real systems, not compliance or policy. What You'll Do: Platform & Runtime Security Design and improve isolation mechanisms for multi-tenant workloads (containers, sandboxing, execution environments) Strengthen boundaries between customers, workloads, and internal systems Identify and mitigate risks in distributed, dynamic compute environments Container &
$175K – $215K/yr
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We’re looking for a Senior Software Engineer to join our Network Engineering team to accelerate building and scaling our innovative systems that support our growing identity platform. In this role, you will build the next-generation infrastructure that underpins all systems at CLEAR. The ideal candidate for this role will approach challenges with an eye toward reliability, simplicity, and scalability. What You'll Do: Develop and maintain a streamlined process for engineers to effortlessly build and deploy scalable and reliable software-defined networking solutions on AWS. Enhance our compute platform (Kubernetes) by integrating new functionalities and features, focusing on AWS networking services and concepts such as VPCs, Route Tables, Security Groups (SGs), ALBs/ELBs, and Route53, as well as implementing Kubernetes networking solutions like service mesh (Istio) to optimize service communication and management. Collaborate across engineering teams to advocate for and implement best practices in observability, utilizing tools like Splunk or Datadog to ensure robust network monitoring. Act as a product owner for our infrastructure, collecting feedback and requirements from engineering teams to address pain points and develop solutions, particularly in the realm of AWS networking and cloud-native design principles. What you're great at: 6+ years of extensive experience in infrastructure and platform development, particularly in software-defined networking and AWS cloud services. Proficient in writing production-grade softwar
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We're building the data infrastructure behind some of the most demanding AI training workloads in the world, and we want sharp, curious people to help us do it. In this role, you'll build and maintain the high-performance data layer our Modeling teams rely on for training and evaluation jobs. As a Software Engineer, Data Infrastructure, you will: Work directly on petabyte-scale storage infrastructure, and the networking and performance challenges that come with it. Collaborate daily with researchers and engineers who are some of the best in the world at what they do. You may be a good fit if you have: 4+ years of experience working on data storage infrastructure Strong command of Python Kubernetes experience, especially on the storage side (Persistent Volumes, CSI drivers, etc.) The ability to transform unstructured data into performant datasets across diverse storage backends including S3, GCS, and POSIX Experience with distributed data processing frameworks such as Apache Beam, Spark, or Flink [Nice-to-have] Familiarity with modern analytics tooling such as BigQuery, Airflow, or dbt Genuine excitement about AI.
From $105K/yr
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We’re looking for an early career Software Engineer to join our Infrastructure team to accelerate building and scaling our innovative systems that support our growing identity platform. In this role, you will build the next-generation infrastructure that underpins all systems at CLEAR. The ideal candidate for this role will approach challenges with an eye toward reliability, simplicity, and scalability. What You'll Do: Develop and maintain a streamlined process for engineers to effortlessly build and deploy scalable and reliable software-defined networking solutions on AWS. Enhance our compute platform (Kubernetes) with new functionalities and features, focusing on AWS networking services and concepts such as VPCs, Route Tables, Security Groups (SGs), ALBs/ELBs, and Route53, optimize service communication and management. Collaborate across engineering teams to advocate for and implement best practices in observability, utilizing tools like Splunk or Datadog to ensure robust network monitoring. Act as a product owner for our infrastructure, collecting feedback and requirements from engineering teams to address pain points and develop solutions, particularly in the realm of AWS networking and cloud-native design principles. What you're great at: 0-2 years of experience in infrastructure and platform development and AWS cloud services. Proficient in Python, with understanding of Kubernetes and container orchestration tools like EKS and ECS. Understand AWS networking services, including VPC design, SGs, NATGWs, ALBs/ELBs, Rout
We are investing in agentic AI and need a Senior AI Engineer to lead the design and delivery of these systems. This is a foundational hire: you will own both the agent-facing workstreams — pipelines, orchestration, conversational interfaces — and the underlying context layer that makes them reliable, including memory management, knowledge graph integration, and retrieval infrastructure. You will work closely with data engineers, project leads, and client stakeholders, and play a key role in shaping how Lynx builds and ships AI solutions at scale. What This Involves: Lead the architecture and delivery of agentic AI systems end-to-end: agents, orchestration, tool use, and multi-step reasoning workflows. Own the context layer: design and implement memory architectures (episodic, semantic, working memory) and integrate GraphRAG and knowledge graph retrieval into agentic pipelines. Build robust RAG systems — including vector retrieval, graph traversal, and hybrid search — and ensure retrieval quality through evaluation frameworks. Translate client requirements into technical designs, presenting approaches and trade-offs to both technical and non-technical stakeholders. Define standards and reusable patterns for agentic AI development that other engineers at Lynx can build on. Set up observability, evaluation, and monitoring pipelines to ensure AI systems perform correctly in production. Requirements: 5–8 years of software or ML engineering experience, with at least 2–3 years building LLM-based or agentic AI systems in production. Deep hands-on experience with agentic frameworks (LangChain, LlamaIndex, AutoGen, CrewAI, or similar) and LLM APIs (OpenAI, Anthropic, etc.). Strong understanding of agent design patterns: ReAct, planning loops, tool use, multi-agent coordination, and memory architectures. Practical experience with GraphRAG or knowledge graph-based retrieval (e.g., Neo4j, Microsoft GraphRAG) and vector databases (Pinecone, Weaviate, Qdrant, etc.). Proficiency in
Become a part of our caring community Most AI engineering jobs are a thin wrapper around a model API. This role is different. We build the platform that transforms millions of clinical documents into trusted, actionable data. Our systems use large language models (LLMs) to read medical records, extract structured facts, answer complex questions with citations back to the source document, and route ambiguous cases to human experts for review. Our users make decisions that impact real healthcare outcomes, so “good enough” is not good enough. Building AI systems that are accurate, reliable, auditable, and scalable is at the core of this role. As a Senior AI Applied Engineer, you will design, build, deploy, and operate production AI systems used at scale within one of the largest health insurers in the United States. You will own solutions end-to-end, from user experience and APIs to model orchestration, evaluation frameworks, infrastructure, and production operations. Why Join Us Build production AI systems where LLMs are in the critical path, not just demos or proofs of concept. Work on extraction, retrieval, agentic workflows, and human-review systems that process real healthcare data at scale. Own projects end-to-end across frontend, backend, AI orchestration, infrastructure, deployment, and operations. Solve challenging problems around accuracy, explainability, traceability, and reliability in regulated environments. Ship quickly in a small, high-impact team that embraces AI-assisted development and rigorous quality standards. Build systems that continuously improve through expert feedback, evaluations, and human-in-the-loop workflows. Key Responsibilities Design, develop, and deploy full-stack AI-powered application
Become a part of our caring community Every large organization is making critical decisions today about how it will leverage AI over the next decade. Few have leaders who can both define that vision and demonstrate its viability through hands-on engineering. This role requires both. We build the platform that transforms millions of clinical documents into trusted, actionable data. Our systems use large language models (LLMs) to read medical records, extract structured facts, answer complex questions with citations to source documents, and route difficult cases to human experts. These capabilities support decisions that impact real healthcare outcomes for members. As a Principal AI Applied Engineer, you will define the technical strategy, architectural standards, and long-term vision for AI-enabled products across the organization. You will influence enterprise-wide decisions regarding AI platforms, model strategies, engineering standards, and technology investments while remaining deeply hands-on in prototyping, experimentation, architecture, and software development. This is the highest-level individual contributor role within the AI Applied Engineering organization. Success requires exceptional technical depth, organizational influence, strategic thinking, and the ability to translate emerging AI capabilities into scalable, reliable, and responsible production systems. Why Join Us Shape the long-term AI architecture and engineering direction for a large enterprise healthcare organization. Influence how AI-enabled products are designed, built, evaluated, deployed, and governed across multiple teams. Drive strategic decisions involving models, vendors, platforms, infrastructure, and shared capabilities. Prototype and validate emerging technologies before the organization invests at scale.</
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for engineers with deep AI/ML and low-level systems experience who want to build the best technical support experience in the world. This isn't a traditional support role — it's an engineering role where you happen to be closest to our customers. You'll split your time roughly 50/50 between working directly with customers and shipping fixes, features, and automation that improve Modal for everyone. When you help a customer debug a training run, you'll also fix the underlying issue in the platform. When you notice ten customers hitting the same friction point, you'll build the tooling or automation that eliminates it entirely. This role is for people who solve problems, not people who answer tickets. The problems you encounter are deeply technical and arise from running some of the most demanding AI workloads in the world. You'll be a member of our eng
From $272K/yr
About Datadog: We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Opportunity: Datadog’s Senior Staff Engineers are technical leaders operating at the forefront of large-scale systems design, building the infrastructure that will support our next five years of growth and beyond. They do this in three major ways: As individual contributors, they bring world-class technical depth to build industry-leading systems in areas such as observability data platforms, distributed query engines, and real-time event streaming at global scale. As technical leaders, they apply broad architectural perspective and deep systems thinking to align design decisions across teams and domains. They work across complex, multi-team problem spaces to define long-term technical direction, drive large-scale initiatives forward, and ensure consistent execution. As engineering stewards, they play a key role in evolving our systems and engineering culture. They actively participate in Datadog’s senior technical community, bringing external insights and internal experience to elevate engineering standards and mentor the next generation of technical leaders. Examples of projects a Senior Staff Engineer may lead include designing and launching a new distributed data storage engine capable of handling hundreds of millions of records per second, building the real-time infrastructure behind a new observability product, or re-architecting a core service to support exponential growth in throughput and complexity. What You’ll Do: Be the technical owner of multiple critical systems or architecture areas, often spanning several t
About Glean: Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With over 100 enterprise SaaS connectors, flexible LLM choice, and robust APIs, Glean gives organizations the infrastructure to govern, scale, and customize AI across their entire business - without vendor lock-in or costly implementation cycles. At its core, Glean is redefining how enterprises find, use, and act on knowledge. Its Enterprise Graph and Personal Knowledge Graph map the relationships between people, content, and activity, delivering deeply personalized, context-aware responses for every employee. This foundation powers Glean’s agentic capabilities - AI agents that automate real work across teams by accessing the industry’s broadest range of data: enterprise and world, structured and unstructured, historical and real-time. The result: measurable business impact through faster onboarding, hours of productivity gained each week, and smarter, safer decisions at every level. Recognized by Fast Company as one of the World’s Most Innovative Companies (Top 10, 2025), by CNBC’s Disruptor 50, Bloomberg’s AI Startups to Watch (2026), Forbes AI 50, and Gartner’s Tech Innovators in Agentic AI, Glean continues to accelerate its global impact. With customers across 50+ industries and 1,000+ employees in more than 25 countries, we’re helping the world’s largest organizations make every employee AI-fluent, and turning the superintelligent enterprise from concept into reality. If you’re excited to shape how the world works, you’ll help build systems used daily across Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub, and many more - deeply embedded where people get things done. You’ll ship agentic capabilities on an open, extensible stack, with the craf
From $244K/yr
We're looking for a Staff Engineer to join the Logs organization at Datadog and help redefine how our customers ingest, query, and derive insights from logs data. In this role, you’ll work closely with Product Managers and customers to drive complex initiatives across ingestion pipelines, search infrastructure, and intelligent log management capabilities - all while pushing the boundaries of what’s possible with AI and distributed systems. You’ll have the opportunity to lead efforts that shape the future of log management. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Partner with Product Managers to define ambiguous product requirements and determine the most impactful solutions for customers Lead technical strategy and execution and design systems surrounding log query performance and ingestion at scale. Explore and prototype new capabilities and collaborate with peers on initiatives spanning AI-powered log management, security and business operations, advanced query capabilities, and external data sources query capabilities. Mentor engineers across levels and contribute to growing a high-performing, collaborative team culture Who You Are: You have deep experience architecting and scaling backend systems, with a strong focus on data-intensive or distributed infrastructure You excel in ambiguous environments, demonstrating a mix of drive, curiosity and pragmatic decision-making You’ve partnered effectively with Product Managers and customers to define product direction and ship impactful features You have expertise in debugging complex systems and optimizing performance across real-time data pipelines You have experience in using AI agents tools in your day-to-day engineering practices You lead by example and enjoy helping others grow through m
Our Mission You call. You wait. You call again. In every other part of your life, you book in seconds. In healthcare, you’re blocked. We’re here to give power to the patient. For nearly 20 years, we’ve built the leading healthcare marketplace - helping tens of millions of people find and book the care they need. Now, we’re going further: building our infrastructure beyond Zocdoc’s marketplace to power access to care wherever patients search, from provider websites and insurance directories to search engines, AI platforms, and more. Healthcare still lacks something every other major consumer industry takes for granted: a seamless way to go from seeking to getting . We don’t want to own the front door to care; there isn't one. We want to make sure all of those doors open when patients are knocking. Fixing healthcare starts with fixing access to it. And we're still just getting started. About the Role We're transforming how healthcare practices interact with Zocdoc, building intelligent systems that understand each practice's needs and guide them toward actions that grow their business. This means personalized homepages, smart recommendations, AI-assisted configuration, and workflows that make Zocdoc essential to daily operations. As Senior Software Engineer, you'll build these systems end-to-end. You'll own features from design through production, work across the stack, and collaborate with Product, Design, and Data Science to ship experiences that matter. What You'll Do Build platform components - including practice profile services, engagement scoring pipelines, recommendation APIs, personalization infrastructure. Ship product features end-to-end - database to API to front-end, owning the full lifecycle. Work with Data Science to integrate ML models build feature pipelines, call model endpoints, instrument feedback loops. Write production-ready code with strong testing, observability, and error handling. Participate in design discussio
Other cities to consider
More places hiring for this role
Get new ai infrastructure system engineer bangalore jobs in New York, United States by email
Daily job updates · Unsubscribe anytime