About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Software Engineer, Distributed Data Systems, you will design, build, and operate some of the largest distributed data systems in the world. You will be responsible for the end-to-end stack to deliver and consume top-quality data for robotics training at exabyte-scale. You’ll manage distributed data pipelines, collaborate closely with researchers to translate requirements into robust systems, and harden pipelines that serve as the backbone for OpenAI’s rapid iteration cycles. We’re looking for engineers who are detail-oriented, have strong experience with distributed systems, and excel at building reliable, large-scale systems in high-stakes environments. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain data infrastructure such as exabyte-scale distributed data processing, data selection, automated labeling, and training data loaders. Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient. Partner with researchers to deeply understand requirements and translate them into production-ready systems. Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation. Deliver the best possible data for training robotics models. You might thrive in this role if you: Have strong experience with distributed systems and large-scale infrastructure with a strong interest in data. Are detail-oriented a
Jobs in United States
Ai Infrastructure System Engineer Bangalore in United States
5,082 active opportunities · Updated October 2026
Showing
15 jobs
Explore current ai infrastructure system engineer bangalore jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
From $132.4K/yr
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Are you passionate about building impactful products for sales and finance teams? Come join the IT Enterprise Systems team at Pinterest where you will be responsible for advancing our sales and marketing systems. What you’ll do: Design, build, and operate full‑stack applications and services on Pinterest’s enterprise infrastructure to support our Sales, Marketing, and Finance teams, from backend services and APIs through integrations and user‑facing workflows. Lead the technical design and implementation of GenAI/ML‑powered services and pipelines that automate and augment enterprise workflows (for example, summarizing sales interactions, enriching account data, or surfacing intelligent recommendations), including clear evaluation frameworks, observability, and validation guardrails. Own the end‑to‑end software development lifecycle for th
From $10K/yr
About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role The Applied AI team at Ramp is at the forefront of leveraging AI to drive innovation across our platform. We are seeking strong full-stack engineers who are proficient in web frameworks, backend development, and infrastructure. You will work on exciting projects such as AI Agents, Retrieval-Augmented Generation, Structured Extraction (we made https://github.com/1rgs/jsonformer ), internal tooling for customer-facing teams, fine-tuning models, and build infrastructure for LLM inference. If you're passionate about working on real production use cases of large language models (LLMs) and want to contribute to groundbreaking AI applications, this role is for you. What You’ll Do Ship full-stack AI projects end to end Build and integrate components for AI infrastructure, supporting production-level inference and fine-tuning Develop and improve engineering processes, tools, and systems to scale AI solutions across Ramp Create tools and internal platforms to enhance the productivity and capabilities of Ramp's AI and engineering teams What You
What we're building Mutiny is the self-improving AI infrastructure for GTM teams to execute faster and close more revenue. Our ambition is to do for revenue velocity what Cursor and Claude Code did for engineering velocity. With Mutiny, everyone in sales and marketing gets a bench of GTM athletes that handle any work across their revenue motion and learn from what's actually moved their deals. In April we re-launched the product as an agent-first platform. Anthropic showcased us as a leader in AI GTM. MRR is growing more than 70% month-over-month, with customers like Uber, Rippling, and Snowflake. We're backed by Sequoia, YC, and Insight, and we're building a generational company. The opportunity Most engineers spend their career making predictable systems faster. You'll spend yours making non-deterministic ones trustworthy. As a senior engineer on our AI product team, you'll architect the Campaign Builder and Agent experiences marketers and sellers open every day to go from idea to personalized assets in minutes. You'll partner directly with product, design, and the founders to define what an agent-first GTM platform should feel like, and your calls on architecture, evals, and guardrails compound across thousands of customer accounts. This role is in person in New York City, five days a week, and we ship weekly. What you'll own The core agent surfaces. Architect and ship the Campaign Builder and Agent experiences end-to-end. Frontend, backend, prompts, evals, the whole stack. Reliability on top of LLMs. Make non-deterministic models feel deterministic at the surface. Build the retries, fallbacks, and orchestration so the customer never sees the failure mode. Evals and guardrails. Define how we measure quality, catch regressions, and keep brand and tone consistent across thousands of customer accounts. Speed and feel. AI products live or die by latency and the loop between intent and output. You'll obsess over both, and use coding agents and agent networks to ship f
About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A
We’re looking for a Software Engineer to architect and build backend systems that enforce data privacy and automate compliance at scale. You’ll work closely with product, infrastructure, security, and legal teams to embed privacy-by-design into our data and access layers. This is a hands-on, high-impact role for an experienced engineer who is passionate about protecting user data while enabling innovation. What You’ll Do Design, build, and operate backend services that enforce policy-driven data access, lifecycle controls, and privacy protections. Develop distributed authorization and identity-aware enforcement mechanisms integrated directly into data services and control planes. Implement auditability, policy hooks, and enforcement observability to ensure compliance is continuously verifiable. Partner with Security, Legal, and Compliance to convert privacy requirements into scalable technical designs and developer-friendly APIs. Harden data platforms and backend services through schema-level controls and data handling constraints by default. Collaborate with infrastructure teams to ensure consistent enforcement across systems while minimizing duplicated implementations. Contribute patterns, libraries, and education that elevate trustworthy data access patterns across the organization. You Might Thrive in This Role If You Have 5+ years of industry experience building and operating backend or infrastructure systems in production. Strong software engineering fundamentals , with fluency in at least one major programming language (e.g., Python, Go, Rust, C++, Java). Experience with distributed authorization, RBAC/ACL systems, encryption-based access, or policy engines. Familiarity with global privacy regulations and their architectural implications. Ability to influence and collaborate with teams across legal, compliance, product, and engineering. A bias toward practical, impactful solutions that balance privacy protections with product needs. Nice to Have Experience wi
About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. This team is central to this mission, setting the core infra strategy and implementing this vision. From site selection to the buildout process, this team sits at the intersection of commercial, technical, strategy, and operations, interacting with teams and executives inside and outside of OpenAI. About the Role We are seeking experienced Data Center Mechanical and Electrical/Power Design Engineers with expertise in designing, operating, and maintaining large-scale data center campuses. The ideal candidate for this role will have extensive background and experience in design and managing critical equipment and facilities, design and operation of MEP (Mechanical, Electrical, Plumbing) systems, and overseeing operational activities from initial phases of Data Center build through delivery and ongoing maintenance. The ideal candidate will have a strong technical background, operational leadership experience, and a proven ability to collaborate with external vendors on critical infrastructure. This role offers the opportunity to lead transformative data center projects with high visibility and impact. If you are passionate about delivering cutting-edge infrastructure solutions, we encourage you to apply. Key Responsibilities Oversee building and MEP design, operation, and maintenance, including reviewing building and MEP drawings and proposals across all project phases. Lead operational activities for large-scale data center campuses, from early design phases through delivery and daily operation. Operate and maintain critical data center facilities and equipment, ensuring reliability and performance. Collaborate with external vendors to select, procure, and manage critical equipment, such as generators, UPS, chillers, and CDUs. Provide technical expertise on all aspects of data center building, equipment, and
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role This is where security meets innovation at enterprise scale. As a staff security engineer, applications at WRITER, you'll be building the security foundations that protect the AI systems powering some of the world's most recognizable brands. You'll work at the intersection of application security, AI infrastructure, and developer enablement—partnering with engineering teams to embed security into every line of code while ensuring our platform remains both powerful and trustworthy. The opportunity is massive: you'll help define how enterprise AI applications are secured, from threat modeling our LLM architectures to building automated security controls that scale across our growing platform. This isn't about saying "no"—it's about finding creative ways to say "yes, and here's how we do it securely." You'll tackle challenges that most security engineers never encounter: securing AI agents, protecting training data pipelines, and designing controls for systems tha
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Join WRITER's security team as a staff detection and response engineer and help protect the AI infrastructure that's transforming how the world works. You'll build sophisticated detection systems that identify attacks targeting our AI platform, training data, and model deployments while creating automated response capabilities that scale with our explosive growth. This isn't just traditional security work – you're defending cutting-edge AI/AGI systems against adversaries who are evolving their tactics as fast as AI itself advances. This role combines hands-on security engineering with strategic thinking to stay ahead of novel threats that don't exist in textbooks yet. You'll be the operational arm of our security function, translating threat intelligence into real-time detections, coordinating incident response across multiple teams, and hunting for sophisticated attacks across GPU clusters and distributed training environments. If you're excited by the challen
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role This is where security meets innovation at enterprise scale. As a security engineer, applications at WRITER, you'll be building the security foundations that protect the AI systems powering some of the world's most recognizable brands. You'll work at the intersection of application security, AI infrastructure, and developer enablement—partnering with engineering teams to embed security into every line of code while ensuring our platform remains both powerful and trustworthy. The opportunity is massive: you'll help define how enterprise AI applications are secured, from threat modeling our LLM architectures to building automated security controls that scale across our growing platform. This isn't about saying "no"—it's about finding creative ways to say "yes, and here's how we do it securely." You'll tackle challenges that most security engineers never encounter: securing AI agents, protecting training data pipelines, and designing controls for systems that didn
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role This is where security meets innovation at enterprise scale. As a security engineer, applications at WRITER, you'll be building the security foundations that protect the AI systems powering some of the world's most recognizable brands. You'll work at the intersection of application security, AI infrastructure, and developer enablement—partnering with engineering teams to embed security into every line of code while ensuring our platform remains both powerful and trustworthy. The opportunity is massive: you'll help define how enterprise AI applications are secured, from threat modeling our LLM architectures to building automated security controls that scale across our growing platform. This isn't about saying "no"—it's about finding creative ways to say "yes, and here's how we do it securely." You'll tackle challenges that most security engineers never encounter: securing AI agents, protecting training data pipelines, and designing controls for systems that didn
From $136K/yr
About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As an ML Platform Engineer at Stitch Fix, you will play a key role in building and maintaining the critical infrastructure that powers machine learning and AI across our organization. You will design, develop, and support scalable, resilient services and frameworks for ML model training and deployment, feature engineering and serving, candidate generation, AI agent deployment and observability, and other core platform capabilities. In this role, you'll contribute to the day-to-day operations of the ML Platform team, ensuring the smooth functioning of existing systems while driving improvements. You’ll collaborate closely with full-stack data scientists, offering consultation and support to help them unlock the full potential of our platform. With significant autonomy, you’ll have the opportunity to shape the future of ML and AI at Stitch Fix. Your ideas and expertise will drive improvements, codify best practices, and influence how we approach machine learning and AI systems at scale. Responsibilities: Collaborate with cross-functional teams, including data scientists, engineers, and business partners, to solve complex distributed systems and business challenges at scale. Be part of a team with high visibility across the organization, driving impactful solutions that make a difference. Share your ideas and help guide the team’s investments toward high-value opportunities. Foster a culture of technical collaboration and contribute to the development of scalable, resilient systems. About You You bring
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Join WRITER's security team as a staff detection and response engineer and help protect the AI infrastructure that's transforming how the world works. You'll build sophisticated detection systems that identify attacks targeting our AI platform, training data, and model deployments while creating automated response capabilities that scale with our explosive growth. This isn't just traditional security work – you're defending cutting-edge AI/AGI systems against adversaries who are evolving their tactics as fast as AI itself advances. This role combines hands-on security engineering with strategic thinking to stay ahead of novel threats that don't exist in textbooks yet. You'll be the operational arm of our security function, translating threat intelligence into real-time detections, coordinating incident response across multiple teams, and hunting for sophisticated attacks across GPU clusters and distributed training environments. If you're excited by the challen
Other cities to consider
More places hiring for this role
Get new ai infrastructure system engineer bangalore jobs in United States by email
Daily job updates · Unsubscribe anytime