About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des
Jobs in United States
Ai Infrastructure System Engineer Bangalore in United States
5,089 active opportunities · Updated October 2026
Showing
15 jobs
Explore current ai infrastructure system engineer bangalore jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Infrastructure Quality team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, general contractors, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. Our work spans from vendor qualification through commissioning, ensuring operational readiness across our global portfolio. About the Role We are seeking an experienced Manufacturing Quality Engineer (MQE) to establish, implement, and manage a manufacturing-focused quality program for datacenter infrastructure. This role will be responsible for vendor oversight, quality assurance, process improvement, and issue resolution for all critical systems. You will lead vendor audits, monitor performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced risks, and operational reliability. By partnering with vendors, construction teams, and internal stakeholders, you will help ensure OpenAI’s datacenters are delivered on time and built to the highest operational standards. Travel Domestic and international travel as needed (estimated 40–60%) to manufacturing sites, datacenter locations, and partner facilities. Key Responsibilities Vendor Oversight & Performance Management Conduct manufacturing evaluation, audits, and improve vendor performance across production, inspection, testing, and delivery phases. Develop and track quality metrics to assess manufacturing performance and identify trends. Partner with vendors to refine processes, training, and quality controls to mitigate risks before shipment. Program Development & Execution Develop and maintain a datacenter-focused m
About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are hiring a Sim Infrastructure Engineer to turn simulation systems into reliable, automated, production-quality pipelines that power model training, evaluation, and hardware-in-the-loop validation. This role owns the automation, orchestration, and tool integration that apply simulation to concrete robotics tasks: building CI/CD for SIL/HIL, presubmit checks, automatic model evaluation, metric computation and reporting, and the runtime infrastructure to run simulations at scale. You will collaborate closely with Sim Realism, Sim Environments, Research, and Ops to make simulation an integrated, reproducible, and measurable part of our ML and robotics workflows. This role is based in San Francisco, CA, and requires in-person 4 days a week. In this role, you will: Build and maintain presubmit checks, continuous integration and deployment pipelines for simulation code, environments, and tasks so simulation artifacts are testable, versioned, and reproducible. Implement end-to-end automation to run model evaluation in sim (SIL) and orchestrate HIL runs; compute realism and task metrics, generate dashboards and alerts, and ensure evaluation is repeatable and auditable. Create robust APIs and connectors so research, training, and data-collection systems can schedule, seed, and evaluate batches of simulations; support RL rollouts, imitation-data collection, and presubmit model checks. Build scheduling, batching and orchestration for running very large numbers of concurrent rollouts (target tens of thousands of rollouts / large RL workloads), sol
From $98K/yr
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity At New Relic, we provide our customers with real-time insights, so they can innovate faster. Our software provides deep observability across the stack, enabling software teams to solve their customer’s problems, accelerate digital transformation, and make DevOps work. You will be at the heart of the teams supporting New Relic’s infrastructure and will work on a team that provides global service mesh and load balancing solutions. We provide these services on-premises, as well as using our multi-cloud infrastructure. We support each other to do our best work through positive communication and continuous improvement. What You’ll Do As a key member of our Infrastructure team, you will design and operate a scalable, resilient ingress data plane that directly impacts the value we provide to our customers. By ensuring the stability and performance of our global service mesh and load balancing solutions, you drive the foundational reliability that the entire New Relic organization depends on to deliver real-time insights. You will leverage advanced automation and infrastructure-as-code to accelerate development speed, allowing our engineering teams to ship safe, incremental changes across a massive fleet with confidence. Your work in evolving our DNS and CDN infrastructure is not just about maintenance; it is about creating a seamless, high-performance environment that enables innovation at scale. Through deep collaboration with Product, Design, and partner platform t
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don't just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. This is a dedicated identity engineering role on our Enterprise Cloud Engineering team. You will own and operate the enterprise identity and access platforms — Okta, Microsoft Entra ID, and modern IGA systems — that govern how every employee, application, service, and AI agent authenticates and is granted access at scale. Our team builds the authentication, authorization, identity governance, and identity lifecycle capabilities that keep Snowflake secure and compliant, partnering closely with Security and Engineering to deliver secure-by-default, Zero Trust access. These identity p
About the Team This team builds and operates the systems that enable OpenAI researchers to run reliable, scalable, and efficient research workflows. The team sits close to research and works across infrastructure, systems, and automation to make sure researchers have the tools and environments they need to move quickly. The work spans software engineering, infrastructure, systems administration, cluster operations, and reliability engineering. As OpenAI’s infrastructure evolves from bespoke bare-metal systems toward more standard, scalable platforms, the team needs engineers who can understand how systems work end-to-end and build the right abstractions without reinventing the wheel. About the Role As a Software Engineer on this team, you will build and operate the infrastructure that supports frontier research and critical research-facing systems. You will work on systems that sit close to the metal, but the role is not limited to classic operations or sysadmin work. We are looking for someone who can reason about networking, bootstrapping, Kubernetes, scalability, automation, and reliability - while also writing software to make these systems better over time. This role is a strong fit for an independent, high-ownership engineer who enjoys reliability-heavy infrastructure work but still wants to build. You do not need to come in as a kernel expert or highly algorithmic optimization engineer, but you should be deeply curious about infrastructure, comfortable debugging complex systems, and excited to support researchers doing novel work. We expect you to: Build and operate reliable infrastructure for research workloads and research-facing services. Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services. Improve cluster bootstrapping, provisioning, automation, and deployment workflows. Debug issues across networking, compute, storage, orchestration, and service reliability layers.
From $196.8K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Sr. Studio Software Engineer for Roblox Studio Platform, you will be a key contributor to the evolution of Roblox Studio, the primary IDE for making massive multiplayer online games on the Roblox platform. Studio provides the mission-critical tools for 3D modeling, animation, and the complete SDLC for millions of developers. We are looking for engineers who thrive on an adventure into the unknown and have experience across various systems, Operating Systems, Game Engines, and Application Frameworks. You’ll tackle projects involving: Core User Features: Architecting application frameworks, windowing systems, and code generation. “AI Native” Features: Pioneering scalable systems that extend to complex, agentic use cases. Foundational Architecture: Driving OS integration, extensibility, and customizability at the deepest levels. Central Backend Systems: Engineering the infrastructure to power a consistent, high-performance UX. Design Evolution: Collaborating with UX designers to translate Studio into a modern, consistent design language. You Will: Design and execute the technical direction to drive the future extensibility and adaptability of the application. Own and deliver complex techn
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Making data driven decisions is key to Plaid's culture. To support that, we need to scale our data systems while maintaining correct and complete data. We provide tooling and guidance to teams across engineering, product, and business and help them explore our data quickly and safely to get the data insights they need, which ultimately helps Plaid serve our customers more effectively. Engineers on Data Infrastructure are domain experts in Data Warehouse, Data Lakehouse, Spark, Workflow Orchestration, and Streaming technologies. We scale our existing data pipelines in a performant and cost efficient way while creating the necessary abstractions to make developing on top of this platform extremely simple for other engineers at Plaid. Responsibilities Contribute towards the long-term technical roadmap for data-driven and machine learning iteration at Plaid Leading key data infrastructure projects such as improving ML development golden paths, implementing offline streaming solutions for data freshness, building net new ETL pipeline infrastructure, and evolving data warehouse or data lakehouse capabilities. Working with stakeholders in other teams and functions to define technical roadmaps for key backe
From $295.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's data infrastructure processes petabytes of data daily, powering analytics, ML, and product decisions for a platform serving 200M+ daily active users. As a Principal Software Engineer in our Data Infra org, you will be the primary technical leader driving the strategic vision, long-term architecture, and massive scalability of our distributed data platforms that power Roblox. You will own and drive the next-generation architecture of our core platforms, which span Kafka, Flink, Spark, Trino, Druid, Airflow and Data Catalog. This role operates under high ambiguity, demanding unparalleled ownership to redefine the limits of infrastructure handling exabyte-scale workloads, and providing a unique opportunity to lead the future evolution of our global data ecosystem. You Will: Define Multi-Year Technical Strategy: Own and drive the end-to-end architectural vision for Roblox's core data platforms spanning Kafka, Flink, Spark, Trino, Druid, Airflow, and Data Catalog systems. Turn multi-year company strategies into concrete, production-grade infrastructure blueprints. Lead Cross-Functional Alignment: Partner closely with executive leadership, platform governance, data science, and product e
From $345K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer on Creator Services Data, you’ll be leading the company’s efforts to build the next generation Data Storage systems to power the millions of experiences on the Roblox Platform. We run the mission critical cloud services, Data Stores , Memory Stores , and Badges , which are crucial for storing game state such as inventory and scores, implementing leaderboards, server lists, and trading, and tracking player progress and achievements. Our team is also responsible for building dashboards to provide insights to Creators using cloud services including Client/Server Performance , Data Stores , and Memory Stores . Finally, our team owns the Roblox Extended Services platform, which provides the capability for large experiences to purchase additional resources for existing services like Data Stores and new services built around compute and generative AI. At its core, this team is focused on solving complex back end distributed systems and storage problems at scale. However, our scope extends to full stack projects spanning all the way from the infrastructure layer, through data storage and data pipelines, microservices, telemetry, game servers,
From $196.8K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. At Roblox, we strive to connect a billion people with optimism and civility, and the Safety organization’s mission is to become the leader in civil immersive online communities. We systematically detect, remove, and prevent problematic accounts, content and behavior, and we make Roblox accounts secure and free from compromise. We cover a broad area of the tech spectrum, including machine learning, classifiers for 3D models, experimentation, automation, detection workflows, and AI-powered text filters. Aligned and partnering with product teams, we use this tool-belt to discover new opportunities, influence and shape the product roadmap and prioritization, build safety products, and measure the impact on our community of users and developers. In doing so, we keep Roblox safe, civil, and inclusive, and we foster positive relationships between people around the world. WHY Safety Data Infrastructure To pro-actively find bad actors and protect good users, Roblox needs to ingest enormous amounts of data and make it usable both to our automated detection systems and for human moderators and customer support agents doing investigations. The Safety Data Infrastructure team is addressing these needs w
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Security Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments—GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Security Software Engineer, you will design and build critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Architect and implement production-grade security services (e.g., auth services, access brokers, secure proxies, key-management infrastructure) that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD. Partner with infrastructure and research engineers to embed security into high-performance compute clusters, enabling rapid model training and deployment without compromising protection. Develop automation and detection tooling to continuously identif
Other cities to consider
More places hiring for this role
Get new ai infrastructure system engineer bangalore jobs in United States by email
Daily job updates · Unsubscribe anytime