Jobs in United States

Ai Infrastructure System Engineer Bangalore in United States

5,082 active opportunities · Updated October 2026

Explore current ai infrastructure system engineer bangalore jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t

AWSKubernetesRestAI
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake is about empowering enterprises to achieve their full potential, and people too. With a culture that’s all in on impact, innovation, and collaboration, Snowflake is the sweet spot for building big, moving fast, and taking technology, and careers, to the next level. About the Role The Cortex Code team is building the future of coding agents for working with data. See our flagship product in action: Cortex Code in Action: Live Demos + AMA . Your work will directly impact how developers and businesses build with data. You'll own the full AI engineering lifecycle: design, prompt/tool engineering, evals, deployment, measurement, and optimization. You'll work with a small, high-powered modeling and infrastructure team. What you will do in this role: Own features end-to-end for Snowflake Cortex Code products. Build agentic workflows, coding harnesses, evaluation pipelines. Build enterprise-grade context engineering: function calling, tool schemas, guardrails, agent teams, and verification/repair. Partner with product and infra: translate customer problems into products and experiments. Collaborate with infrastructure teams to productionize improvements. Work with an elite team of engineers towards building great products Requirements: Bachelor’s degree in Computer Scienc

TypeScriptPythonAIGo
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will help develop and validate next-generation server platforms that power a reliable, high-performing, and cost-efficient infrastructure at scale. You will work across platform bring-up, firmware qualification, hardware validation, fleet integration, and performance optimization to support large-scale production deployments. You Will: Bring-up & Sustaining: Drive key aspects of the hardware development lifecycle, including feasibility studies, hardware bring-up, validation, deployment, and ongoing production support. Platform Optimization: Perform platform integration, performance characterization, and system-level debugging across compute infrastructure, focusing on hardware optimization, driver tuning, and thermal/power efficiency. Hardware Validation: Develop and execute rigorous evaluation and stress-testing strategies for server platforms to ensure reliability and performance under production-scale workloads. Firmware & Fleet Enablement: Support BIOS/BMC firmware qualification, hardware health monitoring, and automation tooling for firmware deployment and lifecycle management. Vendor & Cross-Functi

PythonAWSGitLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa

AWSRestAIGo
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lake using open formats like Apache Iceberg, delivering deep correlation and long-term analytics at dramatically lower cost. A dynamic Knowledge Graph and chat-based AI SRE provide rich context and guided workflows so teams can move from detection to root cause and resolution significantly faster. The Infrastructure team at Observe by Snowflake is responsible for building, scaling, and operating the development and production environments that power our observability platform. We are a small, highly collaborative team with a broad scope, focused on delivering reliable infrastructure while continuously improving the systems that support our engineers and customers. What You’ll Do Design, build, and operate scalable cloud infrastructure in AWS supporting a high-scale observability platform. Improve system reliability, performance, and operational visibility across development and production environments. Develop and maintain CI/CD pipelines and internal tooling to improve developer productivity and deployment safety. Identify and mitigate security risks, and help maintain intern

PythonAWSAzureGCP
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated

PythonKubernetesLinuxMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea

AWSAzureKubernetesCI/CD
V
📍 San Francisco, California, United States· Full-time
✓ High-confidence listing

$165K – $185K/yr

Quick readStrong listing-quality and freshness signals

About VSCO For years, we've helped photographers create their work. Now we're building what comes next. VSCO exists for photographers. Not as a side feature, not as an afterthought, but as the whole point. We build the connected system photographers rely on, and we've spent over a decade earning the trust of a global creative community that takes the craft seriously. Photography is at an inflection point. AI is reshaping what's possible for creative work, and that's where our mission shines. VSCO is building the full photographer's workflow: from creating and editing your work, to delivering it to clients, to running your business. All of it built thoughtfully, with photographers leading the way. We believe the future of photography tools creates more space for creativity, handling the busy work so photographers can focus on the craft. If you care about craft, community, and what technology can unlock for creative people, this is the work. We're a mission-driven and focused company where your work ships quickly, is meaningful, and reaches tens of millions of people worldwide. You'll have a real say in what we build and how we build it. We hire people who don't wait to be asked, naturally connect the dots, care about the quality of what they ship, and believe the best outcomes come from building together. About The Role We’re looking for a Senior Software Engineer, Infrastructure to own the platform that VSCO product and data teams ship on. You’ll join a small infra team that treats AWS, EKS, and GitOps as the default path for new systems, and you’ll spend real time pairing with other teams so - search, Workspace, and data - land in the right Terraform, Helm, and Flux instead of one-off snowflakes. This is a hands-on senior seat. You will design Terraform modules, cut production traffic to EKS, secure traffic with Cloudflare, look for optimizations in our AWS infrastructure, and leave on-call better than you found them. You like production ownership, you can sequence

PythonJavaSQLMySQL
C
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? The Data Infrastructure team at Cohere is responsible for the storage and data movement layer underlying every model training run. We're building the unified storage layer that feeds our training workloads. It needs to serve petabytes of training data and model checkpoints fast enough to keep thousands of GPUs busy across several training clusters. In this role, you’d have an opportunity to build this system from the ground up. You’d be a key contributor, working on a problem few teams have had to solve at this scale. In this role, you will: Design, build, and operate the distributed storage system that feeds model training and evaluation. Run this system multiple on Kubernetes clusters at petabyte scale. Work with researchers and training-infra teams on how jobs actually read and write data, and turn that into throughput, latency, and durability requirements Work through the networking, I/O, and consistency problems of moving large datasets and checkpoints across regions and backends, with GPU idle time and time-to-insight as the measures of success You may be a good fit if you have: Strong storage fundamentals,

PythonKubernetesGitRest
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa

AWSRestAIGo
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $287.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Privacy Infrastructure team, you’ll design and build large-scale systems that operate across Roblox’s platform and data ecosystem. You’ll tackle complex distributed systems challenges, develop reliable and scalable infrastructure, and work across engineering teams to deliver foundational capabilities that serve millions of users. This is an opportunity to take on high-impact technical problems at Roblox scale while helping shape the next generation of our infrastructure. You Will Design, build, and scale reliable infrastructure and platform solutions that protect user data and support the needs of a global platform. Develop foundational capabilities for emerging technologies, including agentic AI workflows, with a focus on safe and responsible data access. Partner with engineering teams across Roblox to integrate scalable data protection capabilities into their systems and development workflows Drive technical strategy and architecture for complex, cross-cutting challenges spanning data, infrastructure, privacy, and security. You Have 7+ years of proven experience as a software engineer. Expertise in Python or Golang; Strong understanding of system

PythonAWSGitRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Platform Analytics team builds the systems OpenAI researchers use to understand the quality and behavior of the models we train including what models are doing, why they behave in a particular way, and how that behavior changes across experiments. Neptune is a core part of this work. It ingests, stores, queries, and visualizes large volumes of metrics from pretraining, post-training, and reinforcement learning. Hundreds of researchers depend on these systems in their daily work to compare experiments, debug unexpected behavior, and decide what to try next. Our scope is broader than metrics. We also build platforms that help researchers analyze samples, traces, evaluation results, and other structured or unstructured data through dashboards, APIs, and increasingly agent-driven workflows. These systems need to remain fast, reliable, and understandable as the scale and complexity of research change quickly. We are not trying to become a consulting team that builds a separate solution for every research project. We work directly with researchers to understand recurring problems, then turn them into reusable infrastructure and platform capabilities that many teams can build on. About the Role We’re looking for a hands-on experienced software engineer who can take ownership of a critical system and drive it from problem definition through production adoption. This person should be able to own a platform such as CacheHouse end to end: define its technical direction, design its data model and storage architecture, integrate it with several research dashboards and workflows, guide one or two engineers, and ensure the system works reliably for its users. The right candidate should already bring the technical judgment, ownership, and execution expected at this level. The primary learning curve should be OpenAI’s stack and research problem space, not learning how to lead a complex engineering effort or deliver a production system. You will work directly with

AWSRestAIC++
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati

PythonSQLAWSLinux
T
📍 Boston, Massachusetts, United States· Full-time
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a SoC Design Verification Engineer to validate the System Management Controller (SMC) and enable seamless multi-chip integration. In this role, you will design and execute tests, build infrastructure, and debug issues across chiplet-based SoCs. You’ll have the opportunity to work with remote mentorship while contributing to the foundation of scalable multi-die systems. This role is hybrid, based out of Toronto, Ontario, Boston, MA or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Proficient in SystemVerilog, SV-UVM, Python, and C/C++ with strong verification skills. Experienced in writing test plans, building infrastructure, and debugging hardware/software flows. Comfortable working with remote mentorship and distributed teams. Familiar with AI-assisted tools like Copilot, Cursor, and Claude to accelerate verification. What We Need Develop and maintain SMC tests and supporting DV infrastructure. Write, execute, and track test plans for chiplet and multi-chip SoC designs. Use C/C++ to develop tests compiled, loaded, and executed directly on the DUT. Triage, analyze, and debug issues in clos

PythonAWSAIC++
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. Job Summary We are looking for an experienced Growth Infrastructure Engineer to build and maintain the technical backbone that enables scalable growth experiments, high-performance data pipelines, and automated systems that drive user acquisition, engagement, and product iteration. This role sits at the intersection of growth, product, and infrastructure — combining deep technical engineering with experimentation and data-driven optimization. You will collaborate with product, data science, and backend teams to ensure that growth initiatives run smoothly and scale efficiently across systems. Key Responsibilities Growth Infrastructure & Systems Design, implement, and maintain scalable infrastructure that supports growth and experimentation needs. Build and optimize analytics pipelines to capture key product and growth metrics (acquisition, activation, retention, etc.). Develop automated workflows for user onboarding, campaign delivery, and performance tracking. Experimentation & Optimization Support A/B testing frameworks and integrate them into production systems. Enable reliable data collection and evaluation for growth experiments. Automate deployment and rollout of growth feature flags and tests. Cross-Functional Collaboration Partner with Growth Product Managers, Data Engineers, and Analysts to define technical requirements for growth initiatives. Translate business goals into technical specifications and system designs. Provide guidance on performance, reliability, and scalability trade-offs. Monitoring & Reliability Implement monitoring and alerting for growth infrastructure services. Troubleshoot production issues and optimize for uptime and performance. Ensure data quality and consistency for report

JavaScriptPythonJavaAWS
🔔

Get new ai infrastructure system engineer bangalore jobs in United States by email

Daily job updates · Unsubscribe anytime