Jobs in United States

Ai Infrastructure System Engineer Bangalore in United States

5,082 active opportunities · Updated October 2026

Explore current ai infrastructure system engineer bangalore jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

MR
📍 Salt Lake City, Utah, United States· Full-time
✓ High-confidence listing

From $56K/yr

Quick readStrong listing-quality and freshness signals

Locations: South Jordan, UT Salary: $56,000 Launch Your Career in Technology Every app, website, payment, and digital service relies on technology running smoothly behind the scenes. When something goes wrong, Production Support Engineers are the people who investigate the issue, restore service, and help prevent it from happening again. If you're curious, analytical, and enjoy solving problems, this is an opportunity to build hands-on experience with cloud platforms, Linux, automation, databases, and large-scale enterprise systems from day one. What Is Production Support? Production Support Engineers keep business-critical applications running reliably in live environments. Think of it this way: Software Engineers build the platform. QA Engineers test the platform. Production Support Engineers keep the platform running when it matters most. Working at the intersection of technology and business, you'll troubleshoot issues, automate processes, and help improve the reliability and performance of systems used by thousands, or even millions, of people every day. If you enjoy solving puzzles, working under pressure, and understanding how large-scale systems work, this could be the perfect place to start your career. What You'll Do As part of a global production engineering team, you'll: Help support large-scale applications and platforms used by leading organizations around the world. Monitor business-critical applications and services to ensure high availability and performance. Investigate and resolve production incidents across applications, infrastructure, databases, and cloud environments. Analyse logs, alerts, and system metrics to identify root causes and prevent recurring issues. Partner with software engineers, infrastructure teams, and business stakeholders to improve system reliabilit

JavaScriptPythonJavaSQL
B
📍 New York, New York, United States· Full-time
✓ High-confidence listing

$192K – $240K/yr

Quick readStrong listing-quality and freshness signals

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Senior Full-Stack Software Engineer on our Task Workflows Platform team, you will build and scale the foundational infrastructure that powers core experiences across Brex. You will work across our extensive suite of platforms - including our multi-channel notifications engine, collaborative commenting system, and dynamic workflow rule builder. You will also play a pivotal role in evolving our task orchestration infrastructure, helping us transition from a centralized, generic experience into highly tailored, p

TypeScriptPythonJavaReact
B
📍 Seattle, Washington, United States· Full-time
✓ High-confidence listing

$192K – $240K/yr

Quick readStrong listing-quality and freshness signals

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Senior Full-Stack Software Engineer on our Task Workflows Platform team, you will build and scale the foundational infrastructure that powers core experiences across Brex. You will work across our extensive suite of platforms - including our multi-channel notifications engine, collaborative commenting system, and dynamic workflow rule builder. You will also play a pivotal role in evolving our task orchestration infrastructure, helping us transition from a centralized, generic experience into highly tailored, p

TypeScriptPythonJavaReact
B
📍 San Francisco, California, United States· Full-time
✓ High-confidence listing

$192K – $240K/yr

Quick readStrong listing-quality and freshness signals

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Senior Full-Stack Software Engineer on our Task Workflows Platform team, you will build and scale the foundational infrastructure that powers core experiences across Brex. You will work across our extensive suite of platforms - including our multi-channel notifications engine, collaborative commenting system, and dynamic workflow rule builder. You will also play a pivotal role in evolving our task orchestration infrastructure, helping us transition from a centralized, generic experience into highly tailored, p

TypeScriptPythonJavaReact
M
📍 New York City, United States· Full-time
✓ High-confidence listingCompany trend -55.6%

From $130K/yr

Quick readStrong listing-quality and freshness signals

About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About the Team The Revenue Marketing team at Mixpanel is responsible for pipeline generation across paid, website, and product-led channels. This role sits within that team as its dedicated engineering owner — accountable for the technical systems that power how people discover, evaluate, and request Mixpanel. You are the only engineer dedicated to this surface, which means you set its technical direction rather than execute against someone else's. You work directly with our marketing teams and partner with Growth Engineering on shared systems. Website design and content are handled by our web design team; your focus is the technical infrastructure, integrations, and systems that sit underneath. About the Role As our Marketing Systems Engineer, you own the technical layer of our marketing engine: the integrations and infrastructure that power our marketing website. The website is our primary lead capture system that turns traffic into pipeline. You own the systems that connect our marketing surface to Salesforce, Customer.io/Hubspot, and our broader GTM stack. You set the technical roadmap for this surface, move fast, measure impact, and treat reliability as a first-class concern rather than a cleanup task.You will also collaborate with Growth Engineering on shared infrastructure including handraiser routing and tracking. Responsibilities Own the technical health of the Mixpanel marketing website: page speed, WCAG compliance, Google Tag Manager, technical SEO, and third-party integrations including Qualified, Optimizely, and TrustArc. Own the integrations between the marketing website and our GTM stack t

JavaScriptPythonJavaAWS
M
📍 San Francisco, United States· Full-time
✓ High-confidence listingCompany trend -55.6%

From $130K/yr

Quick readStrong listing-quality and freshness signals

About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About the Team The Revenue Marketing team at Mixpanel is responsible for pipeline generation across paid, website, and product-led channels. This role sits within that team as its dedicated engineering owner — accountable for the technical systems that power how people discover, evaluate, and request Mixpanel. You are the only engineer dedicated to this surface, which means you set its technical direction rather than execute against someone else's. You work directly with our marketing teams and partner with Growth Engineering on shared systems. Website design and content are handled by our web design team; your focus is the technical infrastructure, integrations, and systems that sit underneath. About the Role As our Marketing Systems Engineer, you own the technical layer of our marketing engine: the integrations and infrastructure that power our marketing website. The website is our primary lead capture system that turns traffic into pipeline. You own the systems that connect our marketing surface to Salesforce, Customer.io/Hubspot, and our broader GTM stack. You set the technical roadmap for this surface, move fast, measure impact, and treat reliability as a first-class concern rather than a cleanup task.You will also collaborate with Growth Engineering on shared infrastructure including handraiser routing and tracking. Responsibilities Own the technical health of the Mixpanel marketing website: page speed, WCAG compliance, Google Tag Manager, technical SEO, and third-party integrations including Qualified, Optimizely, and TrustArc. Own the integrations between the marketing website and our GTM stack t

JavaScriptPythonJavaAWS
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -82%

About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal and VMs dependable & scalable foundations for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Software Engineer, Host Assurance to build and operate the services, APIs, and host software that establish and maintain trust in our compute infrastructure. You will own production software from design and implementation through testing, rollout, observability, and operation. Your work will support capabilities such as machine identity, certificate issuance and enrollment, secure bootstrap, and host attestation across bare-metal and VM environments. Success in this role requires strong technical judgment, the ability to reason across software and host-system boundaries and learn unfamiliar parts of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—inc

AWSRestAIGo
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $293.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox’s daily active users growing at a record pace, we are building the next generation advertising in an immersive metaverse. Our team sits at the core of the Roblox Ads ecosystem, serving billions of impressions while providing the critical infrastructure that connects global brands with our massive community of creators. We are seeking a Principal Software Engineer to drive technical strategy and execution for the Ads Experience team. This team is responsible for the full spectrum of our monetization experience, including Advertiser Experience (Roblox Ads Manager), Brand Experience (Brand Ads & Integrations), and Publisher Experience (Creator Monetization & Commerce). This is a foundational technical challenge beyond standard ad-tech. You will architect the high-scale backbone of our monetization system, driving the technical vision for our Ads Manager, immersive Brand Ads formats, and creator solutions, while pioneering AI integration and balancing a three-sided marketplace economy in real-time. You Will Architect the Roblox Ads Manager: Design and build the end-to-end technical architecture for our flagship advertising platform, ensuring it can handle complex campaign st

TypeScriptReactAWSGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man

PythonSQLAWSGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,

AWSKubernetesCI/CDLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system

PythonSQLAWSLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Fleet team builds core components to enable productive research from small to state of the art scale across OpenAI, with the goal of accelerating progress towards AGI. We frequently collaborate with other teams to speed up the development of new state-of-the-art capabilities. About the Role As we scale up with more researchers and engineers joining OpenAI, we seek a pragmatic and passionate engineer with a strong focus on the development experience for both engineers and scientists. In this role, you will be responsible for building and maintaining systems that allow our research + engineering organization to iteratively develop, test, and deploy new features reliably, with high velocity, and with a frictionless and fast development cycle. You will help oversee and drive to the vision of how we should build, test and deploy software. You will drive the design of our continuous integration pipelines, testing infrastructure, training and support around our build system. Our current environment relies heavily on Python, Rust, and C++, which you will take ownership of and strive to transform into a state of the art development experience for research. Ultimately, your role will be to provide the necessary tools and metrics to support our fast-paced culture and ensure a stable, scalable platform for growth, while also fostering a seamless and low friction experience for OpenAI’s research. This role is based in San Francisco, CA. For a San Francisco role, we use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. You might thrive in this role if you: Have supported large monorepo development and deployment before Are a proficient Python programmer working in large monorepos Are proficient with Docker and Kubernetes Experienced in CI/CD About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boun

PythonAWSDockerKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Software Engineering team is responsible for designing and building the scalable, performant, and secure backend systems that power our products—from early prototypes to large-scale deployments. We collaborate closely with product, hardware, and full-stack teams to ensure our infrastructure enables fast iteration while setting a strong foundation for long-term growth. About the Role As a Backend Engineer , you will design and build services, APIs, and infrastructure that support evolving product needs. You’ll apply a deep understanding of backend systems and maintain enough end-to-end context—from hardware to cloud—to guide technical decisions that best serve the product and team. We’re looking for engineers who thrive in fast-paced, collaborative environments and care deeply about building robust systems that scale. This role is based in San Francisco, CA . We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Architect, build, and maintain high-performance, secure backend systems. Design APIs, data models, and infrastructure to support evolving product needs. Balance near-term development velocity with long-term maintainability and scalability. Collaborate with cross-functional teams to ensure cohesive, end-to-end solutions. You might thrive in this role if you: Have 7+ years of professional software engineering experience, with a focus on backend systems. Have a proven track record of building and scaling systems from early stage to large scale. Are proficient with Python and Go, and familiar with a range of server-side technologies. Have a strong grasp of system design, performance optimization, and security best practices. Can reason about full-stack tradeoffs from hardware through cloud infrastructure. (Nice to have) Have experience with distributed systems and cloud architectures. (Nice to have) Bring a background in instrumentation, analytics, and performanc

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Release Engineer team is responsible for building and maintaining the systems that power software delivery—from CI/CD pipelines and artifact management to release automation and fleet telemetry. We ensure software across bootloaders, firmware, operating systems, and cloud services is built reproducibly, validated rigorously, and released safely at scale. About the Role As a Release Engineer, you’ll design, build, and operate release infrastructure that enables reliable, secure, and traceable software delivery across complex multi-component systems. You’ll partner closely with embedded, cloud, and QA teams to ensure that every build—from development to OTA deployment—is fast, verifiable, and production-ready. We’re looking for engineers who take pride in automation, build reproducibility, and system reliability—and who enjoy building the connective tissue that allows hardware and software to ship together seamlessly. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and operate CI/CD pipelines for multi-component builds (bootloader, firmware, OS images, backend, companion apps) using hermetic toolchains. Define versioning and branching strategies; automate promotions, changelogs, and artifact retention. Integrate unit, integration, and hardware-in-the-loop (HIL) test results; quarantine flaky tests, auto-bisect failures, and block unsafe promotions. Build A/B OTA update flows with verity and health checks; run staged rollouts and canaries; implement safe rollback and roll-forward strategies. Implement code signing for binaries and firmware, generate SBOMs, run vulnerability scanning, and attach build attestations and provenance. Manage dashboards and alerts for build health, promotion latency, failure rates, and fleet update telemetry. You might thrive in this role if you: Have experience building and operating buil

PythonAWSCI/CDGit
TN
📍 New York, NY, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

The mission of The New York Times is to seek the truth and help people understand the world. That means independent journalism is at the heart of all we do as a company. It’s why we have a world-renowned newsroom that sends journalists to report on the ground from nearly 160 countries. It’s why we focus deeply on how our readers will experience our journalism, from print to audio to a world-class digital and app destination. And it’s why our business strategy centers on making journalism so good that it’s worth paying for. About the Role, Mission or Department Overview As a Senior Engineer, Marketing, you'll be embedded with the Marketing team, providing us with technical leadership, consulting, and systems thinking. You'll assemble the orchestration systems, integrations, and AI capabilities that help teams work faster, and in more data‑driven ways. You will lead end‑to‑end design, implementation, and management of the marketing campaign lifecycle system, integrations, internal tools, and learning/experimentation infrastructure. You will report to our Director, Advertising Systems. Responsibilities: You will write high-quality, secure, and well-tested code, contributing to standards and documentation for Marketing infrastructure, data models, and tools You will build the campaign lifecycle orchestration backbone You will build integrations and internal tools that connect project management, collaboration, creative, media, and email/lifecycle platforms You will work with Marketing partners as a technical advisor, translating needs into designs You will integrate AI agents and automations (including LLM-powered flows) into Marketing workflows according to company GenAI guidance You will design and operate data pipelines and models that ingest marketing and performance data into a data layer, supporting experimentation and analytics workflows You will deliver abstractions and APIs that ensure AI systems and our users to create campaign wrap-ups, insights, next-t

TypeScriptPythonReactNode.js
🔔

Get new ai infrastructure system engineer bangalore jobs in United States by email

Daily job updates · Unsubscribe anytime