About the Team The Workload team is responsible for designing and running OpenAI’s LLM training and inference infrastructure that powers frontier models at massive scale. Our systems unify how researchers train and serve models, abstracting away the complexity of performance, parallelism, and execution across vast GPU/accelerator fleets. By providing this foundation, the Workload team ensures that researchers can focus on advancing model capabilities while we handle the scale, efficiency, and reliability required to bring those models to life. About the Role We are looking for an engineer to design and implement the dataset infrastructure that powers OpenAI’s next-generation training stack. You will be responsible for building standardized dataset interfaces, scaling pipelines across thousands of GPUs, and proactively testing performance bottlenecks. In this role, you will collaborate closely with the multimodal researchers, and other infra groups to ensure datasets are unified, efficient, and easy to consume. In this role, you will: Design and maintain standardized dataset APIs, including for multimodal (MM) data that cannot fit in memory. Build proactive testing and scale validation pipelines for dataset loading at GPU scale. Collaborate with teammates to integrate datasets seamlessly into training and inference pipelines, ensuring smooth adoption and a great user experience. Document and maintain dataset interfaces so they are discoverable, consistent, and easy for other teams to adopt. Establish safeguards and validation systems to ensure datasets remain reproducible and unchanged once standardized. Debug and resolve performance bottlenecks in distributed dataset loading (e.g., straggler systems slowing global training). Provide visualization and inspection tools to surface errors, bugs, or bottlenecks in datasets. You might thrive in this role if you: Have strong engineering fundamentals with experience in distributed systems, data pipelines, or infrastructure.
Jobs in United States
Ai Infrastructure System Engineer Bangalore in United States
5,122 active opportunities · Updated October 2026
Showing
15 jobs
Explore current ai infrastructure system engineer bangalore jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Intelligence & Investigations Engineering team builds systems that detect, analyze, and disrupt abuse across OpenAI’s products. We partner closely with the Child Safety team and cross-functional groups to protect users while advancing OpenAI’s goal of developing AI that benefits everyone. About the Role As a Fullstack Engineer focused on child safety, you’ll build data-intensive, AI-powered applications and infrastructure that enable operators and investigators to work effectively and responsibly. You’ll adapt quickly in ambiguous, fast-moving environments to deliver well-crafted, reliable tooling for high-severity safety work. *Candidates should understand this role involves exposure to sensitive and egregious content. In this role, you will: 4+ years of experience as a software engineer Prototype, build, and maintain intelligence systems that detect, triage, and enable efficient human review of possible high severity harm Work hand in hand with operators and investigators, designing and delivering systems that enable them to do their work faster, more accurately, and more safely. Develop across the stack: UIs, services, pipelines, and anything else required to solve the problems we face. Interact with partners across Product Policy, Platform Integrity, Safety Systems, and Research Contribute to the team’s technical strategy, especially for child safety related tools and systems Report on impact in a data-driven fashion You might thrive in this role if you: Have a strong software engineering foundation and enjoy owning systems end-to-end—from infrastructure and data ingestion to frontend tooling Are energized by working at the frontier of AI capabilities, integrating new models and APIs into practical systems Have experience building and operating large-scale data pipelines or search/retrieval systems Are proficient in Python and/or TypeScript, and familiar with tools like Spark, Kafka, Flink, data warehouses, and SQL Take a product-minded ap
$175K – $215K/yr
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We’re looking for a Senior Software Engineer to join our Network Engineering team to accelerate building and scaling our innovative systems that support our growing identity platform. In this role, you will build the next-generation infrastructure that underpins all systems at CLEAR. The ideal candidate for this role will approach challenges with an eye toward reliability, simplicity, and scalability. What You'll Do: Develop and maintain a streamlined process for engineers to effortlessly build and deploy scalable and reliable software-defined networking solutions on AWS. Enhance our compute platform (Kubernetes) by integrating new functionalities and features, focusing on AWS networking services and concepts such as VPCs, Route Tables, Security Groups (SGs), ALBs/ELBs, and Route53, as well as implementing Kubernetes networking solutions like service mesh (Istio) to optimize service communication and management. Collaborate across engineering teams to advocate for and implement best practices in observability, utilizing tools like Splunk or Datadog to ensure robust network monitoring. Act as a product owner for our infrastructure, collecting feedback and requirements from engineering teams to address pain points and develop solutions, particularly in the realm of AWS networking and cloud-native design principles. What you're great at: 6+ years of extensive experience in infrastructure and platform development, particularly in software-defined networking and AWS cloud services. Proficient in writing production-grade softwar
We are hiring a Security Software Engineer to design and implement the hardware-backed security foundations used across OpenAI’s device ecosystem. A central focus of this role is hardening the boundary between our policy systems and the HSMs that protect sensitive cryptographic keys. This boundary determines which operations may be performed, what may be signed, which policies must be satisfied, and how changes to trusted software and policy are authorized. You will develop security-critical software and firmware within, or immediately adjacent to, an HSM trust boundary. Depending on your background, this may include HSM trusted applications, firmware services, cryptographic mechanisms, device drivers, PKCS#11 components, secure-provisioning protocols, or signing-policy enforcement systems. This is a hands-on software-engineering role. You will be expected to design systems, write and review production code, debug across hardware and software boundaries, and carry projects from initial requirements through deployment. It is not an HSM administration, PKI operations, compliance, or architecture-only position. In This Role, You Will Design and implement security-critical software and firmware for HSMs, secure elements, trusted execution environments, and hardware roots of trust. Build and harden the policy-to-HSM boundary responsible for authorizing certificate issuance and cryptographic signing operations. Develop HSM trusted applications, firmware components, host interfaces, device drivers, SDKs, or cryptographic service integrations. Implement or extend cryptographic interfaces such as PKCS#11, OpenSSL providers or engines, platform key-storage APIs, or comparable hardware-security interfaces. Build firmware and software that cryptographically enforces key generation, provisioning, usage, rotation, recovery, and destruction policies. Design and implement HSM-backed certificate authority, code-signing, key-management, and device-identity systems. Develop end-to-end
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We're building the data infrastructure behind some of the most demanding AI training workloads in the world, and we want sharp, curious people to help us do it. In this role, you'll build and maintain the high-performance data layer our Modeling teams rely on for training and evaluation jobs. As a Software Engineer, Data Infrastructure, you will: Work directly on petabyte-scale storage infrastructure, and the networking and performance challenges that come with it. Collaborate daily with researchers and engineers who are some of the best in the world at what they do. You may be a good fit if you have: 4+ years of experience working on data storage infrastructure Strong command of Python Kubernetes experience, especially on the storage side (Persistent Volumes, CSI drivers, etc.) The ability to transform unstructured data into performant datasets across diverse storage backends including S3, GCS, and POSIX Experience with distributed data processing frameworks such as Apache Beam, Spark, or Flink [Nice-to-have] Familiarity with modern analytics tooling such as BigQuery, Airflow, or dbt Genuine excitement about AI.
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides tools that help developers find and fix issues in their applications. The Developer Infrastructure team owns the systems that make every engineer at Sentry more effective at shipping high quality software: developer environments, CI/CD for our open source and closed source codebases, the golden path for building new services, our SDK and library publishing tooling, and the metrics and dashboards that keep it all healthy. Our goal is straightforward: developers should spend their time thinking about the software they build for customers, not the software they use to build it. That means everything from local and cloud development environments, to the CI/CD pipeline that gets code out safely and quickly, to the tooling that catches flaky tests before they cost someone a day. As AI coding agents become part of how engineers work, we're also investing in making sure our environments, CI, and deployment systems support that shift. As a Software Engineer on Dev Infra, you'll help build and scale this infrastructure end to end. In this role, you will Build and maintain the tooling that powers local and cloud-based developer environments Improve CI for our codebases and CD for our deployment pipeline, keeping both fast and reliable as the org and its infrastructure grow Define and evolve the golden path for building new services and libraries at Sentry Contribute to our SDK and library publishing tooling and release processes Build the metrics, dashboards, and flaky test detection that give the org visibility into deployment health and CI reliability Build tooling that gives engineers, and the AI codin
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role: This is a Principal Product Engineering role focused on Money Infrastructure at Replit. You’ll work on the financial backbone that powers how Replit earns money, how builders earn money, and how Agents transact. This role sits at the intersection of engineering, product, and the business. The systems you build directly impact revenue, trust, and some of the most critical user journeys on the platform. Getting them right enables growth, experimentation, and global scale. Getting them wrong creates broken payments, confusing pricing, and lost trust. We’re looking for engineers who can design and scale reliable financial systems while translating complex monetary logic into intuitive, user-friendly experiences for both Replit customers and builders on the platform. We love folks who have a passion for monetizing innovation and being a part of the greater pricing story. You will: Lead the design, architecture, and implementation of Replit’s core money infrastructure, spanning pricing, billing, payments, and monetization. Own and scale the global order-to-cash foundation supporting credit-based subscriptions, usage-based billing, marketplaces, in-app payments, and commerce for Agents. Enable rapid pricing and packaging experimentation across the company by building flexible abstractions and APIs for new SKUs, plans, and monetization models. Build high-converting, localized payment experiences across geographies — thinking globally while enabling users to pay locally. Power builder monetization by creating payment rails for apps, Agents, subscriptions, and new monetization primitives. Partner closely on data specifications with finance, accounting, and data teams to produce accurate, auditable, and reliable f
From $105K/yr
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We’re looking for an early career Software Engineer to join our Infrastructure team to accelerate building and scaling our innovative systems that support our growing identity platform. In this role, you will build the next-generation infrastructure that underpins all systems at CLEAR. The ideal candidate for this role will approach challenges with an eye toward reliability, simplicity, and scalability. What You'll Do: Develop and maintain a streamlined process for engineers to effortlessly build and deploy scalable and reliable software-defined networking solutions on AWS. Enhance our compute platform (Kubernetes) with new functionalities and features, focusing on AWS networking services and concepts such as VPCs, Route Tables, Security Groups (SGs), ALBs/ELBs, and Route53, optimize service communication and management. Collaborate across engineering teams to advocate for and implement best practices in observability, utilizing tools like Splunk or Datadog to ensure robust network monitoring. Act as a product owner for our infrastructure, collecting feedback and requirements from engineering teams to address pain points and develop solutions, particularly in the realm of AWS networking and cloud-native design principles. What you're great at: 0-2 years of experience in infrastructure and platform development and AWS cloud services. Proficient in Python, with understanding of Kubernetes and container orchestration tools like EKS and ECS. Understand AWS networking services, including VPC design, SGs, NATGWs, ALBs/ELBs, Rout
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Luau App Foundations team is responsible for the core infrastructure of the Roblox application. They bridge the gap between app and the performance-heavy Roblox game engine. This team is building the unified stack that powers mission-critical surfaces like Home, Avatar, and Social for millions of concurrent users. They are the ones who make it possible for "Web-style" development efficiency to exist within a high-performance C++ game engine. Why is this role exciting: Technical Pioneer: You will be writing libraries and modules using C++ inside a world class Roblox proprietary Game Engine. Systematic Impact: This is a "Force Multiplier" role. The frameworks and components you build will be used by dozens of other engineering teams to ship their features. Complex Problem Solving: You aren’t just building an app; you’re managing smooth data flow through a client that has to perform perfectly on a $100 Android phone and a $3,000 Gaming PC simultaneously. 0 to 1 Transitions: You will lead the charge in shaping some of the most crucial components of the App written in C++ inside the Game Engine. Key Challenges: Bridging Tech Stacks: The libraries you write sit between a modern UI (written in
About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a Mechanical Engineer to design, build, and own the mechanical side of our robotic actuator dynamometer and test infrastructure. You will create the test stands, couplings, fixtures, load paths, guarding, and serviceable lab hardware that enable rigorous characterization of robotic actuators. This role combines precision mechanical design with hands-on lab work. You will take robotic actuator test infrastructure from requirements and analysis through CAD, fabrication, assembly, commissioning, and iteration, partnering closely with electrical and software engineers to deliver safe, flexible, high-uptime test cells. In this role, you will Own the mechanical architecture of dynamometer and actuator test cells, including frames, bases, load paths, alignment, guarding, and serviceability. Design dynamometer structures, robotic actuator fixtures, load-motor mounts, couplings, shafts, bearings, adapters, and torque-reaction hardware. Translate robotic actuator test requirements into robust mechanical systems for torque, speed, thermal, durability, backdrive, efficiency, and failure testing. Perform first-principles analysis and simulation for stiffness, strength, fatigue, vibration, thermal growth, critical speed, and safety factors. Create precise, repeatable alignment strategies that protect test articles, load machines, sensors, and couplings. Design modular fixturing that supports rapid changeover across actuator and motor variants without compromising measurement quality. Work closely with electrical engineers on cable routing
From $126.4K/yr
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve our production infrastructure. This is a single application for Site Reliability Engineering opportunities across our Infrastructure Platforms department. Rather than asking you to choose the right team or level upfront, we evaluate your skills holistically and match you to the opportunity that best aligns with your experience
From $285.5K/yr
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . As a principal engineer on the Online Systems team, you’ll join a team that powers Pinterest’s most business-critical online systems at massive scale, driving the reliability, efficiency, and evolution behind every core Pinner and Advertiser experience. You'll lead major efforts like multi-region deployment and Kubernetes migration, set the standard for operational excellence, and define the long-term vision for our online serving infrastructure, supporting machine learning and product innovation across the company. This is an opportunity for high-impact technical leadership, broad visibility, and cross-functional influence at the heart of Pinterest’s platform. What you’ll do: Improve reliability, scalability and infra efficiency for Pinterest’s critical online systems across storage and caching, online service and realtime analytics syste
From $10K/yr
About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role The Financial Intelligence team, part of Ramp's Applied AI org, powers the agentic analysis, FP&A, and book-close experiences that over 70,000 businesses rely on to run their finances. We turn messy financial data into a fast, trustworthy, governed data layer that AI agents can work with and build artifacts on top of. This is a frontend-leaning role for an engineer who obsesses over interaction design and data-dense UX, but is just as comfortable designing the API that feeds it. You'll own the surfaces customers actually touch: agentic financial analysis, month-end close, and exploratory data experiences that feel as polished as the best business intelligence tools. If you want your frontend work to sit at the center of real production LLM systems rather than being bolted on at the end, this is the role. What You’ll Do Design and ship customer-facing AI experiences end to end, from React UI through to the APIs and data contracts behind them Build BI-tool-quality interfaces for agentic analysis, FP&A, and book-close workflows:
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a skilled and experienced Senior AI Engineer – AI Platform to join our ClickUp Engineering team. In this role, you will play a critical part in both building the core AI platform and directly applying large language models (LLMs) to deliver intelligent features across ClickUp. You will focus on backend systems that enable scalable, reliable, and secure AI-powered capabilities, while also working hands-on with LLMs to solve real user problems and drive product innovation. Key Responsibilities: Architect, design, and implement scalable AI platform services that support the deployment, orchestration, and lifecycle management of LLMs and other AI models. Apply LLMs and other AI technologies directly to build and enhance ClickUp’s intelligent features, working closely with product and engineering teams to deliver impactful solutions. Build and maintain robust APIs and backend systems that enable seamless integration of AI-powered features into ClickUp’s core platform. Develop infrastructure for model serving, monitoring, logging, and automated evaluation to ensure high reliability and performance of AI services in production. Integrate with multiple LLM providers (e.g., OpenAI, Anthropic, Google) and manage model selection, routing, and fallback strategies for optimal performance and cost. Drive the adoption of best practices in AI privacy, security, and compliance, including data anonymization, secure data handling, and regulatory adherence. Optimize platform performance, scalability, and cost-efficiency, leveraging cloud-native technologies and distributed systems. Stay curre
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a highly skilled Staff AI Engineer – AI Platform to join our ClickUp Engineering team. In this role, you will play a critical part in both building the core AI platform and directly applying large language models (LLMs) to deliver intelligent features across ClickUp. You will focus on backend systems that enable scalable, reliable, and secure AI-powered capabilities, while also working hands-on with LLMs to solve real user problems and drive product innovation. Key Responsibilities: Architect, design, and implement scalable AI platform services that support the deployment, orchestration, and lifecycle management of LLMs and other AI models. Apply LLMs and other AI technologies directly to build and enhance ClickUp’s intelligent features, working closely with product and engineering teams to deliver impactful solutions. Build and maintain robust APIs and backend systems that enable seamless integration of AI-powered features into ClickUp’s core platform. Develop infrastructure for model serving, monitoring, logging, and automated evaluation to ensure high reliability and performance of AI services in production. Integrate with multiple LLM providers (e.g., OpenAI, Anthropic, Google) and manage model selection, routing, and fallback strategies for optimal performance and cost. Drive the adoption of best practices in AI privacy, security, and compliance, including data anonymization, secure data handling, and regulatory adherence. Optimize platform performance, scalability, and cost-efficiency, leveraging cloud-native technologies and distributed systems. Stay current with ad
Other cities to consider
More places hiring for this role
Get new ai infrastructure system engineer bangalore jobs in United States by email
Daily job updates · Unsubscribe anytime