Jobs in United States

Ai Infrastructure System Engineer Bangalore in United States

5,418 active opportunities · Updated October 2026

Explore current ai infrastructure system engineer bangalore jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As the Head of IT at Baseten, you will build, scale, and secure our internal technology function to support our rapid growth. Reporting to our Chief Information Security Officer, you will lead and mentor a team of 5+ IT engineers, leading the charge to transition Baseten from startup-era IT to a highly automated, enterprise-ready IT organization. You will take full ownership of corporate IT infrastructure, Helpdesk operations, corporate identity management, device lifecycles, and vendor procurement. As we scale to support the world’s most dynamic AI companies, you will ensure our internal systems scale seamlessly with our headcount, providing a secure, frictionless, and world-class technology experience for all Baseten employees. RESPONSIBILITIES Team Leadership: Manage, mentor, and grow a team of IT engineers, fostering a high-performance culture focused on technical excellence and end-user satisfaction. Helpdesk Operational Excellence: Build a fast-response support function by establishing clear response SLAs, tracking employee satisfaction metrics, and formalizing on-call and incident response processes. Zero-Touch Automation: Architect and implement automated employee onboarding, offboarding, and role-based access changes through deep integrations across HRIS, MDM, and IAM systems. SaaS Management & Procurement: Establish comprehensive SaaS management processes to eliminate shadow IT, automate access w

Machine LearningAIGoRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Container runtimes were designed for general-purpose software workloads. AI inference is not a general-purpose workload. Running large models at production scale exposes cracks in every layer of the container stack: runtimes unaware of GPU memory constraints, images that take minutes to pull when a model needs to scale to thousands of replicas, and isolation mechanisms that weren't designed for the multi-tenant serving environments that production AI requires. The tools the industry has relied on for a decade weren't built for this, and patching around those limitations at higher layers only goes so far. Baseten owns the entire pipeline, from the moment a developer pushes a model to the moment a request gets a response. That vertical ownership means we can fix these problems at the root. The Runtime Fabrics team is doing exactly that: purpose-building the container runtime and storage layers for AI inference workloads, led by some of the world's top containerd maintainers. As Engineering Manager of the Runtime Fabrics team, you will lead this work, setting technical direction, growing a world-class team of systems engineers, and ensuring the team's output shapes not just Baseten's infrastructure but the open-source container ecosystem at large. If you've contributed to containerd, runc, or related OCI projects and are ready to lead a team solving some of the hardest problems in infrastructure today, we'd love

LinuxMachine LearningAIC++
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $295.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Engineering Manager, Home Infrastructure The Home Infrastructure team builds the mission-critical backend and data systems that power Roblox’s Homepage and Experience Details Page, two of the highest-traffic surfaces on Roblox. These surfaces reach the vast majority of Roblox’s daily active users and are core drivers of discovery, engagement, retention, and platform growth. We are a full-stack product infrastructure team responsible for content distribution across Roblox. Our systems support multiple modes of user interaction, including exploratory browsing, directed discovery, and personalized content recommendations across the many types of content that make up the Roblox ecosystem. This team sits at the intersection of large-scale distributed systems, machine learning-powered personalization, data infrastructure, and product experimentation. We partner closely with Machine Learning, Data Science, Product, Design, Frontend, Ads, Marketplace, Virtual Economy, and other teams across Roblox to build the platforms that help users find the most relevant and engaging content. As Engineering Manager for Home Infrastructure, you will lead a team of Backend and Data Engineers responsible for the e

AWSGitMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. PRODUCT AT BASETEN Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the people who defines it. You'll work directly with our founders and with some of the best systems and infrastructure engineers in the world, and you'll set the standard for what product looks like here. You earn trust by being technical, finding the truth in front of customers, building great cross-functional relationships, and shipping great product experiences. THE ROLE The largest, most demanding Enterprises are starting to run on Baseten and they come with a range of security, compliance, and procurement requirements. Today that readiness is assembled deal-by-deal. You'll own the enterprise-readiness surface end to end and turn it into product: the deployment options customers can choose and buy, compliance posture they can trust, access and security controls their IT teams require, and the billing and spend controls their finance teams expect. What does a complete Baseten Enterprise Product offering look like? RESPONSIBILITIES Drive the Enterprise Readiness customer experience end to end: Partner with GTM and Enterprise Engineering to make "enterprise-ready" a platform-wide capability, not a deal-by-deal scramble. Outcome: readiness becomes a supported, priced product instead of bespoke work asse

ReactMachine LearningAIGo
D
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $100K/yr

Quick readStrong listing-quality and freshness signals

Datadog is seeking curious, driven interns to join our Product Management team and help build products that improve how engineers monitor and understand their systems. As a Product Management Intern, you'll support the product development lifecycle by partnering closely with Engineering, Design, and Product Marketing to bring new ideas and features to life. You'll gain hands-on experience working on products that serve highly technical customers while contributing to meaningful business and user outcomes. Interns are embedded directly within product teams, working on meaningful initiatives alongside full-time Product Managers and contributing to actual product decisions. Our platform processes over 100 trillion events per day across 30,000+ customers in a multi-cloud environment -- giving you direct exposure to large-scale, real-time systems built by engineers, for engineers. It's an environment where you'll develop product thinking, technical communication, and cross-functional collaboration skills by doing the work, not just observing it. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Conduct customer discovery conversations and gather feedback to better understand user needs Drive product initiatives from concept through launch alongside Engineering, Design, and Product Marketing teams Translate customer and business needs into clear product requirements and engineering priorities Analyze customer feedback, product data, and market insights to help inform product decisions Prepare and deliver technical product demonstrations and communication materials Develop technical understanding of Datadog’s observability platform and cloud infrastructure products Who You Are: Targeting a 2028 full-time graduation or start date Pursuing a degree in

GitRestAIGo
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role We're looking for an Engineering Manager to lead a team of highly experienced engineers building the infrastructure that powers Modal's serverless GPU platform. This is a hands-on leadership role — expect to split your time between technical contribution and people management depending on what the team needs. You'll set direction, remove blockers, and build a strong engineering culture as your team tackles hard problems in distributed computing, large-scale data handling, and performance optimization. Who You Are You're an experienced engineering leader who stays close to the work and builds alongside your team when it counts. You earn trust through technical depth, not title. You communicate clearly, help strong engineers move fast without cutting corners, and stay calm and pragmatic under pressure. You care as much about how your team gets to an answer as the answ

JavaLinuxAIC++
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $192K/yr

Quick readStrong listing-quality and freshness signals

Coordination Systems provides foundational distributed systems building blocks for internal Datadog platforms. Our services cover sharding, consensus, resource protection, configuration distribution, and much more. We are looking for a manager to lead the Coordination Systems - Storage team. This team provides essential configuration storage and distribution systems that are depended upon by almost every service and pod at Datadog. We power critical runtime configuration (e.g. feature flags), complex control planes (e.g. dynamic sharding configuration), and much more. Storage is one of four subteams within Coordination Systems. If successful, the candidate will have opportunities to lead other growing and impactful areas such as Resource Protection. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: (Describe role responsibilities here/max 6 bullets) Lead a core team of 5 engineers (distributed, with majority in NYC) Lead ceremonies, prioritize and delegate project Stay hands-on with the code, e.g. isolated features, small remediations, investigation follow ups Stay actively involved in operations, incidents, root cause analysis, etc. Constantly promote a culture of operational excellence, organizing gamedays, conducting operational reviews, staying proactive with reliability Who You Are: (Describe role qualifications here/max 6 bullets) Strong distributed systems skills, able to understand and account for a variety of failure modes, well-versed in end-to-end o11y, validation testing, simulation setup, etc. Worked on platform teams before, providing critical infrastructure to internal stakeholders Experienced in handling significant incidents, both as a responder and follow-up ow

AIRustExcelSEM
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. The Systems Integration team is critical in this mission, turning complex hardware-software development into reliable product signals. We build the shared infrastructure, tooling, and lab environments that let teams test quickly, understand failures, and ship with confidence. About the Role As a Systems Integration Manager , you will lead the team responsible for device validation infrastructure, test automation, developer tooling, and lab operations. This is a player-coach leadership role: you’ll set technical and operational direction, build and develop a team of engineers and lab operations professionals, and stay close to the architecture and hardest systems problems. You will partner closely with device software, OS, firmware, hardware, reliability, QA, and release infrastructure teams to define validation strategy, improve release readiness, and ensure our test environments and quality signals scale with the product. Because this is a new category of devices, you’ll have the rare opportunity to build the validation foundation early—shaping the systems, standards, and operating model that will support products from prototype through launch. We’re looking for a leader who combines strong technical judgment with people leadership, operational rigor, and experience building reliable systems for complex hardware-software products. This role is b

Artificial IntelligenceAI
N
📍 Redmond, United States
✓ Quality checkedCompany trend -13.7%

NVIDIA’s EDA Infrastructure organization builds and operates the systems that support chip development. We are looking for an engineering manager to lead the team responsible for operational processes and platforms across incident management, maintenance, on-call, issue management, and customer-serving readiness. You will own the roadmap and delivery, from defining how teams work to building the tools they use. You will partner with infrastructure and service owners to improve reliability, reduce manual work, and ensure services are ready to support customers. Your team will use automation, AI, and lessons from operational events to drive improvements. What you’ll be doing: Lead a team and own the roadmap for operational processes and platforms, from requirements and delivery through adoption and results. Set technical direction, prioritize work, and guide execution across engineering and operational disciplines. Partner with infrastructure, product, and security teams to establish consistent practices for incident response, maintenance, on-call, issue management, and customer-serving readiness. Hire and develop engineers and technical leads, building a team with clear ownership and accountability. Align priorities across teams, communicate progress and risks, and provide technical leadership during major incidents. What we need to see: <span style="co

Artificial IntelligenceAI
C
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%

From £135K/yr

Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Who are we? Cohere is at the forefront of AI innovation, building cutting-edge language models and AI systems. Our Finance team plays a critical role in supporting our rapid growth and ensuring operational excellence. We value technical expertise, collaborative problem-solving, and a commitment to accuracy and compliance. This role offers the opportunity to make a significant impact on our financial infrastructure while working at the intersection of finance and transformative technology. As we scale our business and navigate an evolving financial landscape, you'll play a pivotal role in driving operational excellence and serving as a key partner on accounting implications of major business decisions. In this role, you will: Own the full lifecycle management of our NetSuite ERP environment, including architecture design, configuration optimization, integration development, and strategic roadmap planning to support General Ledger, AP, AR, Procure-to-Pay, Fixed Assets, Revenue Recognition, Audit, and International Consolidations & Reporting Drive end-to-end process excellence across Order-to-Cash and Procure-to-Pay workflows,

GitAIGoRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. The Systems Integration team is critical in this mission, turning complex hardware-software development into reliable product signals. Lab Operations is the physical backbone of that work: we build and maintain the device fleets, test environments, and hardware-in-the-loop labs that let teams test repeatably, understand failures, and ship with confidence. About the Role As a Lab Operations Manager, Systems Integration , you will own the day-to-day operation of a large-scale consumer device test lab. This is a hands-on operations leadership role: you’ll keep device fleets, test rigs, lab infrastructure, inventory, provisioning, maintenance, and logistics running smoothly so engineers and QA technicians have reliable environments for validation and release testing. We’re looking for someone who is highly organized, technically hands-on, comfortable with consumer electronics and lab equipment, and experienced operating complex physical test environments at scale. Because this is a new category of devices, you’ll have the opportunity to build the lab operating model early—shaping the systems, standards, and workflows that support products from prototype through launch. In this role, you will: Own device fleet and inventory: Manage configuration, deployment, tracking, lifecycle, and accurate asset records for a large fleet of consumer devices and test

Artificial IntelligenceAILogistics
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $295.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox Ads business growing at a rapid rate, we are building large scale ads machine learning infrastructure to deliver effective performance ads to our users, and more business values to our advertisers. We’re looking for an EM to lead a team of exceptional ML infrastructure engineers, build scalable, reliable, and high-performance infrastructure that powers ML systems across our organization. You’ll operate at the scales of hundreds of billions of engagements, and redefine how we deliver performance ads to hundreds of millions of users. You Will: Lead strategic planning and roadmap execution of scalable production-ready ML systems including model training, data pipelines, feature engineering and model inference. Own the architecture, establish engineering best practices of scalability, reliability, and cost-effectiveness of ML infrastructure (e.g., training, serving, feature). Work closely with data scientists, ML engineers, platform teams, and product stakeholders to design, implement, and operate robust ML platforms that accelerate model development and deployment. Recruit, mentor, and grow a high-performing team of ML infrastructure engineers. You Have: 5+ years of experienc

AWSGitMachine LearningAI
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an

KubernetesGitAIGo
P
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, United States· Full-time· Remote
✓ High-confidence listingCompany trend -86.4%
Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is seeking a Staff Software Engineer, Capacity Engineering. The team is responsible for efficiently managing one of the largest-scale cloud-native infrastructures in the world. This role is highly impactful, as efficiency is an ongoing strategic priority for Pinterest. The role has direct visibility across Pinterest Engineering and with Engineering and company leadership. The team is looking for a candidate with a strong background in implementing performance and efficiency projects on large scale distributed systems. In this individual-contributor role you will own and drive performance and efficiency for a core area of Capacity Engineering, partnering with the company-wide efficiency lead and collaborating with performance and efficiency leaders across the organization. What you’ll do: Drive efficiency in large-scale shared environme

PythonJavaAWSKubernetes
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by leading the design of high-performance, scalable and reliable machine learning systems? Do you want to set technical direction and help shape the next generation of AI platforms powering advanced NLP applications? We are looking for a Lead Member of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will provide technical leadership across multiple teams, driving the architecture and strategy for deploying optimized NLP models to production in low latency, high throughput, and high availability environments. You will serve as a key point of contact for customers, leading the design of customized deployments to meet their specific needs, and mentoring engineers to raise the technical bar across the team. You may be a good fit if you have: 8+ years of engineering experience running production infrastructure at a large scale, with a track record of technical leadership Demonstrated experience leading the architecture

AWSAzureGCPKubernetes
🔔

Get new ai infrastructure system engineer bangalore jobs in United States by email

Daily job updates · Unsubscribe anytime