Jobs in United States

Infrastructure Team Manager in United States

1,504 active opportunities · Updated October 2026

Explore current infrastructure team manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. Our mission is to push the frontier of code generation and agentic reasoning, and deploy these capabilities in real-world products such as ChatGPT and the API, as well as in next-generation tools specifically designed for agentic coding. We operate across research, engineering, product, and infrastructure—owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. About the Role As a Performance & Systems Engineer on the Codex team, you will be responsible for whole-system optimization across a complex, evolving stack. Codex spans LLM inference, cloud orchestration, agentic work management, and multiple product surfaces. Your job will be to identify and land high-leverage changes—across infrastructure, modeling, and product layers—that make Codex agents significantly faster and cheaper to serve. We’re looking for generalists who thrive in ambiguity and love chasing performance bottlenecks to ground. This is a high-ownership role where your work will directly improve the experience of millions of users. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Hunt down and address inefficiencies across the Codex system stack, from agent behavior to LLM inference to container orchestration, and beyond. Build tooling to measure, profile, and optimize system performance at scale. Collaborate with researchers and engineers to land high-ROI changes that improve latency and cost. You might thrive in this role if you: Have experience operating across both ML systems and cloud infrastructure. Enjoy diving into messy, ambiguous problems and emerging with clear wins. Think holistically about performance, balancing spee

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s Tax and Trade team sits at the center of OpenAI’s global growth, shaping how cutting-edge AI products, partnerships, and infrastructure scale across borders while navigating complex tax, trade, and regulatory regimes. We operate as strategic operators, not just compliance experts, embedding early in product, finance, policy, and infrastructure decisions to manage risk, unlock incentives, and enable OpenAI to grow responsibly and competitively worldwide. About the Role We’re hiring a Tax Director to help shape some of OpenAI’s most important business, product, infrastructure, and financing initiatives. This role will sit at the intersection of tax, finance, technology, and commercial strategy, advising on the investments, products, partnerships, and operating decisions that enable OpenAI to scale. You will serve as a senior tax business partner to Finance, Legal, Infrastructure, Treasury, Accounting, Product, and commercial teams. Your work will span compute and infrastructure investments, structured financings, commercial contracting, product strategy, new business models, and international expansion. This is a broad, highly commercial role for someone who enjoys solving new problems and influencing business decisions. The ideal candidate combines strong tax judgment with financial acumen, intellectual curiosity, and an understanding of transaction economics, contractual risk allocation, and business operations. You will engage early in strategic initiatives, help teams evaluate possibilities, and develop practical structures that support OpenAI’s growth. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. In this role, you will: Serve as a senior tax business partner across OpenAI, advising on strategic business decisions, product development, commercial models, infrastructure investments, financing initiatives, and international growth. Serve as the lead tax a

AWSGitRestAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $326.1K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Security Software Engineer on the Production IAM team, you will set the technical direction for how identity and access work across Roblox's production infrastructure, from the mTLS-based identity that services use to authenticate to one another, to the privileged access controls that govern how engineers reach production. The team is accountable for Roblox's machine and workload identity platform, its centralized authorization engine, its production access management platform, production PKI and certificate lifecycle, and just-in-time privileged access for engineers. As an individual contributor in Production IAM, you will define multi-year strategy, drive alignment across Roblox Platform, mentor senior and staff engineers, and personally build the hardest parts of these systems. As AI agents become first-class actors in production, you will also help pioneer how they get identity, prove who they are, and receive safely-scoped access. You will Lead the architecture for production identity and access. Define and evolve the end-to-end design for machine, workload, human, and AI-agent identity across our hybrid on-prem and cloud fleet, making secure access invisible when

PythonJavaAWSGit
O
📍 United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. The Industrial Compute team is central to this mission, setting the core infra strategy and implementing this vision. From site selection to the buildout process, this team sits at the intersection of commercial, technical, strategy, and operations, interacting with teams and executives inside and outside of OpenAI. About the Role The Community Engagement Lead will be the primary bridge between OpenAI and the communities in Ohio. This role ensures that OpenAI builds strong, trust-based relationships with local stakeholders, communicates proactively about our projects, and integrates community priorities into our development approach. The role spans engagement, communications, and reputation management, and will partner closely with the Economic Development and Environmental leads. Key Responsibilities Build and maintain relationships with local leaders, community organizations, NGOs, and residents. Develop and execute community engagement strategies for new and existing sites. Represent OpenAI in public forums, hearings, and community events. Partner with the Economic Development Lead on incentive compliance and community benefits. Partner with the Environmental Lead on communicating environmental stewardship and sustainability efforts. Develop proactive communications to address concerns, highlight benefits, and reduce risk of opposition. Monitor community sentiment and advise executives on risks and opportunities. Create a community engagement playbook that can scale across geographies. Qualifications 8+ years in community affairs, public engagement, or corporate communications. Proven track record engaging diverse community stakeholders for large infrastructure or technology projects. Strong public speaking and facilitation skills. Ability to manage sensitive political and reputational issues. Experienc

AWSRestAIGo
O
📍 United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. The Power & Land team owns the energy strategy required to secure reliable, scalable, and economically resilient power for OpenAI’s global data center portfolio. About the Role The Power Trading Lead will own commodity hedging strategy and execution across OpenAI’s data center power portfolio. This role will translate large, dynamic electricity and fuel exposures into practical hedging, procurement, and risk-management strategies that protect infrastructure economics while preserving flexibility for growth. This is an individual contributor lead role and does not have direct reports initially. The role will work across power markets, utility tariffs, retail and wholesale supply structures, natural gas and power hedges, renewable and clean firm products, and portfolio risk analytics to support long-term compute growth. Key Responsibilities Develop and maintain OpenAI’s commodity hedging strategy across electricity, natural gas, and related energy exposures for data center operations and growth. Quantify portfolio exposure by market, site, load shape, tenor, tariff, and supply structure, and translate that exposure into clear hedging recommendations. Evaluate and execute hedging structures including fixed-price supply, forwards, swaps, options, retail supply products, congestion and basis risk mitigation, and related instruments where appropriate. Partner with utilities, suppliers, traders, banks, consultants, and market counterparties to source competitive products and improve risk-adjusted energy economics. Build decision frameworks for when to hedge, how much to hedge, and which risks to retain across different stages of site development, construction, and operations. Coordinate with finance, treasury, legal, procurement, energy regulatory, sustainability, and site-readiness teams to ensure hedging strategy aligns with broa

AWSRestAIGo
O
📍 United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. The Site Readiness & Development team owns the upstream diligence and development work required to convert powered-land opportunities into executable infrastructure options. About the Role The Power Land Developer; Development & Power Projects will support the advancement of powered-land opportunities across the portfolio, reporting to the Development & Power Project Lead. This is an individual contributor lead role and does not have direct reports initially. This role will own day-to-day development execution across approximately 1–3 active sites depending on complexity, including utility applications, interconnection requests, study follow-through, queue milestones, consultant workstreams, schedule and risk management, and readiness handoff for opportunities advancing to deeper diligence, commercial commitment, or execution. The role translates study and utility findings into clear schedule, budget, and risk implications for portfolio decisions. Responsibilities Manage day-to-day advancement of powered-land opportunities from intake through readiness. Own utility application and interconnection follow-through across active powered-land opportunities, including utility inputs, study responses, and queue milestones. Scope and manage external studies and consultant workstreams needed to validate readiness. Track development schedules, dependencies, risk registers, site-readiness milestones, and key budget implications. Translate technical findings into clear schedule, budget, and risk implications and escalate material blockers or trade-offs to the Development & Power Project Lead. Partner with land, environmental, economic development, and community affairs leads to align project advancement. Support expansion planning, infrastructure sequencing, and long-term power-risk mitigation. Define readiness criteria and

AWSRestAIRust
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i

PythonReactNode.jsAngular
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

PythonAWSLinuxRest
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Release Engineer team is responsible for building and maintaining the systems that power software delivery—from CI/CD pipelines and artifact management to release automation and fleet telemetry. We ensure software across bootloaders, firmware, operating systems, and cloud services is built reproducibly, validated rigorously, and released safely at scale. About the Role As a Release Engineer, you’ll design, build, and operate release infrastructure that enables reliable, secure, and traceable software delivery across complex multi-component systems. You’ll partner closely with embedded, cloud, and QA teams to ensure that every build—from development to OTA deployment—is fast, verifiable, and production-ready. We’re looking for engineers who take pride in automation, build reproducibility, and system reliability—and who enjoy building the connective tissue that allows hardware and software to ship together seamlessly. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and operate CI/CD pipelines for multi-component builds (bootloader, firmware, OS images, backend, companion apps) using hermetic toolchains. Define versioning and branching strategies; automate promotions, changelogs, and artifact retention. Integrate unit, integration, and hardware-in-the-loop (HIL) test results; quarantine flaky tests, auto-bisect failures, and block unsafe promotions. Build A/B OTA update flows with verity and health checks; run staged rollouts and canaries; implement safe rollback and roll-forward strategies. Implement code signing for binaries and firmware, generate SBOMs, run vulnerability scanning, and attach build attestations and provenance. Manage dashboards and alerts for build health, promotion latency, failure rates, and fleet update telemetry. You might thrive in this role if you: Have experience building and operating buil

PythonAWSCI/CDGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI's Industrial Compute organization is responsible for planning, delivering, operating, and optimizing the compute infrastructure that powers frontier AI. As OpenAI scales toward becoming an intelligence utility, Industrial Compute coordinates a complex lifecycle spanning infrastructure strategy, capacity planning, provider partnerships, fleet operations, product demand, and financial planning. The organization manages one of the largest and fastest-growing compute footprints in the world, where decisions around capacity allocation, deployment readiness, utilization, reliability, and product demand directly impact product availability, customer experience, and business performance. The Capacity Systems team builds the software platforms, data systems, and automation frameworks that connect these functions into a shared operating model. We transform fragmented planning workflows into scalable systems that enable teams to understand what compute was contracted, delivered, healthy, allocated, and ultimately converted into business and research outcomes. About the Role We are seeking a Capacity Systems Software Engineer to build the platforms and services that power Industrial Compute planning, forecasting, optimization, and operational decision-making. In this role, you will design and develop software systems that connect infrastructure delivery, fleet health, capacity allocation, demand forecasting, deployment readiness, financial planning, and product consumption into a unified system of record. Your work will help OpenAI make better decisions about where compute should be deployed, how capacity should be allocated, and how infrastructure investments translate into business value. You will partner closely with Capacity Planning, Fleet Operations, Infrastructure Engineering, Product, Finance, Supply Chain, and Strategic Sourcing teams to replace spreadsheet-driven workflows with scalable software systems that enable visibility, automation, and dec

TypeScriptPythonJavaSQL
G
📍 Austin, Texas, United States· Full-time
✓ Quality checked

Join the Team at Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore brings together deep expertise to solve complex problems and deliver meaningful progress in AI compute. Ready to raise the standard for supplier quality in advanced AI manufacturing? Apply now to be part of the journey. Job Summary We are seeking a highly motivated individual with experience in managing supply chain quality in a high-tech manufacturing organisation. The ideal candidate will be hands-on, comfortable with ambiguity and able to manage multiple projects and stakeholders simultaneously. Working across our entire supply chain in a technically challenging, fast paced environment, you will ensure the readiness of our global supply chain to ramp successfully as we develop and manufacture the world’s most advanced AI systems and services. The Team The Graphcore Quality Team delivers outstanding customer experience and champions sustainable excellence throughout Graphcore. We are responsible for customer, supply chain and product quality, as well as organisational compliance, governance and assurance. You will be joining a diverse

AIGoExcelSEM
N
📍 Remote, United States· Remote
✓ Quality checkedCompany trend -12.7%

NVIDIA is looking for an experienced software engineer with infrastructure experience to become a senior member of the Cloud Foundations Automation - Development Team. We build and manage the automation ecosystem supporting NVIDIA's GPU Cloud and NVIDIA SuperPod deployments. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most hard-working and dedicated people on the planet working for us. If you're creative and autonomous, we want to hear from you! What you'll be doing: Developing software to enable efficient network design, deployment and day 2 management. Building product focused software solutions, used by internal and external customers. Helping us as we transform our workflows and organization into a centrally orchestrated configuration management framework, operating at scale across geographies. Owning and driving integrations with various service APIs such as Cloud Service Providers, to automate creation of environments and auto populate data sources in turn. Building on open source software, designing and implementing data structures and UI interfaces to automate processes from equipment purchase to device config generation to deployment to operations. Streamlining deployment mechanisms and life cycle operations Developing modern service architectures around streaming data and event pipelines. Working with infrastructure domain experts on true, zero touch deployment solutions and utilizing best of breed high performance computing management solutions. Be a proactive problem solver, looking out for new opportunities to improve our services and customer experience. Communicate readily with your peers across the organization, b

PythonKubernetesAI
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -12.7%

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys

PythonJavaKubernetesLinux
S
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role Supabase's Partnerships function is scaling fast - spanning Technology Partners, Solution Partners, Startups, and Cloud Partnerships. We're looking for a Partner Operations & Systems Lead to own the operational and technical backbone that lets this team move quickly and make decisions with good data. This is a senior, hands-on role. You'll design the systems, processes, and reporting that the whole partnerships org runs on - and, increasingly, you'll build the internal tools yourself. Supabase is investing heavily in AI-assisted development to move faster than a traditional ops build cycle allows, and this role is expected to be a leading example of that inside Partnerships. You'll report directly to the Head of Partnerships and sit alongside our regional and functional leads as a peer, with the mandate to build and enforce operational rigor across all teams. Beyond the infrastructure, we want someone who helps shape where Partnerships places its bets - not just the system that reports on them. What You'll Own Systems & Data Infrastructure Own the partner tech stack end-to-end - CRM partner objects, attribution tooling, and partner-facing portals. Design and maintain partner attribution models and the dashboards leadership uses to evaluate performance. Ensure data hygiene and consistency across partner records, deals, and touchpoints spanning all sub-functions. Process & Program Management Design and run partner onboarding, tiering, and certification programs. Own the RFC / DRI decision-making framework for the partnerships team, including how it's used and refined over time. Run deal-registration and "quarterback" account-ownership processes that keep internal teams

O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s mission is to ensure that general-purpose artificial intelligence benefits all of humanity. The Payments team works across product, engineering, design, and finance to build the financial infrastructure that makes OpenAI’s products accessible to consumers and enterprises around the world. As AI introduces new ways for people and organizations to work, the team is defining how to support and monetize emerging forms of product usage, from usage-based pricing to agentic work. We’re building the foundational systems that help OpenAI products deliver clear, reliable, and scalable payment experiences while ensuring that this powerful technology is deployed responsibly. About the Role In this role, you’ll lead design for one of OpenAI’s most foundational product areas: the payments and monetization infrastructure that supports our consumer and enterprise products. You’ll partner closely with product, engineering, and cross-functional teams to shape how customers understand, manage, and pay for entirely new kinds of AI usage. Your work will extend beyond traditional checkout and billing. You’ll help define the systems, frameworks, and experiences behind durable pay-as-you-go models, Codex usage, and agentic workflows, translating complex business and technical requirements into intuitive experiences. As a product designer in a highly ambiguous and rapidly evolving space, you’ll influence both product strategy and the underlying infrastructure that OpenAI products depend on. This role is based in our San Francisco HQ. We offer relocation assistance to new employees. In this role, you will: Lead the design direction for foundational payments, billing, and monetization experiences across OpenAI’s consumer and enterprise products. Design and ship high-quality, end-to-end product experiences, from early systems and interaction concepts to high-fidelity prototypes and production-ready designs. Shape the infrastructure and product frameworks that support em

AWSRestAIGo
🔔

Get new infrastructure team manager jobs in United States by email

Daily job updates · Unsubscribe anytime