Jobs in United States

Infrastructure Team Manager in United States

1,521 active opportunities · Updated October 2026

Explore current infrastructure team manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI’s 1P Hardware Systems team operates at the intersection of Hardware Engineering, Infrastructure, Supply Chain, Manufacturing, Deployment, and Finance to translate infrastructure demand into an executable systems plan. The Planning Lead connects system demand and deployment timing to site and configuration requirements, power availability, XPU needs, hardware supply commitments, manufacturing capacity, and site readiness. About the Role OpenAI is seeking a 1P Hardware Systems Planning Lead to own the integrated demand, supply, and deployment plan for our 1P AI infrastructure systems. This role will connect infrastructure demand and site power availability to system configurations, XPU requirements, and hardware supply commitments—creating a single, actionable view of whether our deployment plan can be met. You will establish the planning mechanisms that allow teams to see what changed, understand the impact, and act quickly when demand, supply, configuration, site readiness, or power timing moves. You will work across Hardware Engineering, Infrastructure, Supply Chain, Manufacturing, Deployment, Finance, and external partners to identify gaps early and drive recovery plans to closure. This is a highly cross-functional role for someone who combines systems-level thinking, strong analytical judgment, and rigorous program execution. The ideal candidate can turn complex and changing inputs into a clear operating plan, surface the decisions that matter, and drive accountability across teams without relying on formal authority. Success also requires strong communication judgment: the ability to align cross-functional teams, brief leadership at a concise and decision-oriented level, and go deep into the underlying assumptions, dependencies, risks, and recovery plans when needed. In this role, you will: Own the integrated planning process for 1P hardware systems, connecting system demand, deployment timing, site and configuration requirements, power ava

AWSRestAIGo
D
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -93.7%

$192K – $259.8K/yr

Quick readStrong listing-quality and freshness signals

Drata is building the trust layer between great companies - automating compliance, managing risk, and helping organizations prove trust continuously as they scale. We're Dratanauts: a global crew of 600+ professionals united by a culture that rewards integrity, ownership, and raising the bar, no matter where in the world we're working from. Why Join the Drata Team? At Drata, you're not maintaining legacy compliance software - you're building the agentic AI platform defining what trust looks like for the next generation of companies. Here's what makes the work itself worth showing up for: Problems without a playbook: You'll work at the edge of AI and security, building agentic governance, continuous compliance, and real-time trust verification to solve problems that don't have an established answer yet. You're writing it as you go. Real ownership, not just process: Our values center on owning outcomes and raising the bar, not checking boxes. You're expected to have opinions and back them. A seat at the table: Your perspective is unique and valued. Open debate and diverse viewpoints are built into how decisions actually get made here, at every level. Growth at rocketship speed: Drata is scaling fast, which means scope grows fast too. High performers get more ownership, visibility, and experience. A crew, not just coworkers: Dratanauts consistently describe a "come as you are" culture with sharp, curious people—the kind of team that makes hard problems genuinely fun to solve. See what they say here and follow us on LinkedIn for company news, employee stories, and career updates. Job Summary: Drata's AI Platform team builds the production infrastructure that powers AI features across our compliance platform — from MCP servers that make Drata's data available to AI agents, to LLM workflow orchestration that automates SOC 2, TPRM, and policy analysis. You'll own the systems that sit between our AI models and our customers: tool definitions that agents actually understand,

TypeScriptPythonNode.jsAWS
O
📍 United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

AWSRestAIRust
O
📍 United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside our cloud partners, infrastructure providers, and internal engineering teams, we operate hyperscale AI campuses that support the training and deployment of frontier AI models. The Site Operations team serves as OpenAI's on-site operational presence, helping ensure campuses operate safely, efficiently, and in alignment with Industrial Compute standards. We work closely with Hardware Operations, Infrastructure Delivery, Network Operations, Security, Facilities, Construction, and our infrastructure partners to support day-to-day site execution and maintain operational readiness. As Industrial Compute continues to expand globally, Site Operations plays a critical role in ensuring each campus is prepared to support reliable AI infrastructure at scale. About the Role We are seeking a Site Operations Technician to support the daily operation of Industrial Compute campuses. This role acts as OpenAI's on-site technical representative, helping coordinate activities across hardware operations, facilities, construction, logistics, security, and external service providers. You will perform routine site inspections, support asset tracking, coordinate vendor activities, assist with operational readiness, document site conditions, and help ensure infrastructure issues are identified and resolved quickly. The ideal candidate enjoys working in highly technical environments, is detail-oriented, and thrives in fast-paced operational settings where no two days are the same. Key Responsibilities Perform routine walkthroughs of Industrial Compute facilities to verify operational readiness and identify potential issues. Monitor site conditions and report abnormalities involving hardware spaces, network rooms, utilities, logistics areas, and common infrastructure. Support coordination of vendors, contractors, and partner organizations performing work o

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing

PythonAWSKubernetesRest
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -84.1%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Legal team plays a crucial role in advancing our mission by tackling innovative and fundamental legal issues in AI. The team includes professionals from diverse legal fields—technology, AI, infrastructure, privacy, IP, corporate, employment, tax, regulatory, and litigation—who collaborate closely with colleagues across the company. If you are passionate about being a technology lawyer working on cutting-edge challenges, you’ll thrive here. About the Role We are seeking an experienced attorney to serve as the commercial legal lead for Marketing. Based in San Francisco, you will be a primary legal partner to OpenAI’s rapidly growing global Marketing organization, supporting high-impact campaigns, creative production, talent and creator relationships, sponsorships, events, and the agreements and rights that make that work possible. You will work closely with Marketing, Communications, Partnerships, Procurement, Finance, Product, and colleagues across Legal, including Product, Privacy, Regulatory, and IP/Brand, to deliver practical advice at the pace of the business. This is a unique opportunity to help shape OpenAI’s Marketing efforts, negotiate sophisticated transactions, and build scalable legal frameworks for responsible global growth. We operate on a hybrid work model of three days per week in the office and offer relocation support for new employees. In this role, you will: Serve as the commercial legal lead for Marketing, partnering closely with Brand, Creative, Product Marketing, Design, Film and Photo, Performance Marketing, Partner Marketing, Communications, and regional teams from concept through launch. Draft and negotiate a wide range of marketing and entertainment agreements, including agency, production, talent and creator, sponsorship, event, media, content-licensing, marketing-technology, and vendor agreements. Structure and clear the rights needed for campaigns and content, including talent and appearance releases, publicity and

AWSGitRestAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -84.1%
Quick readStrong listing-quality and freshness signals

About the Team The Emerging Products team is a lean, high-output product lab group that builds products at the forefront of model capabilities. We collaborate across all teams within the company, from research and infrastructure to consumer products. The team is responsible for identifying new product opportunities, building them quickly, dogfooding them internally, and then launching the successful products to users. We use data, user research, and analytics to inform our ideas, and make decisions on what experiments are worth iterating, stopping, or scaling. About the Role We’re looking for a senior, product-minded software engineer to own ambiguous 0-to-1 work from idea through prototype, validation, and handoff. This is a full-stack role with a strong frontend and product emphasis: you will build the interfaces and supporting backend systems needed to test new experiences quickly, while making sound architectural choices that enable successful concepts to scale. This role is based in our Mission Bay office in San Francisco. In this role, you will: Build and ship high-quality, product experiments across the full stack. Turn ambiguous user needs and emerging technical capabilities into testable product concepts, using research and metrics to guide iteration. Own technical direction for 0-to-1 projects, balancing speed, reliability, and a clear path from prototype to scalable product. Partner closely with design, product, research, and engineering teams to dogfood, evaluate, launch, and transition successful experiments. You might thrive in this role if you: Have a track record of building and shipping end-to-end products in fast-moving, startup, founder-led, growth, or other high-ownership environments. Bring strong frontend engineering skills and enough backend and systems depth to make sound full-stack architectural decisions. Pair product intuition with evidence, using user research and product data to identify opportunities and make pragmatic tradeoffs. Operat

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. The Power Execution team owns the strategy and execution required to secure reliable, scalable, and economically resilient power for OpenAI’s global data center portfolio. The team sits at the intersection of commercial, technical, policy, legal, and operational work, partnering across OpenAI and with utilities, grid operators, regulators, counterparties, and public-sector stakeholders. About the Role The Energy Regulatory Lead will own energy regulatory strategy and execution for OpenAI’s infrastructure growth. This role will be the primary bridge between the Power Execution team and Public Policy and Government Affairs on energy regulatory matters, ensuring that OpenAI’s external engagement is grounded in project realities and that changing policy and regulatory conditions are translated into actionable infrastructure decisions. This is an individual contributor lead role and does not have direct reports initially. The role combines portfolio-level regulatory positioning with transactional regulatory work: evaluating jurisdictional pathways, supporting utility and energy transactions, coordinating approvals and filings, and helping project teams navigate tariffs, interconnection, load-service requirements, market rules, and regulatory risk from diligence through execution. In this role, you will: Develop and maintain OpenAI’s energy regulatory strategy across priority U.S. markets and, as needed, emerging geographies for infrastructure expansion. Coordinate closely with Public Policy and Government Affairs to shape energy regulatory priorities, engagement plans, messaging, and positions before utilities, public utility commissions, grid operators, state energy offices, and other relevant policymakers. Translate project requirements—load size, timing, reliability, cost, carbon, and expansion needs—into clear regulatory objectiv

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -84.1%

About the Team The Ona team at OpenAI is helping build the software factory for the enterprise. We build infrastructure that enables AI agents to work in secure, customer-controlled cloud environments, with the context, tools, and controls they need to make progress across the software lifecycle—beyond a single developer’s laptop or active session. Our focus is helping enterprises move from experimenting with agents to using them reliably in production. That means solving challenging problems in cloud environments, orchestration, security, and collaboration, while making the experience straightforward for the people directing and reviewing the work. We’re a team that values initiative, close relationships with customers, and exceptional engineering craft. We take ownership, learn quickly, and communicate directly and kindly. About the Role We’re hiring backend-focused Product Engineers across our platform and security product teams. You’ll build infrastructure and customer-facing workflows that let developers and AI agents work reliably in parallel. You’ll work primarily in Go on APIs, complex networking, development environments, and orchestration for long-running tasks. You’ll own outcomes from understanding a user’s problem and choosing an approach through shipping, operating, and improving the solution, working closely with frontend, infrastructure, and security engineers. In this role, you will: Work directly with customers to build developer and security workflows, from getting a project running to investigating findings, reviewing agent-generated changes, and verifying fixes. Build Go services and APIs for provisioning cloud environments, running agents in customer infrastructure, and integrating with source control, CI, and other developer tools. Design reliable orchestration for long-running, parallel work, including durable state, retries, cancellation, and recovery. Build security into execution workflows through clear permissions, credential handling, is

AWSAzureGCPKubernetes
S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.6%

$155K – $400K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Events Analytics Platform (EAP) team is responsible for the infrastructure that powers all of Sentry's time-series data and searching capabilities across billions of events with sub-second latency. We started this initiative by building Snuba, the primary storage and query service for Sentry's event data powered by ClickHouse, and we are now focused on unlocking deeper visibility and reporting across the terabytes of event data our users generate. As a Senior Software Engineer, you will lead efforts to push the boundaries of data visibility at Sentry. You will do this by expanding the capabilities of our search infrastructure, building new capabilities on top of our state-of-the-art storage layer and increasing the performance and integrity of Sentry’s core data services. You will also help shape Infrastructure's technical direction at Sentry and collaborate with Product and other Engineering teams to turn that vision into a reality. If you want to solve the hard problems that come with scaling event data into the petabyte range, this could be the job for you. In this role you will: Expand EAP's ability to deliver data at world-class speed and reliability. Architect and automate services and systems to scale reliably under growing demand. Make architectural trade-offs that balance product requirements with engineering constraints. Maintain and grow the team's code quality initiatives by regularly reviewing code and contributing to design decisions. Lead design and discussions around deliverables the team is working towards. Improve the maintainability and developer experience of the codebases EAP owns. Exa

PythonSQLPostgreSQLRedis
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Recursive Self-Improvement (RSI) team works across research, engineering, product, and infrastructure to build AI systems that accelerate and ultimately conduct high-quality research at OpenAI. We work to automate real research workflows and improve research productivity by building systems and feedback loops, designing evaluations, and training models to develop missing capabilities. Our work spans the full lifecycle of model training, evaluation, and deployment to help researchers move faster and tackle increasingly ambitious problems. About the Role We’re hiring research scientists , research engineers , and AI systems engineers to work on automating research at OpenAI. This role is based in San Francisco, CA. In this role, you will: Design evaluations for research judgment, hypothesis generation and testing, and long-horizon experiment execution. Turn real research workflows and model failures into data and evaluation flywheels. Improve model research capabilities through agent harnesses, synthetic data, RL environments, and model training. Build and maintain safe, reliable integrations between our models and OpenAI’s research infrastructure. Develop research agents, experiment-orchestration systems, and sandboxed runtimes that support real research workflows. Create metrics and economic models to understand RSI’s current and future effects on research productivity, model capabilities, and the safety of internal deployments. This is a high-ownership role for researchers and engineers who thrive in ambiguity, move fluidly between research and implementation, and turn emerging opportunities into rigorous, reliable, scalable results. You might thrive in this role if you: Have research or engineering experience across LLM training, model evaluations, agent systems, synthetic data, research infrastructure, or large-scale distributed systems. Are a strong generalist who can move between open-ended research and practical implementation, turning ambig

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team API Agents builds the shared agent harness, tools, and infrastructure that turn OpenAI’s frontier models into systems that can reliably complete real work. We carry the capabilities behind Codex into a much broader set of products and workflows across software engineering, research, finance, healthcare, enterprise operations, and more. Our work spans search and connected context, computer use, memory, delegation and multi-agent coordination, and safe execution. Sitting at the intersection of Research, Codex, infrastructure, and applied product teams, we build reusable agent capabilities that compound across the ecosystem. About the Role We are looking for an experienced backend software engineer to build the core systems behind the next generation of agents. You will design reliable services and abstractions that help agents find the right context, use tools and computers, retain knowledge, coordinate over long-running workflows, and take action safely. The role combines deep backend and infrastructure work with strong product judgment, with opportunities to work across agent runtimes, orchestration, search, execution environments, identity and permissions, observability, and evaluations. This is software and systems engineering rather than model training: success comes from strong backend fundamentals, high agency, and the ability to turn fast-moving research capabilities into dependable production primitives. In this role, you will: Design, build, and operate the shared agent harness and backend infrastructure that power long-running, high-value workflows across OpenAI and third-party products. Build reusable capabilities across search and connected context, computer use, memory, tool execution, delegation, subagents, and multi-agent orchestration. Establish the foundations agents need to operate safely in production, including secure execution environments, identity and permissions, observability, evaluations, reliability, and cost and latency effi

TypeScriptPythonAWSRest
O
📍 United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI’s Compute Strategy team is responsible for securing and scaling the core resources that power our research and products. We partner across engineering, finance, legal, and operations to identify, negotiate, and execute strategic partnerships that expand OpenAI’s capacity for compute, power, and data center infrastructure. Our mandate spans energy procurement, real estate development, colocation, cloud service providers, silicon and strategic supply chain, and infrastructure financing—ensuring OpenAI can grow with speed, resilience, and cost-efficiency. About the Role We are hiring several Business Development Lead, Compute Strategy positions focused on compute infrastructure. Each hire will bring deep expertise in one or more focus areas while collaborating across the broader infrastructure stack. In this role, you will source opportunities, structure partnerships, and negotiate high-value agreements across OpenAI’s infrastructure ecosystem. You will work directly with external partners and suppliers while collaborating internally with engineering, legal, finance, and operations to ensure we have the resources needed to support state-of-the-art AI systems. This role requires technical fluency, commercial judgment, and disciplined execution. Your work will directly shape how quickly, reliably, and efficiently OpenAI can bring new compute capacity online. Each hire will focus on building partnerships and executing deals in one or more of the following areas: Energy and Power: securing scalable and sustainable energy supply. Land and Real Estate: identifying and securing strategic sites. Colocation : evaluating and contracting for third-party data center capacity. Cloud Service Providers (CSPs): structuring partnerships with hyperscalers and specialized AI cloud providers. Silicon: building semiconductor partnerships to secure advanced silicon and resilient long-term supply. Fiber & Equipment: securing fiber & critical data center equipmen

AWSRestAIGo
🔔

Get new infrastructure team manager jobs in United States by email

Daily job updates · Unsubscribe anytime