About the Team OpenAI is helping build the infrastructure that powers the next generation of artificial intelligence. Through Stargate, we are developing and operating large-scale AI compute campuses that require world-class execution across data center design, construction, commissioning, and operations. The Infrastructure Operations team is responsible for bringing AI infrastructure online and ensuring it operates reliably at scale. We partner closely with hardware, network, deployment, construction, and operations teams to deliver mission-critical environments capable of supporting frontier AI workloads. As our footprint expands, operational excellence becomes increasingly important to ensuring safe, reliable, and efficient campus operations. About the Role We are seeking a Facilities Operations Manager to support the commissioning, operational readiness, and long-term operation of next-generation AI data center campuses. This role sits at the intersection of construction, commissioning, hardware deployment, and facilities operations. You will be responsible for ensuring mission-critical infrastructure is prepared to support hardware deployment, transitioned successfully into production operations, and maintained to the highest standards of reliability and availability. You will lead day-to-day operational execution across electrical, mechanical, controls, and supporting infrastructure systems while partnering closely with commissioning teams, site operators, vendors, and engineering organizations. This role requires a strong blend of technical depth, operational leadership, and cross-functional execution. Key Responsibilities Lead day-to-day operations of mission-critical facility infrastructure across AI compute campuses. Own operational readiness activities supporting new campus deployments and infrastructure expansion. Partner with commissioning teams to transition facilities from construction and startup into steady-state operations. Develop, implement, and
Jobs in United States
Infrastructure Team Manager in United States
1,475 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure team manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
From $156K/yr
As a Product Manager – IaC Detection, you will define, build, and launch capabilities that proactively detect infrastructure issues in code (e.g. Terraform, Helm) before they can be deployed into production and escalate into production incidents. The Infrastructure Monitoring team has pioneered shift-left detection in the industry with Bits Infrastructure Operations , and we’re looking for a Product Manager to expand this capability to a broader set of use cases Customers (and thus developers) are increasingly standardizing on IaC tools to deploy and maintain ever-growing infrastructure in the cloud. At the same time, SREs and Infra teams struggle with an increasing number of production incidents. By shifting-left and identifying high-impact infra changes before they are deployed, we help reduce production incidents, reduce waste, and free up SRE time to focus on value-added tasks. You will own the roadmap to expand IaC detection to a broader set of use cases, including cost detection, blast radius impact, as well as configuration changes on infrastructure powering applications like nginx, postgres and more. You’ll partner closely with Engineering, Design, and customers to build and iterate on the roadmap, build product market fit, drive customer adoption (including internal usage), and focus on coverage and correctness of the AI system. This is an opportunity to lead an initiative at the intersection of AI, infrastructure operations, and autonomous observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the product roadmap for IaC Detection, enabling customers to proactively detect and catch high-impact infrastructure and configuration changes before they are deployed into production and escalate into incidents. Define the end-to
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are hiring a talented Tech Lead Manager (TLM) to lead Snowflake’s System Under Test (SUT) team, responsible for evolving how Snowflake engineers test the Snowflake product locally and at scale in CI. The SUT empowers Snowflake engineers by delivering a reliable, low-latency and cost-efficient developer experience across a high-growth, high-demand surface area. As the TLM for SUT, you will lead a small and highly technical team at the intersection of CI, developer infrastructure, and product engineering. You will set direction, drive execution, and partner broadly across Engineering Systems and product teams to deliver a more reliable, faster, and more maintainable test platform for Snowflake’s engineers. In this role, you will: Lead, coach, and grow the SUT team while creating a high-energy, cohesive environment with strong planning, ownership, and career development. Own the roadmap and execution for SUT rollout across development environments, CI and AI workflows. Drive measurable improvements in startup reliability, latency, and cost, using clear SLOs, dashboards, and operational metrics to guide decisions and raise the bar on execution. Serve as the technical anchor for the SUT domain, shaping architecture and guiding the evolution from legacy systems to a composable
About the Team The Core Services organization builds and runs the mission-critical online services that product teams rely on in production. We own foundational distributed systems and platform capabilities that enable reliable execution, high-performance services, and large-scale file/data needs across our products. This team is distinct from developer infrastructure and data infrastructure—our focus is production service foundations and core runtime services. About the Role We’re hiring an Engineering Manager, Core Services to help lead teams responsible for highly reliable, high-scale distributed systems that sit on the critical path for OpenAI products. Your team will own foundational production systems that OpenAI’s product engineering teams build on. You’ll collaborate closely with product and infrastructure partners to ship reliable services quickly, and help scale systems and teams as OpenAI grows. You’ll partner closely with senior engineering leaders to scale the org, mature operations, and drive major platform initiatives. This role requires strong technical ability. You’ll be responsible for: Managing and growing a high-performing team of infrastructure engineers. Leading teams building and operating large, critical production platforms, including cluster reliability, scaling, and rollout safety. Building and operating mission-critical distributed systems with strong operational rigor (SLOs, incident response, capacity planning, reliability). Setting technical direction for platform foundations such as workflow/orchestration capabilities, large-scale file/blob/storage services, and core service foundations. Partnering with a broad set of stakeholders, including product engineering, adjacent infrastructure teams, and (where relevant) finance/cost partners. Coaching, mentoring, and developing engineers and emerging leaders. You might thrive in this role if you: Have significant experience leading teams that run mission-critical infrastructure in production
From $296K/yr
Datadog’s Cloud Observability group is one of the core data retrieval and processing groups powering our foundational product, Infrastructure Monitoring. The group’s scope includes integration with all major hyperscalers (AWS, Azure, GCP, OCI), as well as both regional and GPU-specific cloud providers. As Director, you will own engineering for all clouds, generating more than 10 million metric points per second, managing ~40 engineers through a team of Engineering Managers. You’ll partner with Senior Directors and product leadership to shape the roadmap, not just execute against it, managing the growth of one of Datadog’s foundational teams. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You'll Do: Own engineering for all of Cloud Observability Manage ~40 engineers through a layer of Engineering Managers; this is a manager-of-managers role Shape the roadmap alongside product leadership rather than simply executing against it — push back on, iterate on, and help author the strategy for your area Drive AI adoption across the engineering org, from tooling and workflows to product features and team practices Navigate cross-team dependencies across the Agent, Telemetry Onboarding, Integrations, Action Platform, and Infrastructure Monitoring. Build and retain engineering talent in NYC, Boston, and Paris, mentor Engineering Managers toward Director readiness, and participate in the on-call rotation Who You Are: You have directly managed Engineering Managers, not just individual contributors You have deep experience with one or more cloud providers, ideally with experience operating large-scale systems in the cloud. You have a solid understanding of cloud economics, as well as how to balance performance and cos
From $192.8K/yr
MongoDB is standing up a dedicated ISV Partnerships function, and this Director role is the leadership seat at the center of that effort. You will own the strategy and execution of MongoDB's ISV partner program across a team of partner managers, covering both industry-specific ISVs in verticals like financial services, healthcare, insurance, retail, and manufacturing, and horizontal AI-native ISVs building platforms and tools that serve builders across industries. This is not a role for someone who wants to manage from a distance. You will set the strategic direction for how MongoDB identifies, activates, and scales ISV partnerships, while staying close enough to the work to directly influence key relationships, shape partner agreements, and coach your team through complex deals. You will be expected to operate as a player-coach: intellectually involved in the most important partnerships, hands-on when the moment calls for it, and consistently raising the bar for what good looks like on your team. This is a greenfield program. You will have real ownership over how it is built, the methodology your team follows, how success is defined and measured, and how the function earns credibility with Sales, Product, and executive leadership. The opportunity is significant for the right leader. This role can be based remotely in the United States. What You'll Do Strategy and Program Leadership Define and own the strategic framework for MongoDB's ISV partner program, including segmentation across vertical and horizontal ISV categories, prioritization criteria, and the GTM approach for each partner tier Build and evolve the program infrastructure: onboarding processes, partner playbooks, co-sell frameworks, success metrics, and the internal operating model that connects ISV partnerships to MongoDB's revenue motion Set the vision for how ISV partnerships contribute to MongoDB's growth, and communicate that vision clearly to executive leadership, cross-functional stakeholders, and
From $280K/yr
Datadog is seeking a Director of Product Management to lead our AI Observability portfolio and shape how organizations build, monitor, and scale AI systems in production. This role leads LLM Observability and helps define the next wave of innovation across GPU Monitoring, Distributed AI Monitoring, and emerging research-oriented tooling such as Model Lab. You will set the vision and strategy for this rapidly growing area, expanding established products while incubating new capabilities that deliver deep visibility into AI infrastructure, model performance, and distributed AI environments. As AI becomes core to modern applications, this team plays a critical role in ensuring customers can deploy and scale AI with confidence. We’re looking for a builder-minded product leader with strong technical depth and hands-on curiosity - someone who has built or worked closely with AI-powered products and understands the realities of production AI. You will lead a team of product managers and partner closely with engineering and design to advance Datadog’s leadership in AI observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the vision and strategy for AI-driven products, ensuring alignment with overall company goals and customer needs. This will include managing our embed program to enhance the capabilities of existing products as well as developing dedicated and independent AI products. Lead and mentor a team of product managers, helping them grow and advance their careers while ensuring the delivery of high-quality, AI-powered features. Collaborate with cross-functional teams including engineering, data science, marketing, and sales to deliver AI product solutions that meet customer needs and business objectives. Identify new opportunities for
From $280K/yr
Datadog is seeking a Director of Product Management for Platforms to lead the internal platforms that power our global engineering organization. This is both a customer and internal platform leadership role, focused on enabling Datadog customers to maximize value with Datadog but also to enable other Datadog products to build, scale, and operate products efficiently and reliably. In this role, you will bring a combination of technical expertise, product management experience, and a deep understanding of platform and shared capabilities to help Datadog grow its leadership position. You will lead a team of product managers and collaborate with senior leadership in product, engineering and design. Your scope includes driving product features shared across all Datadog but also large-scale Datadog’s platform solutions. What You’ll Do Own the vision and strategy for platform products, ensuring alignment with overall company goals and customer needs. Identify new opportunities for innovation and drive them from concept to execution, ensuring they have a measurable impact on customers and the business. Define product roadmaps and manage the prioritization of features and initiatives to ensure the team's efforts are aligned with business goals. Drive Platform Adoption: Partner with engineering and product teams to ensure widespread adoption of shared platforms and shared features, reducing duplication and accelerating delivery. Improve Operational Efficiency: Improve developer velocity, time-to-production, and operational efficiency across Datadog’s engineering ecosystem. Collaborate with cross-functional teams including engineering, design, data science, marketing, and sales to deliver infrastructure product solutions that meet customer needs and business objectives. Define and Track Success Metrics: Define and track platform success metrics, including adoption of platform capabilities, reduction in internal toil, time-to-production improvements, cost efficienc
About the Team OpenAI's Industrial Compute organization builds and operates the infrastructure required to train and serve frontier AI models. The Capacity Planning team connects rapidly changing research and product demand with the compute, networking, storage, power, data center, hardware, and operational resources required to make that demand executable. About the Role We are seeking a Technical Program Manager to build and lead capacity planning across OpenAI's large-scale AI infrastructure. You will translate uncertain workload demand into clear infrastructure requirements, allocation decisions, supply commitments, activation priorities, and long-range capacity strategies. This role sits at the intersection of research, engineering, infrastructure, finance, sourcing, deployment, and operations. You will create the planning models, operating cadences, governance mechanisms, and source-of-truth systems that allow teams to understand what capacity is required, what is available, what is at risk, and what decisions must be made. This is not a finance-only forecasting or reporting role. Success requires technical fluency across the infrastructure stack, strong analytical judgment, and the ability to move consequential decisions forward when requirements, timelines, and supply conditions change quickly. Key Responsibilities Own capacity-planning processes across near-term workload allocation, quarterly execution, and longer-range infrastructure horizons. Translate research, training, inference, and product demand into compute, accelerator, cluster, networking, storage, rack, power, and site requirements. Develop scenarios that make assumptions, confidence levels, constraints, sensitivities, and decision points explicit. Reconcile requested demand against contracted, delivered, installed, activated, and workload-usable capacity. Partner with research and engineering teams to understand workload priorities, technical dependencies, utilization patterns, and changing req
About the Team OpenAI's data and storage infrastructure spans data platforms, online databases, and file/object storage. These systems underpin data ingestion and processing, durable persistence, indexing and retrieval, and product file experiences. As frontier models and agents evolve how they use memory, history and snapshots, the underlying architecture increasingly shapes the capabilities products can deliver—and their latency, reliability, cost and efficiency. About the Role We are looking for a technically deep TPM to independently define and lead multiple programs across data platforms, online databases and storage infrastructure. You will connect model, product and data-consumer requirements to architecture, and work with the relevant engineering teams to take new capabilities through production adoption and repeatable expansion. The design scope is exabyte-scale storage and infrastructure spanning multiple millions of CPU cores. The challenge is not simply forecasting more resources: it is making complete, workload-ready capacity repeatable, with a clear path from product requirements through architecture, deployment and validation. A data pipeline, database query, file operation or execution snapshot can affect whether a product or agent succeeds; you will connect those outcomes to the systems underneath. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate model, product and data-platform needs into precise access patterns, consistency, durability, freshness, availability and scalability requirements. Connect memory, history, retrieval and resumable work to capability and end-to-end latency. Partner with engineering to transform data and storage architecture into repeatable scale units: standardized provisioning, placement, routing, data movement and readiness checks that bring storage, compute and networking online together.
About the Team OpenAI's Industrial Compute organization is building and operating the infrastructure foundation for the next generation of AI. Infrastructure Operations works across facilities, hardware, network operations, incident management, data center engineering, delivery teams, and external partners to bring capacity online safely, understand its operational state, and improve it over time. As OpenAI's data center portfolio grows across first-party and partner-delivered capacity, the organization needs clear goals, trusted data, repeatable processes, and systems that make ownership, risk, readiness, and performance visible. This role will help build the operating mechanisms that allow Infrastructure Operations to scale with rigor. About the Role We are seeking a Technical Program Manager to own the systems, data, reporting, governance, and program-management backbone for Infrastructure Operations. Reporting to the Delivery & Operations Lead, you will translate strategy into executable goals and operating cadences, turn operational needs into software and data solutions, and create the mechanisms that keep a rapidly evolving organization aligned and accountable. This role will also own the current 1P+3P delivery-tracking layer within Operations: milestones, delivery timelines, quantity forecasts, risks, decisions, and executive reporting. You will partner closely with 1P Delivery Program Management, Compute TPMs, Data Center Engineering, construction, commissioning, and operations leaders to ensure that delivery information becomes complete, usable input for readiness, handover, and ongoing operations. You will own program health and the operating system around it: the goals, data definitions, workflows, reporting, decision paths, and follow-through that help functional DRIs execute. The ideal candidate is comfortable in ambiguity, technically fluent enough to implement real systems, and relentless about converting scattered information into durable mechan
About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. In partnership with leading cloud providers, hardware manufacturers, utilities, construction partners, and internal engineering organizations, we are delivering hyperscale AI campuses that power the next generation of frontier AI models. Infrastructure Delivery Operations sits at the center of this effort. Our team develops the operating model that connects infrastructure strategy, supply planning, manufacturing operations, and delivery into a single, integrated system that enables OpenAI to deploy AI infrastructure predictably at scale. We partner across Hardware Engineering, Network Engineering, Capacity Delivery, Hardware Operations, Security, Finance, Strategic Sourcing, and external infrastructure partners to create a single, integrated view of program health. Through governance, operational analytics, executive reporting, and scalable operating mechanisms, we enable leaders to proactively manage risk, optimize capacity, and deliver infrastructure predictably at Industrial Compute speed. About the Role We are seeking a Technical Program Manager, Infrastructure Delivery Operations to drive integrated strategy and delivery across OpenAI's rapidly expanding AI infrastructure portfolio. This role sits at the intersection of infrastructure strategy, New Product Introduction (NPI), supply planning, manufacturing operations, and infrastructure delivery. You will lead highly cross-functional programs spanning engineering, supply planning, manufacturing, logistics, construction, commissioning, and operations, ensuring technical and operational dependencies remain synchronized from planning through production readiness. Beyond driving program execution, you will leverage operational insights to improve capacity planning, infrastructure strategy, and deployment readiness. You will also help operationalize new technologies and suppliers by partnering w
About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to support the next generation of frontier AI systems. Through a combination of strategic partnerships and self-built data center campuses, we are scaling the physical infrastructure needed to deliver compute at unprecedented scale. The Commissioning organization is responsible for ensuring this infrastructure is safely tested, validated, integrated, and transitioned into reliable operations. As the portfolio grows, the team is building common standards, processes, tools, and reporting systems that allow commissioning programs to operate consistently across projects while giving teams and leadership clear visibility into readiness, risk, and execution. About the Role We are seeking a Commissioning Program Manager to build and scale the operating systems behind OpenAI’s infrastructure commissioning programs. You will own the development and continuous improvement of commissioning standards, processes, tools, dashboards, and KPIs across the infrastructure portfolio. You will work closely with commissioning and construction teams to translate field execution needs into practical playbooks, workflows, templates, metrics, and reporting mechanisms that teams can use from construction readiness through testing and turnover. This role sits at the intersection of infrastructure delivery, program management, process design, and data. The ideal candidate understands how complex construction projects operate and can turn fragmented workflows and project data into repeatable systems that improve execution without creating unnecessary administrative burden. Key Responsibilities Develop and maintain commissioning program standards, playbooks, process maps, templates, checklists, stage gates, and acceptance criteria across infrastructure projects. Establish consistent workflows for commissioning planning, construction readiness, QA/QC, issue management, document control, testing evidence
$226K – $285K/yr
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a Supply Chain Program Manager, you will own material readiness and supply chain execution for critical hardware programs spanning custom silicon, systems, memory, storage, networking, and rack infrastructure. You will work cross-functionally with Engineering, Strategic Sourcing, Manufacturing Operations, Finance, Planning, Quality, and external suppliers to develop and execute scalable supply strategies that support aggressive product development and deployment timelines. This role requires deep understanding of hardware supply chains, material planning, NPI execution, supplier management, and operational scaling in constrained and rapidly evolving environments. In this role you will: Material Readiness & Supply Planning - Own end-to-end material readiness across NPI and production phases, including building the necessary framework and processes for enablement. Drive supply planning and execution for long lead-time and constrained commodities including ASICs, HBM, DDR, SSDs, networking, optics, power, thermal, and mechanicals. Build and manage material readiness plans aligned to proto/pre-EVT, EVT, DVT, PVT, and mass production schedules. Monitor supply health, lead times, inventory positions, allocation risk, and capacity constraints. Drive shortage management, allocation mitigation, and recovery planning. Coordinate supply commits, forecast alignment, and supply continuity planning with suppliers and manufacturing partners. Cross-Functional Program Ma
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI compute infrastructure ecosystem. The InfraDev team is central to this mission, setting the strategy and executing the roadmap to scale our supercomputing footprint globally. From site planning to system integration, this team operates at the intersection of commercial, technical, and operational excellence, partnering with leaders across OpenAI and the industry. About the Role We are seeking a Strategic Sourcing Manager who is ready to take on global-scale challenges in AI compute supply and manufacturing — an opportunity to shape the future of supercomputing. As part of the Infrastructure Strategy & Delivery organization supporting Industrial Compute, OpenAI’s next-generation supercomputing platform, you will lead sourcing and strategic supplier engagements for compute infrastructure at hyperscale. This role will drive commercial strategy and supplier accountability across server platforms, accelerators, rack systems, and associated thermal & power delivery components. This is not traditional procurement — it is foundational work enabling OpenAI to deploy compute faster and more efficiently than anyone in the world, while building deeply-integrated partnerships with the global compute supply chain. Key Responsibilities Develop and execute sourcing strategies for the next generation of AI compute infrastructure in partnership with engineering and program leadership. Stay at the leading edge of industry and supply-chain trends to inform category strategy and long-range planning. Initiate, negotiate, and manage commercial agreements across OEM/ODM/JDMs for accelerators, GPUs/CPUs, server platforms, rack systems, liquid-cooling components, PSUs, and supporting mechanical/electrical subsystems. Secure capacity and optionality across a rapidly scaling compute supply chain while mitigating risk and ensuring manufacturing readiness. Partner cross-funct
Other cities to consider
More places hiring for this role
Get new infrastructure team manager jobs in United States by email
Daily job updates · Unsubscribe anytime