About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. The Infrastructure team is central to this mission, setting the core strategy and implementing the vision. From site selection to deployment to operations, this team sits at the intersection of commercial, technical, and operational domains, interacting with experts and executives inside and outside of OpenAI. We design and operate mission-critical facilities that support cutting-edge AI workloads at scale. About the Role We are seeking a Facilities Operations Lead to support the commissioning, deployment, and long-term operation of our next-generation AI data centers. This role bridges the interface between data center construction and hardware landing, ensuring seamless integration of mission-critical infrastructure with hardware deployment timelines. You will define and execute commissioning plans, support infrastructure bring-up, and take ownership of operations and maintenance for cutting-edge, large-scale, AI data centers. You will collaborate closely with design, construction, and hardware teams to define repeatable processes for new data center builds and lead hands-on operations to uphold the performance and reliability of our deployed infrastructure. Key Responsibilities Define and execute sequences of operations, commissioning steps, and bring-up processes for mission-critical data center facilities. Interface with the design and hardware teams to define deployment procedures tailored to each data center and hardware configuration. Oversee installation, commissioning, and operational readiness of large-scale data center campuses. Manage monitoring, maintenance, and quality control of the data center infrastructure, including high-performance liquid cooling systems. Develop on-site operations staffing strategy. Develop and enforce procedures for planed and unplanned downtime and SLAs for critical
Jobs in United States
Data Center Infrastructure Architect in San Francisco
442 active opportunities · Updated September 2026
Showing
15 jobs
Explore current data center infrastructure architect jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui
About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. This team is central to this mission, setting the core infra strategy and implementing this vision. From site selection to the buildout process, this team sits at the intersection of commercial, technical, strategy, and operations, interacting with teams and executives inside and outside of OpenAI. About the Role We are seeking experienced Data Center Mechanical and Electrical/Power Design Engineers with expertise in designing, operating, and maintaining large-scale data center campuses. The ideal candidate for this role will have extensive background and experience in design and managing critical equipment and facilities, design and operation of MEP (Mechanical, Electrical, Plumbing) systems, and overseeing operational activities from initial phases of Data Center build through delivery and ongoing maintenance. The ideal candidate will have a strong technical background, operational leadership experience, and a proven ability to collaborate with external vendors on critical infrastructure. This role offers the opportunity to lead transformative data center projects with high visibility and impact. If you are passionate about delivering cutting-edge infrastructure solutions, we encourage you to apply. Key Responsibilities Oversee building and MEP design, operation, and maintenance, including reviewing building and MEP drawings and proposals across all project phases. Lead operational activities for large-scale data center campuses, from early design phases through delivery and daily operation. Operate and maintain critical data center facilities and equipment, ensuring reliability and performance. Collaborate with external vendors to select, procure, and manage critical equipment, such as generators, UPS, chillers, and CDUs. Provide technical expertise on all aspects of data center building, equipment, and
About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to support the next generation of frontier AI systems. Through a combination of strategic partnerships and self-built data center campuses, we are scaling the power, cooling, electrical, mechanical, and controls infrastructure needed to deliver compute at unprecedented scale. The Commissioning organization is responsible for ensuring this infrastructure is safely tested, validated, integrated, and transitioned into reliable operations. For our self-build campuses, the team operates through a hybrid delivery model: OpenAI provides commissioning leadership, discipline ownership, governance, and project integration, while commissioning partners provide field and test engineering capacity to support inspections, startup, testing, and turnover. About the Role We are seeking a Commissioning Project Lead to own the commissioning strategy and execution for a large-scale, self-build data center project. You will lead the overall commissioning program from early construction planning through startup, functional testing, integrated systems testing, and final turnover. You will establish the commissioning execution plan, integrate commissioning activities into the master project schedule, coordinate multidisciplinary readiness, and lead the vendor commissioning partners providing field and test engineering capacity. This role serves as the primary commissioning interface to project leadership, construction management, contractors, equipment vendors, operations, and commissioning partners. You will be responsible for creating clarity across organizations, identifying readiness and schedule risks early, and ensuring the facility progresses through testing and turnover against clearly defined acceptance criteria. The role will initially support planning and coordination in a hybrid capacity and transition to full-time onsite presence as construction, inspections, startup, testing, and t
About the Team OpenAI’s Industrial Compute team is responsible for building and scaling large-scale compute capacity across first-party data centers, strategic partners, and industrial infrastructure environments. We focus on converting power, land, hardware, and operational execution into reliable compute capacity that can support frontier AI training and inference workloads. This team operates at the intersection of infrastructure delivery, hardware systems, utilities, supply chain, and capacity strategy—ensuring OpenAI can scale compute faster than traditional models allow. About the Role We are seeking a Tokens-as-a-Service (TaaS) Lead to drive the end-to-end conversion of industrial-scale infrastructure investments into usable token capacity for OpenAI workloads. In this role, you will own execution across complex compute programs where raw infrastructure capacity must be transformed into operational GPU throughput. You will coordinate across data center delivery, power, networking, hardware deployment, workload enablement, finance, and external partners to ensure capacity becomes productive tokens as quickly and efficiently as possible. This role is ideal for someone who can bridge physical infrastructure delivery with compute utilization outcomes. Success requires strong systems thinking, elite program leadership, and the ability to drive accountability across internal teams and strategic partners. In this role, you will Lead Tokens-as-a-Service programs across industrial compute environments, including first-party and partner-owned capacity. Convert delivered power, space, and hardware capacity into production-ready token throughput. Build integrated execution plans spanning construction, power energization, rack deployment, networking, cluster readiness, and workload onboarding. Partner with infrastructure engineering, hardware, networking, finance, supply chain, and operations teams. Drive external providers, EPCs, OEMs, utilities, and strategic partners t
About the Team OpenAI's Industrial Compute organization is building and scaling the infrastructure required to support frontier AI. The Infrastructure Strategic Sourcing team connects technical and project requirements to supplier readiness, contracting, purchasing, equipment delivery, and portfolio-level risk visibility across owner-furnished contractor-installed equipment (OFCI), data center networking, rack systems and integration, fiber, cabling, optical interconnects, and related infrastructure. The team partners across Pre-Construction, Design, Construction, Electrical and Mechanical Engineering, Network Engineering, Hardware and Rack Delivery, Strategic Sourcing, Procurement, Legal, Finance, Accounts Payable, Logistics, and external suppliers. We build the operating mechanisms that keep sourcing decisions, purchase execution, long-lead equipment, network and fiber dependencies, rack readiness, and delivery commitments aligned to infrastructure schedules. About the Role We are seeking an Infrastructure Sourcing Operations Lead to own procurement operations across pre-construction, design, construction, and sourcing through purchase order issuance, while maintaining visibility through invoice resolution, production, logistics, delivery, installation, and readiness. The portfolio includes electrical and mechanical OFCI, networking equipment, rack systems and integration, fiber, cabling, optical interconnects, and other infrastructure required to bring capacity online. In this role, you will set priorities, make or escalate decisions that affect cost, supplier relationships, contractual position, and delivery schedules, and define the standards used by execution support for queue management, documentation, tracker maintenance, and recurring reporting. Success requires sound commercial and program judgment, operational rigor, systems thinking, and the ability to turn incomplete information across vendors, tools, and project teams into clear decisions, accountable
About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to support the next generation of frontier AI systems. Through a combination of strategic partnerships and self-built data center campuses, we are scaling the physical infrastructure needed to deliver compute at unprecedented scale. The Commissioning organization is responsible for ensuring this infrastructure is safely tested, validated, integrated, and transitioned into reliable operations. As the portfolio grows, the team is building common standards, processes, tools, and reporting systems that allow commissioning programs to operate consistently across projects while giving teams and leadership clear visibility into readiness, risk, and execution. About the Role We are seeking a Commissioning Program Manager to build and scale the operating systems behind OpenAI’s infrastructure commissioning programs. You will own the development and continuous improvement of commissioning standards, processes, tools, dashboards, and KPIs across the infrastructure portfolio. You will work closely with commissioning and construction teams to translate field execution needs into practical playbooks, workflows, templates, metrics, and reporting mechanisms that teams can use from construction readiness through testing and turnover. This role sits at the intersection of infrastructure delivery, program management, process design, and data. The ideal candidate understands how complex construction projects operate and can turn fragmented workflows and project data into repeatable systems that improve execution without creating unnecessary administrative burden. Key Responsibilities Develop and maintain commissioning program standards, playbooks, process maps, templates, checklists, stage gates, and acceptance criteria across infrastructure projects. Establish consistent workflows for commissioning planning, construction readiness, QA/QC, issue management, document control, testing evidence
About the Team OpenAI's Industrial Compute organization is building and operating the infrastructure foundation for the next generation of AI. Infrastructure Operations works across facilities, hardware, network operations, incident management, data center engineering, delivery teams, and external partners to bring capacity online safely, understand its operational state, and improve it over time. As OpenAI's data center portfolio grows across first-party and partner-delivered capacity, the organization needs clear goals, trusted data, repeatable processes, and systems that make ownership, risk, readiness, and performance visible. This role will help build the operating mechanisms that allow Infrastructure Operations to scale with rigor. About the Role We are seeking a Technical Program Manager to own the systems, data, reporting, governance, and program-management backbone for Infrastructure Operations. Reporting to the Delivery & Operations Lead, you will translate strategy into executable goals and operating cadences, turn operational needs into software and data solutions, and create the mechanisms that keep a rapidly evolving organization aligned and accountable. This role will also own the current 1P+3P delivery-tracking layer within Operations: milestones, delivery timelines, quantity forecasts, risks, decisions, and executive reporting. You will partner closely with 1P Delivery Program Management, Compute TPMs, Data Center Engineering, construction, commissioning, and operations leaders to ensure that delivery information becomes complete, usable input for readiness, handover, and ongoing operations. You will own program health and the operating system around it: the goals, data definitions, workflows, reporting, decision paths, and follow-through that help functional DRIs execute. The ideal candidate is comfortable in ambiguity, technically fluent enough to implement real systems, and relentless about converting scattered information into durable mechan
About the Team OpenAI’s Legal team plays a crucial role in advancing our mission by tackling novel and consequential legal issues in AI. The commercial infrastructure legal team supports the systems and facilities that make advanced AI possible, including compute, first-party silicon, colocation, and first-party data center development. About the Role We’re seeking a senior commercial infrastructure lawyer to lead legal work across our rapidly growing data center development and colocation portfolio. You will advise on complex, high-value transactions involving new data center sites, power, construction, colocation, and related infrastructure arrangements. This role is designed for a lawyer who understands how data centers are developed and operated—not a general real estate practitioner. You will work closely with infrastructure, finance, procurement, and operations teams, while coordinating outside counsel where appropriate, to help projects move quickly and responsibly. San Francisco is preferred, and occasional travel may be required. In this role, you will: Lead commercial legal strategy and risk management for data center development and colocation transactions. Draft, negotiate, and advise on colocation agreements, construction agreements, power-related arrangements, and other contracts supporting data center development. Support first-party projects involving raw land, site and power rights, design and construction, EPC models, third-party build-to-suit structures, and leasebacks. Partner cross-functionally with infrastructure, finance, procurement, operations, and other stakeholders on fast-moving, complex projects. Manage and collaborate effectively with outside counsel across a high volume of infrastructure matters. Develop scalable legal playbooks and contracting approaches for a growing portfolio of data center sites. You might thrive in this role if you have: 12+ years of legal experience, primarily in commercial legal roles. Deep, hands-on experience
$342K – $445K/yr
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Technical Lead to lead deployment and operations for OpenAI’s Silicon & Systems team. This person will become the Directly-Responsible Individual responsible for bringing OpenAI’s custom silicon and associated systems into data center environments, ensuring successful deployment, bring-up, validation, operational readiness, and ongoing reliability at scale. This role sits at the intersection of silicon, systems, infrastructure, data center operations, and software. You will lead a team focused on taking new hardware platforms from lab validation into production data center deployment. You will be responsible for building the operational processes, technical workflows, tooling, and cross-functional alignment required to deploy and operate custom AI hardware reliably in OpenAI’s supercomputing infrastructure. The ideal candidate is both a strong leader and a deeply technical operator. You should be comfortable staying close to the technical details of hardware bring-up, fleet deployment, debugging, system validation, data center integration, and production operations. This role requires strong execution, excellent cross-functional judgment, and the ability to drive clarity in ambiguous, fast-moving environments. In this role, you will: Lead a team responsible for deployment and operations of OpenAI’s custom silicon and systems in data center environments Own the path from hardware bring-up and validation through production deployment, operati
About the Role As a Director, Compute & Infrastructure FP&A, you will own and drive the monthly forecasting process for the Compute & Infrastructure org by partnering with various stakeholders across Finance, Accounting, Tax and Engineering. You will play a critical role in planning and forecasting the company’s largest and most complex cost center ( Compute & Infrastructure ). You will collaborate cross-functionally to develop long-range infrastructure investment plans, evaluate build vs. buy decisions, and ensure capital is deployed efficiently to support rapid growth. You will also provide strategic financial guidance through scenario modeling, ROI analysis, and performance tracking, enabling leadership to make high-stakes decisions under uncertainty. What You’ll Do Own compute financial planning & Forecasting. Build and manage consolidation models for GPU/CPU capacity, storage, networking, and data center investments. Translate infrastructure roadmaps into short- and long-term financial forecasts (LRP, annual planning) Coordinate closely with Corporate FP&A on timelines and process Present insights on a monthly basis to senior management. Drive infrastructure investment decisions. Evaluate build vs. buy, vendor vs. owned infrastructure, and capacity allocation tradeoffs. Develop frameworks for investment trade-offs to guide executive decision making. Build scalable tooling & reporting. Implement stakeholder-facing dashboards to track compute spend, utilization, and efficiency metrics. Improve visibility into unit economics (e.g., cost per training run, cost per inference, cost per customer). Drive forecasting accuracy & accountability. Lead budget vs. actual analysis for compute and infrastructure spend. Identify key cost drivers (utilization, pricing, efficiency gains) and reduce forecast variance. Support close & financial reporting. Partner with Accounting to ensure accurate classification of infrastructure spend (OpEx vs C
About the Team OpenAI Finance is responsible for ensuring the organization is set up for success in pursuit of its mission. The Technical Accounting team plays a crucial role in helping OpenAI navigate complex, judgmental, and rapidly evolving accounting matters with rigor and clarity. We aim to bring both technical excellence and strong business partnership to some of the most novel accounting questions in the industry. About the Role As Senior Manager, Technical Accounting, Compute Infrastructure, you will lead the evaluation, documentation, and operationalization of complex accounting matters related to OpenAI's compute infrastructure, strategic investments, and other non-routine business activities.. This role sits at the intersection of U.S. GAAP technical accounting, infrastructure strategy, financial reporting, controls, and cross-functional execution. Key areas may include cloud compute arrangements, data center and colocation arrangements, lease accounting under ASC 842, power purchase agreements, strategic investments, consolidation evaluations under ASC 810, financial instruments, and other emerging or non-standard arrangements. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead technical accounting analysis for complex, judgmental, and non-routine transactions under U.S. GAAP. Evaluate accounting implications for compute infrastructure arrangements, including cloud compute, data center, colocation, lease, PPA, infrastructure procurement, and related commercial arrangements. Partner with Controllership, Tax, Legal, FP&A, Procurement, Infrastructure, and other cross-functional teams to assess the accounting implications of new products, commercial arrangements, strategic transactions, and business initiatives. Prepare and review technical accounting memoranda, position papers, and other auditor-ready documentation. Translate
$216K – $240K/yr
About the Team OpenAI Finance is responsible for ensuring the organization is set up for success in pursuit of its mission. The Technical Accounting team plays a crucial role in helping OpenAI navigate complex, judgmental, and rapidly evolving accounting matters with rigor and clarity. We aim to bring both technical excellence and strong business partnership to some of the most novel accounting questions in the industry. About the Role As Senior Manager, Technical Accounting, Compute Infrastructure, you will lead the evaluation, documentation, and operationalization of complex accounting matters related to OpenAI's compute infrastructure, strategic investments, and other non-routine business activities.. This role sits at the intersection of U.S. GAAP technical accounting, infrastructure strategy, financial reporting, controls, and cross-functional execution. Key areas may include cloud compute arrangements, data center and colocation arrangements, lease accounting under ASC 842, power purchase agreements, strategic investments, consolidation evaluations under ASC 810, financial instruments, and other emerging or non-standard arrangements. This role is based in San Francisco, CA or remote. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead technical accounting analysis for complex, judgmental, and non-routine transactions under U.S. GAAP. Evaluate accounting implications for compute infrastructure arrangements, including cloud compute, data center, colocation, lease, PPA, infrastructure procurement, and related commercial arrangements. Partner with Controllership, Tax, Legal, FP&A, Procurement, Infrastructure, and other cross-functional teams to assess the accounting implications of new products, commercial arrangements, strategic transactions, and business initiatives. Prepare and review technical accounting memoranda, position papers, and other auditor-ready documentation.
About the Team OpenAI is helping build the infrastructure that powers the next generation of artificial intelligence. Through Stargate, we are developing and operating large-scale AI compute campuses that require world-class execution across data center design, construction, commissioning, and operations. The Infrastructure Operations team is responsible for bringing AI infrastructure online and ensuring it operates reliably at scale. We partner closely with hardware, network, deployment, construction, and operations teams to deliver mission-critical environments capable of supporting frontier AI workloads. As our footprint expands, operational excellence becomes increasingly important to ensuring safe, reliable, and efficient campus operations. About the Role We are seeking a Facilities Operations Manager to support the commissioning, operational readiness, and long-term operation of next-generation AI data center campuses. This role sits at the intersection of construction, commissioning, hardware deployment, and facilities operations. You will be responsible for ensuring mission-critical infrastructure is prepared to support hardware deployment, transitioned successfully into production operations, and maintained to the highest standards of reliability and availability. You will lead day-to-day operational execution across electrical, mechanical, controls, and supporting infrastructure systems while partnering closely with commissioning teams, site operators, vendors, and engineering organizations. This role requires a strong blend of technical depth, operational leadership, and cross-functional execution. Key Responsibilities Lead day-to-day operations of mission-critical facility infrastructure across AI compute campuses. Own operational readiness activities supporting new campus deployments and infrastructure expansion. Partner with commissioning teams to transition facilities from construction and startup into steady-state operations. Develop, implement, and
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You’ll own the hands-on and automation work that brings WAN, fiber, carrier, and cloud-interconnect circuits into service. Partner with network engineers, fiber providers, cloud service providers, colocation teams, and data-center technicians to move each connection from ordered and patched to verified, stable, and ready for handoff. You’ll own Layer 1 troubleshooting and circuit bring-up while building workflows that translate reliable system or model output into precise, approved technician actions, capture field feedback, and drive each connection to a green-port handoff. The right person combines strong physical-networking judgment with practical automation skills: patch-panel and port mappings, optics and light levels, provider coordination, structured operational data, API or scripting workflows, and human-in-the-loop LLM tooling. Responsibilities Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence. Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports. Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching
Other cities to consider
More places hiring for this role
Get new data center infrastructure architect jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime