About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We're looking for an Optical Interconnect System Engineer to design, qualify, and deploy scalable optical connectivity for large-scale AI infrastructure. This role spans fiber-system architecture, optical-mechanical integration, validation, reliability, deployment, and serviceability. You will work with optical, mechanical, electrical, networking, manufacturing, reliability, and data-center teams to translate system needs into practical interconnect solutions. This is a hands-on role for someone who can connect design decisions with installation, qualification, troubleshooting, and long-term operational performance. In this role, you will: Define optical interconnect architectures and requirements across hardware platforms and rack-level systems. Design high-density fiber systems for performance, density, reliability, installation, and serviceability. Lead optical-mechanical integration and cross-functional design reviews. Develop test and qualification plans for optical components, modules, switching platforms, and integrated systems. Own optical loss budgets, routing guidelines, handling requirements, and serviceability criteria. Support system bring-up, deployment, troubleshooting, failure analysis, and reliability improvement. Create reusable design guidelines, interface requirements, and qualification methods. You might thrive in this role if you have: Core experience Experience desi
Jobs in United States
Rack Scale Software Architecture Director in San Francisco
20 active opportunities · Updated October 2026
Showing
15 jobs
Explore current rack scale software architecture director jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team builds next-generation AI-native silicon and systems while working closely with software, research, and manufacturing partners to co-design hardware tightly integrated with AI models. In addition to delivering systems for OpenAI’s supercomputing infrastructure, the team develops the tools, methodologies, and strategic partnerships needed to accelerate hardware innovation. About the Role We’re seeking an experienced Hardware Strategic Sourcing Manager to own sourcing strategy and supplier partnerships for fiber and optical interconnect components across OpenAI’s next-generation AI infrastructure. Reporting to the Head of Partnerships & Strategic Sourcing, you will lead sourcing across fiber cable assemblies, internal optical harnesses, fiber shuffles, optical backplane assemblies, connectorized and standalone passive optical assemblies, fiber-array units (FAUs), fiber-to-chip and coupling interfaces, detachable connectors, optical routing, and assigned optical packaging, assembly, and test services. You will work closely with electrical engineering, optical engineering, systems engineering, mechanical and packaging engineering, quality, rack integration, data-center deployment,manufacturing, supply chain, finance, legal, and program management teams to translate demanding bandwidth, signal integrity, reliability, and scale requirements into resilient supplier partnerships and scalable commercial strategies. Your work will directly support the performance, reliability, manufacturability, and scale of the high-speed optical connectivity required for OpenAI’s next-generation AI systems. In this role, you will: Develop and execute a comprehensive sourcing strategy for fiber and optical interconnect components supporting high-bandwidth AI systems and infrastructure. Own sourcing across optical fiber cable assembli
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a Delivery Director, Capacity programs for our on-premises data center builds and neo cloud (GPU cloud) delivery programs. This is a high-visibility, execution-critical role sitting at the intersection of infrastructure engineering, capacity planning, vendor/partner management, and customer delivery. You will own the end-to-end delivery lifecycle for large-scale compute infrastructure — from initial site/capacity commitments through power, networking, and hardware bring-up, to production-ready GPU/compute capacity landing in the hands of internal teams or customers. You'll be the person who turns ambitious infrastructure roadmaps into predictable, on-time, delivery. RESPONSIBILITIES Own delivery of on-prem infrastructure builds — colocation expansions, power/cooling readiness, rack-and-stack, network fabric bring-up, and hardware acceptance testing — coordinating across colo providers and partners, network engineering, hardware ops, and vendor teams. Drive neo cloud delivery programs — manage capacity delivery from GPU cloud and neo cloud partners (e.g., colocation/bare-metal/GPU cloud providers), including contract milestones, capacity ramps, SLAs, and go-live readiness. Build and maintain master delivery schedules across concurrent, multi-site, multi-vendor programs, integrating power/shell timelines, hardware lead times, logistics, and software/platform readiness into a single critical path.
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role OpenAI's Hardware organization builds supercompute platforms from silicon and boards to full rack-scale systems to power advanced AI workloads. This role owns end-to-end quality for high-speed interconnect hardware across the product lifecycle: early design influence, supplier/contract manufacturer readiness, qualification, ramp, and fleet quality in lab and data center environments. You will be the quality lead for advanced interconnect components and assemblies, including high-speed copper cables, cable cartridges, patch panels, backplane/cable-backplane solutions, high-speed connectors, and related electro-mechanical interfaces. You will partner closely with electrical, mechanical, SI/PI, systems, reliability, operations, and external vendors to prevent escapes and drive rapid, data-driven containment and corrective action. In this role you will: Own quality for advanced interconnect components and assemblies: high-speed connectors, high-speed copper cables, cable cartridges (e.g., cable cassette style assemblies), patch panels & optics, and backplane/cable-backplane interconnect solutions. Drive quality-by-design: participate in design reviews, DFM/DFx, tolerance stacks, material and plating selections, connector mating strategy, strain relief, and assembly methods to reduce variation and field failures. Define and track quality and reliability metrics (DPPM, yield, escapes, RMA/FRACAS trends, Cpk/Ppk where applicable) for interconnects across NPI and m
About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Through strategic partnerships and self-built campuses, we are scaling one of the world's fastest-growing AI infrastructure platforms. The Supply Chain organization ensures critical infrastructure components—from compute systems and networking equipment to integrated rack solutions—are sourced, manufactured, qualified, and delivered with the speed and reliability required to support frontier AI development. We partner closely with Hardware Engineering, Manufacturing Quality Engineering, Infrastructure Delivery, Hardware Operations, Finance, and suppliers worldwide to build a resilient, scalable supply chain capable of supporting rapid infrastructure expansion. As Industrial Compute continues to grow, Supply Chain serves as the operational bridge between engineering innovation and large-scale infrastructure deployment. About the Role We are seeking a Supply Chain Manager to lead strategic execution across sourcing, supplier operations, manufacturing quality, and infrastructure delivery for OpenAI's AI infrastructure portfolio. This role will oversee a multidisciplinary team responsible for strategic sourcing, manufacturing quality engineering, and technical program management while partnering closely with engineering, finance, hardware operations, and deployment teams. You will drive supplier strategy, manufacturing readiness, production planning, quality performance, and operational execution across the full hardware lifecycle. Success requires balancing long-term supplier strategy with day-to-day execution. You'll establish scalable operating mechanisms, strengthen supplier partnerships, manage complex cross-functional programs, and ensure OpenAI can rapidly deploy AI infrastructure without compromising quality, cost, or reliability. This is a people leadership role responsible for developing a high-performing organization while driving operati
About the Team OpenAI’s Industrial Compute team is responsible for building and scaling large-scale compute capacity across first-party data centers, strategic partners, and industrial infrastructure environments. We focus on converting power, land, hardware, and operational execution into reliable compute capacity that can support frontier AI training and inference workloads. This team operates at the intersection of infrastructure delivery, hardware systems, utilities, supply chain, and capacity strategy—ensuring OpenAI can scale compute faster than traditional models allow. About the Role We are seeking a Tokens-as-a-Service (TaaS) Lead to drive the end-to-end conversion of industrial-scale infrastructure investments into usable token capacity for OpenAI workloads. In this role, you will own execution across complex compute programs where raw infrastructure capacity must be transformed into operational GPU throughput. You will coordinate across data center delivery, power, networking, hardware deployment, workload enablement, finance, and external partners to ensure capacity becomes productive tokens as quickly and efficiently as possible. This role is ideal for someone who can bridge physical infrastructure delivery with compute utilization outcomes. Success requires strong systems thinking, elite program leadership, and the ability to drive accountability across internal teams and strategic partners. In this role, you will Lead Tokens-as-a-Service programs across industrial compute environments, including first-party and partner-owned capacity. Convert delivered power, space, and hardware capacity into production-ready token throughput. Build integrated execution plans spanning construction, power energization, rack deployment, networking, cluster readiness, and workload onboarding. Partner with infrastructure engineering, hardware, networking, finance, supply chain, and operations teams. Drive external providers, EPCs, OEMs, utilities, and strategic partners t
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des
About the Team OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical. The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly. About the Role We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity. In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity. This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions. This role is based in San Francisco, CA, with travel as needed. Key Responsibilities Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint Drive readiness to convert contracted compute capacity into schedulable production clusters Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives Build integrated schedules spanning procurement, logistics, installation, st
About the Team OpenAI's Industrial Compute organization builds and operates the infrastructure required to train and serve frontier AI models. The Capacity Planning team connects rapidly changing research and product demand with the compute, networking, storage, power, data center, hardware, and operational resources required to make that demand executable. About the Role We are seeking a Technical Program Manager to build and lead capacity planning across OpenAI's large-scale AI infrastructure. You will translate uncertain workload demand into clear infrastructure requirements, allocation decisions, supply commitments, activation priorities, and long-range capacity strategies. This role sits at the intersection of research, engineering, infrastructure, finance, sourcing, deployment, and operations. You will create the planning models, operating cadences, governance mechanisms, and source-of-truth systems that allow teams to understand what capacity is required, what is available, what is at risk, and what decisions must be made. This is not a finance-only forecasting or reporting role. Success requires technical fluency across the infrastructure stack, strong analytical judgment, and the ability to move consequential decisions forward when requirements, timelines, and supply conditions change quickly. Key Responsibilities Own capacity-planning processes across near-term workload allocation, quarterly execution, and longer-range infrastructure horizons. Translate research, training, inference, and product demand into compute, accelerator, cluster, networking, storage, rack, power, and site requirements. Develop scenarios that make assumptions, confidence levels, constraints, sensitivities, and decision points explicit. Reconcile requested demand against contracted, delivered, installed, activated, and workload-usable capacity. Partner with research and engineering teams to understand workload priorities, technical dependencies, utilization patterns, and changing req
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI compute infrastructure ecosystem. The InfraDev team is central to this mission, setting the strategy and executing the roadmap to scale our supercomputing footprint globally. From site planning to system integration, this team operates at the intersection of commercial, technical, and operational excellence, partnering with leaders across OpenAI and the industry. About the Role We are seeking a Strategic Sourcing Manager who is ready to take on global-scale challenges in AI compute supply and manufacturing — an opportunity to shape the future of supercomputing. As part of the Infrastructure Strategy & Delivery organization supporting Industrial Compute, OpenAI’s next-generation supercomputing platform, you will lead sourcing and strategic supplier engagements for compute infrastructure at hyperscale. This role will drive commercial strategy and supplier accountability across server platforms, accelerators, rack systems, and associated thermal & power delivery components. This is not traditional procurement — it is foundational work enabling OpenAI to deploy compute faster and more efficiently than anyone in the world, while building deeply-integrated partnerships with the global compute supply chain. Key Responsibilities Develop and execute sourcing strategies for the next generation of AI compute infrastructure in partnership with engineering and program leadership. Stay at the leading edge of industry and supply-chain trends to inform category strategy and long-range planning. Initiate, negotiate, and manage commercial agreements across OEM/ODM/JDMs for accelerators, GPUs/CPUs, server platforms, rack systems, liquid-cooling components, PSUs, and supporting mechanical/electrical subsystems. Secure capacity and optionality across a rapidly scaling compute supply chain while mitigating risk and ensuring manufacturing readiness. Partner cross-funct
From $350K/yr
About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. Help Win New Business The Opportunity: We are scaling our dedicated Data Center practice, and we are looking for the person who will lead it. This is a founding commercial role. You will start as a team of one, owning the full sales motion end-to-end, and you will build the team around you as the practice grows. You will define how Flexport goes to market with hyperscalers, hardware OEMs, and the broader ecosystem. You will set the playbook, win the first marquee accounts, and hire the people who scale what you build. Reporting directly to the Regional General Manager, you will operate with a high degree of autonomy and direct access to executive leadership. This role features a 50/50 compensation model (Base + Uncapped bonus, with accelerators), with OTEs of $350k+, designed for strong leaders who take bets on themselves. Why This Role Is Different: The data center logistics market is not a standard freight problem. A single AI compute rack can cost more than $1 million and weigh up to 4,000 pounds. Racks contain Class 9 dangerous goods (lithium-ion batteries) and liquid cooling systems requiring specialized handling. Construction sequencing failures delayed 57% o
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a Rack Power Engineer with deep expertise in high-power conversion and distribution to design, qualify, and support power systems for AI supercomputers. You will own rack power solutions—including power shelves, AC/DC rectifiers, power supply units (PSUs), power management controllers (PMCs), and high-current distribution—from requirements and supplier development through deployment. You will also monitor fleet rack power health, lead debugging and root-cause investigations, and drive improvements into hardware, firmware, and qualification coverage. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own rack power architecture and requirements for high-power AI supercomputing systems, including power budgets, AC input interfaces, DC distribution, redundancy, efficiency, serviceability, and integration with data center infrastructure. Drive the design and supplier development of power shelves, rectifiers, PSUs, PMCs, busbars, connectors, and protection circuits. Review electrical designs and control behavior, and evaluate performance, cost, reliability, and availability trade-offs. Define and execute component, shelf, and rack qualification plans covering load transients, current sharing, hot-swap, startup and shutdown, redundancy failover, fault protection and recovery, thermal limits, and AC disturbances and ride-through
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
About the Team OpenAI's Industrial Compute organization is building and scaling the infrastructure required to support frontier AI. The Infrastructure Strategic Sourcing team connects technical and project requirements to supplier readiness, contracting, purchasing, equipment delivery, and portfolio-level risk visibility across owner-furnished contractor-installed equipment (OFCI), data center networking, rack systems and integration, fiber, cabling, optical interconnects, and related infrastructure. The team partners across Pre-Construction, Design, Construction, Electrical and Mechanical Engineering, Network Engineering, Hardware and Rack Delivery, Strategic Sourcing, Procurement, Legal, Finance, Accounts Payable, Logistics, and external suppliers. We build the operating mechanisms that keep sourcing decisions, purchase execution, long-lead equipment, network and fiber dependencies, rack readiness, and delivery commitments aligned to infrastructure schedules. About the Role We are seeking an Infrastructure Sourcing Operations Lead to own procurement operations across pre-construction, design, construction, and sourcing through purchase order issuance, while maintaining visibility through invoice resolution, production, logistics, delivery, installation, and readiness. The portfolio includes electrical and mechanical OFCI, networking equipment, rack systems and integration, fiber, cabling, optical interconnects, and other infrastructure required to bring capacity online. In this role, you will set priorities, make or escalate decisions that affect cost, supplier relationships, contractual position, and delivery schedules, and define the standards used by execution support for queue management, documentation, tracker maintenance, and recurring reporting. Success requires sound commercial and program judgment, operational rigor, systems thinking, and the ability to turn incomplete information across vendors, tools, and project teams into clear decisions, accountable
Other cities to consider
More places hiring for this role
Get new rack scale software architecture director jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime