We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary We are seeking an accomplished Principal Cloud Storage Engineer to lead the design, engineering, and evolution of our private cloud storage platforms. This role will focus on large-scale storage architecture, data protection, cyber recovery, and resiliency technologies across complex enterprise environments. The ideal candidate will combine deep technical expertise in storage systems with strong leadership, architectural vision, and the ability to influence technical direction across the organization. Key Responsibilities Architect and engineer enterprise storage platforms that ensure data integrity, availability, security, and disaster recovery readiness Design and implement end-to-end storage solutions, including Software Defined Storage, SAN, NAS, and object storage across private cloud and data center environments Drive strategic technology decisions by evaluating emerging products, tools, and standards supporting storage, data protection, cloud, and compute platforms Lead infrastructure initiatives involving storage modernization, data protection, cyber recovery, data migration, and resilience engineering Develop and execute enterprise strategies for backup, recovery, cyber vaulting, and business continuity Create and maintain comprehensive documentation of storage architectures, configurations, policies, and operation
Jobs in United States
Datacenter Liquid Cooling Architect in United States
151 active opportunities · Updated October 2026
Showing
15 jobs
Explore current datacenter liquid cooling architect jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Are you a person who likes to work in a fast-paced organization? NVIDIA is the world leader in Visual Computing. We are passionate about four markets: Gaming, Automotive, Enterprise Graphics and HPC/Cloud Datacenters; in addition to our traditional OEM business. We are well positioned as the ‘AI Computing Company’, and our GPUs are the brains powering modern Deep Learning software frameworks, accelerated analytics, big data, modern data centers, smart cities, and driving autonomous vehicles. We have some of the most forward-thinking and talented people on the planet working for us. If you're forward-thinking, hardworking, driven and if working with extraordinary people across countries sounds interesting, this job is for you. We are now looking for a Human Resources Business Partner to provide HR support onsite in Santa Clara, CA for our Worldwide Field Organization in a dynamic and collaborative environment. This is a global organization, and we are looking for someone to be passionate about supporting and building strategies to enable NVIDIA to achieve success. You’ll partner with a cross-functional group of subject matter experts to design and execute strategies for how we staff, onboard, develop, motivate, retain and organize work. You will need excellent communication skills, critical thinking and planning ability, and the agility to function in a fast paced and innovative environment. What you'll be doing: This position will be an integral enabler of the mission of our Field organization. In this position you will work with the senior leaders and leadership teams within NVIDIA organizations to develop and execute the HR strategies that champion organizational and people effectiveness. You will think strategically as well as roll up your sleeves and dive deep into practical application. You must understand business priorities and translate them into an HR ag
NVIDIA is seeking an a PCB Library Engineer to join our PCB Design Infrastructure team. In this role, you will help develop and maintain the PCB library assets used across NVIDIA's Data Center, AI, Networking, Automotive, and Graphics products. Working alongside experienced PCB designers, library engineers, mechanical engineers, manufacturing engineers, and component engineers, you will create and validate component footprints, schematic symbols, mechanical components, panel definitions, and other critical design assets that enable successful product development. This position provides an excellent opportunity to build expertise in PCB design, manufacturing, component engineering, and design automation while supporting some of the most advanced computing platforms in the world. What you'll be doing: Develop PCB footprints, padstacks, schematic symbols, and mechanical library content using Cadence PCB design tools. Review component datasheets, package drawings, and engineering specifications to create accurate design libraries. Support library verification, release, and documentation processes. Partner with PCB design, mechanical engineering, manufacturing engineering, and operations teams to resolve library-related issues. Learn and apply industry standards, including IPC requirements, DFM, DFA, and DFT principles. Support quality initiatives to ensure library content is accurate, manufacturable, and scalable. Participate in continuous improvement and automation efforts within the library environment. Develop a strong understanding of PCB fabrication, assembly, and component technologies. What We Need to See: BS degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, Manufacturing Engineering, or a related field or equivalent experience. <p
About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to support the next generation of frontier AI systems. Through a combination of strategic partnerships and self-built data center campuses, we are scaling the physical infrastructure needed to deliver compute at unprecedented scale. The Commissioning organization is responsible for ensuring this infrastructure is safely tested, validated, integrated, and transitioned into reliable operations. As the portfolio grows, the team is building common standards, processes, tools, and reporting systems that allow commissioning programs to operate consistently across projects while giving teams and leadership clear visibility into readiness, risk, and execution. About the Role We are seeking a Commissioning Program Manager to build and scale the operating systems behind OpenAI’s infrastructure commissioning programs. You will own the development and continuous improvement of commissioning standards, processes, tools, dashboards, and KPIs across the infrastructure portfolio. You will work closely with commissioning and construction teams to translate field execution needs into practical playbooks, workflows, templates, metrics, and reporting mechanisms that teams can use from construction readiness through testing and turnover. This role sits at the intersection of infrastructure delivery, program management, process design, and data. The ideal candidate understands how complex construction projects operate and can turn fragmented workflows and project data into repeatable systems that improve execution without creating unnecessary administrative burden. Key Responsibilities Develop and maintain commissioning program standards, playbooks, process maps, templates, checklists, stage gates, and acceptance criteria across infrastructure projects. Establish consistent workflows for commissioning planning, construction readiness, QA/QC, issue management, document control, testing evidence
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You’ll own the hands-on and automation work that brings WAN, fiber, carrier, and cloud-interconnect circuits into service. Partner with network engineers, fiber providers, cloud service providers, colocation teams, and data-center technicians to move each connection from ordered and patched to verified, stable, and ready for handoff. You’ll own Layer 1 troubleshooting and circuit bring-up while building workflows that translate reliable system or model output into precise, approved technician actions, capture field feedback, and drive each connection to a green-port handoff. The right person combines strong physical-networking judgment with practical automation skills: patch-panel and port mappings, optics and light levels, provider coordination, structured operational data, API or scripting workflows, and human-in-the-loop LLM tooling. Responsibilities Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence. Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports. Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching
About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We're seeking a Security Engineer to join our First-Party Hardware team. In this role, you will own the end-to-end security foundation for OpenAI's first-party AI hardware systems, working across hardware security, embedded security, system security, and practical deployment at data center scale. You will partner with silicon, hardware, firmware, infrastructure, manufacturing, operations, and security teams to define and deliver system-level device trust. This includes boot integrity, device identity, provisioning, attestation, management-plane security, storage encryption, debug controls, firmware update and recovery, RMA, and decommissioning. You will be accountable for turning threat models into requirements, requirements into implementation, and implementation into validation evidence that can support launch decisions. Location: San Francisco, CA (Hybrid: 3 days/week onsite) Relocation assistance available. In this role, you will: Own security requirements, threat models, validation strategy, and launch-readiness evidence for first-party hardware platforms from early design through production deployment. Design and review secure boot, measured boot, roots of trust, platform firmware resilience, firmware signing, recovery, and anti-rollback strategies across heterogeneous devices. Own device identity, provisioning, enrollment, attestation, certificate lifecycle, and key-management requirements across manufacturing and data center bring-up. Harden management
About the Team The Finance Platform & Technology team at OpenAI builds and scales the future-proof systems and data architecture that power our core financial operations. We enable business agility, compliance, and operational excellence across quote-to-cash, procure-to-pay, inventory, and asset management for both B2B and B2C. Our focus is on modernizing workflows through strategic integrations, scalable automation, and seamless data flows empowering smarter decisions, reliable reporting, and sustainable growth as OpenAI evolves About the Role We are looking for a Supply Chain Transformation Architect to redesign and modernize our end-to-end supply chain operations supporting robotics, consumer hardware, and data center infrastructure. This role sits at the intersection of process, systems, and data. Your primary focus will be transforming supply chain processes across planning, procurement, manufacturing, logistics, and fulfillment—then enabling those processes with the right systems architecture, data foundation, and AI-driven automation. You will help move the organization from manual, reactive operations to intelligent, data-driven supply chain execution. In this role you will: Lead End-to-End Supply Chain Transformation Evaluate and redesign core supply chain processes across demand planning, supply planning, procurement, manufacturing operations, logistics, and fulfillment. Identify operational bottlenecks, fragmented workflows, and manual processes that limit scalability. Build standardized process frameworks and operating models that support rapid scaling of hardware programs. Drive Operational Excellence Implement structured supply chain practices such as: S&OP / Integrated Business Planning Supply risk management Inventory optimization Supplier collaboration frameworks Logistics visibility and execution models Establish operational KPIs and governance to improve predictability, responsiveness, and resilience. Architect the Digital Supply Chain Tra
About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role OpenAI's Hardware organization builds supercompute platforms from silicon and boards to full rack-scale systems to power advanced AI workloads. This role owns end-to-end quality for high-speed interconnect hardware across the product lifecycle: early design influence, supplier/contract manufacturer readiness, qualification, ramp, and fleet quality in lab and data center environments. You will be the quality lead for advanced interconnect components and assemblies, including high-speed copper cables, cable cartridges, patch panels, backplane/cable-backplane solutions, high-speed connectors, and related electro-mechanical interfaces. You will partner closely with electrical, mechanical, SI/PI, systems, reliability, operations, and external vendors to prevent escapes and drive rapid, data-driven containment and corrective action. In this role you will: Own quality for advanced interconnect components and assemblies: high-speed connectors, high-speed copper cables, cable cartridges (e.g., cable cassette style assemblies), patch panels & optics, and backplane/cable-backplane interconnect solutions. Drive quality-by-design: participate in design reviews, DFM/DFx, tolerance stacks, material and plating selections, connector mating strategy, strain relief, and assembly methods to reduce variation and field failures. Define and track quality and reliability metrics (DPPM, yield, escapes, RMA/FRACAS trends, Cpk/Ppk where applicable) for interconnects across NPI and m
About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary CVS Health is seeking a highly skilled and visionary Sr. Manager to lead and manage our enterprise Cisco Application Centric Infrastructure (ACI) environment and the transition to Cisco Nexus Dashboard Fabric Controller (NDFC). This role sits within the Network Engineering team and is pivotal to driving the modernization and scalability of our next-generation data center fabric. This role will help set and drive the network technology strategy for ISTS, ensuring that ISTS delivers on our mission to transform technology and provide an agile, cost optimized and resilient set of network infrastructure services to meet the evolving needs of all the CVS Health lines of business. The ideal candidate will have experience supporting complex enterprise network environments and demonstrate deep expertise in Cisco ACI and NDFC technologies. As a strategic leader, you will be responsible for setting the direction, leading a team of engineers, and ensuring the operational integrity, performance, and evolution of our fabric-based data center architecture. This position will also mentor staff members in an effort to develop excellent enterprise networking skills. Key Responsibilities Lead the deployment, administration, and lifecycle management of the Cisco ACI environment Oversee the strategic transition to and ongoing management of the Cisco
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. At Micron, we are transforming how the world uses information to enrich life. The High Bandwidth Memory (HBM) Design team develops industry-leading memory solutions that enable advances in Artificial Intelligence, high-performance computing, graphics, and data-center applications. Our engineers collaborate across global teams to deliver innovative memory architectures and semiconductor technologies that power next-generation computing systems. As an HBM Design Engineer Intern, you will work alongside experienced memory engineers and gain hands-on experience in semiconductor design, simulation, verification, and analysis. This internship provides exposure to industry-standard design methodologies, EDA tools, and cross-functional collaboration throughout the product development lifecycle. You will also have opportunities to apply AI-Assisted and AI-Enabled workflows to improve engineering productivity, debug efficiency, and design quality. Responsibilities Assist with the design, simulation, analysis, and verification of HBM memory and logic circuits using industry-standard semiconductor design tools. Support RTL development, circuit implementation, timing analysis, power analysis, and functional validation activities. Develop scripts, automation solutions, and AI-Assisted workflows to improve design productivity, debug efficiency, and design-flow quality. Collaborate with Design, Verification, Physical Design, CAD, and Product Engineering teams on technical reviews, debug activities, and project deliv
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. At Micron, we are transforming how the world uses information to enrich life. The High Bandwidth Memory (HBM) Design team develops industry-leading memory solutions that enable advances in Artificial Intelligence, high-performance computing, graphics, and data-center applications. Our engineers collaborate across global teams to deliver innovative memory architectures and semiconductor technologies that power next-generation computing systems. As an HBM Design Engineer Intern, you will work alongside experienced memory engineers and gain hands-on experience in semiconductor design, simulation, verification, and analysis. This internship provides exposure to industry-standard design methodologies, EDA tools, and cross-functional collaboration throughout the product development lifecycle. You will also have opportunities to apply AI-Assisted and AI-Enabled workflows to improve engineering productivity, debug efficiency, and design quality. Responsibilities Assist with the design, simulation, analysis, and verification of HBM memory and logic circuits using industry-standard semiconductor design tools. Support RTL development, circuit implementation, timing analysis, power analysis, and functional validation activities. Develop scripts, automation solutions, and AI-Assisted workflows to improve design productivity, debug efficiency, and design-flow quality. Collaborate with Design, Verification, Physical Design, CAD, and Product Engineering teams on technical reviews, debug activities, and project deliv
We are seeking a highly skilled and hard-working Senior Test Developer / test engineer to join our multifaceted Enterprise Software QA team. This role offers an outstanding opportunity to leave your mark on the design, construction, optimization and testing of large-scale infrastructure for various foundational NVIDIA unified cloud services and data center offerings. If you are a dedicated engineer with strong expertise in cloud infrastructure and distributed systems and want to apply your skills with AI tools, this role could fit you perfectly. You will thrive in an exciting, innovative environment. What you'll be doing: Work with development teams on test plans for all layers of SW stack for cloud infrastructure, execution, reviews, failure analysis and assessing overall quality and risk. Work with customer PMs on software issues including technical feedback from OEMs and CSPs. Develop key benchmarks to track execution and deploy process improvements to improve efficiency Leverage AI skills to expedite the test scope, test plan, execution and automation workflows. Lead NVIDIA Cloud and Data Center bring up activities which will involve validation, reporting, working with engineering to debug issues, providing design input at times, adding coverage in different areas. Design, develop and maintain CI/CD pipelines for continuous testing in cloud environments when needed. Perform performance, scalability, and reliability testing of cloud services. Implement and maintain test environments in cloud platforms such as AWS, Azure, or Google Cloud. Supervise the infrastructure to alert on significant events, ensuring the highest level of system performance and reliability. Work with various different partner teams to ensure availability of clusters to test on and take the lead in resolve all issues. Working with tea
Other cities to consider
More places hiring for this role
Get new datacenter liquid cooling architect jobs in United States by email
Daily job updates · Unsubscribe anytime