Jobs in United States

Datacenter Liquid Cooling Architect in United States

151 active opportunities · Updated October 2026

Explore current datacenter liquid cooling architect jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

DC
📍 New York, New York, United States· Full-time
✓ High-confidence listing

From $131K/yr

Quick readStrong listing-quality and freshness signals

Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production

PythonAWSAzureCI/CD
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for an experienced System Level Test Engineer to join our Product Test and Diagnosis Department (PTD). In this role, you will lead the development and deployment of System Level Test (SLT) solutions for next-generation AI processors. Working closely with cross-functional teams, you will contribute to the design and implementation of SLT hardware, software, automation, and characterization solutions that support silicon bring-up, yield learning, manufacturing readiness, and production deployment. The ideal candidate will possess strong technical depth in semiconductor test and validation, a passion for solving complex engineering challenges, and a strong focus on product quality and manufacturability. The Team The Product Test and Diagnostics team’s role is to detect and manage hardware defects that arise from the manufacture and use of our products. This covers chips, boards and finished systems and takes place both in the manufacturing sites and in the field. Responsibilities and Duties Lead development and deployment of SLT hardware and software solutions supporting silicon bring-up, charac

PythonAISEMHR
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Manufacturing Test Engineering Manager Position Summary We are seeking an experienced Manufacturing Test Engineering Manager to lead the development and execution of the end-to-end manufacturing test strategy for next-generation AI server platforms and datacenter infrastructure. This role is responsible for defining and driving the manufacturing test architecture from L6 board assembly through L11 rack-level integration and final system validation , ensuring world-class product quality, manufacturability, and production scalability. This leader will manage a team of 3–5 Manufacturing Test Engineers while partnering closely with Hardware, Firmware, Platform, Validation, Quality, Operations, and Joint Design Manufacturing (JDM) partners. The role owns the manufacturing test strategy, test coverage, factory test infrastructure, manufacturing capacity planning, and continuous improvement of manufacturing quality. Key Responsibilities Manufacturing Test Strategy Define and own the end-to-end manufacturing test strategy from L6 board assembly through L11 rack integration . Develop standardized manufacturing test methodologies that optimize quality, throughput, cost of test, and scalability across multiple products and JDM sites. Establish manufacturing test standards, best practices, and engineering processes that support high-volume server manufacturing. Technical Leadership & People Management Lead, mentor, and develop a team of 3–5 Manufacturing Test Engineers supporting multiple hardware programs. Establish team priorities, allocate resources, and ensure successful execution of manufacturing test deliverables. Foster a culture of technical excellence, accountability, collaboration, and continuous improvement. Serve as the primary technical escalation point for manufacturing test and production issues. Cross-Functional Engineering Collaboration Partner with Hardware, Platform, Firmware, Validation, Reliability, Quality, and Operations teams to ensure manufac

PythonLinuxAIExcel
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for an experienced Silicon Test Engineer to join our Product Test and Diagnosis Department (PTD). This is a pivotal role and will involve building a team of engineers to develop System Level Test (SLT) capability within the company. Working closely with a cross-functional team you will implement SLT tests for a family of next generation AI Processors. The ideal candidate should have a focus on quality and demonstrate a good understanding of the importance of production test on the success of a product. T hey will have a proven Functional Test or ATE Test Engineering background, and will have a pragmatic, hands-on and flexible approach to a fast-changing environment. The Team The Product Test and Diagnostics team’s role is to detect and manage hardware defects that arise from the manufacture and use of our products. This covers chips, boards and finished systems and takes place both in the manufacturing sites and in the field. Responsibilities and Duties Managing a team of

PythonGitAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa

AWSRestAIGo
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA is seeking a Senior Technical Program Manager to join the CSP Engagements team, focused on deep technical engagement with hyperscale cloud service providers for NVIDIA’s next‑generation datacenter systems such as Vera Rubin NVL72. This role is intended for experienced systems and embedded software leaders—including software engineering managers, technical leads, or senior architects—who have led datacenter server and platform software programs and can operate as a trusted technical partner to hyperscale CSP engineering teams. As a member of the CSP Engagements team, you will act as the primary technical engagement leader between NVIDIA’s system software organizations and CSP platform, system software, and AI teams, ensuring alignment, readiness, and successful large‑scale deployment of NVIDIA‑based datacenter solutions. What you will be doing: Lead deep technical engagements with hyperscale CSPs as the primary NVIDIA point of contact for system software, firmware, and platform readiness for NVIDIA datacenter products. Partner directly with CSP system software, firmware, and infrastructure engineering leaders to align on software architecture, bring‑up plans, deployment readiness, and production requirements for NVIDIA‑based server and rack‑scale platforms. Represent CSP technical priorities internally, advocating for customer requirements and tradeoffs across NVIDIA’s system software, firmware, hardware, silicon, and product teams are aligned to customer needs, timelines, and constraints. Own the end‑to‑end CSP engagement lifecycle, from early technical alignment and pre‑production readiness through large‑scale deployment, escalation management, and sustained production support. Drive bi‑directional technical communication: translating CSP system‑level requirements into actionable focus areas for NVIDIA engineering teams, while clearly communicating N

LinuxArtificial IntelligenceAI
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

We are developing advanced multi-rack, multi-tenant AI/ML datacenters with NVIDIA GB200, and upcoming GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team to focus on the cloud-native stack for datacenter products like GB200. In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex scheduling challenges across racks, tenants, and clouds as part of the CSP engagements team. What you’ll be doing: Perform deep-dive debugging of multi-rack, multi-tenant clusters: scheduler behavior, container runtime issues, device-plugin crashes, RDMA/IB fabric anomalies, etc. Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom micro-services that expose new GPU capabilities. Drive joint architecture reviews and “whiteboard” sessions with CSP and internal platform teams; convert findings into RFCs and upstream pull requests. Create reproducible testbeds (Helm/Ansible/Terraform) that mirror customer environments; automate validation and benchmark suites. Deliver technical collateral-design docs, how-to guides, demo scripts-and present at customer on-sites, KubeCon, and SlurmUG. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Strong source-level expertise in Kubernetes internals (scheduler, CRI/CNI/CSI, operators) and Slurm (federation, power-save, plugins). Hands-on experience integrating next-gen GPUs (Blackwell/GB200/GB300) or comparable accelerators into containerized clusters. Proven track record debugging large-scale, cloud-native stacks across ne

PythonKubernetesArtificial IntelligenceAI
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA is seeking a Senior Firmware Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware development with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you will be doing: Design and develop firmware solutions for manageability and observability of data center servers. Actively participate in hardware bring-up activities, OOB firmware development, protocol stacks (Redfish, PLDM, MCTP, NSM) and hardware-software co-design for Cloud Service Provider deployments. Debug and troubleshoot NVIDIA GPU firmware issues, power management, performance, and thermal control problems for data center deployments, providing active support to CSPs. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Deep expertise in embedded firmware, server management controllers, and hardware bring-up with proven track record of shipping production BMC solutions Strong knowledge of DMTF protocols (Redfish, IPMI, PLDM, MCTP, SPDM), telemetry frameworks, and out-of-band management architectures Expert-level skills in C/C&

Artificial IntelligenceAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa

AWSRestAIGo
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA Architecture Modeling group is looking for Architects, Functional Modeling Engineers, and Simulation experts to join various architecture efforts across GPU/ SOC Architecture teams. A key part of NVIDIA's strength is to innovate in parallel computing fields, delivering the highest performance in the world for high-performance computing. We are constantly looking for ways to improve our SoC and Systems architecture and maintain our leadership. In this position, you will be working with other world-class architects on modeling, analysis and validation of chip & system architectures and features that advance the state of art in performance and efficiency. What you'll be doing: Modeling and analysis of SoC & Systems algorithms and features, across datacenter, automotive, and client products Build and deliver platforms for SOC's that enable left shift for the SW teams aligned with project milestones Work closely with the SOC architects and guide modeling teams to deliver high-quality functional models that involve SOC+GPU use cases Collaborate with our EDA partners to align on customer-facing technologies Develop tests, test plans, and testing infrastructure for new architectures/features and code coverage analysis and reporting Ensure alignment between the various modeling teams at NVIDIA, GPU modeling teams, and modeling teams overseas What we need to see: Master’s or PhD in Computer Science, Electrical Engineering, Computer Engineering, or a related relevant field (or equivalent experience) with 5+ years of relevant work experience. Strong programming ability: C++, C along with a good understanding of build systems (CMAKE, make) , toolchains (GCC, MSVC) and libraries (STL, BOOST) Computer Architecture background with experience in modelling wit

PythonDockerAIJenkins
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA's Silicon Co-design Group (SCG) sits at a rare intersection: we own the full product development lifecycle, from early architecture definition through silicon bringup to product release. Our ArchDev team is the hub for silicon and system-level feature development, driving tradeoff analysis, system integration, and POR alignment across the entire organization. If you want to see your work go from whiteboard to world-class silicon, this is where that happens. What You'll Be Doing: Architect and integrate system-level performance and power management features, controllers, and policies to optimize product efficiency across datacenter and client products . Build feature roadmaps to address low-power, low-noise, and performance-per-watt product needs through prototyping, use-case analysis, and cost/benefit trade-offs. Partner with architecture, ASIC, board/platform, software/firmware, and marketing teams to drive design decisions and debug complex issues. Track industry trends and market needs and translate them into forward-looking roadmaps that keep NVIDIA's products ahead of the curve. Lead debug efforts, develop workarounds, and support bringup , validation, manufacturing, and customer escalations. What We Need to See: <

PythonLinuxAI
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

The Senior RMA Customer Experience Manager will work with the key customer account team leads to ensure customers have timely case updates. The position has to decide on fulfilment priority based on Service Level Agreements and track regional buffer stock. The manager needs to work with account leads to track the impact of open RMAs on customer datacenter capacity, so as to make allocation decisions. What you’ll be doing: Customer Satisfaction: Manage RMA cases and lead all aspects of replacement part fulfillment to clear the open RMA backlog per Service Level Agreements. Work with account teams to prioritize backlog in case of supply constraints. Case Management: Attend customers meeting, provide accurate and up-to-date information on case status, track of customer feedback and develop strategies to improve the overall customer experience. Metrics and Improvement: Generate case management metrics and analyze data to identify process improvement opportunities. Publish weekly status updates on open RMAs – supply commits, ready for pickup, shortage alerts. Collaboration: Work closely with quality and manufacturing teams to effectively communicate RMA updates and statuses to customer account teams. Facilitate seamless information flow between teams to provide exceptional customer support. Uptime and Buffer Stock: Track key customer installations uptime and work with planning team to maintain buffer stock near customer sites to proactively address potential supply shortages. Supply Allocation: Collaborate with production and repair factory planners to map out supply allocations for the open RMA backlog. Optimize resource allocation for timely resolution. Serve as Lead: Trainer and support for other case managers on program, process, customer interface and working internally as new programs develop for NPI and implemented into our global customer service team. What we need to see:

AISupply ChainCustomer Service
T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our Tensix team is building custom AI compute cores, RISC-V CPUs, and chiplet-based architectures for datacenter, edge, and automotive AI. Design Verification Engineers on this team validate compute IP and subsystems and build scalable DV infrastructure to keep verification fast, automated, and production-grade. This role is hybrid, based out of Toronto, ON, Austin, TX or Belgrade, Serbia. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced in modern verification methodologies with strong SystemVerilog skills and exposure to structured testbench development. Comfortable working from block-level to system-level verification and reasoning about microarchitecture behavior from specs and waveforms. Proficient in Linux environments with Python, or Bash scripting for automate builds, parse logs, manage CI pipelines. Skilled in coverage-driven verification and confident debugging across RTL, testbench, and workload scenarios. Motivated by AI hardware and eager to learn verification strategies for new architectures and domains. What We Need Contribute to verification of Tensix IP and subsystems from early planning through tape-out, owning cov

PythonAWSCI/CDLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

PythonAWSLinuxRest
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology and amazing people. Today, we're harnessing the boundless possibilities of AI to build the next era of computing. An era in which our GPU acts as the brain of computers, robots, and self-driving cars that can understand the world. Accomplishing unprecedented goals calls for imagination, inventiveness, and exceptional talent from around the world. As a NVIDIAN, you'll be immersed in a diverse, encouraging environment where everyone is inspired to do their best work. Join our team and discover how you can build a lasting impact on the world. NVIDIA designs the silicon behind AI, accelerated computing, and graphics. Power and thermal architecture decisions sit behind every watt of performance and every degree of thermal headroom! We are the Silicon Co-Design Group (SCG). We are hiring a Principal Architect to scout the research and industry landscape, identify emerging system-level co-design ideas in power and thermal, and drive them across teams into product differentiation across NVIDIA's roadmap. This role shapes system, platform, and datacenter feature/behavior, and partners with teams across Nvidia. SCG scope spans across architecture, design, software, operations, platforms, and productization. What you'll be doing: Architect next-generation system, platform, and datacenter-level power and thermal co-design solutions. Scan internal research, academia, standards bodies, and silicon, memory, packaging, and platform partners for what is emerging. Build the product differentiation case for each candidate idea, performance, power, reliability, schedule, cost, and brainstorm what is worth pursuing. Lead end-to-end co-design from concept to product. Drive alignment across architecture, VLSI, softw

Machine LearningAI
🔔

Get new datacenter liquid cooling architect jobs in United States by email

Daily job updates · Unsubscribe anytime