Jobiba hiring network

Datacenter Liquid Cooling Architect Jobs

15 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current datacenter liquid cooling architect jobs. Use filters to narrow by work mode, employment type, experience and date posted.

T
Tenstorrent
📍 Toronto• $100K – $500K/yr
1 day ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As a Datacenter Liquid Cooling Architect, you will define, design, and architect next-generation liquid cooling infrastructure for Tenstorrent’s large-scale AI training and inference clusters. You will partner with systems engineering, mechanical engineering, software, and cross-functional design teams to develop chassis-, rack-, and cluster-scale cooling solutions, including CDU integration, telemetry and control, leak detection, and resilient operating strategies. This role will help shape reliable AI datacenter architectures and deployments for both internal and external customers. This role is on-site, based out of Toronto, Canada, Austin, Texas or Santa Clara, California. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A datacenter and system thermal design professional with 10+ years of experience architecting cooling infrastructure for complex computing environments. An experienced liquid cooling architect who can design chassis- and rack-scale solutions for large AI training and inference clusters. A systems thinker who understands how mechanical, electrical, software, facility, and systems engineering decisions come toge

F
Flexport
📍 San Francisco• Full-time• From $350K/yr
1mo ago

About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. Help Win New Business The Opportunity: We are scaling our dedicated Data Center practice, and we are looking for the person who will lead it. This is a founding commercial role. You will start as a team of one, owning the full sales motion end-to-end, and you will build the team around you as the practice grows. You will define how Flexport goes to market with hyperscalers, hardware OEMs, and the broader ecosystem. You will set the playbook, win the first marquee accounts, and hire the people who scale what you build. Reporting directly to the Regional General Manager, you will operate with a high degree of autonomy and direct access to executive leadership. This role features a 50/50 compensation model (Base + Uncapped bonus, with accelerators), with OTEs of $350k+, designed for strong leaders who take bets on themselves. Why This Role Is Different: The data center logistics market is not a standard freight problem. A single AI compute rack can cost more than $1 million and weigh up to 4,000 pounds. Racks contain Class 9 dangerous goods (lithium-ion batteries) and liquid cooling systems requiring specialized handling. Construction sequencing failures delayed 57% o

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. The Infrastructure team is central to this mission, setting the core strategy and implementing the vision. From site selection to deployment to operations, this team sits at the intersection of commercial, technical, and operational domains, interacting with experts and executives inside and outside of OpenAI. We design and operate mission-critical facilities that support cutting-edge AI workloads at scale. About the Role We are seeking a Facilities Operations Lead to support the commissioning, deployment, and long-term operation of our next-generation AI data centers. This role bridges the interface between data center construction and hardware landing, ensuring seamless integration of mission-critical infrastructure with hardware deployment timelines. You will define and execute commissioning plans, support infrastructure bring-up, and take ownership of operations and maintenance for cutting-edge, large-scale, AI data centers. You will collaborate closely with design, construction, and hardware teams to define repeatable processes for new data center builds and lead hands-on operations to uphold the performance and reliability of our deployed infrastructure. Key Responsibilities Define and execute sequences of operations, commissioning steps, and bring-up processes for mission-critical data center facilities. Interface with the design and hardware teams to define deployment procedures tailored to each data center and hardware configuration. Oversee installation, commissioning, and operational readiness of large-scale data center campuses. Manage monitoring, maintenance, and quality control of the data center infrastructure, including high-performance liquid cooling systems. Develop on-site operations staffing strategy. Develop and enforce procedures for planed and unplanned downtime and SLAs for critical

awsrestai
View job →

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the hardware platforms that support the next generation of AI systems. The Opportunity As a Systems Engineering Intern, you will contribute to projects that combine hardware, firmware, and software engineering for advanced AI compute platforms. You will work with experienced engineers on subsystem design, laboratory testing, system validation, automation, and performance analysis. The internship provides hands-on experience with modern hardware development and system-level engineering. You will own clearly defined technical tasks with guidance from the team and document your methods, results, and conclusions. What You Will Do Support the design and testing of CPU and high-speed input and output subsystems for advanced compute platforms. Run laboratory tests and measurements to help evaluate performance, power, signal behavior, and reliability. Contribute to system-level validation by creating scripts and tools that streamline testing, data collection, and analysis. Explore emerging input and output technologies, including PCIe 6.0 and 800G Ethernet, and learn how they support advanced computing workloads. Assist with investigations into platform power, cooling, and energy efficiency, including liquid-cooling systems for high-performance processors. Use power meters, oscilloscopes, logic analyzers, or comparable lab equipment under appropriate supervision. Analyze test results, identify unexpected behavior, and work with engineers to reproduce and investigate issues. Collaborate across hardware, firmware, software, m

pythonartificial intelligenceai
View job →

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the hardware platforms that support the next generation of AI systems. The Opportunity We are looking for a recent graduate or early-career engineer to join the Hardware Platform Development team as a Graduate Systems Engineer. You will contribute to the design, integration, validation, and performance analysis of advanced AI compute platforms. You will work with experienced hardware, firmware, software, mechanical, thermal, and systems engineers throughout the development lifecycle. The role combines subsystem engineering, hands-on laboratory work, test automation, data analysis, troubleshooting, and clear technical documentation. Start: September, 2027 Location: Austin, Texas, USA What You Will Do Contribute to the design, integration, and testing of CPU and high-speed input and output subsystems for advanced compute platforms. Take ownership of defined engineering tasks from requirements and test planning through execution, analysis, and technical review. Develop system-level validation plans, procedures, scripts, and tools that improve test coverage, repeatability, data collection, and analysis. Evaluate platform performance, power, signal behavior, reliability, and interoperability using laboratory measurements and system data. Investigate emerging input and output technologies, including PCIe 6.0 and 800G Ethernet, and assess their use in advanced computing systems. Support platform power, cooling, and energy-efficiency investigations, including liquid-cooling systems for high-performance processor

pythonartificial intelligenceai
View job →

Mechanical and Thermal Laboratory Technician Position Summary Graphcore is a globally recognized leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data center hardware that provide the specialized processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Austin, which will play a central role in Graphcore's work building the future of AI computing. Responsibilities The Mechanical and Thermal Laboratory Technician is a hands-on technical role supporting the development and validation of advanced AI hardware systems for data center environments. Working as part of a cross-functional engineering team, this individual will be responsible for executing mechanical and thermal laboratory testing, supporting product validation activities, prototype fabrication and assisting with troubleshooting and root-cause analysis of complex hardware systems. Requirements Associate degree in Mechanical Engineering Technology or a related technical field preferred. Equivalent combinations of education, training, and relevant experience will be considered, including experienced non-degreed candidates or candidates with degrees in unrelated disciplines. Minimum of 5 years of experience working in mechanical laboratories, machine shops, test labs, or similar technical environments. Experience with server hardware platforms and data center equipment. Knowledge of Direct Liquid Cooling (DLC) systems and their implementation in server environments. Experience operating forklifts, pallet jacks, and other material-handling equipment. Key Responsibilities Execute mechanical and thermal test plans to validate hardware designs a

aisemtraining
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
27 days ago

About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. In this role you will: As a Hardware Test Engineer, you will work on Machine Learning/AI hardware system projects to craft the solutions for current and future data center deployments. You will bring a strong understanding of hardware system testing, excellent project management skills, and the ability to collaborate across multiple teams to ensure efficient lab operations. You will be responsible for designing, implementing, and executing comprehensive test plans that ensure the reliability, performance, and scalability of our supercomputing hardware systems. You will develop detailed test plans and methodologies tailored to hardware components, including processors, memory modules, custom accelerators and interconnects. You will collaborate with hardware design, manufacturing, firmware teams and vendors to identify, analyze, and resolve issues affecting hardware, power, thermal and high-speed interconnects. You will perform in-depth debugging on the hardware system Excellent analytical skills to diagnose hardware issues, troubleshoot problems, and propose solutions. Ability to interpret complex test data, identify trends, and draw meaningful conclusions. High-speed links, with a focus on SerDes (Serializer/Deserializer) technology to assess signal integrity, error rates, and overall link performance. You will collaborate with the lab manager to maintain the equipment and hardware systems, including oscilloscopes, thermal test chambers, liquid cooling systems, and other mea

REMOTEpythonawsrest
View job →
O
1mo ago

About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des

awsrestai
View job →

About the Team: OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. Our Stargate program develops and deploys massive, state-of-the-art data center campuses in partnership with industry leaders today—and through future OpenAI infrastructure projects tomorrow. We design for scale, speed, and reliability, and we need experienced technicians who can translate network blueprints into physical reality. About the Role: We are seeking a Senior Data Center Networking Technician who thrives in fast-moving build environments and is eager to roll up their sleeves during active datacenter deployments. Your first assignment will focus on the physical bring-up of network infrastructure at a large partner-operated campus, collaborating with partner teams and their delivery vendors to achieve agreed performance and reliability targets. As that campus reaches steady state, you will transition to lead network deployment for future OpenAI data center projects, defining standards and guiding implementation across multiple locations. Candidates must be able to sit onsite in Abilene, Texas 5 days per week Key Responsibilities Serve as OpenAI’s technical lead technician during the current campus build, partnering with internal engineers and external contractors on design reviews, installation plans, and acceptance criteria. Spend significant time on the data-center floor performing inspections, assisting with cable routing/termination when needed, conducting fiber testing (OTDR, power levels, continuity), and resolving installation challenges in real time. Troubleshoot and optimize cabling routes, patching, and equipment turn-up to ensure clean, reliable handoff to network operations. Contribute to design discussions and peer reviews for structured cabling and physical network layouts, providing practical field feedback to engineering teams. Develop repeatable engineering standards, as-built do

pythonawslinux
View job →

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Infrastructure Quality team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, general contractors, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. Our work spans from vendor qualification through commissioning, ensuring operational readiness across our global portfolio. About the Role We are seeking an experienced Manufacturing Quality Engineer (MQE) to establish, implement, and manage a manufacturing-focused quality program for datacenter infrastructure. This role will be responsible for vendor oversight, quality assurance, process improvement, and issue resolution for all critical systems. You will lead vendor audits, monitor performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced risks, and operational reliability. By partnering with vendors, construction teams, and internal stakeholders, you will help ensure OpenAI’s datacenters are delivered on time and built to the highest operational standards. Travel Domestic and international travel as needed (estimated 40–60%) to manufacturing sites, datacenter locations, and partner facilities. Key Responsibilities Vendor Oversight & Performance Management Conduct manufacturing evaluation, audits, and improve vendor performance across production, inspection, testing, and delivery phases. Develop and track quality metrics to assess manufacturing performance and identify trends. Partner with vendors to refine processes, training, and quality controls to mitigate risks before shipment. Program Development & Execution Develop and maintain a datacenter-focused m

awsrestai
View job →
O
1mo ago

About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to support the next generation of AI at unprecedented scale. Through a combination of strategic partnerships and self-built campuses, we are developing and operating large-scale data center infrastructure across power, cooling, networking, compute, construction, and site operations. The scale and complexity of this infrastructure introduces a broad range of environmental, health, and safety considerations across site development, design, construction, equipment deployment, commissioning, and ongoing operations. EHS is a critical part of how we build infrastructure that is safe, resilient, compliant, and capable of operating at scale. About the Role We are seeking an EHS Lead to establish and drive environmental, health, and safety strategy across OpenAI’s rapidly expanding compute infrastructure portfolio. This role will develop the EHS framework for large-scale data center development and operations, partnering closely with engineering, construction, infrastructure delivery, facilities, operations, security, legal, environmental, and external development partners. The EHS Lead will help ensure that safety and environmental considerations are embedded into projects from early design and site development through construction, commissioning, and operations. The role will establish standards and operating mechanisms, assess and mitigate risks, oversee EHS performance across internal teams and third-party partners, and provide technical leadership on complex or high-consequence safety issues. Success requires the ability to operate strategically while maintaining strong technical depth and executional rigor in fast-moving, highly complex infrastructure environments. In this role, you will: Develop and own EHS strategy, standards, programs, and operating mechanisms across OpenAI’s data center and compute infrastructure portfolio. Establish scalable EHS requirements for site de

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

awslinuxrest
View job →
N
23 hrs ago

The Silicon Co-Design Group (SCG) sits at the crossroads of architecture, design, marketing, operations, and productization. Our work spans early architecture through final product delivery across Datacenter, Gaming, Robotics, Automotive, and Embedded markets. We work closely across functions to deliver chips that change what is possible. NVIDIA’s Silicon Co-Design Group is hiring a Chip Lead to serve as the technical lead for one of our most consequential silicon programs! This is not project management, and it is not a senior IC role. You are the person the program partners with on its hardest technical questions — when the right answer is not obvious, and when leadership needs a single technical perspective to align on direction. You are accountable for the technical integrity of the chip end-to-end. You guide co-design feature integration, help resolve the toughest multi-functional bugs, and serve as the project-specific custodian of the qualification playbook. The program runs cleaner because you are on it! What you’ll be doing: You will partner across design, validation, software, and manufacturing to keep the program’s technical narrative clear and on track. Day to day, you will: Serve as the single technical point of contact for multi-functional decisions, issues, and trade-offs. Co-lead program-level feature integration from chip to system, surfacing inter-function dependencies and guiding them to resolution. Help resolve the program’s hardest multi-functional bugs by translating ambiguous, multi-team symptoms into root-cause closure on areas such as HBM, power and thermal, high-speed I/O, and packaging. Steward the qualification playbook. When the playbook does not fit a situation, guide the mitigations and capture the lessons as reusable methodology for other SSG programs. Shape the program’s technical narrative by surfacing key risks, trade

aiproject management
View job →
🔔

Get new datacenter liquid cooling architect jobs by email

Daily job updates · Unsubscribe anytime