Are you ready to contribute to world-class innovation and push the boundaries of what's possible? At NVIDIA, you'll have the opportunity to be part of a team that is driving groundbreaking impacts across various markets. As a Thermal Solutions Development Engineer, you will play a pivotal role in our Silicon Codesign Group, transforming thermal solution concepts into lab-ready builds and beyond. What you will be doing: Build thermal solutions for engineering characterization and validation of next-gen GPU/SOC products, ensuring flawless delivery from concept to lab. Drive end-to-end development and deployment of thermal solutions, collaborating with internal teams and external vendors on build requirements, prototype evaluation, test system integration, and software automation. Improve thermal design processes by incorporating feedback and findings, developing workflow and maintaining our world-class standards. Work closely with system architects, chip and board designers, and software/firmware engineers in a dynamic and high-energy environment to bring industry-defining products to market. Apply AI-enabled approaches and AI tools to accelerate design iteration, test planning, and characterization/validation triage (e.g., requirements/spec summarization, experiment prioritization, log/telemetry summarization, anomaly/outlier detection), improving cycle time, coverage, and traceability while validating outputs against physics, specs, and lab measurements. Partner with AI/tooling teams as the thermal domain SME to define use-cases, success criteria, and evaluation methods; provide feedback to improve tool reliability and usability. What we need to see:
Jobs in United States
Senior Thermal Solutions Design Engineer in United States
1,941 active opportunities · Updated October 2026
Showing
15 jobs
Explore current senior thermal solutions design engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a “learning machine” that constantly evolves by adapting to new opportunities that are hard to solve, that only we can address, and that matter to the world. This is our life’s work, to amplify human creativity and intelligence. Make the choice to join us today. Design-for-X Engineering at NVIDIA works on groundbreaking innovations involving crafting creative solutions in AI for Chip Design and AI for Predictions in various use cases in manufacturing testing on some of the industry's most complex semiconductor chips. What you'll be doing: As a senior member in our team, you will work on innovating in the DFT Power, Thermal & Voltage Noise Methodology areas. This will include working on groundbreaking low power & thermal solutions for our manufacturing tests to be enabled at conditions that push the boundaries for our datacenter GPUs. You will work with multi-functional teams including Product Development & Power Architecture, implementing brand-new methodologies on hard-to-solve problems for improving our outgoing quality of chips. You will work on post-silicon data analysis for power to architect the next-gen solutions. In addition, you will help develop and deploy DFT methodologies for our next generation products using Applied ML & Gen AI solutions. You will also help mentor junior engineers on test designs and trade-offs including cost and quality. What we need to see: BSEE (or equ
NVIDIA is seeking a Senior Firmware Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware development with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you will be doing: Design and develop firmware solutions for manageability and observability of data center servers. Actively participate in hardware bring-up activities, OOB firmware development, protocol stacks (Redfish, PLDM, MCTP, NSM) and hardware-software co-design for Cloud Service Provider deployments. Debug and troubleshoot NVIDIA GPU firmware issues, power management, performance, and thermal control problems for data center deployments, providing active support to CSPs. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Deep expertise in embedded firmware, server management controllers, and hardware bring-up with proven track record of shipping production BMC solutions Strong knowledge of DMTF protocols (Redfish, IPMI, PLDM, MCTP, SPDM), telemetry frameworks, and out-of-band management architectures Expert-level skills in C/C&
Define the physical limits of what a chip can do, then break them! NVIDIA's Silicon Co-Design Group sits at the crossroads of architecture, silicon, systems, and manufacturing, where first principles thinking translates directly into product outcomes at scale. As GPU power density and thermal limits approach physical boundaries, the gap between system architecture needs and manufacturable, testable capabilities widens with every generation, and this role exists to close it. You wouldn't be maintaining existing solutions, you'd be identifying where physical constraints will erode NVIDIA's competitive edge before anyone else sees it coming, then building the cross-domain responses that ensure the next generation leapfrogs those limits entirely. The problems here don't have known answers yet: Every generation of NVIDIA silicon pushes closer to the limits of physics. The Co-Design Architect role isn't about optimizing within known bounds; it's about redrawing those bounds entirely. If you want your work to shape what's physically possible in the world's most demanding computing products, this is where that work happens. What you'll be doing: Power, Thermal & Packaging Limits Analysis: Gather and analyze the hard constraints of packaging, power delivery, and thermal dissipation against roadmap demands, predicting when those constraints will reduce competitiveness and acting before they do. Cross-Domain Solution Development: Develop solutions that span silicon features, packaging innovation, firmware and hardware co-design, and DFX updates, ensuring future products are not only performant but testable to the highest quality and reliability standards. Packaging Technology Evaluation: Assess emerging packaging technologies for their voltage/frequency, noise, reliability, and testability implications, determining which are worth adopting an
Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debug compl
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a best-in-class family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from a diverse group of backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a senior validation lead engineer to lead at-scale rack validation efforts for next-generation AI hyperscale systems. This role focuses on post-silicon system validation across the full lifecycle, ensuring functional, electrical, and thermal performance meets product objectives. You will own end-to-end blade and rack validation including planning, development, execution, and debug while collaborating across firmware, systems, and hardware teams. The Team The Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Responsibilities and Duties Lead post-silicon validation of AI compute blades and racks including test planning, development, and automation. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness. Own execution against program achievements and report validation progress and risks. Triage test failures, collect debug data, and collaborate on root cause analysis. Track
Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu
NVIDIA has been redefining computer graphics, desktop gaming, and enhanced computing capabilities for more than 25 years. Today, we are tapping into the unlimited potential of AI to define the next era of computing. As a NVIDIAN, you will work on problems that sit at the boundary of architecture, silicon, firmware, software, and production, where strong judgment matters as much as technical depth. We're the Silicon Design for Productization (DFP) Team, within the broader Silicon Co-Design Group, and we turn power and thermal design into executable productization methodology. Power and thermal are among the most complicated problems we work on at NVIDIA because they sit at the intersection of architecture, workload behavior, silicon variation, firmware policy, platform constraints, and product goals. Small decisions here have an outsized impact on performance, efficiency, reliability, bring-up speed, and ultimately what the product can deliver in the field. We define how features move from concepts to bring-up, characterization, validation, and release. In this role, you will help us build that bridge. We're looking for an engineer who reasons from first principles, flourishes with ownership in a fast-paced environment, and uses AI with sound judgment. What you’ll be doing: Lead the effort across multi-functional teams to keep the program’s power and thermal productization strategy clear, executable, and on track. Create methodology and silicon test plan based controller designs and architecture, including characterization process, debug tools, fuse/firmware settings and lab requirements. Drive resolution for challenging silicon issues through structured hypotheses, measurement plans, and root-cause closure. Steward the Power and Thermal playbook when the existing productization methodology
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company”. We are looking to grow our company, and grow our teams with the smartest people in the world. What you’ll be doing: You will work with ground breaking technologies for the Tegra SoC and various NVIDIA embedded platforms Implement power and thermal management software features in Linux Kernel and user space Collaborate with power architects, hardware and software engineers on platform power estimation and optimization Optimize the software stack to improve performance, efficiency, and responsiveness for edge AI and robotics use cases. Focus on improving compute and memory utilization, reducing latency and power consumption, and tuning system-level performance to deliver reliable and scalable AI workloads across demanding real-world edge environments. What we need to see: MS in CS, CE, EE, Systems Engineering or related software/hardware engineering major, or equivalent experience 8+ years of software development experience with a significant focus on Linux Excellent C programming/debugging skills within Linux kernel and user space software Background with working on embedded systems and ARM processor specific System-level debugging experience and problem-solving skills Excellent communication skills Ways to stand out from the crowd: Understanding of the Linux power and thermal management features (schedule
The SCG Architecture team is hiring a Senior Power Integrity Co-Design Engineer to architect and deliver di/dt mitigation across silicon, package, board, and platform. This role bridges architecture, silicon, and platform — translating product noise targets into shipped specifications, and feeding silicon findings back into the next generation's build. Success in this role requires strong systems thinking and a willingness to accept ambiguity. It also requires the ability to apply AI as a force multiplier while maintaining rigorous engineering judgment. What you'll be doing: Architect voltage-noise mitigation across the full stack — silicon, package, board, platform — and own the codesign trade-offs between them. Co-design noise features with Speed, Power, Reliability, Circuit Design , Power-Arch, ASIC, and platform teams. You're the connective tissue across the codesign web. Work with other team members to define product-level voltage noise targets, drive them to closure, and sign them off at shipment. Build and take ownership of the Sim-to-Si correlation methodology for noise. You know when a model is lying and when silicon is. Model and prototype next-gen noise features — transient sense, droop response, mitigation IP, and codify them so every future program inherits them. Lead show-stopper noise bugs during bringup. The critical issues stop with you. Drive architecture-level codesign tradeoffs across V/F Power Noise Reliability Thermal (Noise-Variation) and (Noise-to-Closure) boundary work, where the highest-leverage innovation lives. What we need to see: BS / MS / PhD in EE, CE, or related (or equivalent experience). 5+ years in silicon power integrity, voltage noise, or PDN. Deep expertise in at least one of
Senior Low Observables (LO) Mission Systems Integration Engineer Company: The Boeing Company The Senior LO Mission Systems Integration Engineer will lead design, integration, analysis, test, and production activities for low observable materials, structures, apertures, sensors, and radomes for mission systems. The role requires strong technical leadership, independent execution of complex tasks, advanced use of computational electromagnetic (CEM) solvers, and the ability to produce and defend technical results with minimal oversight. Key Responsibilities Lead design, modeling, and analysis efforts for LO materials, coatings, structures, and integrated mission systems to meet signature reduction and performance requirements. Independently apply CEM solvers (e.g., FEKO, HFSS, CST, WIPL‑D, xFDTD, or equivalent) to model antennas, apertures, radomes, and LO treatments; perform design optimization and trade studies. Drive integration of LO treatments with mechanical, thermal, and sensor/antenna system constraints; identify and mitigate producibility, testability, and sustainment risks. Plan, execute, and oversee laboratory and field tests (anechoic chamber measurements, antenna/radome characterization, RCS measurements); lead data collection, reduction, and validation activities. Develop and validate predictive models, reconcile simulations with measurements, perform uncertainty quantification, and provide actionable recommendations to engineering and program leadership. Produce high-quality technical documentation: test plans, test reports, technical memoranda, design reviews, and customer briefings. Mentor and provide technical oversight to junior engineers and technicians; ensure adherence to engineering processes, qual
We're looking for a Senior Mechanical Engineer for Midjourney Medical — from precision electromechanical assemblies up to the large-scale structures and mechanisms that hold everything together and make it move. This role spans the full range of scale: one week you might be refining a compact transducer mount, the next you're architecting a structural frame or designing the motion system that positions hardware within it. This is a hands-on role on a small cross-functional team. You'll take problems from whiteboard sketch to working hardware yourself — designing in CAD, prototyping in the shop, testing, iterating, and integrating with research teams, electrical and software engineers along the way. We move fast and expect you to drive your own work: identifying what needs to happen next, making sound engineering calls without waiting for permission, and shipping hardware that works. What you'll do Own parts of the mechanical systems end to end — structures, mechanisms, and electromechanical integration Design large-scale structures: frames, weldments, enclosures, and support systems, with attention to stiffness, weight, manufacturability, and serviceability Design mechanisms: linkages, motion stages, actuation systems for precise, reliable positioning and articulation Integrate transducers, electronics, cabling, and thermal management into large physical systems Perform engineering analysis (tolerance stack-ups, structural/FEA, mechanism kinematics) to de-risk designs before committing to hardware Build and test relentlessly — 3D printing, rapid prototyping, and hands-on fabrication are core to how you'll validate designs Select parts and materials for system-level integration, balancing performance, cost, and lead time Drive projects independently from concept through validation, and collaborate closely with a small cross-disciplinary team to hit project goals What we're looking for Bachelor's degree in Mechanical Engineering or a related field 5+ years of experien
From $243.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will play a key role in enabling our mission to deliver a reliable, high-performing, and cost-efficient infrastructure that powers the world’s play. In this specialized role, you will be the technical lead for our GPU and AI accelerator ecosystem. You will be responsible for the full lifecycle of GPU hardware, from initial architectural evaluation and firmware qualification to large-scale fleet integration and performance tuning. You will ensure that Roblox’s massive-scale rendering and ML workloads run on the most optimized and stable hardware possible. You Will: Architect & Prototype: Prototype next-generation GPU-accelerated hardware platforms, ensuring seamless integration between high-density compute nodes, high-speed interconnects (NVLink/PCIe Gen5/6), and system firmware. GPU Optimization: Drive the integration, performance testing, and debugging of GPUs in our fleet, focusing specifically on hardware-level optimizations, driver tuning, and thermal/power management. Validation & Certification: Develop and execute rigorous evaluation and stress-testing strategies for GPU-heavy server platforms to ensur
From $243.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will help develop and validate next-generation server platforms that power a reliable, high-performing, and cost-efficient infrastructure at scale. You will work across platform bring-up, firmware qualification, hardware validation, fleet integration, and performance optimization to support large-scale production deployments. You Will: Bring-up & Sustaining: Drive key aspects of the hardware development lifecycle, including feasibility studies, hardware bring-up, validation, deployment, and ongoing production support. Platform Optimization: Perform platform integration, performance characterization, and system-level debugging across compute infrastructure, focusing on hardware optimization, driver tuning, and thermal/power efficiency. Hardware Validation: Develop and execute rigorous evaluation and stress-testing strategies for server platforms to ensure reliability and performance under production-scale workloads. Firmware & Fleet Enablement: Support BIOS/BMC firmware qualification, hardware health monitoring, and automation tooling for firmware deployment and lifecycle management. Vendor & Cross-Functi
The Silicon Co-Design Group (SCG) sits at the crossroads of architecture, design, marketing, operations, and productization. Our work spans early architecture through final product delivery across Datacenter, Gaming, Robotics, Automotive, and Embedded markets. We work closely across functions to deliver chips that change what is possible. NVIDIA’s Silicon Co-Design Group is hiring a Chip Lead to serve as the technical lead for one of our most consequential silicon programs! This is not project management, and it is not a senior IC role. You are the person the program partners with on its hardest technical questions — when the right answer is not obvious, and when leadership needs a single technical perspective to align on direction. You are accountable for the technical integrity of the chip end-to-end. You guide co-design feature integration, help resolve the toughest multi-functional bugs, and serve as the project-specific custodian of the qualification playbook. The program runs cleaner because you are on it! What you’ll be doing: You will partner across design, validation, software, and manufacturing to keep the program’s technical narrative clear and on track. Day to day, you will: Serve as the single technical point of contact for multi-functional decisions, issues, and trade-offs. Co-lead program-level feature integration from chip to system, surfacing inter-function dependencies and guiding them to resolution. Help resolve the program’s hardest multi-functional bugs by translating ambiguous, multi-team symptoms into root-cause closure on areas such as HBM, power and thermal, high-speed I/O, and packaging. Steward the qualification playbook. When the playbook does not fit a situation, guide the mitigations and capture the lessons as reusable methodology for other SSG programs. Shape the program’s technical narrative by surfacing key risks, trade
Other cities to consider
More places hiring for this role
Get new senior thermal solutions design engineer jobs in United States by email
Daily job updates · Unsubscribe anytime