NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated
Salary not disclosed
Check market pay for comparable Senior GPU Memory Architect roles before applying.
Role overview
Job description
NVIDIA is seeking a world-class computer architect to contribute to the development of future high-performance computing systems, with a focus on enhancing the power-constrained performance of the hardware. Ideal candidates will have a strong track record of understanding and analyzing memory systems architecture to improve performance per watt (perf/W) and performance per millimeter (perf/mm). A broad perspective across the field of computer architecture and depth in the area of power, performance, and area (PPA) analysis is highly desirable. NVIDIA has pioneered programmable GPUs and the CUDA language and is a world leader in high-performance computing technology, with aggressive plans for future processors. This position offers the opportunity to have a real impact in a fast-moving, technology-focused company.
What you will be doing:
Develop innovative high-performance processor and system architectures, focusing on the memory system and energy efficiency.
Develop architecture and micro-architecture features to improve the state-of-the-art in GPU memory systems, optimizing along the axes of perf/W, perf/mm, and perf/$.
Develop and enhance architecture prototype models for power and noise analysis.
Participate in performance and power simulation of features to analyze, define, and improve energy per byte.
Analyze benchmarks, application workloads, and performance/power simulation and emulation results to identify areas for architecture optimizations.
Debug power, performance, and functional issues with high-level models, RTL simulation and emulation, silicon, and systems.
Collaborate with outside partners on system infrastructure.
What we want to see:
10+ yrs of experience in CPU/GPU architecture, memory systems design with a focus on energy efficiency in the system.
Bachelors in Electrical Engineering, Computer Science, or related field (or equivalent experience).
Expertise in power analysis and modeling in pipeline-based architectures.
Good understanding of power analysis tools like Ansys Power Artist, Synopsys PTPX, etc.
Experience with C, C++, and some scripting languages.
Ways to stand out from the crowd:
MS/PhD in Electrical Engineering, Computer Science, or related field
Experience with hardware design languages such as Verilog or VHDL
Prior experience in developing and optimizing algorithms for power efficiency
Demonstrated ability to innovate and drive continuous improvement in system architecture
Excellent communication skills, both written and verbal, for effective collaboration with internal teams and external partners.
You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.What they are looking for
Skills & requirements
Hiring company
Nvidia
Explore this employer's active roles, salary signals and company profile on Jobiba.
Keep exploring
Similar active roles
Fresh roles matched to this title and market.
$114.8K – $183.6K/yr
From $243.3K/yr
🔔 Get job alerts
New Senior GPU Memory Architect jobs in Ca, Santa Clara, United States, straight to your inbox.
No spam · Unsubscribe anytime