NVIDIA is seeking a world-class computer architect to contribute to the development of future high-performance computing systems, with a focus on enhancing the power-constrained performance of the hardware. Ideal candidates will have a strong track record of understanding and analyzing memory systems architecture to improve performance per watt (perf/W) and performance per millimeter (perf/mm). A broad perspective across the field of computer architecture and depth in the area of power, performance, and area (PPA) analysis is highly desirable. NVIDIA has pioneered programmable GPUs and the CUDA language and is a world leader in high-performance computing technology, with aggressive plans for future processors. This position offers the opportunity to have a real impact in a fast-moving, technology-focused company. What you will be doing: Develop innovative high-performance processor and system architectures, focusing on the memory system and energy efficiency. Develop architecture and micro-architecture features to improve the state-of-the-art in GPU memory systems, optimizing along the axes of perf/W, perf/mm, and perf/$. Develop and enhance architecture prototype models for power and noise analysis. Participate in performance and power simulation of features to analyze, define, and improve energy per byte. Analyze benchmarks, application workloads, and performance/power simulation and emulation results to identify areas for architecture optimizations. Debug power, performance, and functional issues with high-level models, RTL simulation and emulation, silicon, and systems. Collaborate with outside partners on system infrastructure. What we want to see: 10+ yrs of experience in CPU/GPU architecture, memory systems design with a focus on energy efficiency in the system. Bachelor
Jobs in United States
Senior Gpu Memory Architect in United States
1,941 active opportunities · Updated October 2026
Showing
15 jobs
Explore current senior gpu memory architect jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology and amazing people. Today, we're harnessing the boundless possibilities of AI to build the next era of computing. An era in which our GPU acts as the brain of computers, robots, and self-driving cars that can understand the world. Accomplishing unprecedented goals calls for imagination, inventiveness, and exceptional talent from around the world. As a NVIDIAN, you'll be immersed in a diverse, encouraging environment where everyone is inspired to do their best work. Join our team and discover how you can build a lasting impact on the world. NVIDIA's Silicon Co-Design Group (SCG) leads the full product development lifecycle, from early architecture definition through silicon bringup to product release. The ArchDev team is the hub for silicon and system-level feature development, driving tradeoff analysis, system integration, and POR alignment across the entire organization. This is where ideas become chips, and chips become products that define the state of the art — and we're building that future with some of the most motivated engineers in the industry. We're looking for a Senior Memory Systems Engineer to own HBM and LPDDR integration in sophisticated SoCs. This role covers the full stack, including silicon, package, embedded software, testing, and product development. The engineer will resolve the toughest system-level memory challenges throughout the process, building solutions that hold up at scale. What you'll be doing: HBM & LPDDR System Integration and Bringup: Drive HBM and LPDDR system integration, bringup, characterization, and debug for next-generation SoCs — taking memory subsystems through the full arc from first silicon to production-ready at scale. Full-Stack Memory Closure: <s
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company”. We are looking to grow our company, and grow our teams with the smartest people in the world. What you’ll be doing: You will work with ground breaking technologies for the Tegra SoC and various NVIDIA embedded platforms Implement power and thermal management software features in Linux Kernel and user space Collaborate with power architects, hardware and software engineers on platform power estimation and optimization Optimize the software stack to improve performance, efficiency, and responsiveness for edge AI and robotics use cases. Focus on improving compute and memory utilization, reducing latency and power consumption, and tuning system-level performance to deliver reliable and scalable AI workloads across demanding real-world edge environments. What we need to see: MS in CS, CE, EE, Systems Engineering or related software/hardware engineering major, or equivalent experience 8+ years of software development experience with a significant focus on Linux Excellent C programming/debugging skills within Linux kernel and user space software Background with working on embedded systems and ARM processor specific System-level debugging experience and problem-solving skills Excellent communication skills Ways to stand out from the crowd: Understanding of the Linux power and thermal management features (schedule
$114.8K – $183.6K/yr
Job Title Senior Software Engineer - Image Reconstruction (C++/CUDA) Job Description Build the GPU-native engine that turns raw CT physics into life-saving images inside Philips scanners worldwide , wringing maximum performance from constrained hardware at the frontier of C++, CUDA, and AI alongside world-class physicists. Your role: Help r e-architect an entire CT image-reconstruction pipeline, from raw detector physics to the final clinical image, as a fully GPU-native, massively scalable platform in C++ and CUDA. Work shoulder-to-shoulder with physicists, algorithm architects, and platform engineers, translating complex signal and image-processing models into high-performance GPU implementations that balance image quality against compute cost. Collaborate with CT platform teams around the globe to design, build, test, and deploy an industry-leading reconstruction software platform. Push the frontier of performance engineering and applied AI by profiling, parallelizing, optimizing memory use across caches and shared memory, and pioneering learned models that replace expensive physics with fast, accurate equivalents. This hybrid role is based in Cleveland, OH, with three flexible days in the office and two remote days each week . Regular travel is not typically </sp
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each brings together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We are searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will be instrumental in ensuring our data center solutions deliver industry-leading performance for accelerated computing workloads. What you will be doing: Design and execute comprehensive performance benchmarking strategies for our data center platforms and products Characterize real-world AI training, inference, and HPC workloads at scale Define, track, and report key performance indicators (throughput, latency, efficiency, scaling) Build automation tools and frameworks for performance monitoring and analysis Identify and analyze performance bottlenecks across compute, memory, network and storage subsystems Work closely with architecture, hardware,
From $243.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will play a key role in enabling our mission to deliver a reliable, high-performing, and cost-efficient infrastructure that powers the world’s play. In this specialized role, you will be the technical lead for our GPU and AI accelerator ecosystem. You will be responsible for the full lifecycle of GPU hardware, from initial architectural evaluation and firmware qualification to large-scale fleet integration and performance tuning. You will ensure that Roblox’s massive-scale rendering and ML workloads run on the most optimized and stable hardware possible. You Will: Architect & Prototype: Prototype next-generation GPU-accelerated hardware platforms, ensuring seamless integration between high-density compute nodes, high-speed interconnects (NVLink/PCIe Gen5/6), and system firmware. GPU Optimization: Drive the integration, performance testing, and debugging of GPUs in our fleet, focusing specifically on hardware-level optimizations, driver tuning, and thermal/power management. Validation & Certification: Develop and execute rigorous evaluation and stress-testing strategies for GPU-heavy server platforms to ensur
At NVIDIA, we push the boundaries of computing innovation. Our ASIC Verification Engineers focus on developing the world’s top SoCs and GPUs. Joining us as a Senior ASIC Verification Engineer - GPU means working on modern technology powering consumer graphics and AI applications. This position is ideal for those passionate about technology and eager to impact computing’s future. What you'll be doing: As a key member of our ASIC Verification team, you will verify the design and implementation of the industry's leading GPUs. You will be responsible for verifying the ASIC build, architecture, golden models, and micro-architecture using advanced verification methodologies such as UVM or equivalent. Understand the design and implementation of your unit/cluster/chip, define the verification scope, develop the verification infrastructure, and verify the correctness of the design. Collaborate with architects, designers, and pre- and post-silicon verification teams to accomplish your task. What we need to see: Bachelor's Degree in EE, CS, or CE or equivalent experience. 5+ years of relevant experience. Experience in verification using random stimulus along with functional coverage and assertion-based verification methodologies. Experience with design and verification tools (VCS or equivalent simulation tools, debug tools like Verdi, Indago, GDB). Expertise in System Verilog or similar HVL. Strong debugging and analytical skills. Perl and C/C++ programming language experience desirable. Strong communication skills and the ability & desire to work as a great teammate are huge pluses. Experience in crafting test bench environments for unit and system level verification. #LI-Hybrid Your base salary will be det
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. We are looking for a motivated Performance engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! This is an outstanding opportunity for someone with HPC and performance background to advance the state of the art in this space. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters. Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available Triage and root-cause performance issues reported by our customers Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information <li
NVIDIA is leading groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU -- our invention -- serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables groundbreaking creativity and discovery, and powers inventions that were once considered science fiction, including artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We build communication libraries like NCCL, NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co-design next-gen data center platforms and scalable communications software. DL and HPC applications have a huge compute demands and already run at scales of up to tens of thousands of GPUs. GPUs are connected with high-speed interconnects (e.g. NVLink, PCIe) within a node and with high-speed networking (e.g. InfiniBand, Ethernet) across nodes. Efficient and fast communication between GPUs directly impacts end-to-end application performance. This impact continues to grow with the increasing scale of next generation systems. This is an outstanding opportunity to advance the state-of-the-art, break performance barriers, and deliver platforms the world has never seen before. Are you ready to build the new and innovative technologies that will help realize NVIDIA's vision? What you will be doing: Investigate opportunities to improve communication performance by identifying bottlenecks in today's systems. Design and implement new communication technologies to accelerate AI and HPC workloads. Explore innovative solutions in HW and SW for our next generation platforms as part of co-design efforts involving GPU, Networking, and SW architects. Build proofs-of-concept, conduct experiments,
NVIDIA’s Silicon Co-Design Group is the team that gets every GPU, SoC, and CPU silicon program from first power-on to high-volume production. We are hiring a Senior Manager to lead our Test, Manufacturability, Reliability & Quality (TMRQ) organization. This is not a coordination role . Your work decides if a product can be built at scale and trusted in the field. These include production test development (SLT, BLT), control run flow, system reliability stress (HTOL), platform- and board-level manufacturing issue closure, and field diagnostic test development. You lead a team of individual contributors and a first-line manager at the layer where silicon, platform, and software collide with manufacturing reality. Decisions you make show up in yield curves, production ramp , and customer escapes. You are the leader the program turns to when a build is stuck, a control run is fallout-heavy, or a field return points back at silicon . The exceptional hire also uses AI deliberately — with demonstrated workflow impact and the judgment to know where it compresses real work and where it introduces risk. What you will be doing: Keep programs moving. Own the technical execution and velocity of SLT, BLT, Board/Chip/Rack CR, and system reliability stress (HTOL) across every GPU, SoC, and CPU silicon program. Close the hardest multi-functional failures. Resolve Vmin and binning escapes, performance shortfalls, and power anomalies by driving root-cause across design, methodology, DFT, ATE, package, software/firmware, and manufacturing — and own the WARs and productized fixes through to confirmation. Give leadership the clarity to act. Convert raw integration signals — CR fallout, BLT/SLT yield, SHTOL/CHTOL data, RMA trends, customer escalations — into decision-ready options that enable executive leadership to act with confidence on POR, QS/PS gates, a
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent Senior Engineering Manager to lead a large firmware engineering organization delivering end-to-end manageability firmware for NVIDIA's next generation Data Center Compute Systems. This role owns HGX product line and OpenBMC-based management firmware and MCU firmware components in data center platforms, including architecture, execution, quality, reliability, telemetry, and customer readiness. We are seeking an experienced senior leader with strong technical depth, broad system perspective, and a proven ability to lead large teams through complex product cycles. This role is onsite in Santa Clara, CA, USA. If you're creative and autonomous, we want to hear from you! What you'll be doing: Lead a large firmware engineering organization delivering OpenBMC based firmware and MCU firmware for next-generation Data Center Compute Systems. Own HGX platform as a lead for Firmware and System software readiness working across the organization. Define and drive the long-term firmware roadmap, balancing architectural innovation with product execution and delivery milestones. Drive architecture strategy across BMC, MCU, platform software, manageability, health management, and data center firmware interfaces. <spa
NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a “learning machine” that constantly evolves by adapting to new opportunities that are hard to solve, that only we can address, and that matter to the world. This is our life’s work, to amplify human creativity and intelligence. Make the choice to join us today. Design-for-X Engineering at NVIDIA works on groundbreaking innovations involving crafting creative solutions in AI for Chip Design and AI for Predictions in various use cases in manufacturing testing on some of the industry's most complex semiconductor chips. What you'll be doing: As a senior member in our team, you will work on innovating in the DFT Power, Thermal & Voltage Noise Methodology areas. This will include working on groundbreaking low power & thermal solutions for our manufacturing tests to be enabled at conditions that push the boundaries for our datacenter GPUs. You will work with multi-functional teams including Product Development & Power Architecture, implementing brand-new methodologies on hard-to-solve problems for improving our outgoing quality of chips. You will work on post-silicon data analysis for power to architect the next-gen solutions. In addition, you will help develop and deploy DFT methodologies for our next generation products using Applied ML & Gen AI solutions. You will also help mentor junior engineers on test designs and trade-offs including cost and quality. What we need to see: BSEE (or equ
NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a “learning machine” that constantly evolves by adapting to new opportunities that are hard to solve, that only we can address, and that matter to the world. This is our life’s work, to amplify human creativity and intelligence. Make the choice to join us today. Design-for-Test Engineering at NVIDIA works on groundbreaking innovations involving crafting creative solutions for DFT architecture, verification and post-silicon validation on some of the industry's most complex semiconductor chips. What you'll be doing: As a senior member in our team, you will work with pre-silicon and post-silicon data analytics - visualization, insights and modeling. Design and uphold sturdy data pipelines and ETL processes for the ingestion and processing of DFX Engineering data from various origins Lead engineering efforts by collaborating with cross-functional teams (execution, analytics, data science, product) to define data requirements and ensure data quality and consistency You will work on hard-to-solve problems in the Design For Test space which will involve application of algorithm design, using statistical tools to analyze and interpret complex datasets and explorations using Applied AI methods. In addition, you will help develop and deploy DFT methodologies for our next generation products using Gen AI solutions. You will also help mentor junior engineers on test designs and trade-offs including cost and quality. What we need to see: BSEE (or equivalent experience) with 5+, MSEE with 3+, or PhD wi
NVIDIA Architecture Modeling group is looking for Architects, Functional Modeling Engineers, and Simulation experts to join various architecture efforts across GPU/ SOC Architecture teams. A key part of NVIDIA's strength is to innovate in parallel computing fields, delivering the highest performance in the world for high-performance computing. We are constantly looking for ways to improve our SoC and Systems architecture and maintain our leadership. In this position, you will be working with other world-class architects on modeling, analysis and validation of chip & system architectures and features that advance the state of art in performance and efficiency. What you'll be doing: Modeling and analysis of SoC & Systems algorithms and features, across datacenter, automotive, and client products Build and deliver platforms for SOC's that enable left shift for the SW teams aligned with project milestones Work closely with the SOC architects and guide modeling teams to deliver high-quality functional models that involve SOC+GPU use cases Collaborate with our EDA partners to align on customer-facing technologies Develop tests, test plans, and testing infrastructure for new architectures/features and code coverage analysis and reporting Ensure alignment between the various modeling teams at NVIDIA, GPU modeling teams, and modeling teams overseas What we need to see: Master’s or PhD in Computer Science, Electrical Engineering, Computer Engineering, or a related relevant field (or equivalent experience) with 5+ years of relevant work experience. Strong programming ability: C++, C along with a good understanding of build systems (CMAKE, make) , toolchains (GCC, MSVC) and libraries (STL, BOOST) Computer Architecture background with experience in modelling wit
Other cities to consider
More places hiring for this role
Get new senior gpu memory architect jobs in United States by email
Daily job updates · Unsubscribe anytime