NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data
Jobs in United States
Senior Gpu Memory Architect in United States
1,941 active opportunities · Updated October 2026
Showing
15 jobs
Explore current senior gpu memory architect jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior IT Auditor passionate about SOX and internal examination. The role supports the Senior Manager of IT SOX Compliance to strengthen NVIDIA’s internal control environment over financial reporting (ICFR). It involves collaborating with IT and accounting/finance groups to evaluate and build efficient business and IT controls that mitigate financial reporting risk. What you'll be doing: Plan, scope, and implement the end-to-end SOX 404 lifecycle, including walkthroughs, risk-control matrices, testing, deficiency evaluation, remediation, and reporting. Partner with process owners to assess risks in new and changing business processes,-identify the key IT controls, and build out narrative, flowchart and control description. Drive continuous improvement and automation by seeking opportunities to streamline, standardize, and automate controls, reducing operational friction while maintaining control effectiveness. Provide guidance and training to control owners to promote awareness and understanding of SOX compliance requirements. Serve as the primary point of contact for external auditors, and ensure a seamless, efficient audit process. Coach
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. At NVIDIA, we're not just transforming the world of computer graphics and AI; we're setting the stage for the future of autonomous driving. As a Lead Safety Architect, you will be at the forefront of our autonomous vehicle technology, ensuring its safety at scale. You will collaborate with the most innovative engineers and technologists to integrate safety measures into our latest DRIVE products. This role is paramount in achieving and exceeding NVIDIA's high safety standards, making your work both exciting and impactful! What you’ll be doing: Representing NVIDIA’s functional safety strategy and architectures to the customer Working closely with customers to understand their functional safety requirements and system architectures and feeding those back into the development teams Assisting customers to safely integrate and validate our products in their systems and vehicles Supporting customer facing safety collateral Tailoring functional safety platforms and safety analyses for strategic customers You will be working closely with safety management, solution architects, sales and technical marketing teams to deliver state of the art products
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. Join NVIDIA's NIM team and be part of an exceptionally ambitious project in Santa Clara, CA! As a Senior Software Engineer, NIM Tools, you will have the remarkable opportunity to build a groundbreaking model customization and deployment lifecycle platform from inception. This isn't just another feature team—you will be defining the structure for a new product surface accessed by ISVs and CSPs internationally. Your work will empower customers to take models from selection through fine-tuning, evaluation, deployment, and compliance flawlessly. What you'll be doing: Compose and build the fine-tuning handoff pipeline, including LoRA adapter repackaging, re-quantization, and re-validation into NIM. Develop the evaluation harness, ensuring models meet our high standards. Implement the observability and attestation layer to produce auditable compliance artifacts. Work in close partnership with ISVs and CSPs to roll out NVIDIA NIMs on a large scale. Define and improve durable platform APIs, steering clear of one-off integrations. Ensure flawless completion of projects through strict attention to detail and proven methodologies. Wha
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which the GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As a Developer Technology Engineer you will be at the forefront of innovation, working with leading industry partners and exciting OSS projects to accelerate RTX & DGX SoC performance for agentic use. This role offers an outstanding opportunity to collaborate with world-class talent and make a significant contribution to the next era of enterprise and consumer AI. What you'll be doing: Own key engagements with our fast-paced developer ecosystem partners, optimizing their applications to deliver outstanding end-to-end performance on RTX and DGX SoCs. Take ownership of performance optimization for compute-intensive CPU, AI, and 3D graphics workloads, working with domain experts to identify bottlenecks and implement effective solutions. Work across system software, GPU driver, architecture, and NVIDIA Research teams to influence next-generation, high-performance SoC platforms by bringing real-world workflows and actionable insights from partner and customer needs. Provide technical guidance and mentorship to junior engineers while contributing to an inclusive and high-performing team environment. What we need to see: BS or MS degree in Computer Science, Engineering, Mathematics or related degree (or equivalent experience) 5+ years of respective work experience as software developer Proficiency in C/C++, Python, software
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build cloud-native data and storage services for hybrid and multi-cloud infrastructure, including dataset discovery, ingestion, governance, checkpointing, observability, and low-latency access. Develop scalable cloud-native services and APIs that support exabyte-scale, high-performance GPU training and inference workflows. Work closely with product managers, internal AI teams, platform teams, and partner engineering teams to understand requirements and turn them into reliable production systems. Collaborate with SRE, operations, and support teams to improve service reliability, performance, observability, on-call readiness, and operational scale. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, and verification. What we need to see: BS in Computer Science, Information Systems, Computer Engineering, or equivalent experience, with 5+ years of software engineering experience. Strong foundation in algorithms, data structures, distributed systems, and practi
NVIDIA is seeking a Senior Firmware Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware development with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you will be doing: Design and develop firmware solutions for manageability and observability of data center servers. Actively participate in hardware bring-up activities, OOB firmware development, protocol stacks (Redfish, PLDM, MCTP, NSM) and hardware-software co-design for Cloud Service Provider deployments. Debug and troubleshoot NVIDIA GPU firmware issues, power management, performance, and thermal control problems for data center deployments, providing active support to CSPs. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Deep expertise in embedded firmware, server management controllers, and hardware bring-up with proven track record of shipping production BMC solutions Strong knowledge of DMTF protocols (Redfish, IPMI, PLDM, MCTP, SPDM), telemetry frameworks, and out-of-band management architectures Expert-level skills in C/C&
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform & node designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each bringing together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We’re searching for a highly motivated, technical leader to design, drive, and operationalize rack-scale factory and deployment flows for next-generation data center products. The ideal candidate will combine deep systems expertise, decisive technical leadership, and a passion for building reliable, debuggable, and scalable manufacturing and deployment solutions. What you’ll be doing: Lead and drive rack-scale/L11 flows for factory and initial data center deployment. Design and implement end-to-end factory workflows, including firmware flashing sequences, security provisioning, and deployment of software mitigations. Collaborate with data center architects, ODMs, and OEMs to define factory and data center requirements that ensure efficient and reliable production ramp. Champion reliability, debuggability an
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G
As one of the technology industry's most desirable employers, NVIDIA has been redefining accelerated computing, computer graphics and leading the Artificial Intelligence revolution. NVIDIA's innovation is fueled by its great technology—and amazing people. We are seeking a Senior Silicon and System Product Lead to influence, innovate and take our next generation products to the market. As part of the Silicon Solutions Team, we architect and deliver groundbreaking system solutions that integrate all aspects of the system from silicon design, software design to operations and final deployment in multiple market segments that NVIDIA serves. This position offers an unique opportunity to collaborate with multiple organizations in the company and grow your career in a high impact role. We need a passionate, hard-working and creative individual to lead the products all the way from market analysis to delivering the features on the final product. What you'll be doing: Drive product performance and power targets, trade-off features/configurations and provide innovative solutions to complex silicon and system level problems. Evaluate new market segments and use cases; translate market requirements to engineering problem statements and metrics. Innovate Performance, power, yield and quality optimizations and features for the world’s fastest power-shipping products in the GPU and SoC market segments spanning gaming, automotive, datacenter and DL/AI. Develop methodologies and requirements for multi-functional teams to drive silicon and system product features to production. Incorporate productization feedback to improve the next generation. Lead the team for feature requirements and schedule from architecture to silicon phase of projects. Work alongside system architects, designers, marketing teams, chip and board designers, software/firmware engineers, HW/S
The Silicon Co-Design Group (SCG) sits at the crossroads of architecture, design, marketing, operations, and productization. Our work spans early architecture through final product delivery across Datacenter, Gaming, Robotics, Automotive, and Embedded markets. We work closely across functions to deliver chips that change what is possible. System Integration sits at the intersection of all of them. It is the layer where every architecture, design, software, and manufacturing decision meets reality. When something breaks late in a program, it usually breaks here first. We are hiring a Senior Manager to lead this team in the US and partner closely with teams globally. Your work will sit on the critical path of every NVIDIA silicon program, and the bar you set for system integration is the bar we ship to! The two hardest, highest-leverage problems in this seat: Find critical silicon issues earlier — often before software is production-ready. Left-shifting post-silicon coverage is the highest-value thing System Integration can do. Standing up wide-area testing as a repeatable capability is a core part of the role. Keep programs on milestone when upstream dependencies slip. Validation plans collide with reality every program! The team needs new strategies, not just contingency plans, to keep moving when software, firmware, or methodology slip. You will design and run those strategies. What you’ll be doing: Plan and execute post-silicon feature integration, PVT validation, and wide-area testing across NVIDIA’s GPU, SoC, and CPU programs. Build wide-area and in-system test as a repeatable capability that shifts post-silicon coverage left, so issues surface before we are production-ready. Lead resolution of the most complex system-level issues, RMAs, and HW/SW interaction problems with creative workarounds and focused lab experimentation. Deve
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe
We are developing advanced multi-rack, multi-tenant AI/ML datacenters with NVIDIA GB200, and upcoming GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team to focus on the cloud-native stack for datacenter products like GB200. In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex scheduling challenges across racks, tenants, and clouds as part of the CSP engagements team. What you’ll be doing: Perform deep-dive debugging of multi-rack, multi-tenant clusters: scheduler behavior, container runtime issues, device-plugin crashes, RDMA/IB fabric anomalies, etc. Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom micro-services that expose new GPU capabilities. Drive joint architecture reviews and “whiteboard” sessions with CSP and internal platform teams; convert findings into RFCs and upstream pull requests. Create reproducible testbeds (Helm/Ansible/Terraform) that mirror customer environments; automate validation and benchmark suites. Deliver technical collateral-design docs, how-to guides, demo scripts-and present at customer on-sites, KubeCon, and SlurmUG. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Strong source-level expertise in Kubernetes internals (scheduler, CRI/CNI/CSI, operators) and Slurm (federation, power-save, plugins). Hands-on experience integrating next-gen GPUs (Blackwell/GB200/GB300) or comparable accelerators into containerized clusters. Proven track record debugging large-scale, cloud-native stacks across ne
Other cities to consider
More places hiring for this role
Get new senior gpu memory architect jobs in United States by email
Daily job updates · Unsubscribe anytime