Jobs in United States

Senior Memory System Engineer in United States

1,941 active opportunities · Updated October 2026

Explore current senior memory system engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

N
📍 Santa Clara, United States
✓ High-confidence listingExact matchCompany trend -8%
Quick readExact title match for your search

NVIDIA is now looking for a Senior Memory System Engineer to join our ASIC Memory Subsystem team! As a Senior Systems Engineer at NVIDIA, you'll join a group of hardworking engineers to develop and architect innovative Memory Solution for Tegra SoCs. In this position, you'll make a real impact in a multifaceted, technology-focused company. You will work with memory controller/PHY and Platform / System architect, Firmware, SI/PI, Memory suppliers to design and architect cutting edge, high speed and lower power memory technology for NVIDIA CPUs and SOCs. What You Will Be Doing: Analyze future DDR/LPDDR/HBM technologies to determine optimum performance, power, function and RAS in memory for Next generation SOC and Systems. Collaborate with ASIC Architects, Designers, Software and Firmware SW/FW teams to drive memory technology and associated requirements for memory controllers. Define Memory module, Package, and PCB layouts appropriate to the system workloads Debug and bring up memory evaluation / validation and failure issues on memory technology. Collaborate with DRAM suppliers and industry partners on to develop memory and memory related component technology. What We need to see: Bachelor's degree or master’s degree in Electrical Engineering, Computer Engineering (CE), or a related field (or equivalent experience) 10 years of proven track record in DRAM design, module design, or memory sub system design. Deep understanding and strong fundamental of memory design, features, ECC algorithm, SI and PI (Training algorithm) in DDR, LPDDR, and HBM. Strong understanding of memory sub system level interaction with Cache, Memory controller and PHY. Experience in the design, bring-up and validation for memory failure analysis Experience with Python, C/C

N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology and amazing people. Today, we're harnessing the boundless possibilities of AI to build the next era of computing. An era in which our GPU acts as the brain of computers, robots, and self-driving cars that can understand the world. Accomplishing unprecedented goals calls for imagination, inventiveness, and exceptional talent from around the world. As a NVIDIAN, you'll be immersed in a diverse, encouraging environment where everyone is inspired to do their best work. Join our team and discover how you can build a lasting impact on the world. NVIDIA's Silicon Co-Design Group (SCG) leads the full product development lifecycle, from early architecture definition through silicon bringup to product release. The ArchDev team is the hub for silicon and system-level feature development, driving tradeoff analysis, system integration, and POR alignment across the entire organization. This is where ideas become chips, and chips become products that define the state of the art — and we're building that future with some of the most motivated engineers in the industry. We're looking for a Senior Memory Systems Engineer to own HBM and LPDDR integration in sophisticated SoCs. This role covers the full stack, including silicon, package, embedded software, testing, and product development. The engineer will resolve the toughest system-level memory challenges throughout the process, building solutions that hold up at scale. What you'll be doing: HBM & LPDDR System Integration and Bringup: Drive HBM and LPDDR system integration, bringup, characterization, and debug for next-generation SoCs — taking memory subsystems through the full arc from first silicon to production-ready at scale. Full-Stack Memory Closure: <s

R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $196.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Game Engine team at Roblox works on the systems that power the experiences at the heart of the metaverse. Our code is the technical foundation in our client, editor, and our simulation servers. The Game Engine team is broadly split into departments — Audio-Video-Communication, Avatar, Core AI, Digital Matter, Productivity, and Systems. This position is for the Systems Runtime pod. The Runtime pod builds and maintains the foundational C++ components that power the entire Roblox engine and Studio stack. We own the runtime layer — the performance, efficiency, and usability of the core primitives that every other engineering team relies on. Think of Runtime as the “standard library” and execution engine for Roblox's C++ world. Our ownership centers on three pillars: Concurrency System — our task scheduler and fiber runtime power how work is parallelized across CPU cores, letting teams scale features across platforms while keeping code understandable and debuggable. Memory System — we own the engine's memory allocator stack, tracking pipelines, and observability tooling to make memory behavior predictable and prevent leaks and fragmentation. Profiling and Observability — we build and evolve

AWSGitRestAI
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company”. We are looking to grow our company, and grow our teams with the smartest people in the world. What you’ll be doing: You will work with ground breaking technologies for the Tegra SoC and various NVIDIA embedded platforms Implement power and thermal management software features in Linux Kernel and user space Collaborate with power architects, hardware and software engineers on platform power estimation and optimization Optimize the software stack to improve performance, efficiency, and responsiveness for edge AI and robotics use cases. Focus on improving compute and memory utilization, reducing latency and power consumption, and tuning system-level performance to deliver reliable and scalable AI workloads across demanding real-world edge environments. What we need to see: MS in CS, CE, EE, Systems Engineering or related software/hardware engineering major, or equivalent experience 8&#43; years of software development experience with a significant focus on Linux Excellent C programming/debugging skills within Linux kernel and user space software Background with working on embedded systems and ARM processor specific System-level debugging experience and problem-solving skills Excellent communication skills Ways to stand out from the crowd: Understanding of the Linux power and thermal management features (schedule

G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxArtificial IntelligenceAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxAIC++
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are looking for a highly experienced RTL engineer to own critical on- and off-chip interconnect components for our custom AI accelerator platform. You will drive the microarchitecture and RTL implementation of scalable on-chip communication fabrics connecting high-bandwidth compute, memory, and I/O subsystems as well as purpose-built off-chip interfaces and protocols needed to enable custom computing at scale. This is a senior, hands-on engineering role with broad technical ownership. You will drive design from requirements through the full silicon lifecycle, from architecture definition and performance analysis through RTL implementation, verification closure, physical design convergence, bring-up, and production readiness. You will plan and oversee the work of junior engineers and help drive and develop productive engineering relationships with external partners and help manage partner execution. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the microarchitecture, RTL design, and delivery of major SoC interconnect components, including network-on-chip fabrics, switches, routers, bridges, protocol adapters, arbiters, and traffic-management logic as well as off-chip protocol bridges and interfaces. Drive third party engagements to develop novel networking and interface protocols and silicon IP while ensuring high quality and de

AWSRestAIRust
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable &#34;flight recorder&#34; for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as &#34;Hardware Fault,&#34; &#34;Software Bug,&#34; or &#34;Environment Issue.&#34; This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define &#34;signals of impending failure,&#34; enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6&#43; years in systems programming. Experience building automated

PythonKubernetesLinuxMachine Learning
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Chemical and Slurry Systems Engineer within Micron’s Global Facilities and Construction Team, you will deliver engineering expertise supporting the planning, design, construction, operation, and maintenance of facilities systems across Micron’s global manufacturing network. You will participate in capacity and scenario planning, evaluate impact to facilities infrastructure, and serve as a technical domain expert throughout all project phases. Your work ensures environmental, safety, regulatory, and code compliance while driving system reliability and operational excellence. Responsibilities Develop safe, reliable designs for Chemical and Slurry systems and ensure compatibility of materials of construction with all chemicals used. Create, review, and maintain corporate equipment specifications, engineering standards, and design guidelines. Advise sites on system capacity, load projections, and scenario‑driven impacts. Provide technical feedback on infrastructure additions or modifications to support manufacturing and technology roadmap changes. Supply Micron standards to design partners and conduct timely reviews of design packages, construction documents, and engineering work. Support project cost and schedule development and validate scope alignment with collaborator requirements. Coordinate with Facilities, Design, Construction, Procurement, and Manufacturing to ensure clear communication and project alignment. Provide technical guidance during construction,

AIProcurementRecruitment
M
📍 Boise, ID - Main Site, United States
✓ Quality checkedCompany trend -75%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Smart Manufacturing Product Owner at Micron Technology, Inc., you will be responsible to enable innovation and new frontier technology for Smart manufacturing to define, drive and deliver end to end smart manufacturing solution, coordinated across functions of the business. The team will look into applying industry-leading the best methodologies in automation, AI and machine learning to improve Micron’s product development, business and administrative processes across the company. You are joining a team where 'we' matter and hold high standards to achieve perfection. Key Responsibilities Lead the development and implementation of AI-driven Smart Manufacturing solutions to improve product yield, quality, test coverage, and system-level performance across Micron's semiconductor operations. Partner with Test Solutions Engineering, Product Engineering, Global Quality, and multi-functional collaborators to find opportunities and drive digital transformation initiatives. Collaborate with Data Scientists, Machine Learning Engineers, Solution Architects, and Data Engineers to define requirements, prioritize solutions, and deliver end-to-end projects. Translate business needs into detailed business requirements, user stories, acceptance criteria, and data requirements while ensuring alignment with enterprise standards and governance. Develop arguments for proposed solutions, including value quantification, analysis of return on investment and net present value, success metrics, and impl

Machine LearningArtificial IntelligenceAIRecruitment
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali

AWSRestAIRust
S
📍 Bellevue, Washington, United States· Full-time
✓ High-confidence listingCompany trend -92.9%
Quick readStrong listing-quality and freshness signals

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake is now building a world class OLTP service on Postgres. We’re hiring Senior Postgres Engineers to help build the new Snowflake Postgres service. This is an exciting opportunity to help build a large scale, multi-cloud Postgres offering with access to an existing customer base. AS A SENIOR POSTGRES ENGINEER YOU WILL: Develop orchestration layer/ control-plane of large scale databases Work with AWS, Azure, and GCP APIs Build High Availability and Disaster Recovery solutions Tune Postgres to operate at scale for some of the largest datasets in the world Secure and ensure customer data is protected Work alongside some of the brightest minds in the industry and help redefine the space by inventing at various layers of the software stack — from broad distributed systems to low-level optimizations that fundamentally change the performance characteristics of the system to cater to our customer needs. OUR IDEAL SENIOR POSTGRES ENGINEER WILL HAVE: 7+ years industry experience designing, building and developing large scale systems in production Experience building and maintaining distributed, highly available, fault tolerant services Excellent understanding of low level operating systems concepts including multi-threading, memory management, networking, storage, performance,

PythonJavaPostgreSQLAWS
E(
📍 San Francisco Bay Area, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. The residency You own one hard problem end to end. You write the proposal, build the system, design the evaluation, ship behind a gate, and finish with a write-up of what turned out to be true, including the parts that didn't work. You'll sit in the production codebase with a senior mentor and real production data. Recent residents have shipped self-improving harnesses, inference-cost work, agent memory, and eval infrastructure. Your project gets scoped with you, not handed to you. The problem space The loop we care about: production traces become data, data becomes training and evaluation, and better agents produce better traces. Projects live somewhere on that loop. Harness and inference-time work. Context engineering, tool and skill design, orchestration, and deciding where extra inference compute actually pays. Self-improvement loops run behind hard fences. Post-training for agents. SFT on curated trajectories, preference optimization, RL on real agent tasks. Reward design where outcomes are verifiable, process vs. outcome supervision, distilling frontier behavior into cheaper models. Environments and rewards. Turning enterprise workflows into training and eval environments: fi

PythonAIGo
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

We're building the platform that lets long-running autonomous agents operate safely inside NVIDIA's enterprise. These are not assistants on a developer's laptop. They are fleets of agents deployed in the cloud, running continuously at scale on shared accelerated compute. They take on real work across enterprise systems, so people get far more done than they could before. This role defines the constructs that agents are built from: the blueprints they start from, the tools, skills, and plugins that power them against enterprise data, the runtime safety harness that keeps them in bounds, and the connections into credential management, sandbox, memory, and observability. The team designs and ships these building blocks so that agent developers across the company can stand up a new agent, wire it in, and run it for days or weeks. Security and safe execution come out of the box, not something each team has to get right on its own. Today an agent runs inside a single harness. Claude, Codex, and open-source agent harnesses each work differently underneath, with their own execution model, tool interface, and telemetry shape. The platform smooths over those differences, so a single skill, safety policy, or trace works the same no matter which harness is running. We want to enable agents that act on a person's behalf, governed and secure, continuously evaluated and self-improving. These agents coordinate and hand work off to each other, with identity and policy following every hop. They route and tune themselves across harnesses from live eval signals, and get better from their own production telemetry instead of waiting on a human to retrain them. Have you run agents on a harness like Claude or Codex and hit the walls that show up when they run for real, for days, against live systems — and wanted them to learn from it on their own? We're building the platform that solves those problems once, for every team. What you'll b

PythonAIFinanceHR
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's Cache team is building a next-generation caching solution designed to deliver sub-millisecond average latency, horizontal scalability, and high efficiency—all at a drastically lower cost. Our ultimate vision is to shape a caching infrastructure capable of supporting 1 billion Daily Active Users while reducing costs by 90%. We are turning hours of onboarding and capacity expansion into seconds, freeing service owners entirely from managing cluster lifecycles. As a Senior Engineer on the Cache team (part of the Infra Storage org), you will innovate and operate large-scale, in-house distributed systems to solve Roblox's ever-growing caching challenges. You will report directly to the Engineering Manager for the Cache team. (Check out our recent engineering blog post here to learn more about the team's latest work!) You will: Lead the architectural transition to a next-generation, multitenant caching service built on ValKey, ensuring strict data, resource, and failure isolation for all tenants. Drive systemic optimizations to mitigate head-of-line blocking, manage hot keys, and maximize CPU and memory utilization across physical machine clusters. Design and build robust frameworks to a

RedisAWSKubernetesGit
🔔

Get new senior memory system engineer jobs in United States by email

Daily job updates · Unsubscribe anytime