NVIDIA is now looking for a Senior Memory System Engineer to join our ASIC Memory Subsystem team! As a Senior Systems Engineer at NVIDIA, you'll join a group of hardworking engineers to develop and architect innovative Memory Solution for Tegra SoCs. In this position, you'll make a real impact in a multifaceted, technology-focused company. You will work with memory controller/PHY and Platform / System architect, Firmware, SI/PI, Memory suppliers to design and architect cutting edge, high speed and lower power memory technology for NVIDIA CPUs and SOCs. What You Will Be Doing: Analyze future DDR/LPDDR/HBM technologies to determine optimum performance, power, function and RAS in memory for Next generation SOC and Systems. Collaborate with ASIC Architects, Designers, Software and Firmware SW/FW teams to drive memory technology and associated requirements for memory controllers. Define Memory module, Package, and PCB layouts appropriate to the system workloads Debug and bring up memory evaluation / validation and failure issues on memory technology. Collaborate with DRAM suppliers and industry partners on to develop memory and memory related component technology. What We need to see: Bachelor's degree or master’s degree in Electrical Engineering, Computer Engineering (CE), or a related field (or equivalent experience) 10 years of proven track record in DRAM design, module design, or memory sub system design. Deep understanding and strong fundamental of memory design, features, ECC algorithm, SI and PI (Training algorithm) in DDR, LPDDR, and HBM. Strong understanding of memory sub system level interaction with Cache, Memory controller and PHY. Experience in the design, bring-up and validation for memory failure analysis Experience with Python, C/C
Jobiba hiring network
Senior Memory System Engineer Jobs
15 active opportunities · Updated for September 2026
Fresh results
15 shown
Explore current senior memory system engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Senior Manager, Package PE, Product and System Engineering (High Bandwidth Memory) — Fab 10A, Singapore. Apply via Workday.
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Game Engine team at Roblox works on the systems that power the experiences at the heart of the metaverse. Our code is the technical foundation in our client, editor, and our simulation servers. The Game Engine team is broadly split into departments — Audio-Video-Communication, Avatar, Core AI, Digital Matter, Productivity, and Systems. This position is for the Systems Runtime pod. The Runtime pod builds and maintains the foundational C++ components that power the entire Roblox engine and Studio stack. We own the runtime layer — the performance, efficiency, and usability of the core primitives that every other engineering team relies on. Think of Runtime as the “standard library” and execution engine for Roblox's C++ world. Our ownership centers on three pillars: Concurrency System — our task scheduler and fiber runtime power how work is parallelized across CPU cores, letting teams scale features across platforms while keeping code understandable and debuggable. Memory System — we own the engine's memory allocator stack, tracking pipelines, and observability tooling to make memory behavior predictable and prevent leaks and fragmentation. Profiling and Observability — we build and evolve
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company”. We are looking to grow our company, and grow our teams with the smartest people in the world. What you’ll be doing: You will work with ground breaking technologies for the Tegra SoC and various NVIDIA embedded platforms Implement power and thermal management software features in Linux Kernel and user space Collaborate with power architects, hardware and software engineers on platform power estimation and optimization Optimize the software stack to improve performance, efficiency, and responsiveness for edge AI and robotics use cases. Focus on improving compute and memory utilization, reducing latency and power consumption, and tuning system-level performance to deliver reliable and scalable AI workloads across demanding real-world edge environments. What we need to see: MS in CS, CE, EE, Systems Engineering or related software/hardware engineering major, or equivalent experience 8+ years of software development experience with a significant focus on Linux Excellent C programming/debugging skills within Linux kernel and user space software Background with working on embedded systems and ARM processor specific System-level debugging experience and problem-solving skills Excellent communication skills Ways to stand out from the crowd: Understanding of the Linux power and thermal management features (schedule
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are looking for a highly experienced RTL engineer to own critical on- and off-chip interconnect components for our custom AI accelerator platform. You will drive the microarchitecture and RTL implementation of scalable on-chip communication fabrics connecting high-bandwidth compute, memory, and I/O subsystems as well as purpose-built off-chip interfaces and protocols needed to enable custom computing at scale. This is a senior, hands-on engineering role with broad technical ownership. You will drive design from requirements through the full silicon lifecycle, from architecture definition and performance analysis through RTL implementation, verification closure, physical design convergence, bring-up, and production readiness. You will plan and oversee the work of junior engineers and help drive and develop productive engineering relationships with external partners and help manage partner execution. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the microarchitecture, RTL design, and delivery of major SoC interconnect components, including network-on-chip fabrics, switches, routers, bridges, protocol adapters, arbiters, and traffic-management logic as well as off-chip protocol bridges and interfaces. Drive third party engagements to develop novel networking and interface protocols and silicon IP while ensuring high quality and de
Job Summary Reporting to the Memory Validation leadership team, the Senior Silicon DDR/HBM Validation Engineer will be responsible for the bring-up, validation, characterization and debug of advanced memory subsystems used in next-generation AI compute platforms. The role will focus on DDR and HBM technologies, working closely with silicon design, firmware, characterization, platform and systems teams to ensure robust memory subsystem functionality, performance and reliability. The successful candidate will take ownership of significant validation activities, contribute to debug and root-cause analysis efforts, and help improve validation methodologies, automation and infrastructure. The Team The Memory Validation team sits within the Validation organisation and is responsible for the bring-up, validation, characterization and debug of memory subsystems across Graphcore silicon and platform products. The team supports DDR and HBM validation activities throughout the product lifecycle, from first silicon through production readiness. Engineers work closely with architecture, RTL, firmware, characterization, systems and platform teams to ensure memory technologies meet functionality, performance, reliability and performance objectives. Responsibilities and Duties Execute validation and bring-up activities for DDR and HBM memory subsystems Verify memory bring-up software, firmware and scripts against defined project requirements Debug firmware, hardware and system-level issues and contribute to root-cause analysis activities Analyse system logs, validation data and characterization results to identify failures and performance issues Perform PHY characterization and analog-level analysis during stress testing and validation activities Develop and execute functional, stress, performance and corner-case validation tests Perform signal integrity, voltage, frequency and timing measurements using laboratory instrumentation Char
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role The AI team is building cutting-edge solutions that bring the power of AI directly to edge devices while seamlessly integrating with cloud infrastructure. We are looking for a Senior Software Engineer to design and develop high-performance, scalable services to support AI workloads across edge and cloud environments. What You Might Do Design, build, and maintain services that power AI-driven applications, ensuring scalability and performance. Develop APIs and microservices that facilitate seamless integration between cloud-based AI models and edge devices. Optimize data pipelines and storage solutions for real-time AI inference and processing. Implement security and privacy best practices for distributed AI systems. Work closely with AI researchers, infrastructure engineers, and frontend developers to deliver end-to-end AI-driven solutions. Build and optimize an agent orchestration runtime that enables tool use, memory management, and multi-step reasoning across LLMs, APIs, and edge-connected systems. Develop robust logging, monitoring, and alerting systems to ensure system reliabilit
About VSCO For years, we've helped photographers create their work. Now we're building what comes next. VSCO exists for photographers. Not as a side feature, not as an afterthought, but as the whole point. We build the connected system photographers rely on, and we've spent over a decade earning the trust of a global creative community that takes the craft seriously. Photography is at an inflection point. AI is reshaping what's possible for creative work, and that's where our mission shines. VSCO is building the full photographer's workflow: from creating and editing your work, to delivering it to clients, to running your business. All of it built thoughtfully, with photographers leading the way. We believe the future of photography tools creates more space for creativity, handling the busy work so photographers can focus on the craft. If you care about craft, community, and what technology can unlock for creative people, this is the work. We're a mission-driven and focused company where your work ships quickly, is meaningful, and reaches tens of millions of people worldwide. You'll have a real say in what we build and how we build it. We hire people who don't wait to be asked, naturally connect the dots, care about the quality of what they ship, and believe the best outcomes come from building together. About The Role VSCO is hiring a Senior Staff Engineer to own Reflex, our production GPU image-processing engine, and the path it takes into Studio Pro on iOS and macOS. This is a critical technical leader. You will deepen a stack that is already in customers’ hands: new capabilities (RAW, large-image export, desktop), production quality (color correctness, memory, performance), and the contracts that let product engineers ship on top of the engine without becoming graphics programmers. You will also lead company-wide engineering mentorship and force-multiplication — tools, workflows, and AI-assisted development that make the rest of the team more effective. In This
Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. About the General Agents Team The General Agents team, part of Scale’s Enterprise organization, builds robust general agents for customer use cases and applications. The team sits at the intersection of frontier agent development and real-world deployment, translating state-of-the-art reasoning and agentic capabilities into reliable, production-grade systems that drive real economic value. Our agents are scalable systems built around recurring enterprise problem domains, with a strong emphasis on generalization, extensibility, and deployment across many customers. About the Role As a Senior/Staff Machine Learning Engineer (MLE) on the General Agents team, you’ll play a critical role in designing, building, and deploying production-ready AI agents that solve high-impact enterprise problems. You will work across the full agent lifecycle—from model and system design to evaluation, deployment, and iteration—bridging cutting-edge agentic techniques with the constraints and requirements of real customer environments. You will: Design and implement end-to-end agent systems that combine LLM reasoning, tool use, memory, and control logic to solve recurring enterprise use cases. Build scalable, reliable agent architectures that can be deployed across many customers with varying data, tools, and constraints. Develop evaluation frameworks, datasets, environments, and metrics to measure agent performance, reliability, and business impact in production settings. Collaborate closely with product managers, customers, data annotators, and other engineering teams to translate enterprise requirements into robust agent designs. Productionize frontier agent techniques (e.g.,
EnCharge AI is a leader in advanced AI hardware and software systems for edge-to-cloud computing. EnCharge’s robust and scalable next-generation in-memory computing technology provides orders-of-magnitude higher compute efficiency and density compared to today’s best-in-class solutions. The high-performance architecture is coupled with seamless software integration and will enable the immense potential of AI to be accessible in power, energy, and space constrained applications. EnCharge AI launched in 2022 and is led by veteran technologists with backgrounds in semiconductor design and AI systems. Senior Emulation Engineer Location: India - Remote Job Description: At EnCharge AI, we are building the next generation of AI compute silicon — purpose-built for high-performance, low-power, and scalable AI inference. As an Emulation Engineer, you will play a critical role in validating complex AI accelerator architectures on emulation platforms before tape-out. This position is ideal for someone passionate about bridging the gap between hardware and software in fast-paced, deep tech environments. Responsibilities: • Set up and maintain Siemens Veloce emulation and prototyping platforms • Adapt SoC designs for Emulation and Prototyping • Develop and debug emulation testbenches and system-level environments • Support pre-silicon validation, power/performance analysis, and early software bring-up. Participate in silicon bring-up and validation. • Collaborate with design and verification teams to isolate design issues and accelerate debug. • Optimize performance of the emulation workloads and reduce turnaround time. • Work with firmware/software teams to enable use of emulators for OS and driver testing. Required Background: • BS/MS/Ph.D. in EE, CS, or related field with 7+ years of SoC design experience. • Experience with emulation platforms (Veloce, Palladium, or ZeBu) and FPGA-based prototyping systems (proFPGA, HAPS, or Protium) • Experience with emula
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role HP IQ is seeking a Senior Power Systems Engineer to drive end to end system architecture, modeling, and optimization for wearable platforms. The candidate will build power models, define system use cases and operating modes, partner with product and engineering teams to meet aggressive power targets, and define manufacturing tests and field telemetry for NPI. The ideal candidate brings strong systems thinking, hands-on experience with low power architectures, and the ability to build models and telemetry to drive tradeoffs from SOC level to board level to maximize user runtime. What You Might Do: Power architecture responsibility for a wearable device covering SoC states, sensor, audio, and memory architecture tradeoffs at board level. Battery, PCM, and charging profile definition and validation Design power experiments, test plans, and measurement & telemetry mechanisms across firmware and hardware to optimize power and performance. Build and maintain system-level power models to model expected power consumption for user profiles. Map power models by subsystem, correlating
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role HP IQ's Connectivity team is seeking an Embedded Firmware Engineer with strong hands-on experience across RTOS and embedded Linux platforms. You'll bring up new hardware, develop and support firmware across core device subsystems, and help scale products from prototype to fleet deployment. A key part of this role is enabling on-device intelligence — bringing AI models, sensing algorithms, and local processing to lightweight, power-constrained devices at the edge. What You Might Do Design, develop, and debug firmware across RTOS and embedded Linux platforms. Lead hardware bring-up — board bring-up, driver integration, and firmware support through to production. Develop and maintain firmware for subsystems such as connectivity, power and other sensors. Integrate lightweight model inference and sensing/proximity algorithms within tight compute, memory, and power budgets. Support fleet-scale deployment: OTA updates, field diagnostics, and post-launch sustainment. Collaborate with hardware, systems, and QA teams to troubleshoot system-level issues and drive root-cause fixes to closure. Ess
About VSCO For years, we've helped photographers create their work. Now we're building what comes next. VSCO exists for photographers. Not as a side feature, not as an afterthought, but as the whole point. We build the connected system photographers rely on, and we've spent over a decade earning the trust of a global creative community that takes the craft seriously. Photography is at an inflection point. AI is reshaping what's possible for creative work, and that's where our mission shines. VSCO is building the full photographer's workflow: from creating and editing your work, to delivering it to clients, to running your business. All of it built thoughtfully, with photographers leading the way. We believe the future of photography tools creates more space for creativity, handling the busy work so photographers can focus on the craft. If you care about craft, community, and what technology can unlock for creative people, this is the work. We're a mission-driven and focused company where your work ships quickly, is meaningful, and reaches tens of millions of people worldwide. You'll have a real say in what we build and how we build it. We hire people who don't wait to be asked, naturally connect the dots, care about the quality of what they ship, and believe the best outcomes come from building together. About The Role VSCO is hiring a Senior Software Engineer to be a primary contributor on Reflex, our production GPU image-processing engine, and significantly guide the path it takes into Studio Pro on iOS and macOS. You will deepen a stack that is already in customers’ hands: new capabilities (RAW, large-image export, desktop), production quality (color correctness, memory, performance), and the contracts that let product engineers ship on top of the engine without becoming graphics programmers. In This Role, You Will Be a primary contributor on the Reflex imaging engine: DAG execution, WGSL operators, GPU resource/memory budgets, and color management (linear workin
Get new senior memory system engineer jobs by email
Daily job updates · Unsubscribe anytime