资讯🔥7.0
Micron Explores Near-GPU NAND Flash to Run Bigger LLMs
📌 概要
美光正探索开发'近GPU NAND'架构,将高耐久NAND闪存模块置于GPU附近,而非远端存储池。该方案定位于HBM/DRAM与常规存储之间,以较低密度换取更快I/O速度、带宽和读取时间,帮助GPU直接访问数百GB存储,从而运行更大的LLM模型。
⚡ 关键要点
- ▸美光探索'近GPU NAND':将高耐久NAND置于GPU封装或PCB附近
- ▸定位介于HBM/DRAM与传统存储之间,牺牲密度换取I/O速度与带宽
- ▸目标是让GPU直接访问数百GB存储池,支撑更大LLM的运行
Micron is reportedly exploring the idea of developing high-endurance NAND Flash modules that would be positioned closer to the GPU, rather than being located far away in the storage pool behind multiple protocols. The company is said to be investigating ways to bring NAND Flash closer to handle workloads that don't require the latency and bandwidth of traditional HBM or regular DRAM modules. This architecture, currently referred to as "near-GPU NAND," would serve as a middle ground between density and durability. This concept would allow GPUs to access a pool of NAND Flash that has a lower density than conventional TLC or QLC NAND, but instead focuses on improving I/O speeds, bandwidth, and read times.
For instance, this would enable a GPU to access a potential storage pool with hundreds of gigabytes of storage directly on the GPU package or PCB, similar to what DRAM like HBM or GDDR7/LPDDR6 already does. This tier would sit between GPU memory and regular system storage, acting as a very fast buffer, which would be ideal for the inference of massive LLMs. At that point, LLMs wouldn't be memory-bound and could run on systems with fewer GPUs, but with more storage and memory tiers to compensate for the required inference size. Additionally, this NAND would be significantly cheaper than HBM or regular DRAM and would be sized according to system specifications. Images below are HBF slides from SK hynix and Sandisk.
Read full story
For instance, this would enable a GPU to access a potential storage pool with hundreds of gigabytes of storage directly on the GPU package or PCB, similar to what DRAM like HBM or GDDR7/LPDDR6 already does. This tier would sit between GPU memory and regular system storage, acting as a very fast buffer, which would be ideal for the inference of massive LLMs. At that point, LLMs wouldn't be memory-bound and could run on systems with fewer GPUs, but with more storage and memory tiers to compensate for the required inference size. Additionally, this NAND would be significantly cheaper than HBM or regular DRAM and would be sized according to system specifications. Images below are HBF slides from SK hynix and Sandisk.
Read full story