Micron Explores Near-GPU NAND Flash for Supporting Larger-Scale AI Models
1 day ago / Read about 0 minute
Author:小编   

Micron is developing a high-endurance 'Near-GPU NAND' flash memory module, which is planned to be deployed inside the GPU package or on the PCB, close to the GPU. This module strikes a balance between storage density and endurance. Although its storage density is lower than that of traditional TLC or QLC NAND, it enhances I/O speed and bandwidth. Acting as a high-speed buffer layer between video memory and regular storage, it enables the GPU to directly access a storage pool of hundreds of GBs, making it suitable for running large-scale AI large language models and potentially breaking the memory capacity limitations for LLM inference.