According to Taiwan's DigiTimes, NVIDIA is considering adjusting the high-bandwidth memory (HBM) solution for its next-generation Rubin Ultra AI accelerator, changing from the original 12-layer stack to an 8-layer configuration. This move aims to address rising memory costs, enhance cost-effectiveness per unit bandwidth, and improve overall system economics. After the adjustment, the memory capacity per GPU will decrease from 288GB to 192GB, a reduction of 33%, while bandwidth remains unchanged or slightly increases, significantly improving the cost-performance ratio per GB of bandwidth.
