According to Cailian Press on August 7, 2026, NVIDIA is testing multiple versions of its Rubin Ultra GPU, with some versions featuring lower HBM memory capacities than originally planned, in response to supply chain bottlenecks and a shortage of high-end memory driven by surging AI demand. This move reflects the ongoing pressure on data center construction costs to be passed downstream. Sources familiar with the matter revealed that NVIDIA has tested at least three Rubin Ultra GPU variants with reduced HBM specifications over the past few weeks. Semiconductor industry research firm TrendForce stated that in addition to the initially planned 12-layer HBM4e configuration, NVIDIA is now evaluating three alternative options: 8-layer HBM4e, 12-layer HBM4, and 8-layer HBM4, with the final product specifications yet to be determined. TrendForce believes that this specification adjustment testing stems from two key factors: first, the anticipated industry-wide DRAM shortage in 2027 is expected to constrain wafer capacity available for HBM production; second, uncertainties remain regarding the testing and validation progress of 12-layer HBM4e, as well as the yield rates during mass production. Meanwhile, given that the supply shortage of LPDDR5X memory chips is expected to persist until 2027, NVIDIA has decided to halve the SOCAMM memory module capacity in its Vera Rubin superchip modules. Other AI hardware vendors are similarly adjusting memory configurations based on supply conditions. TrendForce noted that several cloud service providers are considering reducing the HBM capacity of their next-generation, in-house-developed AI-specific chips. The firm forecasts that total HBM bit shipments will grow by 50%–60% year-on-year in 2027, though this growth will still fall short of meeting market demand. Consequently, TrendForce believes that HBM suppliers will retain pricing power throughout 2027, while AI chip vendors will face dual pressures of supply shortages and rising costs. Semiconductor industry analysis firm SemiAnalysis pointed out that the hardware cost savings NVIDIA achieves from reduced HBM memory usage will be redirected toward switches and optical interconnect components. SemiAnalysis believes that NVIDIA is leveraging hierarchical storage and optical interconnect technologies to alleviate the pressures of high HBM memory prices and tight supply. However, industry insiders noted that while NVIDIA can partially offset the performance degradation from reduced memory specifications by enhancing GPU computing power and optimizing interconnect bandwidth, overall system deployment costs and cluster construction complexity will increase. For major cloud providers like Microsoft, Meta, Amazon, and Google, which continue to ramp up AI capital investments, this suggests further upward pressure on future data center construction costs.
