Analyst Ming-Chi Kuo has disclosed that Nvidia has re-launched the Rubin CPX project—a significantly revamped artificial intelligence (AI) inference prefilling acceleration GPU initiative. Production is slated to begin in the first quarter of 2027. The newly designed Rubin CPX boasts computational capabilities nearly on par with the standard Rubin GPU, but with enhanced prefilling performance and a power consumption rate of 2300W. It comes equipped with 168GB of HBM4 memory.
From an architectural standpoint, eight Rubin CPX units are grouped together to form a compute tray, and eight such trays are combined to create a rack module. This setup is horizontally scalable using Spectrum-6 all-copper Ethernet and incorporates an independent MGX ETL rack design. Horizontal scaling across rack modules is facilitated through the Spectrum-6+OSFP optical solution. Nvidia suggests pairing the new Rubin CPX with the standard Rubin in a 1:1 ratio, with Ethernet-based RDMA connections linking the two rack types. This configuration enables the product to manage prefilling workloads with greater deployment flexibility and cost efficiency. Additionally, the compute tray's 1.34TB HBM memory capacity caters to the demands of most long-context prefilling and KV Cache establishment tasks.
