Aliyun Unveils New-Generation, In-House Developed Storage: Supercharging AI Training Across Millions of Cards, Slashing AI Storage Expenses by 69%
8 hour ago / Read about 0 minute
Author:小编   

At the 2026 Apsara Conference, Aliyun introduced CPFS, a cutting-edge, high-performance storage solution tailored for next-generation AI training clusters. This storage system boasts throughput in the hundreds of terabytes per second (TB/s) and tens of millions of input/output operations per second (IOPS), with a single file system capable of scaling up to 100 petabytes (PiB). It significantly accelerates the training process for models that utilize millions of cards, slashing the average model startup time by half, boosting peak computing power utilization by 30%, reducing AI storage costs by up to 69%, and doubling the capacity to handle business operations.

CPFS leverages a fully in-house developed architecture that segregates data and metadata services, allowing a single file system to manage trillions of files efficiently. It facilitates unified management of training data, seamless scaling in sync with GPU clusters, as well as intelligent lifecycle management and tiered cold-hot storage functionalities.

Additionally, Aliyun unveiled KVCacheStore, a KVCache acceleration engine designed for large-scale model inference. Functioning as a novel G3.5 storage layer, it adeptly handles massive caches, delivering a throughput of 40GB/s per compute node and achieving millions of queries per second (QPS). A single instance of KVCacheStore can support KV storage on a scale of hundreds of billions, with measured cache hit rates showing an improvement of over 20%.

Aliyun is committed to fostering collaborative innovation across computing, networking, and storage domains to deliver high-performance, scalable, and cost-effective storage solutions for AI training and applications.