Moonshot AI Unveils Innovative Attention Architecture: Kimi Linear
2025-10-31 / Read about 0 minute
Author:小编   

At the heart of the Kimi Linear architecture lies the Kimi Delta Attention (KDA) linear attention module. This ingenious module capitalizes on the finite-state memory inherent in recurrent neural networks, employing a sophisticated gating mechanism to do so. The Kimi Linear model excels in task execution, delivering a remarkable surge in efficiency. When juxtaposed with full attention models, the KDA module slashes Key-Value (KV) cache consumption by a staggering 75%, while simultaneously amplifying decoding throughput sixfold during the processing of long texts on a million-scale. It functions as a seamless 'plug-and-play' substitute for full attention architectures, bolstering both performance and efficiency.