On September 10, DeepSeek officially launched DeepSeek V4.1 Flash, the smallest model in its new model architecture series. This model features native multimodal visual understanding capabilities. By significantly reducing the KV Cache size, its demand for HBM has decreased to 1/4, and its demand for SSD has dropped to 1/8 compared to the previous generation. In Agent usage scenarios, this improvement significantly reduces cache hit costs, thereby substantially cutting the usage expenses for Agent-type tasks.
