DeepSeek officially announced on September 10, 2026, the launch of its groundbreaking V4.1 Flash model. This latest iteration boasts native multimodal visual comprehension abilities, underpinned by a robust 552 billion-parameter Mixture of Experts (MoE) architecture and a Causal-Encoder-Decoder framework. Despite its compact design, featuring only 8 billion input activations and 16 billion output activations, the V4.1 Flash model manages to slash operational costs relative to its counterparts of similar scale, all while surpassing the V4 Pro in terms of performance metrics.
One of the standout features of this model is its significantly reduced KV Cache size, which, in turn, leads to a notable decrease in the demands placed on High Bandwidth Memory (HBM) and Solid State Drives (SSDs). This optimization further drives down the expenses associated with Agent tasks, making the V4.1 Flash model an even more attractive proposition.
In a move to streamline accessibility, the API call designation has been updated to "deepseek-flash." Consequently, both the V4 Flash series and the V4 Pro have been rerouted to the V4.1 Flash model, ensuring a seamless transition for users. Notably, Tencent's suite of products, including WorkBuddy, CodeBuddy, and OpenCode, have already embraced this cutting-edge model, which is now available as open-source software. The new pricing structure for the V4.1 Flash model came into effect on the same day as its announcement, September 10, marking a new chapter in cost-effective and high-performance AI solutions.
