DeepSeek has officially rolled out the DeepSeek V4.1 Flash, a Mixture-of-Experts (MoE) model boasting a staggering 55.2 billion parameters. As the most compact member of its latest model architecture lineup, this model introduces a groundbreaking Causal-Encoder-Decoder framework. It incorporates 8 billion input-activated parameters and 16 billion output-activated parameters. Its performance not only matches but exceeds that of several state-of-the-art flagship models, including the esteemed DeepSeek V4 Pro.
In terms of pricing, DeepSeek has made a strategic downward adjustment to the V4.1 Flash, while maintaining its peak-valley pricing mechanism. These new rates are now officially in place. Moreover, the model's KV Cache efficiency has undergone a remarkable enhancement. It now requires only one-fourth of the High Bandwidth Memory (HBM) and one-eighth of the Solid State Drive (SSD) space compared to its predecessor, the V4 Flash. Simultaneously, its context length capability has been expanded to 1M, with decoding costs remaining virtually unchanged, regardless of the increase in context length.
The API for DeepSeek V4.1 Flash is currently operational, while older iterations like the V4 Flash have been discontinued. The V4 Pro will undergo a phased retirement starting September 14, with requests seamlessly redirected to the new model and billed at the updated rates. It's worth noting that DeepSeek is also actively fostering adaptation within the open-source community, having concurrently released DeepSeek Harness v0.1.5. Applications such as Tencent WorkBuddy and OpenCode have fully embraced and integrated this model. Furthermore, DeepSeek has open-sourced DeepJIT, a lightweight, cross-backend xPU kernel JIT compilation library, further cementing its commitment to innovation and collaboration.
