DeepSeek-V4 Series Unveils Its Pioneering Open-Source Multimodal Vision Model
1 day ago / Read about 0 minute
Author:小编   

On August 31, 2026, DeepSeek made a significant stride in AI development by officially releasing the source code of its inaugural experimental multimodal model within the V4 series—DeepSeek-V4-Flash-Vision-Exp. This groundbreaking release, licensed under the MIT framework, is constructed upon the robust V4-Flash architecture, augmented with a sophisticated vision module. This enhancement endows the model with comprehensive image comprehension abilities, empowering it to accurately describe visual content, recognize text within screenshots, analyze intricate charts, and perform a myriad of other tasks.

Supporting a wide array of image formats, including JPEG, PNG, GIF, and WebP, DeepSeek-V4-Flash-Vision-Exp demonstrates versatility and adaptability. Boasting a staggering 305 billion parameters, the model not only maintains parity with V4-Flash in executing pure-text Agent tasks but also showcases remarkable enhancements in visual understanding Agent benchmark evaluations. Its multimodal Agent capabilities are on par with those of Opus-4.8, marking a significant advancement in the field of artificial intelligence.