DeepSeek has open-sourced its first experimental multimodal model, DeepSeek-V4-Flash-Vision-Exp, on the Hugging Face platform. Licensed under the MIT protocol, the model boasts 305 billion parameters and is built on the V4-Flash-0731 base architecture, integrating a visual encoder and Aligner module. Developers can access complete development resources related to the model. Unlike traditional multimodal models, DeepSeek-V4-Flash-Vision-Exp focuses on enhancing multimodal Agent capabilities, enabling it to parse visual information such as webpage screenshots and collaborate with tools to complete complex tasks. Within just ten days of open-sourcing, seven quantized versions of the model have been released, lowering the barrier for local deployment. In terms of commercialization, DeepSeek has adopted a 'services-first, ecosystem-follows' approach, launching paid API access services first and open-sourcing the model weights ten days later. Notably, the model's visual module does not compromise its text processing capabilities; in fact, it has shown improvements across multiple test metrics, with some even surpassing Claude Opus 4.8, though there is still room for improvement in the NL2Repo project. Officials state that its multimodal Agent capabilities are approaching industry benchmarks. DeepSeek is building a comprehensive Agent solution, with this open-sourced model serving as a critical component. Its design emphasizes deep synergy between components and scenario adaptation, contrasting sharply with the modular assembly paths prevalent in the industry.
