DeepSeek has just rolled out a cutting-edge experimental multimodal visual comprehension model, DeepSeek-V4-Flash-Vision-Exp. This innovative model is now accessible through DeepSeek's API platform, where users can effortlessly activate it by setting specific parameters. When it comes to performance, the model's prowess in handling pure text rivals that of the official DeepSeek-V4-Flash iteration. Meanwhile, its capabilities as a multimodal agent are on par with Opus-4.8, especially shining in visual understanding agent benchmark assessments. It accommodates three distinct invocation formats, facilitating seamless integration with a diverse array of agent tools and supporting the input of both images and text. Moreover, the model provides three approaches for image input. Additionally, the Files API featured on the platform empowers users to reference images using a file_id, thereby minimizing redundant uploads and optimizing bandwidth usage.
