Qwen3-VL Open-Source Lightweight Models Propel Multimodal Tech Forward
2025-10-15 / Read about 0 minute
Author:小编   

News coming from the ModelScope Community reveals that on October 15, 2025, Alibaba Tongyi made a significant move. It introduced two new Dense architecture models, Qwen3-VL-8B and Qwen3-VL-4B, to the Qwen3-VL family and generously made them open-source.

These two models are real space-savers, demanding less video memory while still packing all the punch of the Qwen3-VL series. For each model size, there are two distinct versions available: Instruct and Thinking.

Qwen3-VL-8B shines brightly in a multitude of evaluations. Whether it's tackling STEM problems, answering Visual Question Answering (VQA) queries, performing Optical Character Recognition (OCR), comprehending videos, or handling Agent tasks, it outperforms competitors like Gemini 2.5 Flash Lite and GPT-5 Nano. In fact, its performance is getting dangerously close to that of the previous generation's ultra-large-scale model, Qwen2.5-VL-72B.

On the other hand, Qwen3-VL-4B has its sights set on end-side applications. It offers a higher cost-effectiveness ratio, making it an ideal choice for deploying intelligent terminals that require AI-powered visual understanding capabilities.

Through innovative architectural design, these two models have successfully tackled the "see-saw" problem that has long plagued small models - the trade-off between visual and textual capabilities. They've achieved a harmonious improvement in both visual precision and textual robustness.

Currently, both models can be accessed on the ModelScope Community and Hugging Face platforms. What's more, they come with support for the FP8 version, making them even more versatile and accessible.

  • C114 Communication Network
  • Communication Home
7 X 24 Track global technological trends
Hot Topic