Alibaba Unveils Its Next-Gen Foundation Model Architecture: Qwen3-Next
2025-09-12 / Read about 0 minute
Author:小编   

Alibaba has officially rolled out its cutting-edge foundation model architecture, Qwen3-Next, and has also made the Qwen3-Next-80B-A3B series models, which are built upon this architecture, available as open-source resources. When pitted against the MoE (Mixture of Experts) model structure employed in Qwen3, Qwen3-Next has brought about a number of pivotal enhancements.

These enhancements encompass the incorporation of a hybrid attention mechanism, which allows the model to better focus on relevant information; the utilization of a highly sparse MoE structure, improving computational efficiency; the implementation of a range of training-stabilizing and user-friendly optimization strategies, ensuring smoother and more effective training processes; and the adoption of a multi-token prediction mechanism, which significantly boosts inference efficiency.

The Qwen3-Next-80B-A3B-Base model is a powerhouse, boasting a staggering 80 billion parameters, out of which 3 billion are actively engaged in processing. This model delivers performance that is on par with, or in some cases, even surpasses that of the Qwen3-32B dense model. What's more, the training cost associated with the Qwen3-Next-80B-A3B-Base model is less than one-tenth of that required for the Qwen3-32B model, making it a highly cost-effective solution.