Ali Tongyi Unveils Qwen3.8-Flash Model: Featuring 125 Billion Parameters and a Training Cost Merely 1/9 of Its Predecessor
2 day ago / Read about 0 minute
Author:小编   

The Ali Tongyi Qianwen team has recently rolled out the multimodal Mixture of Experts (MoE) model, Qwen3.8-Flash, and has made the weights of Qwen3.8-Flash-Next, regarded as a prototype for the Qwen4 series, available as open-source. The primary Qwen3.8-Flash model is equipped with a staggering 125 billion parameters, complemented by 51 billion N-gram embeddings, and activates 6 billion parameters for each token. It inherently supports a context window of 262,144 tokens, with the potential to extend up to 1 million tokens. This model has undergone enhancements in four crucial domains: attention mechanisms, residual connections, embedding layers, and optimization algorithms. Remarkably, its training cost is a mere one-ninth of that of Qwen3.7-Plus. Moreover, it exhibits significantly improved coding and office capabilities, securing top positions in 8 out of 14 benchmark tests. Qwen3.8-Flash is set to be accessible on the Qianwen AI platform, providing API services. Additionally, the model weights have been open-sourced on Hugging Face and ModelScope, with an FP8 quantized version also being made available.

  • C114 Communication Network
  • Communication Home
7 X 24 Track global technological trends
Hot Topic