On August 26th, Ali QianWen introduced the Qwen3.8-Flash model, which incorporates a multimodal Mixture of Experts (MoE) architecture. This model serves as an early sneak peek into the forthcoming Qwen4 architecture. The full production version of Qwen3.8-Flash will be accessible soon through the Qwen Cloud API, with a pricing structure set at $0.16 for every 1 million input tokens and $0.47 for every 1 million output tokens. Impressively, the model encompasses a total of 125 billion parameters, including 51 billion N-gram embedding parameters. However, only 6 billion parameters are activated per token, guaranteeing exceptional cost-efficiency.
