QianWen Unveils Large-Scale Speech Synthesis Model: Qwen-Audio-3.0-TTS
1 day ago / Read about 0 minute
Author:小编   

On July 20, 2026, Alibaba's QianWen project made waves in the tech world by launching its advanced large-scale speech synthesis model, Qwen-Audio-3.0-TTS. This cutting-edge model comes in two versions: a real-time interactive Flash version, boasting a swift first-packet delay of just 300 milliseconds, and a high-quality generation Plus version for superior audio output. The new iteration of the model showcases remarkable enhancements in several key areas, including fine-grained label control, adherence to freestyle instructions, multilingual and dialect support, and robustness in complex acoustic environments. It is capable of supporting 16 languages and 20 dialects, and comes complete with a premium voice library. Moreover, the audio output quality has been upgraded to an impressive 48KHz, ensuring crystal-clear sound for users.