On September 23, Qianwen officially rolled out the Qwen-Audio-3.1 series of advanced large-scale speech models. This significant upgrade thoroughly refines three fundamental models: speech recognition, speech synthesis, and real-time speech interaction. Additionally, it introduces two groundbreaking models—the audio creation model, Qwen-Audio-3.1-TTS-Next, and the audio understanding model, Qwen-Audio-3.1-ASR-Next. With these enhancements, the five new speech models now collectively establish a comprehensive audio capability framework that spans "understanding, generation, interaction, and creation." In a move to alleviate financial burdens for users, Qianwen has reduced the prices of all Qwen-Audio speech models. Specifically, TTS prices have been slashed by approximately 70%, Realtime prices by about 85%, and ASR prices by an impressive 95%.
