Volcano Engine Unveils Doubao Speech Model 2.0: A Dual Leap in Semantic Understanding and Emotional Expression
2025-10-16 / Read about 0 minute
Author:小编   

On October 16, 2025, Volcano Engine made an official announcement, introducing the Doubao Speech Synthesis Model 2.0 (Doubao - Seed - TTS 2.0) and the Voice Replication Model 2.0 (Doubao - Seed - ICL 2.0). Built upon the innovative architecture of the Doubao large language model, these two models mark a significant transformation in speech technology, evolving from mere "text - to - speech reading" to "emotion - infused expression grounded in understanding."

The speech synthesis model is capable of comprehending multi - round dialogue contexts. It can precisely convey tone, pauses, and emotional fluctuations, offering users fine - grained control over aspects like speech speed and vocal timbre through specific instructions.

The Voice Replication Model 2.0 has taken a big step forward. Not only can it replicate vocal timbres in a matter of seconds, but it also now boasts emotional interpretation capabilities. This makes it highly adaptable to a wide range of scenarios, including novel narration and dialogue interaction.

After undergoing specialized optimization for educational settings, the model has achieved an impressive 90% accuracy rate in reading complex formulas and symbols across all subjects, from elementary to high school levels. This performance far exceeds the industry average.

At present, both models are accessible via the Volcano Engine speech console. They are already serving a diverse clientele, including OPPO and Yangcong Xueyuan, and are being applied in various scenarios such as dialogue assistants and educational support.

  • C114 Communication Network
  • Communication Home
7 X 24 Track global technological trends
Hot Topic