As per Xiaomi Technology, the next-gen Kaldi team from Xiaomi Group's AI Lab has recently introduced the ZipVoice series of text-to-speech (TTS) models. These models, built on the Flow Matching architecture, encompass ZipVoice—a zero-shot, single-speaker TTS synthesis model—and ZipVoice-Dialog, a zero-shot dialog TTS synthesis model.
ZipVoice effectively tackles the prevalent challenges of bulky parameter sizes and sluggish synthesis speeds that plague existing zero-shot TTS models. Meanwhile, ZipVoice-Dialog overcomes the stability and inference speed limitations that currently hinder dialog TTS synthesis models.
