Taobao’s TaoMate-H3 Model Opens Up: Enabling ‘Minute-Level’ Continuous Audio-Visual Creation Through Three-Step Generation
1 day ago / Read about 0 minute
Author:小编   

TaoMate-H3 is an innovative audio-visual joint streaming generation model, crafted by Alibaba’s TaoLive AIGC team. It builds upon the foundation of the MiniMax H3 model, incorporating a unique three-step LoRA (Low-Rank Adaptation) inference process alongside autoregressive generation technology. This combination allows the model to progressively generate content in small, manageable segments, eliminating the need to wait for the entire video to be produced before viewing.

The model excels at creating videos that feature dialogues, singing voices, or ambient sounds, all based on textual descriptions. It supports minute-level continuation, meaning it can seamlessly extend audio-visual content in real-time, and offers multiple output resolutions in both landscape and portrait formats. This capability significantly cuts down on the time required to produce the initial segment of content, providing invaluable technical support for live streaming, virtual character performances, and similar scenarios.

Developers can now access the inference code and LoRA weights for TaoMate-H3 on GitHub and Hugging Face. These resources are available for download, enabling developers to experience, further develop, and conduct research on the model.

TaoMate-H3 boasts three core capabilities that set it apart:

  • Three-Step Streaming Generation: This approach allows for the gradual and efficient creation of audio-visual content, enhancing the overall generation process.
  • Joint Audio-Visual Generation: The model seamlessly integrates audio and visual elements, producing cohesive and engaging multimedia content.
  • Minute-Level Continuous Continuation: It supports the real-time extension of content, ensuring a smooth and uninterrupted viewing experience.

The model’s versatility makes it suitable for a wide range of applications, including e-commerce product explanations, character singing performances, anime dialogue creation, and ambient effect generation.

The inference process of TaoMate-H3 comprises four key stages:

  1. Segmented Content Organization: Content is organized into manageable segments for efficient processing.
  2. Audio Guidance Preparation: Audio cues are prepared to guide the generation process, ensuring synchronization between audio and visual elements.
  3. Joint Streaming Generation: The model generates audio and visual content simultaneously, in a streaming fashion.
  4. Decoding Output: The generated content is decoded and outputted in the desired format.

By leveraging cutting-edge technologies such as three-step LoRA, Self Forcing autoregressive training, audio-visual KV caching, continuous time encoding, and audio track guidance, TaoMate-H3 achieves efficient and stable generation. Tests conducted on specified hardware demonstrate that TaoMate-H3 significantly enhances generation efficiency compared to the original MiniMax H3 model, with a markedly shorter ready time for the initial segment.

The TaoMate-H3 project adheres to relevant open-source protocols, ensuring transparency and accessibility for the developer community. The TaoLive AIGC team is committed to ongoing optimization of the model’s speed and plans to gradually open-source more streaming audio-visual generation models in the future, fostering innovation and collaboration in the field.