Tencent has launched the groundbreaking AudioGenie training-free multi-agent system, marking a significant advancement in the realm of multi-modal to multi-audio (MM2MA) generation. This innovative system is capable of synthesizing a diverse array of audio content, including sound effects, speech, and music, from multi-modal inputs like videos, texts, and images. It effectively addresses the challenge posed by the scarcity of high-quality paired data. AudioGenie employs a sophisticated two-tier architecture, comprising a generation team and a supervision team. Experimental results have demonstrated its outstanding performance, paving the way for new applications in cross-modal audio generation.
