On August 14, MiniMax, an AI company, introduced its latest innovation—the MiniMax-Music3 music generation model. This cutting-edge tool has the capability to create entire songs, extending up to 5 minutes in length, based solely on lyrics and music descriptions provided. It leverages a sophisticated hierarchical autoregressive framework, which integrates a Global LLM boasting 8 billion parameters and a Local LLM with 600 million parameters. The Global LLM takes charge of grasping the overarching semantics and structural dynamics of songs, whereas the Local LLM hones in on the meticulous restoration of fine-grained acoustic details. The model's output is in stereo WAV format, sampled at 32kHz with a 16-bit depth, ensuring the preservation of musical themes and rhythms throughout lengthy audio sequences. It adeptly handles various song structures, including intros, verses, and beyond.
