On September 11, Xiaomi made an official announcement regarding the release and open-sourcing of Xiaomi-CocktailASR-1. This is a large-scale, industrial-grade model specifically crafted for target speaker voice recognition, aimed at overcoming the notorious "cocktail party" problem. The model adopts an end-to-end Large Language Model (LLM) architecture. By utilizing a snippet of reference audio from the target speaker as a voiceprint cue, it can precisely isolate and transcribe the target user's voice, even amidst the chaos of complex environments where numerous individuals are speaking simultaneously.
