To address the evaluation challenges inherent in long-term speech conversation scenarios, a collaborative team from the University of Melbourne and the University of New South Wales has developed the Multi-Session Speech Memory Benchmark, known as VoxMem. This benchmark adopts a two-dimensional classification approach, focusing on 'acoustic evidence × memory operations', and encompasses 15 distinct examination combinations. Notably, all questions remain unchanged across varying history lengths, ranging from 8K to 64K.
VoxMem boasts an extensive dataset, comprising 3,196 evaluation instances and 34,743 speech sessions, amounting to approximately 177 hours of audio content. The evaluation results from 15 prominent audio large models reveal that, at a context length of 32K, none of the models attained an overall accuracy rate surpassing 40%. The models exhibited a significantly better performance in recalling speech semantics compared to their ability to retain native audio information, such as speaker identity, paralinguistic cues, and environmental sounds. Moreover, the models displayed remarkably low accuracy in tracking alterations in tone and background sounds.
As the history length increased, the memory accuracy for all categories of information experienced a decline. Presently, the speech memory bottleneck in audio large models primarily stems from challenges in information identification, localization, and associative binding, rather than from reasoning operations.
