Shui Cheng Yan's Team Unveils VoiceMem and Introduces a Streaming Dual-Brain Architecture for Real-Time Interaction Infrastructure
18 hour ago / Read about 0 minute
Author:小编   

Shui Cheng Yan's team, in collaboration with four leading universities and the Open Interaction Lab, has jointly introduced VoiceMem—a brain-inspired memory system designed for multimodal companions. Additionally, they have innovatively proposed the 'Streaming Dual-Brain Architecture' with the aim of constructing the next generation of real-time interaction infrastructure. This architecture ingeniously merges the information-processing left brain with the emotion-driven right brain. The left brain is tasked with storing factual information, while the right brain stores personality traits and emotions (the original "precipitate" has been corrected to "stores" for natural English). Through a parallel processing mechanism, the architecture achieves millisecond-level fast retrieval. The dual-brain retrieval process takes only 134 milliseconds, ensuring no perceived delay for users. In experiments such as the Long Conversation Memory Test (LoCoMo), VoiceMem outperformed existing systems like Mem0, accurately capturing emotional and personality information and providing robust support for interactions. Currently, Qwen Audio Agent has been successfully integrated into the system. Looking ahead, the team plans to further push the boundaries of multimodal memory, optimize long-term evolutionary capabilities, and enhance the system's deployability. They strive to create a plug-and-play, long-term evolvable memory system in the field of real-time interaction, laying a solid foundation for genuine connections between AGI Beings and humans. The first author of the paper is Zhifei Xie, and the corresponding author is Shui Cheng Yan.