Fudan University Unveils Human-Centric Audio-Visual Tracking Benchmark for Complex Environments
2 day ago / Read about 0 minute
Author:小编   

The CVL Laboratory at Fudan University has introduced a specialized audio-visual tracking benchmark, named AVTrack, which is specifically designed for complex scenarios centered around human activities. The associated research paper has been accepted for presentation at the prestigious ICML 2026 conference. Moreover, the project's homepage, along with the paper, codebase, and dataset, have all been made accessible to the public. AVTrack primarily tackles the core challenges of audio-visual correspondence and long-term identity maintenance within dynamic and intricate settings. The benchmark comprises 871 video clips, each averaging 54.0 seconds in length, and offers a total of 3,120 pixel-level instance trajectories with consistent cross-frame identities. These clips span six diverse source categories, including TV series and vlogs.
As a dedicated testing benchmark, AVTrack categorizes eight particularly challenging scenarios, such as visual occlusion and audio-visual inconsistency. The complexity of these scenarios far surpasses that of existing datasets in the same domain. The research team conducted tests on representative VIS (Video Instance Segmentation) and AVIS (Audio-Visual Instance Segmentation) methods, as well as on the advanced Gemini 2.5 Pro model. The results revealed that current methods struggle to perform effectively in complex scenarios, and the audio-visual comprehension abilities of general large models do not readily translate into stable pixel-level speaker tracking capabilities.
To address these challenges, the team proposed a modular model called AVTracker, which achieves optimal performance on the HOTA (Higher Order Tracking Accuracy) metric through a local-to-global three-stage design. However, scenarios involving audio-visual inconsistency and visual occlusion continue to pose significant difficulties for all methods tested. This has prompted the team to delve deeper into these challenging scenarios and explore the underlying mechanisms involved.
The overarching goal of AVTrack is to analyze model performance and identify shortcomings in complex audio-visual fusion scenarios. By doing so, it aims to drive further exploration into models' continuous understanding capabilities for speakers in dynamic and ever-changing environments.