To address the issue of multimodal large models "understanding but failing to act correctly" in spatial intelligence, related research has introduced the VA-Bench evaluation benchmark. This benchmark constructs 280 foundational scenarios covering 14 categories of robotic manipulation tasks and establishes a complete interactive closed loop of observation-reasoning-action-correction. Models are required to autonomously identify information gaps, adjust perspectives, and output precise action parameters. Test results show that leading multimodal models excel in target recognition and localization as well as operational semantic understanding, achieving scores close to full marks. However, their success rate in completing entire tasks is only around 50%. Their capability shortcomings primarily focus on three areas: fine-grained spatial execution, online error correction, and dual-arm coordination. Comparative experiments indicate that the ability to actively perceive and complete information is more crucial than simply providing multi-view images. Additionally, model performance significantly declines in tasks involving spatial layout transfer and long-sequence combinations. The VA-Bench evaluation benchmark clearly measures the specific capability boundaries of current multimodal large models in transitioning from spatial understanding to reliable physical actions.
