Shanghai AI Lab Unveils IWR-Bench, Exposing Flaws in AI Video-to-Web Transformation
2025-10-21 / Read about 0 minute
Author:小编   

As reported by AIBASE, in October 2025, the Shanghai Artificial Intelligence Laboratory, working hand - in - hand with Zhejiang University, introduced the world's inaugural video - to - web evaluation benchmark, IWR - Bench. This move effectively bridged the gap in the dynamic interaction evaluation for AI front - end development. This benchmark sets a challenging task for models: they are required to reconstruct webpage interactions using a combination of 'video + static resources'. It encompasses a variety of scenarios, including the popular 2048 games and airline ticket booking processes. The models are then evaluated based on two key metrics: the visual fidelity score (VFS) and the interaction functionality score (IFS). When 28 mainstream models were put through the rigorous evaluation process, the results were quite revealing. GPT - 5, a well - known model, only managed to score a total of 36.35 points. Its IFS was a mere 24.39%, and its VFS stood at 64.25%. These figures clearly lay bare the AI's limitations in comprehending dynamic logic.