On October 20, 2025, Meituan's LongCat team officially rolled out the VitaBench evaluation benchmark tailored for large-scale model agents. This particular benchmark zeroes in on three everyday scenarios that occur with high frequency: ordering food delivery, dining out at restaurants, and traveling. It has established an interactive evaluation setting that encompasses 66 tools and has crafted cross-scenario tasks. For the very first time, it quantitatively dissects task complexity across three dimensions: in-depth reasoning, tool utilization, and user interaction. In doing so, it brings to light the capability boundaries of current agents when faced with real-world scenarios.
