On October 20, 2025, Meituan's LongCat team made an official announcement, launching VitaBench—a benchmark specifically designed for assessing large-model intelligent agents. This benchmark is meticulously crafted to mirror real-world situations. It leverages common scenarios such as food delivery orders, dining out at restaurants, and traveling as prime examples to build an interactive evaluation setting. This setting encompasses 66 tools and incorporates cross-scenario comprehensive tasks.
For the first time, the team conducted a quantitative analysis of intelligent agent tasks from three key perspectives: deep reasoning capabilities, tool utilization efficiency, and user interaction quality. Their research uncovered a striking finding: the success rate of currently leading reasoning models in handling complex cross-scenario tasks stands at a mere 30%. This starkly highlights the substantial gap between the capabilities of existing intelligent agents and the practical demands of real-life applications.
VitaBench is now entirely open-source, with the overarching goal of serving as a foundational resource to support the research, development, and practical deployment of intelligent agents in real-world contexts.
