As reported by 36Kr, Guillermo Rauch, the CEO of the American startup Vercel, took to social media to share that internal agent testing frameworks had been employed to assess the real-world task performance of multiple mainstream large language models. The findings were striking: Kimi K2, an open-source model crafted by China’s Moonshot team, demonstrated a remarkable fivefold increase in speed and a 50% boost in accuracy compared to GPT-5 when deployed in agent application scenarios. This stellar performance not only eclipses that of GPT-5 but also outshines other internationally renowned closed-source models, such as Claude Sonnet4.5, sparking widespread acclaim within the tech community.
