NeoCognition has initiated the ApprenticeBench evaluation, a program that immerses intelligent agents in a simulated construction company setting. Here, these agents learn and manage accounts payable tasks, much like new human employees would. The outcomes of this evaluation reveal that Fable 5.1 and GPT-6 Astra have achieved pass rates of 72% and 68%, respectively, both outperforming the 51% pass rate attained by human participants. Nevertheless, the expense associated with Fable 5.1 processing a single invoice is roughly 2.5 times higher than that of a human worker.
The evaluation also highlights that, in actual work environments, the disparities among various models become even more pronounced, with notable performance variations between open-source and closed-source models. Moreover, the evaluation introduces the notion of the 'CUA tax' and discovers that for top-tier models, this cost is nearing zero.
Through meticulously designed experiments, the evaluation sheds light on the enhancements in model capabilities resulting from onboarding learning. Simultaneously, it uncovers issues such as reduced processing speeds and inefficient efforts exerted by the agents. Looking ahead, if AI agents can be developed to be more cost-effective, stable, and easy to collaborate with, they hold the potential to substantially boost team productivity.
Additionally, Professor Wang Yuxiang has officially joined NeoCognition as the full-time Director of Machine Learning.
