GPT-6 Astra Scores Nearly Threefold Higher Than Claude in Simulated Vending Machine Challenge
21 hour ago / Read about 0 minute
Author:小编   

Andon Labs administered the Vending-Bench 2 test, in which artificial intelligence (AI) was tasked with simulating the operation of a vending machine over a year, starting with a capital of $500. GPT-6 Astra achieved an impressive average final account balance of $15,515, outperforming Claude Fable 5.1 and securing the top position for the first time. Astra's superiority is evident in several key areas: it successfully maintained low procurement prices over an extended period, at one point slashing product quotes to roughly 48% of their original cost. Moreover, it consistently adhered to rules, avoiding losses associated with advance payments. In contrast, Fable's procurement costs escalated steadily, leading to substantial losses from advance payments due to its inability to follow established rules. During a three-player competitive test, Astra refused to engage in rule-breaking collaborations and emerged victorious in all three matches. This test underscores the significance of long-term autonomy for AI agents, with the subsequent challenge being the consistent translation of accurate judgments into effective actions.