GPT-6 Outperforms on ARC-AGI-3, Prompting ARC to Introduce New Evaluation Centered on Innovation
2 day ago / Read about 0 minute
Author:小编   

GPT-6 Astra has attained an impressive 99.9% score on the semi-private test set of ARC-AGI-3, effectively maxing out the test set and rendering the original ARC series evaluation framework outdated. This outcome necessitates a complete overhaul of the existing questions. François Chollet, the visionary behind the ARC Prize, has disclosed plans for the upcoming generation of evaluations: ARC-AGI-4 is slated for release in the first quarter of next year, with a specific focus on evaluating AI's proficiency in continuous learning and curriculum learning over extended periods. Additionally, ARC-AGI-5, often referred to as the "ultimate human exam," is currently in the preparatory phase, with its core objective centered on the capacity to "invent."

Models such as GPT-6 Astra have substantially enhanced their scores by refining evaluation engineering techniques, a development that has ignited debates regarding the fairness of the current testing framework. ARC-AGI-3 assesses AI's exploration, modeling, and other capabilities through gamified levels. However, due to the constrained and deterministic nature of the testing environment, its saturation does not equate to the attainment of Artificial General Intelligence (AGI).

The ARC series evaluations may reach their conclusion by the sixth or seventh generation, signaling the arrival of AGI when the disparity in learning efficiency between humans and AI becomes immeasurable. Given the swift progression of AI technology, it is imperative that evaluation questions are continually updated. Should open-ended invention tasks also prove incapable of differentiating between human and machine capabilities, then the notion of "no more tests to take" will undoubtedly become a focal point in AGI discussions.