GPT-6 Astra Cracks 5 Erdős Problems; Other 4 AI Models Draw Blanks
3 hour ago / Read about 0 minute
Author:小编   

Epoch AI, in collaboration with the University of Manchester, has unveiled a research paper detailing the results of the FrontierMath Erdős (FME) benchmark test. This test was crafted as a direct response to Terence Tao's critique of AI's mathematical prowess. It leverages 68 unsolved mathematical conundrums posed by Erdős to sidestep the issue of data contamination, mandating that AI models generate formal proofs utilizing the Lean proof assistant. All participating models underwent evaluation under identical conditions.

In this rigorous examination, five leading AI models vied for supremacy, with only GPT-6 Astra managing to secure a non-zero score. It achieved a modest yet noteworthy 3% success rate, while its four counterparts registered zero successes. During the standard test phase, GPT-6 Astra adeptly solved two problems. Furthermore, in an extended test with a more lenient computational budget, it tackled an additional three, culminating in a total of five problems resolved. Among these triumphs were Erdős's inaugural significant problem from 1931 and the Erdős-Sós conjecture.

The paper also candidly acknowledges the test's constraints, including the potential for underestimating AI's capabilities due to the high costs associated with formalization and the absence of an assessment metric for AI's capacity to pose questions. As it stands, 63 out of the initial 68 problems remain stubbornly unsolved, beckoning future exploration and innovation.