GPT-6 Astra Solves the Ultimate Challenge: FrontierMath Tier 4 Fully Conquered
4 hour ago / Read about 0 minute
Author:小编   

In 2026, GPT-6 Astra achieved a groundbreaking feat by successfully cracking the final, most formidable problem in FrontierMath Tier 4. This milestone marked a significant turning point, as all problems within this research-grade test—previously considered the ultimate test of mathematical prowess for large models—had now been solved by AI at least once. Consequently, Epoch AI announced that Tier 4 had reached a state of saturation. The FrontierMath test made its debut in 2024, with the primary objective of creating a mathematical benchmark that could withstand the rapid advancements in AI capabilities. Initially, it comprised three tiers: Tier 1, Tier 2, and Tier 3. However, as the capabilities of AI models continued to soar, a more challenging Tier 4 was introduced in 2025. The problems in this test were meticulously crafted by esteemed experts, including mathematics professors, to ensure their complexity and rigor. In its early stages, only a mere three problems were solved. In June 2026, FrontierMath unveiled version 2, which, after careful corrections and the removal of certain problems, ultimately retained 43 problems in Tier 4. Subsequently, the performance of various AI models continued to improve, with GPT-6 Astra standing out by achieving an impressive accuracy rate of 97.6%. Nevertheless, FrontierMath did not rest on its laurels; it introduced a new test featuring unsolved research problems, along with the FrontierMath Erdős Challenge, which demands formal verification proofs. Out of the 68 problems presented in the latter, Astra managed to solve only two, highlighting the ongoing challenges and the limitless potential for further exploration in the field of mathematical AI.