On September 3 (local time), OpenAI made an announcement about the GPT-6 Astra model. However, the release did not proceed as planned. Shortly after its publication, the announcement was swiftly removed, and the webpage remained inaccessible for a considerable duration. The official explanation pointed to a content management system failure and an internet outage, stressing that the withdrawal had no connection to the benchmark test scores. Nevertheless, following the reposting of the blog, numerous evaluation metrics underwent several rounds of modifications. Initially, Astra's hallucination rate was reduced from 4.2% to 2%, only to be subsequently revised back to 4.2%. Additionally, the hallucination rate of GPT-5.6 Sol and its internal cybersecurity evaluation scores also experienced changes. Concurrently, certain scores for models under Anthropic exhibited similar fluctuations, decreasing and then increasing. Some data adjustments even took place prior to the initial publication of the blog post. OpenAI claimed that these adjustments were made to ensure the data most accurately reflected the model's available performance, while also acknowledging that factors such as testing conditions could affect the results. However, this explanation has sparked doubts among AI experts regarding whether OpenAI engaged in 'benchmark manipulation.' This incident once again brings to the forefront the long-standing controversy over the accuracy of benchmark tests in the AI industry, with companies like Meta having encountered similar controversies in the past. Industry insiders are advocating for standardized practices that clearly define changes in evaluation conditions. Benchmark test scores are of paramount importance for AI companies. Yet, data fluctuations and external skepticism may pose challenges for clients and investors in assessing model suitability and could also tarnish OpenAI's market reputation.
