Analysis Firm Charges Google and Meta with Benchmark Manipulation; Both Parties Stand Firm
7 hour ago / Read about 0 minute
Author:小编   

The analysis firm SemiAnalysis recently released an article revealing that Google's Gemini 3.8 Flash and Meta's Muse Spark 1.3 achieved outstanding scores on the Terminal-Bench 2.1 leaderboard, securing the 2nd and 4th positions, respectively. However, when evaluated on the Terminal-Bench 4.0 leaderboard, which integrates new anti-cheating measures and updated questions, the performance of these two models saw a marked decline. Gemini 3.8 Flash tumbled to the 12th position with a score of 19.1, while Muse Spark 1.3 scored 33.3 and also experienced a significant drop in ranking. In response to these allegations, Meta's Chief AI Officer refuted the claims, highlighting that GPT-5.6 Sol exhibited an even larger performance disparity yet was not accused of benchmark manipulation. The officer emphasized that Muse Spark 1.3 prioritizes cost-effectiveness. SemiAnalysis further uncovered sophisticated strategies for benchmark manipulation within the industry: major companies no longer train directly on the original questions but instead purchase training data that closely mimics publicly available benchmark questions. The related industrial supply chain is already well-established, with average quarterly contract prices ranging from $300,000 to $500,000, involving a multitude of companies. Additionally, Datacurve, the operator of the DeepSWE leaderboard, also sells relevant training data, sparking concerns about potential conflicts of interest. SemiAnalysis contends that public benchmarks will ultimately be "gamed through" and advocates for the development of high-quality private benchmarks. However, this approach could also lead to public leaderboards becoming mere marketing tools, depriving developers of a fair and transparent public standard for evaluating models.