GPT-5 Testing Under Scrutiny: Accusations of Cheating and Strategically Avoiding Hard Questions
2025-08-12 / Read about 0 minute
Author:小编   

At the GPT-5 launch event, OpenAI found itself at the center of controversy due to a chart that was disproportionately scaled. Later revelations revealed that in the SWE-bench Verified test, OpenAI completed only 477 out of 500 questions yet secured a high score of 74.9%. In contrast, Anthropic's Claude Opus 4.1 achieved a score of 74.5% on the full set of 500 questions. SemiAnalysis highlighted that the 23 questions skipped by OpenAI could potentially undermine the fairness of the results. Adding to the skepticism, the SWE-bench Verified test set itself was designed by OpenAI, raising concerns about potential bias in the testing methodology. Moreover, in the IOI 2025 competition, OpenAI's internal model achieved impressive results, but this was not the publicly available version, further fuelling discussions about testing standards and marketing tactics. These revelations have intensified public doubts regarding the transparency and fairness of OpenAI's testing practices.

  • C114 Communication Network
  • Communication Home
7 X 24 Track global technological trends
Hot Topic