The official programming ability test of GPT-5 has sparked widespread controversy due to the revelation that the SWE-bench Verified subset utilized in the test comprises only 477 questions, as opposed to the original 500. SWE-bench, a comprehensive metric for assessing a model's autonomous programming capabilities, originally featured 500 questions in its Verified subset. However, OpenAI has unilaterally omitted 23 questions from the test. Should these omitted questions be awarded a default score of zero, GPT-5's score would drop below that of Claude Opus 4.1, with a mere 0.4% separating the two. Notably, OpenAI had previously disregarded certain questions during the release of GPT-4.1, citing incompatibility with their infrastructure for running solutions. This action has once again fueled concerns regarding the potential manipulation of GPT-5's programming ability evaluation results.
