On September 29, Anthropic, the creator of Claude, unveiled a report highlighting that Zhipu GLM-5.3 has exhibited formidable capabilities in exploiting vulnerabilities. However, it also points out notable deficiencies in its security safeguards, which could potentially make it easier for malicious actors to launch attacks. During the ExploitBench evaluation, GLM-5.3 successfully crafted complete end-to-end programs for exploiting vulnerabilities in 50 out of 410 tries, attaining a success rate of roughly 12%—a figure nearing the 14% success rate achieved by Claude Mythos Preview. Furthermore, the security constraints of GLM-5.3 can be circumvented through relatively straightforward techniques. For example, by concocting justifications like 'red team exercises,' attackers achieved a 64% success rate. This rate surged to 92% when pre-filled reasoning content was utilized, and even reached a perfect 100% after directly tweaking the model weights to eliminate the refusal mechanism.
