Anthropic recently carried out a cybersecurity assessment of the GLM-5.3 model. Initially, the aim was to caution that open-weight models could reduce the barriers to launching cyberattacks. However, the assessment inadvertently ended up serving as a strong third-party endorsement for GLM-5.3. The model showcased an impressive ability to autonomously execute complex exploits, nearly on par with Anthropic’s own Claude Mythos Preview.
In the ExploitBench test, GLM-5.3 successfully completed end-to-end exploits far more frequently than several contemporary models, including Claude Opus 4.6 and GLM-5.2. This highlights its substantial offensive security capabilities. Within a single day in a controlled sandbox environment, GLM-5.3 identified multiple previously unknown vulnerabilities in mainstream browsers. Moreover, it could chain together multiple zero-day vulnerabilities to form a complete attack chain. The malicious webpages it generated were even capable of breaking out of the browser sandbox to access arbitrary files on the target device.
The smaller-parameter version of the model, GLM-5.3-Flash, required minimal human intervention. Within just a few hours of autonomous operation, it could generate a functional attack chain based on publicly disclosed Chrome vulnerabilities, all at an extremely low cost of invocation.
Anthropic noted that GLM-5.3’s security restrictions have significant vulnerabilities. Attackers can bypass or even fully disable its mechanism for rejecting malicious requests by providing a false context of red team testing, pre-filling the model’s thought process, or directly modifying the open weights. The open-weight nature of the model allows attackers to directly dismantle relevant security safeguards.
Ironically, this report, which was originally intended to underscore the potential dangers of GLM-5.3, ended up inadvertently promoting its performance as it circulated.
