Anthropic: GLM-5.3 Safeguards Easily Bypassed, Enabling Autonomous Cyber Exploits
Key point
Anthropic reports that Zhipu AI's GLM-5.3 model, while capable of autonomous cyber exploits comparable to Claude Mythos Preview, has safeguards that are easily bypassed or removed, significantly increasing risks to malicious actors.
Details
Anthropic's analysis of Zhipu AI's GLM-5.3 reveals that the model possesses strong capabilities for autonomously building end-to-end cyber exploits, similar to Anthropic's Claude Mythos Preview. However, unlike Claude models which are released with strict safeguards or limited access, GLM-5.3 is released as an open-weight model with safeguards that can be easily bypassed or removed. Anthropic found that attackers could bypass GLM-5.3’s safeguards between 64% and 100% of the time using simple techniques in simulated tests, whereas safeguarded Claude models refused all such requests.
Benchmark Performance
In ExploitBench tests targeting Chrome V8 vulnerabilities, GLM-5.3 developed end-to-end exploits in 50 of 410 attempts, compared to 56 of 410 for Claude Mythos Preview. In Anthropic's internal Binary Exploitation benchmark, GLM-5.3 achieved full control flow hijacks in 4% of tasks, versus 6% for Claude Mythos Preview. Previous models like Claude Opus 4.6 and GLM-5.2 scored 0% in these benchmarks, indicating a significant capability threshold has been crossed.
Real-World Attack Scenarios
In human-in-the-loop tests, researchers used GLM-5.3 to discover and chain multiple undisclosed vulnerabilities in a popular web browser's JavaScript engine within one day, creating a webpage capable of reading arbitrary files from a visitor's computer. Using the smaller GLM-5.3-Flash variant, researchers combined known vulnerabilities to build a reliable exploit chain for ARM64 targets that bypassed Pointer Authentication Code (PAC) hardening. This process required 20 minutes of human attention and 8 hours of model processing, costing $20.40 via the Zhipu API.
Safeguard Bypass and Removal
Although GLM-5.3 refuses harmful requests by default, these safeguards are easily circumvented. Anthropic demonstrated that 'abliteration' (a technique to remove refusal mechanisms) could be performed on the open-weight model for approximately $4,400 (2,200 GPU hours). This reduced refusal rates to about 3% on JailbreakBench, 2% on HarmBench, and 12% on StrongREJECT, while maintaining cyber capabilities. Even without abliteration, simple techniques like deceptive prompting or prefilling thinking tokens increased model participation in harmful tasks to 64% and 92% respectively, compared to 0% for Claude models which cannot be ablated or prefilled in this manner.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.