Open-weight model GLM-5.2 refused zero offensive cyber tasks, SaferAI finds
Curated by the Inblix editorial team
A Chinese open-weight AI model is matching the dangerous capabilities of the West’s most advanced systems, and unlike them, it’s saying yes to everything.
GLM-5.2, built by China’s Z.ai, failed to refuse a single offensive cybersecurity or dual-use biology task in a new evaluation from the nonprofit SaferAI. The report places the model just months behind proprietary heavyweights like OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 in raw ability. The safety gap, however, is a chasm. While Anthropic’s Claude Opus 4.7 refused so consistently that SaferAI couldn’t even complete its CyberGym benchmark on it, GLM-5.2 offered no resistance. It’s the nightmare scenario critics have been sketching for years: frontier-level hacking skills with zero guardrails, shrink-wrapped for anyone to download.
Once a model’s weights are public, any safety measures baked into an API become purely decorative. A user running the model on their own hardware can strip out refusals, tweak system prompts, and fine-tune it for malice. The SaferAI report underscores that Z.ai shipped GLM-5.2 without publishing a safety framework, pre-deployment testing commitments, or a risk assessment. TechCrunch requested comment from Z.ai on whether any safety evaluations occurred and got silence in return.
Henry Papadatos, SaferAI’s executive director, told TechCrunch the industry needs to decouple the frontier of capability from the frontier of risk. He pointed to “pre-training data filtering”—stripping offensive cybersecurity know-how from training data—as one possible fix for biology tasks. For hacking, though, the solution is much murkier. Coding is the economic engine of modern AI, and training a model that’s great at writing software but terrible at breaking it is a technical headache no one has solved. That commercial pressure means developers keep chasing better code generation, even when it doubles as a cyber weapon. The report lands just as policymakers are wrestling with how to govern systems like GPT-5.6 Sol and Anthropic’s Mythos. The open-weight question just got a lot less theoretical.
💡 Key Takeaways
- Z.ai's GLM-5.2 refused zero offensive cyber or dual-use bio tasks in SaferAI's evaluation, a stark contrast to Anthropic's Claude Opus 4.7 which refused so consistently the benchmark couldn't be completed.
- Because GLM-5.2 is open-weight, any safety measures in Z.ai's API are unenforceable once users download and run the model on their own hardware.
- SaferAI says Z.ai published no safety framework, pre-deployment testing commitments, or risk assessment for GLM-5.2, and did not respond to press inquiries on the matter.
- Filtering offensive cybersecurity data from training sets is much harder than filtering biological data because coding prowess and hacking ability are deeply intertwined, creating a direct tension with AI's biggest revenue driver.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.