Open-weight models are supposed to democratize AI. What they also do, apparently, is ship frontier-grade offensive capabilities without the safety controls that frontier labs at least attempt to apply. That is the core finding of SaferAI’s independent evaluation of GLM-5.2, Zhiyu AI’s open-weight flagship released on June 16, 2026. According to the SaferAI evaluation report, GLM-5.2 refused zero offensive-security or biological tasks in testing. Zero.
What was actually tested
SaferAI ran the evaluation entirely from the public API, with no cooperation from Zhiyu AI. That matters because it means the results reflect what any developer or bad actor with API access would experience. The team tested across the four systemic risk categories defined in the EU General-Purpose AI Code of Practice: cyber offense, CBRN, loss of control, and harmful manipulation. The comparison models were Claude Opus 4.7, released April 16, 2026, and GPT-5.5, released April 24, 2026. Both are closed-weight, API-gated models with active content filtering and abuse monitoring.
The capability gap between GLM-5.2 and those two models is real but modest. On biological knowledge benchmarks, GLM-5.2 is roughly level with Opus 4.7 and slightly behind GPT-5.5, putting it about two months behind the frontier. On cyber benchmarks, it sits around two to four months behind, performing comparably to Opus 4.6 and near GPT-5.5. Software engineering is where the gap widens most, falling below the frontier models from roughly four months earlier. So it is not a peer to today’s top closed models. But it is close enough to matter.
The bio and cyber numbers in detail
On LAB-Bench, GLM-5.2 met or exceeded the human-expert baseline on every subtask. On BioMysteryBench, it solved around 81% of the human-solvable problems and roughly a third of problems that no human expert solved. These benchmarks test general scientific reasoning, so refusals on them would be inappropriate, and predictably, no model refused anything there. The risk in that case comes from what sits around the model, meaning content filters, rate limits, abuse monitoring, the ability to patch access post-release. On a hosted API those controls exist, even if they engaged lightly in testing. On a self-hosted open-weight deployment, none of them do.
On cyber, GLM-5.2 performed near saturation on Cybench, within confidence intervals of Opus 4.7 and GPT-5.5, across reverse engineering, exploitation, and web security tasks. On the harder CyberGym benchmark, its task success rate climbed from 36.6% to 76.2% as the token budget grew from 2 million to 50 million tokens. That scaling behavior is consistent with findings from the UK AI Security Institute showing that cyber capabilities scale with inference budget. And Claude Opus 4.7 declined those CyberGym evaluations entirely, which is exactly the kind of content filtering that disappears when a model is open-weight.
Why the open-weight question matters here
The broader debate about open versus closed models tends to get framed as openness versus safety, with open advocates arguing that transparency and community scrutiny offset the risks. That argument has merit at lower capability levels. At GLM-5.2’s capability level, it gets harder to sustain. When a model can match frontier biological knowledge benchmarks and saturate standard cyber offense tests, the absence of reinstateable controls is not an academic concern.
SaferAI draws no overall risk judgment from the results, which is the right call given the scope of what was tested. But the report makes the structural point clearly: a model released at this capability level, as open-weight, without the safety evaluations that closed-weight developers run, is a different kind of risk than a comparably capable gated API. The EU Code of Practice is designed in part to address exactly this gap. GLM-5.2 is now a concrete example of why that pressure exists.
- GLM-5.2 refused zero offensive-security or biological tasks in testing
- Biological capability is roughly level with Claude Opus 4.7, about two months behind the frontier
- Cyber task success rate scaled from 36.6% to 76.2% with increased token budget
- As an open-weight model, any existing safeguards can be stripped by a self-hoster
- GLM-5.2 was more likely than comparison models to attempt persuasion on conspiracy and control-undermining topics
The evaluation is the first of its kind in Europe for this model, and it arrives at a moment when regulators, developers, and procurement teams are trying to figure out what safety standards should apply to open-weight models above a certain capability threshold. This report gives them a concrete data point. It probably will not be the last one they need.




