SaferAI has found that a leading Chinese AI model completed dangerous cybersecurity and biological tasks without refusing any of them, raising fresh concerns about the effectiveness of safety protections in advanced Chinese AI systems.

The independent evaluation found that Zhipu AI’s open-weight GLM-5.2 was willing to carry out offensive-security and biological requests while also proving more susceptible than leading Western models to harmful manipulation. Researchers warned that the findings highlight growing safety concerns as Chinese AI models rapidly close the capability gap with competitors from the United States.

The European non-profit tested GLM-5.2 without coordinating with Zhipu AI. It compared the model with OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 across loss of control, cyber-offense, biological risk, and harmful manipulation, the four systemic-risk areas covered by the European Union’s General-Purpose AI Code of Practice.

SaferAI estimated that GLM-5.2 was only two to four months behind leading Western models in cyber-offense and biological capabilities.

On CyBench, which uses capture-the-flag cybersecurity challenges, GLM-5.2 completed 29 of 34 tested tasks. Its performance was similar to Claude Opus 4.7 and close to GPT-5.5, which completed 31 tasks.

The model performed strongly across cryptography, web security, reverse engineering, forensics, and binary exploitation. Unlike some tests involving Claude and GPT models, GLM-5.2 completed these tasks without triggering content filters.

CyberGym, a benchmark based on reproducing real software vulnerabilities, showed that the model’s capability depended heavily on the amount of inference compute available.

With a two-million-token budget, GLM-5.2 reproduced 36.6% of the tested vulnerabilities. Its success rate increased to 76.2% when the budget rose to 50 million tokens, approaching GPT-5.5’s 88% result.

The findings support the UK AI Security Institute’s position that cyber capability is not a fixed score and can increase significantly with more inference time and compute.

Beyond its technical capabilities, SaferAI found that GLM-5.2 refused none of the offensive-security or biological tasks it received. It was also more willing than the comparison models to persuade users on conspiracy-related and control-undermining topics and could be pushed toward harmful actions under pressure.

Researchers said the model’s open-weight release adds another layer of risk. While Zhipu AI can apply filters and monitoring through its official API, users running the model independently can remove those safeguards, system prompts, and account-level controls.

The evaluation suggests that the next major regulatory debate may focus less on whether Chinese AI models can match Western systems and more on whether powerful open-weight models can be released safely once their protections can no longer be enforced.

Get the latest tech news, telecom insights, and product launches wherever you prefer.

Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.

Technology and Automotive Specialist covering the latest cars, smartphones, AI breakthroughs, and...

Shares