As policymakers grapple with how to govern increasingly powerful AI systems such as OpenAI’s GPT-5.6 Sol and Anthropic’s Mythos, a Chinese open-weight model has quietly closed the gap with the industry’s frontier leaders—raising urgent questions about safety and regulation.
The New Frontier: Open-Weight Models Rise
GLM-5.2, the open-weight AI model developed by China’s Z.ai, now trails OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 by only a few months on cyber and bio capabilities, according to a new report from AI safety nonprofit SaferAI. While the capability gap narrows, the divide between what these models can do and how safely they are deployed is expanding—a trend that safety experts find deeply concerning, especially as 2026 brings more open-weight releases and regulatory debates intensify.
Safety Evaluations Reveal a Stark Disparity
SaferAI’s evaluation, conducted via Z.ai’s public API, found that GLM-5.2 refused none of the offensive cyber or dual-use biology tasks it was given. By contrast, Claude Opus 4.7 “refused so consistently that SaferAI could not complete CyberGym on it at all.” CyberGym is a benchmark for evaluating cybersecurity capabilities; OpenAI used it in the assessment preceding last month’s Hugging Face breach.
This disparity underscores a warning critics have sounded for years: open-weight models could place highly capable AI in the hands of malicious actors, with no enforceable mechanism to police usage once weights are downloaded. As open-weight systems near the capabilities of proprietary frontier models, the debate has shifted from whether they can compete to how society manages the risks of their widespread release.
“The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly,” said Henry Papadatos, executive director of SaferAI, in an interview with TechCrunch.
The Enforcement Problem: Safeguards That Don’t Travel
While Z.ai could implement safety measures on its hosted API, those protections become unenforceable once someone runs the model on their own hardware. Users can strip safeguards, fine-tune weights, or alter system prompts with ease. In contrast, frontier developers like OpenAI and Anthropic rely on classifiers, refusal training, and API-level controls to mitigate dangerous cyber and bio assistance—yet even these are far from foolproof.
Jailbreaks routinely bypass protections on deployed models. Far.ai, an AI safety nonprofit, identified hundreds of universal jailbreaks—reusable keys that succeed on most harmful requests—in frontier models like xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro. According to Far.ai, jailbreaks succeed when attackers combine manipulation techniques such as roleplaying, authority impersonation, fake conversation history, and follow-up prompts to exploit weak points in a model’s defenses.
However, the safeguards that work for closed models do nothing for open-weight ones, which are designed to run on any infrastructure with any set of protections—or none at all. This reality demands a new approach.
Toward Safe Open-Weight AI
Papadatos advocates for a nuanced strategy: “The objective should clearly be that the good capabilities—the safe ones—are accessible to anyone, and then we try to remove the bad ones, even in an open source fashion.”
One technique he highlighted is pre-training data filtering, which could reduce a model’s exposure to harmful knowledge before it is ever released. Combined with targeted post-training alignment and community-driven red-teaming, such measures could help open-weight models deliver benefits without creating unacceptable risks.
As 2026 unfolds and open-weight models continue their rapid ascent, the AI community faces a defining challenge: how to balance openness and innovation with safety and accountability. The gap between capability and safety is not inevitable—but closing it will require deliberate design, robust evaluation, and global cooperation.
via TechCrunch AI
