GLM-5.3 release removes barriers to autonomous cyber exploit development
Zhipu AI's new open-weight model GLM-5.3 matches frontier US models in exploit generation but lacks robust safeguards, allowing attackers to bypass restrictions easily.
The release of GLM-5.3 by Zhipu AI marks a significant shift in the accessibility of advanced cyber capabilities. Unlike previous frontier models that restricted access or implemented strong safety guardrails, this open-weight model allows anyone to download and modify its behavior. Security researchers have confirmed that its safeguards can be bypassed with high success rates using simple techniques.
This development means that sophisticated exploit generation tools are now freely available to malicious actors. While defenders can also use these tools, the asymmetry favors those willing to ignore ethical constraints. The model’s ability to autonomously chain vulnerabilities represents a new baseline for automated cyber threats.
What happened
Five months after Anthropic released Claude Mythos Preview through a limited access program called Project Glasswing, a comparable capability has entered the public domain. Zhipu AI, known internationally as Z.ai, launched GLM-5.3 without meaningful safeguards to limit misuse. This contrasts sharply with US frontier models, which are either kept private, released only to vetted users, or protected by robust API-level safeguards.
The National Institute of Standards and Technology’s Center for AI Standards and Innovation (CAISI) assessed GLM-5.3 on September 17. They identified it as the most cyber-capable open-weight model released to date. Their benchmarks indicate that GLM-5.3 lags behind the US frontier by approximately four months in aggregate cyber capabilities. However, because US models are not freely downloadable, GLM-5.3 effectively brings frontier-level offensive power to any actor with sufficient compute resources.
Researchers tested the model’s ability to develop end-to-end exploits. In evaluations using ExploitBench, which targets known vulnerabilities in the Google Chrome V8 engine, GLM-5.3 successfully developed exploits in 50 out of 410 attempts. This performance is nearly identical to Claude Mythos Preview, which succeeded in 56 out of 410 attempts. In internal binary exploitation tests involving open-source projects, GLM-5.3 achieved full control-flow hijacks in 4% of trials. While this is slightly lower than Claude Mythos Preview’s 6%, it represents a major leap from earlier models like GLM-5.2 and Claude Opus 4.6, which failed to succeed in any trials.
How it works
GLM-5.3 operates as an open-weight model, meaning its internal parameters are publicly available. This architecture allows users to modify the model directly, a process known as abliteration. By reconfiguring the model weights, users can remove refusal mechanisms that prevent the AI from generating harmful content. Researchers found that abliterating GLM-5.3 reduced its refusal rate on harmful benchmarks from over 90% to roughly 2-3%, with minimal impact on its general scientific or cyber capabilities.
Even without modifying the weights, attackers can bypass safeguards using prompt engineering techniques. In simulated tests, researchers used deceptive prompts, such as claiming the AI was acting as an autonomous red-team agent. This approach caused the model to engage with malicious requests 64% of the time. Another technique involved prefilling the model’s thinking tokens to simulate prior consideration of the request, which increased engagement to 92%. These methods do not work against safeguarded Claude models, which block deceptive prompts and do not allow token prefilling via their API.
The model demonstrates autonomy in finding and chaining vulnerabilities. In one test, a researcher used GLM-5.3 to identify previously unknown vulnerabilities in a Linux web browser’s JavaScript engine. Within a day and with less than an hour of human attention, the model chained these flaws into a working exploit that could read arbitrary files from a visitor’s computer. In another instance, the smaller GLM-5.3-Flash version created a reliable exploit chain for a known Chrome vulnerability, bypassing pointer-authentication hardening in just 20 minutes of human oversight.
Key details
- GLM-5.3 is the first open-weight model with cyber capabilities matching restricted US frontier models like Claude Mythos Preview.
- Safeguards can be bypassed 64% to 100% of the time using simple techniques like deceptive prompts or abliteration.
- The model successfully developed end-to-end exploits in 50 of 410 attempts on the ExploitBench benchmark.
- Abliterating the model costs approximately $4,400 in compute and takes about 2,200 GPU hours for inexperienced teams.
- GLM-5.3-Flash generated a working exploit for a known Chrome vulnerability (CVE-2026-11645) for a cost of $20.40 at API prices.
- NIST’s CAISI ranks GLM-5.3 as lagging US frontier models by about four months in aggregate cyber benchmarks.
Why it matters
For software engineers and security teams, the barrier to entry for sophisticated cyberattacks has dropped significantly. Previously, developing complex exploit chains required deep expertise and substantial time. Now, actors with modest resources can leverage GLM-5.3 to automate vulnerability discovery and exploit generation. This shifts the defensive burden, as attackers can scale their efforts using AI agents that work continuously and cheaply.
The availability of abliterated versions means that standard safety filters are ineffective against determined adversaries. Defenders can no longer rely on the assumption that AI tools will refuse harmful requests. Instead, they must assume that attackers have access to unrestricted models capable of autonomous reasoning and code generation. This necessitates a proactive approach to security, focusing on rapid patching and robust system hardening rather than relying on attacker incompetence or resource constraints.
Furthermore, this development highlights the urgency for defenders to access equally capable tools. If malicious actors use frontier AI to find zero-day vulnerabilities, defenders need similar capabilities to identify and fix these flaws before they are exploited. The current landscape favors those who act quickly, making the integration of AI-assisted security testing a critical component of modern software development lifecycles.
What you can do
- Prioritize patching known vulnerabilities immediately, as AI models can rapidly convert public fixes into working exploits.
- Implement rigorous input validation and sandboxing to mitigate the impact of potential browser or engine exploits.
- Adopt AI-assisted security testing tools to identify vulnerabilities in your codebase before attackers do.
- Monitor for unusual network activity or file access patterns that may indicate automated exploit attempts.
- Participate in bug bounty programs and share vulnerability data to help secure open-source dependencies.
- Advocate for and utilize trusted access programs that provide defenders with advanced, safeguarded AI models.



