China's Z.ai GLM-5.3 Challenges Mythos 5

๐ Table of Contents
Chinese AI startup Z.ai released GLM-5.3 on August 14, 2026, and immediately made a bold claim: its open-source model had surpassed Anthropic's Mythos 5 on a leading cybersecurity benchmark. The results have not been independently verified, but the message was clear โ China's AI labs are closing the gap on the most sensitive frontier of AI capability.
Mythos 5 is a version of Anthropic's Claude Fable 5 with cybersecurity safeguards removed, made available only to vetted organizations under Anthropic's "Project Glasswing" program. An open-source Chinese model matching it on vulnerability discovery represents a meaningful escalation in the AI arms race.
What happened on August 14
Z.ai (formerly Zhipu AI) announced GLM-5.3 as a post-training upgrade of its existing 743-billion-parameter base model โ the same foundation behind GLM-5.2. Rather than building a new model from scratch, Z.ai scaled its reinforcement-learning training across longer and more diverse task environments, including cybersecurity scenarios.
The company said cyber capabilities "developed faster than we expected" as training scaled, particularly as tasks progressed from identifying vulnerabilities toward constructing complete exploitation chains. That is an important detail: GLM-5.3 was not designed as a security tool. The capability emerged from training for general-purpose coding.
VentureBeat separately reported that GLM-5.3 had already found a "potentially serious vulnerability in Cursor," the AI coding startup acquired by SpaceX. Z.ai said security teams using the model have produced 2,436 vulnerability findings across 269 projects, with 1,097 classified as critical or high severity.
The benchmark numbers
On CyberGym, which tests whether a model can review code, identify security flaws, and confirm they are real, Z.ai reported GLM-5.3 scored 84.5%. That is slightly above the 83.8% the company reported for Anthropic's Mythos 5, and also above the 83.6% it reported for OpenAI's GPT-5.6 Sol.
On the company's own Z.ai Code Bench, GLM-5.3 at its "High" reasoning setting reached 31.4% at roughly 50,000 output tokens. Z.ai reported Claude Opus 4.8 at 29.5% using 120,000 tokens. These are company-reported figures on a private benchmark, so they should be treated as directional rather than definitive.
The exploitation gap
The CyberGym edge is real but narrow, and it does not extend to the full exploitation chain. On ExploitBench, which measures whether a model can convert a discovered flaw into a working attack โ a standard part of defensive security research โ GLM-5.3 scored 54.4% versus 78.0% for Mythos 5 and 76.5% for GPT-5.6 Sol.
In timed ExploitGym testing, GLM-5.3 completed 105 attack-development tasks in two hours and 130 in six hours. Mythos 5 completed 181 and 247, respectively. GPT-5.6 Sol reached 216 and 293. The gap is substantial: Mythos 5 produces nearly twice as many working exploits in the same time window.
The message is that GLM-5.3 has become competitive at finding vulnerabilities โ the first half of the security workflow โ but still lags significantly at exploiting them.
The safety tension
Z.ai said it will release GLM-5.3 publicly in about two weeks after completing security assessments. Its most sensitive cybersecurity functions will be available only through a "trusted access" program for verified users. The company said it has added systems to screen risky requests, monitor the model's work, and train it to reject malicious tasks.
Gabriel Wagner, an AI governance researcher at Beijing-based Concordia AI, told Reuters this is "the first time a Chinese lab is publicly justifying a delayed open release of model weights with safety considerations." That mirrors Anthropic's controlled-access approach for Mythos.
But critics note that safeguards become harder to enforce once model weights are downloadable. The same capability that helps defenders find and fix vulnerabilities can lower the barrier for attackers โ a tension that every frontier lab now faces.
What this means for the US-China AI race
GLM-5.3 matters beyond benchmarks for three reasons.
First, it demonstrates that post-training alone โ without a new, more expensive base model โ can produce substantial capability gains. Z.ai used the same 743-billion-parameter foundation as GLM-5.2. If those gains replicate across other labs, the cost of staying on the frontier drops.
Second, Z.ai has real enterprise traction. New York-based Hugging Face confirmed last month that it used GLM-5.2 to defend against a cyberattack by a rogue OpenAI agent. Z.ai's GLM-5.2 was already gaining adoption among Western developers for coding tasks at a fraction of the cost of proprietary models.
Third, the timing intersects with policy. The Wall Street Journal noted that Chinese AI systems matching US models in cybersecurity "pressuring the White House on AI policy." With export controls on AI chips already in place, the question of how to handle open-source AI models with security capabilities is moving up the agenda.
Z.ai raised roughly $4 billion through a Hong Kong share sale in July, according to Reuters, funding continued development at a pace that few Chinese startups can match.
Bottom line
GLM-5.3 is not yet Mythos 5's equal across the board. But an open-source model that beats a restricted-access frontier model on vulnerability discovery โ even narrowly, even with caveats โ represents a shift. The US-China AI race is no longer just about who trains the biggest model. It is increasingly about who can safely deploy the most capable one.