Z.ai released GLM-5.3 on 14 August 2026, and the company says every capability gain over its predecessor came from post-training alone: the base model is identical to GLM-5.2. Z.ai describes the approach directly: “Scaling post-training is all we did for GLM-5.3.”

Z.ai says GLM-5.3 improved 50 percent over its predecessor on the company’s internal Z.ai Code Bench. It also claims the open-weights lead on Agents’ Last Exam and, separately, on Terminal Bench 3.0. Those numbers come from Z.ai’s own tables; the release does not include independent benchmark results.

The more consequential figure sits in a different column. On CyberGym, a benchmark for vulnerability discovery, GLM-5.3 posts 84.5, edging out Mythos 5’s 83.8 and GPT-5.6 Sol’s 83.6. On ExploitBench, which measures turning a found flaw into a working exploit, GLM-5.3 reaches 54.4, more than double GLM-5.2’s 24.4, though it still trails Mythos 5’s 78.0 by a wide margin.

Z.ai’s own read on that gap is the sharpest framing available: “The pattern across the three is consistent: the further up the exploitation chain a benchmark sits, the larger the gain from GLM-5.2 and also the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where we are furthest behind.” Elsewhere in the post, the company adds that the trajectory outran its own forecasts: “As we scaled post-training, cyber capability developed faster than we expected.”

Z.ai backs the benchmark claim with field results. Working with security teams in China against real codebases, the model surfaced 2,436 vulnerabilities across 269 projects once specialists had reviewed, filtered and deduplicated the results, 1,097 of them rated medium to high severity. The company has published a Z.ai Security Disclosure Ledger tracking the findings: 53 disclosed publicly, 2,383 still under embargo. Severity breaks down to 107 critical, 990 high, 1,286 medium, and 53 low. The oldest flaw traces back to 1981, and the average vulnerability sat undiscovered for 26.6 years.

These are open weights. Z.ai says it will release them publicly two weeks after launch, once safety evaluation and hardening are complete: “We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.” A model that leads a public vulnerability-discovery benchmark, and more than doubled its own exploit-generation score, is going out to anyone who can download it, on a two-week clock Z.ai set itself.

On coding efficiency, GLM-5.3 at High reasoning effort reaches 31.4 percent on Z.ai Code Bench using roughly 50,000 output tokens, ahead of Claude Opus 4.8’s 29.5 percent at 120,000 tokens by Z.ai’s own account. The company is candid about where it still lags: GLM-5.3 “remains behind Claude Fable 5, which reaches 39.5% at Max effort.” One operational note for developers: GLM-5.3’s API no longer supports disabling the reasoning step entirely, so integrations that set thinking.type to “disabled” need to update that call or their requests will fail.

Security teams evaluating open-weights coding models should treat the CyberGym and ExploitBench numbers as the headline, not the footnote, and budget review time before the two-week release window closes.

Z.ai published these figures and quotes in its own GLM-5.3 release post, dated 14 August 2026.