Anthropic published a security analysis on 29 September concluding that GLM-5.3, a model from Zhipu AI, known outside China as Z.ai, can build working cyberattacks and that its built-in refusals are easy to defeat. Anyone can download the model. That combination, Anthropic argues, is what makes it different from every other model at this level of skill.

Everything below comes from Anthropic’s own write-up, and Anthropic sells a rival product. The company says its tests ran in sandboxed environments against offline targets it set up itself, and it says its findings broadly match an assessment that the US government’s NIST Center for AI Standards and Innovation published on 17 September.

On ExploitBench, which measures how well a model can attack known flaws in the V8 engine inside Google Chrome, Anthropic reports that GLM-5.3 built a complete working exploit on 50 of 410 attempts. Its own Claude Mythos Preview managed 56. On a separate internal test built from open-source projects, GLM-5.3 fully hijacked a program’s control flow in 4 percent of trials against 6 percent for Mythos Preview. Anthropic says earlier models, including Claude Opus 4.6 and GLM-5.2, scored zero.

In one human-led session, a researcher gave GLM-5.3 a sandboxed Linux build of a popular browser. Over roughly a day, and with limited supervision, Anthropic says the model turned up several previously unknown bugs inside the JavaScript engine of that browser, then linked them into a webpage that reads files off a visitor’s machine. A second session used the smaller GLM-5.3-Flash on a known Chrome flaw. Anthropic says it produced a working attack chain for an ARM64 target after 20 minutes of human attention and eight hours of machine work, which would have cost $20.40 at Zhipu’s API prices.

The safeguard results are the sharper claim. Anthropic tested three ways around GLM-5.3’s refusals, using an overtly malicious request in a simulated environment. A role-play prompt, telling the model it was a red-team agent on an exercise, got it to engage 64 percent of the time. Prefilling its reasoning so it appeared to have already agreed got 92 percent. The third method, “abliteration,” is not a prompt trick on the shipped model. It edits the downloadable weights to strip out refusals, and the modified build engaged 100 percent of the time. Anthropic says its own team, with no prior experience, did the edit in about 2,200 GPU hours, roughly $4,400 in compute, and refusal rates on public benchmarks dropped from above 90 percent to between 2 and 12 percent depending on the test. General science scores did not change.

Anthropic says none of these tricks worked on its own Claude models. That is partly by design of the delivery method: Claude runs behind an API that gives users no way to prefill its reasoning, and its weights are not public, so nobody can abliterate it. The company also notes that several developers published abliterated GLM-5.3 versions within days of release. The comparison flatters a closed model, and independent testers have not yet replicated either side of it.

Anthropic draws two conclusions. First, it expects state and non-state attackers to use models like this one. Second, it wants defenders to get the strongest tools available, and says vetted defenders can already use Claude Mythos 5.1 through its trusted-access programs. It also calls on governments to run safety testing on capable models, including GLM-5.3’s successors.

The practical bind for security teams is timing. Attack-grade capability is now something a small team can run on rented GPUs, while the best defensive models remain gated behind vetting. Any organization that patches on a monthly cycle should expect the gap between a public fix and a working exploit to shrink toward hours.

Reported by Anthropic (research post) on 29 September 2026.