Z.ai opened API access to GLM-5.3 on Tuesday, and the Chinese startup did something unusual for a new flagship: it did not raise the price. Developers pay the same $1.40 for a million input tokens and $4.40 for a million output tokens that GLM-5.2 charged. Cached input runs $0.26 per million tokens, and Z.ai is temporarily waiving the storage fee on that cache.

The API rollout arrives one day after AI Insiders covered GLM-5.3’s launch, when the model topped a cybersecurity benchmark and Z.ai reaffirmed plans to open its weights. That story was about capability. This one is about what capability actually costs to run, and the two numbers do not agree.

Artificial Analysis, which tracks model cost against benchmark performance, scores GLM-5.3 at 60 on its Intelligence Index. That ties Kimi K3 as the top-scoring open-weights model and lands seven points above GLM-5.2’s mark. On its own, a higher score at an unchanged price would read as a straightforward upgrade.

It is not, because Artificial Analysis also measures what a finished task costs, not just what a token costs. The firm puts a GLM-5.3 task at roughly $0.68 on that index, up from about $0.44 for GLM-5.2, a difference of nearly 55 percent even though the per-token prices did not move. The gap traces to verbosity: Artificial Analysis found GLM-5.3 writes longer responses than its predecessor, so a task that once consumed a given number of output tokens now consumes more of them. A flat rate card does not produce a flat bill when the model itself uses more tokens to say the same thing.

That distinction will get lost in most coverage of this release, which will lead with the unchanged $1.40 and $4.40 figures and stop there. Anyone comparing GLM-5.3 to GLM-5.2 on a cost basis needs the per-task number, not the per-token one.

Against frontier rivals, GLM-5.3 still prices as a bargain. Combine one million input tokens with one million output tokens and the model totals $5.80. Kimi K3 asks $18.00 for the same mix, Claude Opus 5 asks $30.00, and GPT-5.6 Sol asks $35.00. Cheaper options exist too: GPT-5.6 Luna runs $0.20 for input and $1.20 for output at OpenAI, and Gemini 3.7 Flash runs $0.75 and $3.75 at Google through the end of 2026, when that rate is set to roughly double.

Access is still partial. Developers on Z.ai’s GLM Coding Plan can currently reach GLM-5.3 only through an OpenAI Chat Completions-compatible endpoint. Z.ai has said it plans to release the model’s weights openly, but it has not set a date or named a license.

Teams weighing a move from GLM-5.2 to GLM-5.3 should benchmark cost on completed tasks, not on the rate card. A workload heavy on long completions will land closer to the 55 percent premium Artificial Analysis measured than to the price parity Z.ai is advertising.

Reporting from VentureBeat (Carl Franzen, August 18, 2026).