Alibaba’s newest flagship, Qwen3.8 Max, pulled even with Claude Opus 4.8 this week, tying at a score of 56 on the Intelligence Index that Artificial Analysis publishes. Kimi K3 still leads the group at 57, and it gets there for roughly a quarter less money per completed task than Qwen needs. That gap between an identical score and a very different bill is the part worth paying attention to.

The improvement itself is genuine. Qwen3.7 Max, Alibaba’s prior release, sat at 46 on the same index, so the new version gained ten points in a single generation and moved ahead of Zhipu AI’s GLM-5.2, which scored 51. On GDPval-AA, a separate benchmark built around realistic work tasks, Qwen3.8 Max climbed 468 Elo points to reach 1,739, enough to pass Kimi K3’s 1,685. Only Anthropic’s Claude Opus 5 scores higher on that measure, at 1,852.

Here is the catch. Artificial Analysis found that Qwen3.8 Max needed 64 reasoning steps to complete an average task, compared with 14 for the previous generation, and because each step resends the full conversation so far, total input volume ballooned roughly fifteenfold. Alibaba did lower its list prices: input now runs $2.00 per million tokens versus $2.50 before, output sits at $6.00 versus $7.50, and cached tokens cost $0.25 versus $0.50. None of that mattered once the step count exploded. A single completed task on the Intelligence Index now runs $1.14, more than double the $0.53 that Qwen3.7 Max cost. Kimi K3 finishes the same task for $0.86 and still scores a point higher, and GLM-5.2 comes in cheapest of the four at $0.57.

This is the number buyers should actually be watching. Per-token pricing tells a purchaser almost nothing about what a model costs to run once it starts taking dozens of extra steps to reach an answer, and a rate card that looks 20 percent cheaper can still leave a team paying double once the model’s own behavior multiplies the token volume behind the scenes. Cost per finished task is the metric that captures that reality, and it is also the metric almost no lab volunteers in its launch materials. Teams comparing frontier models on sticker price alone are reading the wrong column.

Qwen3.8 Max also lost ground on two accuracy-adjacent tests. AA-LCR, a test of long-document retrieval accuracy, slipped two points from the previous version. AA-Omniscience, which scores whether a model admits uncertainty rather than fabricating an answer, dropped ten points. Raw accuracy on that test held near 31 percent, but the hallucination rate climbed from 23 percent to 40 percent, meaning the model now guesses instead of declining to answer far more often than its predecessor did.

Teams evaluating Qwen3.8 Max for production workloads should benchmark on total task cost and step count, not the published per-million-token rate, before committing budget to it over Kimi K3 or GLM-5.2.

Reported by The Decoder on August 6, 2026.