Nathan Lambert, who writes the AI research newsletter Interconnects, argues that the gap between Chinese and American frontier labs looks smaller than it is for a simple reason: release policy, not capability, sets the benchmark clock. His case centers on Z.ai, the Beijing lab formerly known as Zhipu AI, whose new GLM-5.3 model landed close to Western frontier scores on agentic coding tests this week.
Lambert’s central claim is about timing. He writes that Z.ai’s time to release is likely measured in days, while OpenAI and Anthropic typically hold models back for months of internal testing before public launch. If that gap is real, the leaderboard is not purely a capability ranking. It is also a measure of how long each lab is willing to sit on a finished model before shipping it.
That claim deserves more scrutiny than repetition. Lambert has no visibility into OpenAI’s or Anthropic’s unreleased internal models, so his assertion that those companies are “far better” internally than what the public sees is an inference, not a measurement. It rests on the reasonable but unverifiable assumption that months of pre-release testing correspond to hidden capability gains rather than safety work, product polish, or simple caution. Readers should treat it as an informed guess about systems nobody outside those labs can currently evaluate.
Lambert is more careful on the benchmaxxing question, and the nuance matters. He does not conclude that Z.ai is gaming its scores. His read is that Z.ai cares somewhat more about public benchmark placement than OpenAI or Anthropic do, partly because a strong showing on aggregators such as the Artificial Analysis Intelligence Index feeds directly into fundraising and morale. But he draws a line between that kind of attention and outright test-set optimization, and says GLM-5.3 does not look “fried” the way a heavily benchmaxxed model would.
The most underreported piece of Lambert’s argument is about the supply chain behind these models rather than the models themselves. He reports that the reinforcement learning data industry is expanding rapidly inside China, and that American data companies selling RL environments and labelling work to Chinese labs are a significant driver of that growth. In his account, Chinese labs can buy access to some of the same RL environments that train Western frontier systems, then release the resulting model faster than the environment’s original American customer does. Lambert is explicit that this is based on sources and industry rumor rather than confirmed transaction data, and he flags large error bars on how big or consequential that market actually is. He stops short of drawing a policy conclusion from it, and the piece does not call for export controls or restrictions on that data trade.
Lambert also credits Z.ai’s own history: the GLM line traces back to a Tsinghua University research group founded in 2019, giving the lab roughly the same multi-year runway that OpenAI and Anthropic have had. He frames the post-training approach behind GLM-5.3 as a real technical achievement rather than a shortcut, distinct from his skepticism about the release-timing comparison.
For operators benchmarking vendors, the practical takeaway is to treat public leaderboard rankings as a snapshot of what each lab has chosen to ship, not a full accounting of what each lab has built. Teams evaluating GLM-5.3 against a Western model still in limited release should ask what internal testing that Western model has yet to clear, since Lambert’s own argument implies the comparison may not be apples to apples for long.
Nathan Lambert, Interconnects, August 14, 2026.