OpenAI shipped GPT-6 Astra, its newest flagship model, and Zvi Mowshowitz spent the following days assembling reactions to it from researchers, developers and OpenAI’s own staff for his newsletter, Don’t Worry About the Vase. His own conclusion: Astra is likely the most capable model released to date for what he calls “ambitious projects,” excelling at 3D generation, game creation, computer use and coordinating multiple subagents on long tasks. He is far more measured on the parts of the job that matter for daily work. Regular coding, in his assessment, is “not a quantum leap” over Sol, OpenAI’s prior flagship, and for back-and-forth discussion and editing he still reaches for Fable 5.1, Anthropic’s current top model.

Mowshowitz frames the debate over whether Astra qualifies as AGI as the first one he has found worth taking seriously, though he personally does not think it clears that bar. He also flags a claim he attributes to OpenAI itself: in his words, the company has already given a partial, “soft announced” acknowledgment that a more capable model exists internally, well short of any formal statement of its abilities or a release date.

The benchmark picture splits along a line familiar to anyone who tracks self-reported AI results. Sam Altman, OpenAI’s chief executive, put out launch-day figures crediting Astra with a perfect 100% on ExploitBench, alongside a 98% mark on FrontierMath Tier 4 and 99.9% on ARC-AGI 3, all figures the company reported itself. Mowshowitz treats the ExploitBench number as a warning sign rather than a selling point: a perfect score, in his view, points to data contamination that OpenAI’s own system card acknowledges. Francois Chollet, ARC-AGI’s creator and an independent party to the release, offers a more grounded figure: Astra scores 66% on the benchmark’s standard harness, and close to 100% only with a custom, continuous conversation harness that costs roughly $360 per game to run. Chollet calls it a major breakthrough in interactive reasoning without endorsing the AGI label himself.

A second, more technical finding troubles Mowshowitz more than any leaderboard. Researchers including Neel Nanda found that Astra scores 159 on Epoch’s Capabilities Index with its reasoning disabled, only 10 points below its 169 score with full chain-of-thought visible. Fable 5.1 drops 35 points under the same test. A separate study found Astra’s results on serial reasoning tasks jump sharply once meaningless filler tokens replace visible reasoning in the prompt, evidence the model can carry out cognition it never writes down. That combination matters because chain-of-thought monitoring, the tool alignment researchers rely on to check what a model is actually doing, depends on the model narrating its own reasoning. Astra appears to need that narration less than any model before it.

On price, Astra and Fable 5.1 both list headline rates of $10 per million input tokens and $50 per million output tokens, but Fable’s cached input costs $0.25 per million against Astra’s $1, a gap that compounds across long agentic runs. Mowshowitz’s own read on user preference comes from a poll of his readers: a swing of roughly 15 to 16 percent moved from Anthropic to OpenAI as a primary daily model, concentrated among casual users rather than the power users who still favor Fable.

For teams renewing model contracts this quarter, the practical takeaway is not to pick a winner. Mowshowitz’s own recommendation, repeated through the piece, is to run hard problems through both models and compare notes, since Astra and Fable 5.1 now diverge more on style and reliability than on raw capability.

Zvi Mowshowitz published this analysis on Don’t Worry About the Vase (Substack) on September 12, 2026.