A cheap proof checker, not raw intelligence, is why AI just cracked ten open problems in mathematics, according to Zvi Mowshowitz. OpenAI disclosed this week that an internal build of its next model, Astra, produced Lean-certified solutions to ten unresolved problems, among them a disproof of Connes’s rigidity conjecture and new bounds on sphere packing, for roughly $2,000 in inference cost. AI Insiders covered that announcement yesterday. Mowshowitz’s read of what it means is the more interesting question.
His framing: ordinary computers have been superhuman at arithmetic for decades without anyone calling it intelligence. Mowshowitz argues language models have now reached the same status for advanced mathematics, and were already ahead of typical human specialists in coding and offensive cybersecurity. The common thread, in his account, is not raw capability. It is that all three domains hand a machine a cheap, mechanical way to check whether an answer is correct.
That is the load-bearing step in his argument. A Lean certificate settles whether a proof holds in seconds, without a human reading a line of it. Mowshowitz’s claim is that a meaningful share of AI research and development shares that property: speed up a training run or shrink a model and you can measure the change directly. If verification transfers that cleanly, the capability that cracked these ten problems should also let a lab accelerate work on its own systems, not just external domains. For that inference to hold, the parts of AI research that actually gate progress, choosing what to try next, judging whether a result generalizes, would have to be as cheaply checkable as a Lean proof. That is the condition his argument needs and does not fully establish.
“Verifiable” spans an enormous range. A Lean proof is checked in seconds by a program with no judgment at all. Whether a research direction was worth pursuing, whether a training run’s gains generalize, whether an architecture change helps once scaled up, are also technically verifiable, but only slowly, expensively, and partly by taste. Mowshowitz concedes something close to this when he compares math to coding and cyber, writing that no formula can tell you whether code is good, “only that it passes its unit tests or that you captured the flag.” If the bottleneck inside frontier AI research is the slow, expensive kind of verification rather than the instant kind, the transfer from formalized proofs to research self-improvement is far weaker than the framing suggests.
Mowshowitz goes further than the public evidence supports in two places. He predicts that whoever gets traction on genuine AI research self-improvement loops will find itself in “an overwhelmingly strong position.” He also guesses that Anthropic’s own specializations matter “at least as much and probably moreso” than OpenAI’s head start in reinforcement learning. Both are informed hunches he offers as such, not conclusions drawn from Astra’s results. Ten formalized proofs demonstrate that verifiable domains fall faster than unverifiable ones. They do not demonstrate that any lab has closed a loop on self-improving research.
The framing deserves to be taken seriously, because it explains a pattern rather than just describing a result: math, coding, and cyber keep falling first, and cheap verification is the common factor across all three. But the strong version of the claim, that the same dynamic hands one lab a decisive research advantage, assumes away the hardest part of AI research, which is deciding what is worth trying before anyone knows if it worked. Until a lab shows progress on the slow-verification problems, architecture judgment, data curation, whether a scaled-up run will pay off, the math proofs remain evidence about math, not evidence about who wins the research race.
Operators evaluating lab claims about AI-driven research acceleration should ask which kind of verification is doing the work: the instant, mechanical kind that produced these ten proofs, or the slow, judgment-heavy kind that actually gates what ships in a frontier model. Those are different problems, and only one of them has been solved.
Zvi Mowshowitz published this analysis on his blog on 3 August 2026.