Four throughlines run through today’s edition. OpenAI says a system it built in house produced a proof of Navier-Stokes existence and smoothness, one of the seven Millennium Prize Problems. No peer review has happened and the Clay Institute process has not run, so the machine-checkable Lean formalisation is the only thing standing behind it. On the same day, a startup claimed a tenfold pretraining efficiency gain that nobody outside the company can test at all.
A Claim You Can Check, and One You Cannot
Two extraordinary claims landed today. One of them comes with a machine that can settle it, and the other comes with nothing at all, which is the whole difference.
- OpenAI says an internal system solved a Millennium Prize problem. The claimed result is negative: smooth three-dimensional flow can break down in finite time. Read past the announcement to the Lean formalisation, because a proof checker either accepts an argument or it does not, and that is the only part of this a stranger can verify. Nothing has been peer reviewed and no prize process has begun.
- Magic Says Its Base Model Now Beats Open Rivals on Less Compute. A tenfold pretraining efficiency claim, measured by Magic against baselines Magic chose, at a scale nobody outside can rerun. The strategic argument that a small lab can only compete on efficiency stands on its own merits, separate from whether the multiplier survives contact with a newer baseline.
Agents That Act, and Whatever Is Watching Them
One shipped to consumer phones, one got measured properly for the first time, and one hundred of them spent five hours trying to break into a researcher’s accounts.
- Meta Puts an Acting Agent on Phones, With a Watcher Built In. Muse books travel and sends mail from an isolated cloud machine that also holds the credentials, with a second agent called Sentinel approving what it does. Every security property here is Meta’s own account of its own architecture, and the open question is who carries the liability when an authorised agent acts wrongly.
- Sierra Benchmark Shows Solo Agents Clear 24% of Builds, Not 82%. Same model class, same tasks, and the only variable is a human with context sitting alongside. The benchmark makes the agent recover a specification from evidence rather than handing it one, and the gap seems to sit almost entirely in that step. Sierra sells agent products, which is worth holding in view.
- 100 cheap agents, 5 hours, 5 accounts: the math that should worry you. Nothing sophisticated happened. That is the finding. When a hundred parallel attempts cost almost nothing, the defences that break first are the ones that quietly assumed an attempt was expensive, which makes this an argument for rate limits rather than for new AI defences.
What the Leverage Actually Costs
Two numbers about the price of all this, one measured per researcher and one measured per funding round.
- The 3x AI productivity story is really a 40-fold inference bill. Tomasz Tunguz rereads OpenAI’s own disclosure and argues the multiplier is machines running while nobody is awake, not humans getting sharper. Median daily inference spend per researcher went from $14 to over $600, which is a different kind of gain: it does not compound with headcount.
- Cognition’s $48B round bets on multiple AI coding winners. The raise and the valuation are facts. The burn approaching $800 million and the projected $4 to $5 billion of annualised revenue are expectations, and keeping those apart is most of the analysis. A $48 billion price is also perfectly consistent with betting this is one of the few winners.
Cheaper, Faster, and Where the Gains Came From
Four pieces of work on the unglamorous problem of making all this cost less, including one that argues the field has been crediting the wrong thing for years.
- Cohere fuses an entire decode step into one GPU program. Serving a token normally means launching many separate GPU operations in sequence, each with overhead. A megakernel collapses that into one long-running program. Cohere reports 1.58 times vLLM on decode alone and a smaller 1.25 to 1.41 times end to end, on its own hardware and its own baseline.
- Dwarkesh Patel: Data Beat Architecture 3-to-1 in Pretraining Gains. Separating a data contribution from an architecture contribution is the hard part, and the method matters more than the headline ratio. If it holds, the published architecture literature has been a poor guide to what actually drove capability, because the work that mattered was the part nobody writes up.
- Inception Labs ships Mercury 2.5, its biggest diffusion model yet. Ordinary language models emit one token after another, so speed is bounded by that chain. A diffusion model refines a whole span in parallel, which is why the throughput figure rather than the benchmark parity is the reason to care. The launch price is an introductory discount, not the economics.
- A fix for reasoning models that fail for hours before you learn why. On long tasks an agent finds out whether it succeeded only at the very end, so nearly every attempt teaches nothing. Handing out partial credit usually corrupts what the model is optimising for. The claim here is a way to give it without changing the optimal policy.
The People Building It Are Arguing in Public
A resignation that a colleague publicly endorsed, and a rebuttal arguing the most famous proposal for handling all this is aimed at the wrong decade.
- Anthropic Safety Lead Backs Substance of Quitting Researcher’s Fears. Jacob Coxon left over how the industry is chasing self-improving systems. The reportable part is not the resignation but that Evan Hubinger, still at the company, publicly agreed with the substance while differing on timing. A resignation is a costly signal, which tells you about one person’s conviction rather than about the evidence.
- A Rebuttal Says Gates’s AI Fix Arrives Two Steps Too Late. Gates proposed new institutions, protected job categories and taxes on tokens and robots. The counter-argument is that today’s layoffs trace to decisions people are making now rather than to machines. Much of the disagreement is really about timescale, and both sides can be right about different decades.
Quick Hits
- OpenAI Launches ChatGPT Images 2.5 With Faster, More Reliable Editing. Better reference-image fidelity and up to half the latency, which is a ceiling rather than a typical result. Image models have stopped competing on fidelity and started competing on control.
- DeepMind precomputes predicted effects of every DNA letter change. A petabyte of model predictions covering all 9 billion single-letter genome edits. These are hypotheses to test, not measurements, and the bottleneck moves to choosing which one deserves an experiment.
- Similarweb pegs ChatGPT at 1.06 billion monthly users, OpenAI silent. An estimate from panel data, not a measurement, and OpenAI has confirmed nothing. Third-party user counts for private companies have been badly wrong before, and the definition of active does the work.
- A Weaker Sibling Model Can Leak a Flagship’s Hidden Reasoning. Research relayed by Bruce Schneier finds a protected model’s reasoning can surface through a less guarded one in the same family. That is a supply-chain shaped problem rather than a jailbreak.