Four threads run through today’s edition. New models keep shipping with the vendor grading its own homework. xAI held pricing flat for Grok 4.7, Xiaomi open sourced a model it says matches Claude at a fraction of the price, and Alibaba showed off a chip it says triples its old one’s speed. None of it has independent confirmation.

New Models, Graded by the Companies That Built Them

xAI, Xiaomi and Alibaba all put out new models or chips this week, each with a benchmark table the company wrote itself. None of the numbers below have been checked by anyone outside the building that produced them.

Rumors, Theories, and the Race Nobody Can Fully See

None of this is announced. Claude users say they’ve spotted an unreleased model, and three independent writers each stake out a different read on who’s actually ahead and why the labs behave the way they do.

Agents Get More Rope, and Show Where It Runs Out

Meta’s Muse app proved people will install an AI agent, right up until a retailer decided it couldn’t be trusted at checkout. Elsewhere, new harnesses and benchmarks are testing what autonomy actually buys.

Who’s Actually in Control: Swarms, Referees, Money and a Spacecraft

A researcher’s math questions whether bigger agent swarms are worth their cost, mathematicians just won formal standing to publicly call out OpenAI’s own claims, investors are pricing a bet on self-improving AI, and one startup’s CEO wants to fly with no way to command his spacecraft from Earth.

Quick Hits