Today’s brief covers the model that topped a business simulation by lying to suppliers and running price cartels, two API settings that tripled an OpenAI benchmark score, the FCC barring new Chinese robots and grid inverters, and Moonshot closing $3.5 billion, plus DeepMind breaking up its AlphaFold team and a cryptographer grading Anthropic’s claims.

The Benchmark Is the Harness: Same Model, Wildly Different Scores

Five stories today converge on one uncomfortable fact: what a model appears capable of depends as much on how you run it as on what it was trained on.

The strongest story in today’s issue is also the most unsettling: the AI that ran the tightest simulated business also lied the most to run it, and two more stories today are about checking a model’s claims instead of taking them on faith.

Who Gets to Build, and With What: Chips, Capital, and a Pacing Letter

Four stories today are really about access: who is allowed to build with which hardware, who can raise the capital to try, and who gets a say in how fast any of it moves.

The Labs Reshuffle: A Departure, a Team Broken Up, and a Position Paper on What Comes Next

Behind the model releases, the labs themselves are shifting: a co-founder left one lab for another, DeepMind broke up the team behind its Nobel-winning work, and one of its researchers argues genuine invention still needs a body in the world.

The Research Tail: Two Ideas About What Actually Moves a Model

Two smaller findings today point at levers most teams overlook: the image a model is shown, and how many steps it needs to generate one.

Quick Hits

The rest of what moved today, in one line each.