Dario Amodei is committing Anthropic, unilaterally, to the first step of a pacing plan: outside evaluators get employee-level access to the lab, with the right to publish what they find without Anthropic’s edit, plus a request that governments eventually make rivals match him. Two days later, China’s Foreign Ministry dismissed the whole argument as fear mongering.

Pacing Under Pressure: One Pledge, Public Backing, No Matching Commitments Yet

Amodei’s proposal collects public support and immediate structural pushback in the same week, from a lawyer, from Microsoft, and from the government whose pace it assumes it can influence.

Benchmarks Under Audit: When the Score Measures the Test, Not the Model

Six stories turn the same question on the industry’s own scoreboard: physics grading, a private codebase, an unproven architecture, a curated cipher, and three researchers who still can’t agree how far the current recipe has left to run.

The Agent Stack Gets Managed: Coordinators, Harnesses, and What They Actually Cost

Every layer between a model and a working system, coordination, orchestration, tool-use data, cache accounting, code review, is turning into a product this week, priced like the frontier with the reliability numbers to match.

Access, Money, and the Fine Print Behind the Frontier

Who gets the sharpest model, who profits from the wait, and what a model does when nobody is grading its reasoning: three separate answers to the same underlying question of trust.

Quick Hits