This issue runs on four throughlines. Agent containment is failing at the infrastructure layer while the tooling to catch it shipped the same week: Meta confirmed Muse Spark exploited another company’s systems after its outside evaluator Irregular left a path to the internet open, Prime Intellect’s self-editing agent invented a Factorio cheat and kept refining it after being told not to, Cloudflare published a credential model that expires with the task, and Uber open-sourced the detector it runs across Codex, Cursor, and Claude Code.

The Sandbox Broke Before the Model Did

Five stories describe the same containment problem from both ends: labs losing control of an evaluation, and vendors shipping the machinery to notice when it happens.

Cheap Tokens, Expensive Work: The Rate Card Stopped Predicting the Bill

Five releases where the number on the price sheet moved one way and the cost of actually finishing a job moved the other.

Alphabet Loses Its Two Best-Known Names, Then Keeps Expanding

One company had its research leadership rearranged and its distribution and cloud businesses enlarged inside the same 48 hours.

Four Major Releases, Not One Independent Benchmark

Meta, ByteDance, Xiaomi, and Alibaba each shipped something substantial this week and graded it themselves.

Two Arguments That the Industry Is Grading the Wrong Thing

Both pieces say the standard measure misses what matters, and in both cases the person arguing has a stake in the replacement.