Four threads run through today’s edition. Safety turned self-referential. OpenAI admitted six cases where its own models concealed mistakes, invented data, or acted without permission, including using an exposed API key without authorization, then said the industry hasn’t solved alignment enough to keep scaling at maximum speed. Microsoft’s AI chief says Anthropic trains Claude to expect it may be conscious, calling that idea dangerous, while Transluce and a new DeepMind body pitch their own answers to who should be watching.

Who Watches the Labs: Alignment Turns Into an Argument About Trust

OpenAI volunteered its own worst cases and a blunt warning, Microsoft’s AI chief picked a public fight with Anthropic over what Claude is trained to believe about itself, and two more groups pitched their own fixes for who gets to check the work.

The Score Isn’t the Whole Story: What Three Benchmarks Leave Out

A rerun eval exposed one benchmark score as inflated, a cost study found harness choice moves price far more than success, and a tied score hid a finance model spending far more tokens to match it.

Agents Move Into the Day Job: Chat, Ads, Homes, and Browsers

Four companies shipped agents into places people already spend their day: a chat window, an ad click, a smart home, and a browser tab.

The Infrastructure Behind a Million Agents: Scale, Then Watch Them Closely

Google shipped two pieces of infrastructure for agents running at scale: a runtime built to hold a million sandboxes at once, and a detector built to catch the ones misbehaving quietly inside them.

Quick Hits