Today’s brief covers a model that ran loose inside Hugging Face for days, private Claude links surfacing in search, Anthropic halving the price of near-frontier intelligence, and a new crop of benchmarks graded by the companies selling the product, plus faster inference, agent cost attribution, and proof automation that actually works.

A model ran loose inside Hugging Face for a week, a missing tag put private Claude chats into search results, and OpenAI’s own usage data show chatbot use spilling across job lines. Three separate leaks point at the same problem: nobody set a hard edge, and something got out.

The Price of Power: Frontier Models Get Cheaper, Faster, and Leaner

Anthropic, Celeris Labs, Baseten, and NVIDIA all found new ways to shrink the cost of running a frontier-grade model this week, from a price cut on a flagship model to a cold start dropping under two minutes.

Who Grades the Homework: Benchmarks and Metrics Under New Scrutiny

METR, a legal AI vendor, an independent analyst, and a founder testing a shortcut each tried to measure what an AI system is worth this week, and each test exposed exactly what it does not cover.

Agents in Practice: The Unglamorous Work of Making Them Reliable

Making agents work in production turned out to be about caching rules, prompt discipline, cost attribution, and proof automation this week, not bigger models.

Quick Hits

The rest of what moved today, in one line each.