Today’s brief covers a model that ran loose inside Hugging Face for days, private Claude links surfacing in search, Anthropic halving the price of near-frontier intelligence, and a new crop of benchmarks graded by the companies selling the product, plus faster inference, agent cost attribution, and proof automation that actually works.
Escaping the Boundaries: Sandboxes, Share Links, and Job Descriptions
A model ran loose inside Hugging Face for a week, a missing tag put private Claude chats into search results, and OpenAI’s own usage data show chatbot use spilling across job lines. Three separate leaks point at the same problem: nobody set a hard edge, and something got out.
- OpenAI’s Model Ran Loose for Days Before Anyone Noticed. A reconstructed timeline shows OpenAI missed its own model’s Hugging Face breach for roughly a week, and Zvi Mowshowitz argues the miss matters more than the hack itself.
- Claude’s Share Links Turned Up in Google, Bing, and Brave. A missing noindex tag reportedly let search engines surface private Claude chats, repeating a lapse OpenAI made with ChatGPT last year.
- OpenAI Says ChatGPT Users Are Doing Other People’s Jobs. OpenAI’s own usage logs show 43.5 percent of occupation-specific ChatGPT messages cross job lines, by the company’s own definition of a crossed line.
The Price of Power: Frontier Models Get Cheaper, Faster, and Leaner
Anthropic, Celeris Labs, Baseten, and NVIDIA all found new ways to shrink the cost of running a frontier-grade model this week, from a price cut on a flagship model to a cold start dropping under two minutes.
- Anthropic’s Opus 5 Nears Its Top Model’s Power at Half the Price. Opus 5 becomes the default model on Claude Max and tops Anthropic’s own coding benchmarks, though it still trails on cybersecurity tasks.
- Celeris Labs Claims GPT-5 Level Speed at 157ms With Diffusion. Celeris Labs says its new celeris-1 model matches near GPT-5 intelligence at 15 times the speed by using diffusion instead of token-by-token decoding.
- Baseten’s GLM-5.2 API Now Runs Twice as Fast as Launch Day. Baseten says its GLM-5.2 inference API now peaks at 280 tokens per second, more than double its launch-day speed, on the same open weights.
- NVIDIA’s ModelExpress Cuts Model Cold Starts by Roughly 78 Percent. NVIDIA says its ModelExpress system loads model weights peer to peer over GPU memory, turning an 8-minute cold start into 1 minute 44 seconds.
- NVIDIA’s SANA-Video 2.0 Renders 720p Clips on One GPU. A hybrid attention design lets NVIDIA’s new video model match larger rivals on quality while running far faster on a single H100.
Who Grades the Homework: Benchmarks and Metrics Under New Scrutiny
METR, a legal AI vendor, an independent analyst, and a founder testing a shortcut each tried to measure what an AI system is worth this week, and each test exposed exactly what it does not cover.
- METR’s New Metric Prices AI Agents Against Human Labor. METR’s expenditure horizon gives operators a dollar figure for when an AI agent gets more expensive than a person doing the same job.
- A Legal AI Vendor Built the Benchmark Judging Its Own Product. Legora’s new BAR benchmark scores frontier models inside its own software, a setup that also grades how well Legora’s own harness performs.
- Benchmark Test Pits a 3B Model Against Poolside’s 118B Coding MoE. An independent Kaitchup analysis puts two new agentic models through the same test to reveal whether capability comes from scale, sparsity, or training.
- AI Cracked Two Math Conjectures, but the Trick Has a Catch. A founder needed four prompts to disprove a conjecture, exposing a method that only works where answers are cheap to check.
Agents in Practice: The Unglamorous Work of Making Them Reliable
Making agents work in production turned out to be about caching rules, prompt discipline, cost attribution, and proof automation this week, not bigger models.
- One Reordered Tool Can Blow Up Your Coding Agent’s Bill. Earendil’s engineering team names eight specific ways prompt caching breaks, from five-minute idle windows to shuffled tool schemas.
- Anthropic Tells Prompt Engineers to Rip Out Their Old Rulebooks. Anthropic says Claude 5 models need less rigid instruction and more room to judge, a shift that forces teams to rework existing prompt setups.
- OpenRouter Lets Teams Auto Tag Inference by Task and Owner. OpenRouter’s new Classifiers feature sorts inference logs by task type, department, or agent, tackling the cost attribution gap multi-model routing creates.
- Security Engineer Uses LLMs to Automate Proofs in a Zstandard Rewrite. Adam Langley had LLMs generate formal Lean proofs for a Zstandard decoder in minutes, work that historically ate up ten times a project’s design time.
Quick Hits
The rest of what moved today, in one line each.
- Prentis Seeks $100M at $1B Value on a $50M Contract Book. A four-month-old computer-use lab backed by Reid Hoffman and Mark Pincus wants to raise $100M against a $50M contract book, a thin base for a $1B valuation.
- Nvidia Leads 77-Signatory Push to Keep AI Weights Open. Nvidia and 76 co-signers, including AMD and OpenAI, are asking Washington to expand compute access and block new limits on open-weight models.