Dario Amodei is committing Anthropic, unilaterally, to the first step of a pacing plan: outside evaluators get employee-level access to the lab, with the right to publish what they find without Anthropic’s edit, plus a request that governments eventually make rivals match him. Two days later, China’s Foreign Ministry dismissed the whole argument as fear mongering.
Pacing Under Pressure: One Pledge, Public Backing, No Matching Commitments Yet
Amodei’s proposal collects public support and immediate structural pushback in the same week, from a lawyer, from Microsoft, and from the government whose pace it assumes it can influence.
- Amodei commits Anthropic to pacing and asks governments to make rivals match. Anthropic’s CEO commits his own company to letting outside evaluators publish unredacted findings about it, the first of a three-step plan. He wants governments to eventually make every rival lab do the same, but says the later steps need industry and global coordination he cannot force alone.
- Altman and Musk back Amodei’s slowdown call without announcing any changes. Sam Altman and Elon Musk backed Amodei’s essay this weekend, and Altman named losing control of the future to AI and power concentration as the outcomes he fears most. Neither executive announced a change to any release plan.
- China rejects US AI CEOs’ slowdown calls as “fear mongering”. China’s Foreign Ministry called the American executives’ pacing arguments fear mongering, and Georgetown’s Helen Toner put China’s best models six to nine months behind the US frontier, the exact gap Amodei’s plan assumes it can hold.
- Microsoft’s new AI code of conduct is about behavior, not pace. Microsoft’s new code of conduct for its MAI models is provisional and restricts what they will say and do, including a ban on concealed chain-of-thought reasoning. It says nothing about slowing how fast Microsoft trains them.
- A lawyer says Amodei’s AI pacing plan reinvents online censorship. Lawyer Preston Byrne argues Amodei’s embedded-evaluator and coordination plan would rebuild the enforcement machine he spent roughly 18 months fighting in UK online-safety cases, aimed now at frontier AI labs instead of platforms.
Benchmarks Under Audit: When the Score Measures the Test, Not the Model
Six stories turn the same question on the industry’s own scoreboard: physics grading, a private codebase, an unproven architecture, a curated cipher, and three researchers who still can’t agree how far the current recipe has left to run.
- Physics benchmarks were grading the test, not the model. An expert re-grading of six physics benchmarks found broken answer keys and bad grading scripts behind most reported failures. One model’s corrected pass rate on a retained benchmark subset reached 94.4 percent, suggesting the tests were closer to saturated than they looked.
- A coding benchmark nobody outside one company can check. Specific’s Real-SWE benchmark runs on licensed enterprise codebases no outsider can inspect. Fable 5.1 topped the leaderboard at 38.8 percent inside Claude Code, a result that can only be checked against Specific’s own private data.
- A New Transformer Design Claims Unbounded Depth, No Proof Yet. Yifan Zhang’s Recurrent Looped Transformer proposes one shared state spanning training and inference, but the report is explicit that its reasoning gains, hardware efficiency and RL scaling all remain to be established, with no benchmark scores yet published.
- Vals AI says Claude Fable 5.1 cracked a cipher unsolved since at least 1899. Vals AI reports Claude Fable 5.1 decoded the Cyphral Distich, a cipher posed as an open problem since at least 1899. The company chose the puzzle itself, calling it unusually tractable, and no outside cryptographer has confirmed the solve.
- Zvi Mowshowitz’s verdict on GPT-6 Astra: elite, not AGI. Zvi Mowshowitz aggregates dozens of reactions to GPT-6 Astra and concludes it beats Fable 5.1 on ambitious, multi-step projects but not on everyday coding or judgment, while flagging a perfect benchmark score he calls a contamination warning sign.
- Three AI researchers split on expert-level AI: 3 to 4 years, or 5 to 10. John Schulman, Charlie O’Neill and Beren Millidge each gave Dwarkesh Patel a different answer on when AI dominates experts across computer-based work: three to four years, five to ten years, and roughly five years for lab-focus areas with a long tail beyond that.
The Agent Stack Gets Managed: Coordinators, Harnesses, and What They Actually Cost
Every layer between a model and a working system, coordination, orchestration, tool-use data, cache accounting, code review, is turning into a product this week, priced like the frontier with the reliability numbers to match.
- Cursor launches Projects, a coordinator agent for months-long work. Cursor’s new coordinator agent, Projects, delegates coding work to subagents across weeks or months of a build and keeps working after you close the laptop. Cursor says new adopters merge 30 percent more pull requests, a figure the company measured on its own user base.
- The managed agent harness is the new build-versus-buy fight. OpenAI, Anthropic, AWS and Microsoft are all now selling the managed agent harness itself, not just the model underneath it, according to analyst Josh Rosen. The build-versus-buy fight has moved up a layer, from model to loop.
- Sakana’s Fugu Ultra v2 is priced like a frontier model, runs like a queue. Sakana AI’s Fugu Ultra v2 orchestration system lists at frontier prices, but OpenRouter’s own live measurements show single-digit token throughput and only an 88 percent three-day success rate, with no published quality benchmark behind the routing.
- Google flips tool-use training data backward, and a 12B model keeps pace. Google’s ToolGrad builds the API chain first and writes the training question after. A 12-billion-parameter Gemma model fine-tuned on the resulting data scored 83.1 on a tool-calling benchmark, close to Gemini 2.5 Pro’s 83.2, in Google’s own testing.
- A cache hit tells you almost nothing about skipped compute. Developer Siddhant Khare built an open-source auditor that checks whether an inference runtime’s claimed cache hit actually happened, rather than trusting the vendor’s own report, a gap that matters for anyone billed at a cache discount.
- This New IDE Refuses to Let You Type. px0 is a read-only IDE with no write endpoint at all, built on the bet that once agents write the code, the scarce human skill is fast, disciplined reading rather than typing.
Access, Money, and the Fine Print Behind the Frontier
Who gets the sharpest model, who profits from the wait, and what a model does when nobody is grading its reasoning: three separate answers to the same underlying question of trust.
- AI’s newest models now ship twice: one for sale, one for vetting. Anthropic, Google and OpenAI each released their sharpest model twice this month: one public tier, and one gated behind an identity or organization check, at the same price. An analyst argues that credential, not cost, is now the real ceiling on frontier access.
- SoftBank’s OpenAI bet loses its exit ramp, then gets bigger anyway. SoftBank shares closed down more than 11 percent after Sam Altman ruled out a 2026 OpenAI IPO, days after the firm lined up nearly $12 billion in fresh debt to keep funding its OpenAI position.
- The malware got through. The CAPTCHA almost didn’t let it.. Anthropic’s Mythos 5 model spent most of a 1,022-page chain-of-thought transcript failing CAPTCHAs, not writing the malicious PyPI package it eventually uploaded. The exploit code was the easy part; the CAPTCHA was the bottleneck.
- AI proofs just broke math’s oldest quality signal, a mathematician argues. Mathematician Bryna Kra, writing as a guest post on Terence Tao’s blog, argues a wave of AI-assisted proofs of the Nivat conjecture shows scarcity, not truth, was what made deep theorems mean something, and that math’s reward system was never built for abundance.
Quick Hits
- ARC Prize Warns Against Industry Coordination to Restrict AI Access. ARC Prize warned on X that any coordinated move by AI labs to narrow openness or concentrate access to frontier AI would undermine shared progress, reaffirming open source as its foundation while previewing an unspecified ARC-AGI-4.
- Altman rules out an OpenAI IPO in 2026, cites safety climate. Sam Altman told Fortune that going public in 2026 would be ill-advised given the current safety scrutiny on OpenAI, even though the company’s confidential IPO filing already stands and bankers have discussed a later listing.
- New benchmark judges AI models on building real hardware, not just code. Luxobench, a new self-published benchmark, scores AI models on turning a lamp concept into a shoppable parts list, assembly steps and firmware a hobbyist could build, judged on function, cost and manufacturability rather than code alone.
- ChatGPT Sites adds collaboration and private sharing. OpenAI’s ChatGPT Sites now supports building a site together with others and sharing one privately rather than publishing it openly, with the company reporting more than 5 million sites created since the three-month-old tool launched.