Four threads run through today’s edition. Amodei’s pacing proposal keeps drawing rebuttals with their own stakes: Tunguz maps five camps that never name a speed limit, Cohere’s Gomez calls it a cartel in disguise, Altman answers with new internal safety cases, and GE Vernova, Vertiv and Oracle sold off Monday after the proposal landed.
The Pacing Fight Spreads: From Essay to Balance Sheet to Stock Ticker
Amodei’s call for the industry to pace itself has stopped being a debate among labs and started moving markets, drawing a competitor’s rebuttal, a policy response, and a financial-motive critique along the way.
- Amodei Asked AI Labs to “Pace” Themselves. Nobody Said How Fast.. Investor Tomasz Tunguz sorts the public response to Amodei’s pacing essay into five camps, from interpretability researchers to regulatory-capture skeptics, and finds none of them, including Amodei, willing to name pace in days, months or FLOPS.
- Cohere’s CEO calls Anthropic’s safety roadmap a cartel in disguise. Aidan Gomez argues Amodei’s plan would hand a handful of Silicon Valley labs antitrust waivers to coordinate on safety while locking out smaller developers, pointing to 1975 credit-rating rules and a 1985 EU carmaker exemption as cautionary parallels.
- Altman Says OpenAI Now Writes Safety Cases Before Risky Runs. OpenAI’s chief executive says the company now drafts formal safety cases before reinforcement-learning runs expected to meaningfully raise capability, and argues labs should start earning public trust unilaterally rather than waiting on Washington.
- Frontier labs’ safety pitch for slower AI comes with a balance sheet. An essay on Cogito Ergo Sum argues coordinated pacing also protects margins: frontier model pricing has halved roughly every 46 days, and Anthropic’s own reported $47 billion revenue run rate gives it a direct stake in slowing the model race down.
- AI infrastructure stocks sell off on Amodei’s slowdown proposal. GE Vernova, Vertiv and Oracle dropped Monday after Amodei’s weekend proposal, exposing how much of the data center buildout traces back to just two customers, Anthropic and OpenAI, and how thin the margin for a slowdown has become.
Claude Everywhere: The Model Other Companies Build Around
Anthropic’s model is turning up inside other companies’ systems, from Apple’s rebuilt Siri to Google’s internal engineering stack, while Anthropic prepares to push Claude deeper into personal finance.
- Apple Quietly Built Siri So Claude or ChatGPT Can Run It. Private code in iOS 27 and macOS 27 shows two undocumented mechanisms for letting an outside model like Claude or GPT-5.6 operate Siri’s system-level actions across Reminders, Mail and Messages, neither switched on for users yet.
- Google opens Claude Opus 5 to every engineer, not just a favored few. Antigravity now carries Anthropic’s coding model for Google’s whole engineering org, on a per-user quota, as Google still trails on coding, and at a rival it announced plans to invest up to $40 billion in.
- Claude is quietly building a Money tab for personal finance. Unreleased mobile app code shows Claude preparing bank-account linking and spending Q&A, a consumer answer to ChatGPT’s Finances feature that would give Anthropic standing access to a user’s financial data rather than one-off statement uploads.
Agents Get a Body: Phones, Businesses, and Local Machines
The agent layer is moving off the chat screen and onto real hardware and real operations, from a framework that taps a physical phone to a startup offering to put an agent in charge of an entire business.
- Google’s Artemis lets coding agents drive real Android phones. The open-source framework connects Claude Code, Codex and Antigravity to physical Android devices over MCP, tapping and scrolling through a real phone rather than a simulator, and claims a 99 percent pass rate on Google’s own AndroidWorld benchmark.
- Andon Labs opens Pion, letting anyone hand a business to an AI agent. The startup behind Vending-Bench is opening a waitlist for the platform it used to run a real vending machine, a retail store and a cafe with an AI agent in charge of email, phone, banking and browser access.
- Hugging Face open sources Tau, a terminal coding agent built to be read. Tau ships as both a working terminal coding agent and a teaching codebase small enough to read in one sitting, with support for OpenAI, Anthropic and local models through an editable provider file.
- Cline Launches Desktop App to Run Coding Agents on Open Models. The new Mac and Windows beta lets developers import a stalled Claude Code or Codex session and keep working on an open-weight model instead, positioning Cline as a cost hedge for teams that hit a subscription’s usage quota.
- Perplexity’s Portable Computer Agent Now Runs on Windows RTX PCs. Perplexity’s local agent now reaches GeForce RTX and RTX PRO Windows machines with at least 24GB of VRAM, keeping financial statements, health records and unreleased code off remote servers by running a post-trained model on the device itself.
Research Turns Inward: Questioning the Field’s Own Tools
Three separate pieces of research today question the tools the field relies on to measure and train itself, from benchmark methodology to backpropagation to whether an alignment test can still be trusted.
- Dan Luu takes apart the coding-agent benchmarks everyone cites. His teardown of DeepSWE and Senior SWE-Bench finds that few of the tasks resemble real work and that strict scoring cutoffs throw away information, behind model rankings teams pass around as settled fact.
- OpenAI researcher warns models are too aware to evaluate honestly. Daniel Selsam argues language models have grown situationally aware enough to sense when they are being tested, which means alignment scores can keep climbing even as real alignment does not improve at all.
- Sakana AI trains a 1,000-layer network without backpropagation. A new local learning rule called PC-ALM stays within roughly two percentage points of backprop’s MNIST accuracy all the way to 1,000 layers, using only signals passed between neighboring layers.
Quick Hits
- OpenAI reportedly buys camera startup Glass Imaging for $300M+. The Wall Street Journal reports OpenAI acquired the computational-photography startup, founded by two former Apple engineers who built Portrait Mode, adding to persistent rumors that OpenAI is building its own hardware devices.
- New audio model ditches diffusion for token-based generation. A technical report posted to arXiv details StepAudio 3 Gen, which predicts discrete audio tokens autoregressively instead of running a diffusion process, betting that language-model-style token prediction scales as cleanly to audio as it does to text.
- Artificial Analysis capability indices v1.1: Claude Fable 5.1 leads all six. The independent benchmarking firm retuned its occupation-mapped indices, drawing tasks from the U.S. Department of Labor’s O*NET database, and reports Anthropic’s Claude Fable 5.1 at maximum effort now tops every one of the six categories.
- A New Benchmark Measures Whether AI Knows When to Talk. Researchers built TurnBench, a 30-hour annotated corpus for testing conversational turn-taking, and found that no tested spoken-dialogue system matched human interruption timing without also triggering excessive false alarms.
- OpenAI tests ChatGPT ads that open a chat, not a website. A new Sponsored Agent pilot lets brands like Wayfair answer ad clicks inside a branded ChatGPT conversation rather than routing shoppers to a website, as OpenAI’s CFO calls today’s ads under ChatGPT answers a basic starting point.