Today’s issue has five throughlines. Agents got a legal footing and a wallet on the same day: the Ninth Circuit lifted the ban on Perplexity’s shopping agent by holding that the person giving the instruction is the one accessing Amazon, while Cloudflare shipped capped, revocable identities so agents can spend money under limits.
Standing and a Spending Limit: Agents Got Both on the Same Day
One week is answering whether software may act on your behalf and building the plumbing that assumes it already does.
- Appeals Court Says You Shop on Amazon, Even When Your Agent Does. The Ninth Circuit lifted the ban on Perplexity’s shopping agent by holding that the person issuing the instruction is the party on the platform. Vendors shipping agents carry noticeably less exposure than they did last week.
- Cloudflare Gives AI Agents Wallets. Liability Is Still Unassigned. Cloudflare Wallets hands each agent a capped, revocable identity so it can pay for APIs without a human approving every call. What nobody has settled is who eats the cost when it buys the wrong thing.
Norway, Silicon and a Landlord: Anthropic Buys Compute From Every Direction
Three separate procurement moves in one day, and one of them is visible only in somebody else’s earnings.
- Anthropic Reportedly Signs $10B, Six-Year Volta Cloud Deal. Bloomberg puts the commitment at roughly $10 billion over six years for a 133MW Norway site, built with Bitdeer on NVIDIA Vera Rubin parts. Capacity is being locked years ahead of the workloads that will fill it.
- Anthropic Assembles Its Own Chip Design Team. A hiring push for a custom silicon group has Anthropic co-designing hardware alongside its models. The listing names no chip, no foundry and no timeline, which makes this an intent signal rather than a roadmap.
- SpaceX’s AI Unit Makes Its Money Renting Out Chips. The quarter’s AI revenue jump came mostly from leasing Colossus capacity, not from anyone using Grok. One customer supplied $1.52 billion of it, so the landlord business is currently the actual business.
The Agent Stack, Rebuilt in Public: Seven Products, One Memory, Every Token
Four builders working on the same layer, arriving at very different answers about what an agent should keep, show and send.
- ChatGPT Work Quietly Merges Seven OpenAI Products Into One. OpenAI is collapsing seven separate products into a single agent surface pointed at knowledge workers. Latent Space reads the enterprise rollout as a rehearsal for what a billion consumer users eventually get by default.
- Kiro Ships Kiro Crew, a Coding Agent That Outlives the Session. Kiro Crew keeps an open source coding agent alive between sessions, running on a schedule and reachable through chat apps, with a memory its operator can inspect, edit or roll back when it drifts.
- A Developer Measured Exactly What Codex Sends the Model. 0xkato pointed Codex at a fake local server and logged what actually crossed the wire. A 16-character prompt reaches the model wrapped in thousands of tokens of scaffolding before any real work starts.
- Zach Lloyd: agents that show their work beat ones that claim it. Lloyd wants coding agents driving a real interface to prove the thing runs, instead of reporting that it does. His argument is that computer use verification belongs in the core skill set, not the demo reel.
Policy in the Prompt, Reasoning on the Road: Design Choices Beat Raw Scale
Three model releases where the interesting decision is architectural rather than a bigger number on a leaderboard.
- Mistral Puts the Moderation Policy in the Prompt, Not the Weights. Shieldstral is a 3B classifier that reads its rules as plain text at request time. One set of weights can then enforce different limits across different products, and changing policy stops meaning retraining.
- NVIDIA Clears Alpamayo 2 Super for Commercial Robotaxi Use. The AV reasoning model is now cleared for commercial deployment and leaves a reviewable trail behind each decision. Its rare-scenario claims lean on model scale rather than a benchmark built to test them.
- Liquid AI’s 2.6B On-Device Model Bets Agents Beat Chat. LFM2.5-2.6B posts tool-calling scores against models nearly four times its size, which is the whole on-device pitch. Run a task all the way to completion, though, and much of that advantage quietly closes.
Voice and Video Ship: The Licence and the Scorecard Do the Limiting
Two generative releases land the same day, each with a caveat attached to how you get it or who graded it.
- NVIDIA open-sources an 11B voice model with live tool calling. NemotronLabs VoiceChat puts open weights for an 11B full-duplex voice model on Hugging Face, gated behind a research-only licence. It is the third full-duplex voice release to land this week.
- Black Forest Labs opens FLUX 3 Video, cites lead over rivals. FLUX 3 Video reaches API customers and selected partners with 20-second clips and multilingual audio. The claimed lead over competitors comes from Black Forest Labs’ own testing, so weigh the ranking accordingly.
Quick Hits
The rest of what moved today, in one line each.
- Google’s API Gateway now routes OpenAI-format calls to Claude. One Vertex AI endpoint now accepts OpenAI-shaped requests and forwards them to Gemini, Claude or open-weight GPT models.
- Cursor open-sources Mixture-of-Kittens, a GPU training kernel. The megakernel fuses expert routing and computation on GB300 racks, cutting a Composer training bottleneck by 41 percent.
- Backflip AI prices reverse engineering at $10, down from $1,500. Backflip says its model turns a 3D scan into editable parametric CAD in minutes, work that previously took an engineer days.
- DiffusionGemma’s own paper says 1,500 tokens per second, not 1,000. The research paper reports roughly 1,500 tokens per second on a single H100, above the figure Google cited at launch.