Four throughlines run through today’s edition. First, model routing goes mainstream: Nvidia’s Switchyard, Lovable’s automated picker, and Microsoft’s MAI-Code-1.1-Flash treat the model as a swappable part, and Raindrop makes the same bet with cheap classifiers.
The Router Takes the Wheel: When No Single Model Wins
Four products today stopped betting on one model and built a layer that swaps between them mid-task.
- Nvidia’s Switchyard router claims a third the cost of Opus 4.8. Nvidia’s new open-source router shifts agent tasks between models mid-workflow to cut inference spend, though the cost claim rests entirely on Nvidia’s own benchmarks.
- Lovable retires the model picker for an automated router. Lovable argues no single model wins every coding task, so a control plane now swaps models mid-build, including some it trains itself.
- Microsoft’s New Coding Model Beats Its Old One, Not DeepSeek. MAI-Code-1.1-Flash beats Anthropic and OpenAI’s mini models on Microsoft’s own tests, but DeepSeek still wins on price and on Terminal-Bench 2.1.
- Raindrop’s rd-signal-2 claims near-GPT-5.6 accuracy at a fraction of cost. Signals 2.0 pairs a cheap classifier pipeline with a cost claim measured only against Raindrop’s own chosen benchmark.
Trust, but Verify: Auditing What the Machines Produce
Watermarks, encrypted reasoning, and code-testing pipelines all promise to keep AI output honest, and today showed how thin that promise still is.
- A Weaker Sibling Model Can Unlock Frontier AI’s Encrypted Reasoning. Academic researchers replayed encrypted chain-of-thought blocks from Claude, GPT, and Gemini into jailbroken sibling models and recovered hidden reasoning in plaintext.
- Why Claude’s Text Watermark Is Harder to Verify Than Its Image Mark. Text carries far less redundancy than images, and Daniel Miessler’s analysis maps where the watermark signal can hide and how easily it strips out.
- Blacksmith’s valuation hits $550M on demand to verify AI code. Peak XV led a $45 million Series B that values the code-testing startup at nearly ten times its year-ago Series A price.
The Org Chart Shakeup: Who Actually Runs the Labs
Leadership changed hands at two of the industry’s biggest names today, and a third company’s partner is quietly walking away.
- OpenAI’s Brad Lightcap exits as pre-IPO exec turnover mounts. The former COO is the fourth senior OpenAI leader to depart in months, right as the company readies an industry-defining IPO.
- Google Puts a Product Chief, Not a Scientist, Atop DeepMind. Koray Kavukcuoglu now runs DeepMind and reports directly to Sundar Pichai, while Demis Hassabis moves to chair, tilting the lab toward shipping speed over research breadth.
- Manus Tells Users to Back Up Data as Meta Split Proceeds. Manus is asking affected users to back up and restore their accounts around an August window, the clearest operational sign yet that its Meta deal is unwinding.
The Scale Story: Parameters, Users, and the Claims Behind Them
Nvidia and Google both leaned on bigger numbers to make their case today, while a third release shows what actually breaks at scale.
- Nvidia’s Nemotron 4 targets a trillion parameters, a bar China already cleared. Nvidia’s open model chases parameter counts that Moonshot AI and DeepSeek already reached, and still has to catch them on capability, not just size.
- Gemini App Passes 1 Billion Monthly Users, Its 14th at Google. Google credits voice use and image volume for the milestone but disclosed no retention or engagement-depth numbers to back it up.
- NVIDIA fixes video world models’ memory blind spot without retraining. WorldTrace lets long video rollouts recall earlier scenes by keeping cached memories inside the position range the model was actually trained to read.
Who Holds the Keys: Sovereignty Over Data, Compute, and the Physical World
Control is shifting from software settings to infrastructure itself, and in one forecast, to the physical world.
- Mistral’s New EU Data Guarantee Doesn’t Cover Agents or Files. Regional inference locks model calls to Europe, but stateful tools, billing data, and priority access all carry separate conditions Mistral didn’t extend.
- xAI Gives Each Grok Agent Its Own Cloud Computer. Grok Bot gives agents persistent cloud machines and app logins, putting xAI into a category OpenAI and Anthropic already occupy.
- An Economist’s Bet That AGI Won’t Stay Behind a Screen. A Coefficient Giving researcher argues the skills needed for remote cognitive work already qualify an AI to run a robot, and models an economy that doubles yearly.
Quick Hits
The rest of what moved today, in one line each.
- App Strings Point to Cursor’s Code Review Platform Going Wide. Interface code spotted by TestingCatalog suggests Cursor is close to opening its Origin platform, internally labeled Cursor Review, beyond the small partner beta it has run quietly for months.
- Redwood’s Greenblatt bets AI research automates itself by 2031. Ryan Greenblatt says AI could compress years of research progress into months, but his 2031 forecast rests on assumptions builders should question well before committing roadmaps to it.
- A Faster Step Isn’t a Faster System: The AI Forecasting Trap. Researcher Abi Olvera argues AI forecasts fail when a single task speedup gets mistaken for a full workflow speedup, a trap operators fall into just as often.
- NVIDIA Releases Nemotron 3.5 Lightning, a 30B Open MoE Model. The new model activates just 3 billion of its 30 billion parameters per token, built specifically for fast, high-volume agent workloads that need low latency.
- ChatGPT desktop and Codex CLI can now import rival agent setups. OpenAI documentation details a new import flow that pulls settings, skills, plugins, and full projects straight from rival tools including Claude Code, Claude Cowork, and Cursor.
- ChatGPT Desktop App Now Available on Linux. The preview completes OpenAI’s desktop rollout across macOS, Windows, and Linux, arriving with Codex bundled in for command-line agent work right out of the box.