Three throughlines run through today’s edition. The frontier’s economics keep shifting. DeepSeek priced V4-Pro at $0.87 per million output tokens, Microsoft shipped its third in-house model to stop paying OpenAI and Anthropic for inference, and Ramp data show Anthropic’s own Fable 5 holding just six percent of enterprise token share.
The Price War at the Frontier: Cheap Weights, Costly Contracts
Three data points from today show the same pressure from different directions: falling token prices, in-house model builds, and soft enterprise adoption of the priciest model on the market.
- DeepSeek Prices V4-Pro at $0.87 per Million Output Tokens. DeepSeek’s benchmark claims are its own, but its July token share trailing only Anthropic is the number that should worry rivals pricing above a dollar.
- Microsoft Ships MAI-Thinking-1, Its Own Reasoning Model for Copilot. Microsoft’s third in-house model this year swaps into Copilot products where it previously paid OpenAI and Anthropic for inference, cutting a line item it controls.
- Fable 5’s steep price meets a soft enterprise market. Ramp spend data show Anthropic’s flagship model captured just six percent of enterprise token share, and one economist calls that ceiling, not a ramp.
Today’s Round of Releases: Open Weights and the Long Game
New models keep landing at both ends of the market: a dense frontier release built for agent work, and an open-weight giant published for anyone to deploy.
- xAI Releases Grok 4.6, Aimed at Long, Multi-Step Agent Work. xAI says the model matches GPT-5.6 Sol on a composite benchmark index and is offering double usage limits for the first week to pull developers in.
- Qwen3.8’s Weights Land: 2.4T Parameters, 95B Active Per Token. Alibaba’s Qwen team published the full model card confirming the architecture, benchmarks and deployment paths it had only teased in July.
- A developer’s month with Grok 4.6: fast, polished, still needs steering. Eric Zakariasson ran xAI’s new model as his daily driver and says the real win is tempo, paired with a prompting habit that forces the model to check its own work.
The Ghost in the Machine: Agents Still Need a Supervisor
Five stories today converge on the same constraint: shipping an agent is getting easier, but judging whether it did the job right is not.
- Claude in Chrome turns its side panel into an agent workspace. Anthropic says the browser extension now carries an agent’s skills, connectors and history across desktop, web and mobile, turning a sidebar into a persistent workspace.
- A field guide to how much SQL access your agent should get. Developer Pamela Fox lays out four ways to wire an agent into Postgres, ranked by how much damage a single bad guess can do to the database.
- Specula’s Bug-Hunting Agent Impresses a Reviewer, but Not Fully. Murat Demirbas credits Specula’s automated bug hunting in a close read of the paper, but notes it never proves whether per-module fixes secure a whole system.
- Agent adoption’s real constraint is verification, not capability. A VC’s framework for hiring agents like employees argues that judging their output, not extending their capability, is the part enterprises are actually stuck on.
- Three Engineers, Hundreds of PRs: The Review Loop That Makes It Work. Pandas creator Wes McKinney says a three-person team merges hundreds of pull requests weekly by centering adversarial review, not autonomous agent loops, in the process.
Who Captures the Value: Infrastructure, Code and the Usage Gap
Three stories today trace where the money actually lands, from chip financing to a Stockholm coding startup to OpenAI’s own usage data.
- Nvidia’s Real Product Is a Synthetic Hyperscaler, Not Just GPUs. Clark Tang argues Nvidia’s fleet software and financing platforms have assembled a hyperscaler’s economics, turning capital lock-in into a moat rather than a risk.
- Lovable Doubles Its Valuation to $13.3B in Eight Months. The Stockholm coding startup priced its new round at roughly 27 times June revenue, a bet that growth keeps outrunning the multiple investors just paid.
- OpenAI’s enterprise data shows a widening gap, in its own metric. New OpenAI research finds top firms generate 8.3 times more output tokens per user than laggards, a consumption figure that doubles as a sales pitch.
Automating the Lab: AI Research Turns on Itself
Three stories today ask how far AI has already moved into the work of building and judging AI, and who still gets a vote on the direction.
- Milestones for AI Automating AI Research Are Falling Ahead of Schedule. A researcher survey found several markers labs set for automated AI research already crossed, years earlier than the interviewees themselves had expected.
- A Fields Medallist maps where LLMs excel at math, and where not. Timothy Gowers argues OpenAI’s math wins are real but narrow, and offers a test for that boundary that does not require trusting any published benchmark.
- Hinton, Li and Ng Split on Tactics but Agree AI Must Stay Open. At Ai4 in Las Vegas, the three researchers offered different, occasionally clashing reasons for resisting concentrated control over how AI development proceeds.
Quick Hits
The rest of what moved today, in one line each.
- Microsoft’s MAI-Image-2.6 climbs to No. 2 on Arena leaderboard. Microsoft says its newest image model gained 79 Elo points overall and now ranks above image models from Meta, Google and xAI on the Arena leaderboard.
- Google Quietly Builds an Agent Console Inside AI Studio. Screens spotted inside AI Studio suggest Google is quietly adding a dashboard to manage the Cloud agents it launched back in May, replacing what had been a set of scattered configuration screens for operators.