Four throughlines run through today’s twenty-one stories: Anthropic and the Pentagon, a harder market to forecast, a faster generative-video clock, and a rough day for benchmarks.
Anthropic and the Pentagon: A Court Win, Then a Hiring Push
One ruling and one job posting show Anthropic fighting the Pentagon in court while still trying to sell it software.
- Judge Rules Pentagon’s Anthropic Risk Label Unlawful. A California federal judge found the Defense Department’s supply-chain risk designation against Anthropic was retaliatory and procedurally defective, and ordered it lifted, though a second suit over the same label is still pending in Washington.
- Anthropic Posts a $700,000 Job to Rebuild Its Pentagon Pipeline. A new Head of National Security Sales listing shows Anthropic pushing back into defense contracting even before its Pentagon standoff is fully resolved.
Generative Video: The Race to Render in Real Time
Three releases today chase the same goal: video generation fast and controllable enough to use like a tool, not wait on like a render.
- fal Ships H3 Max, Renders Video in Under 3 Seconds. fal is selling its post-trained MiniMax H3 variant at half price for a week to move developers onto the faster endpoint.
- SGLang’s lossless MiniMax-H3 video inference gain on H200s: 1.95x. The real number is the lossless 1.95x speedup on Nvidia’s H200s: a separately quoted 6.24x figure trades away video quality to get there.
- Google gives its Gemini video model editing controls, not just clips. Omni 1.1 Flash adds scene extension, keyframe interpolation, and 4K upscaling to the Gemini API, pushing output toward directable footage instead of one-shot clips.
The Money Behind AI: Forecasts, Funding, and a $125 Billion Miss
Five stories today turn on the gap between what forecasters expected the AI economy to do and what it is actually doing.
- Sell-Side Nvidia Forecasts Have Been Wrong by $125 Billion. A year ago Wall Street pegged Nvidia’s FY2028 revenue near $310 billion. The company’s own guide now points toward $700 billion.
- Anthropic and OpenAI’s Combined Revenue Run Rate Hits $100B. Epoch AI pegs the two labs’ annualized run rate at roughly $100 billion, a growth pace with no real precedent in software history.
- Forecasters Bet Semiconductor Stocks Cool While Software Heats Up. A panel of experts and superforecasters expects chip stocks to lag their own trend line while software stocks beat theirs, with wide uncertainty on both.
- DeepSeek Nears a Funding Round at a $74 Billion Valuation. The Hangzhou lab known for cheap models is reportedly raising billions to fund compute as it eyes a Shanghai listing.
- a16z Launches $1.1 Billion Fund to Bet on AI’s Physical Layer. Machine Age shifts a16z’s focus from software margins to the chips, power, and buildout underneath AI models, a departure from the firm’s usual software bets.
Evaluation Gets Harder: Benchmarks, Blind Tests, and a Broken Guardrail
Three stories today complicate the question of how much to trust a benchmark score or a safety guardrail.
- Claude Opus 5 Tops New Science Benchmark at Just 30% Resolution. Stanford researchers built the eval from real scientific workflows, and every model tested resolved fewer than a third of the 70 tasks.
- DeepMind pilots what it calls the first double-blind AI evaluation. A cryptographic box hides benchmark prompts from Gemini Flash Lite, targeting the contamination problem that inflates published AI scores.
- Researcher Cracks Claude Code Auto Mode Defense Most of the Time. Simon Willison flags a Johann Rehberger exploit that beat Anthropic’s default coding-agent guardrail in most attempts.
Enterprise Tooling: Anthropic and Cohere Court Developers
Two moves today aim squarely at enterprise developers rather than consumers.
- Anthropic Opens a Standard for Agents to Run Lab Hardware. The Model Hardware Standard lets AI agents drive microscopes, liquid handlers, and robotic arms, and Anthropic is testing it with a small circle of partners before any wider release.
- Cohere prices document parsing at $1.50 per 1,000 pages. Parse, a vision language model for enterprise document extraction, undercuts hyperscaler OCR pricing in an already crowded field.
Quick Hits
- OpenAI’s GPT-5.6 discounts show up as usage, not just price. OpenRouter’s routing data ties steep discounts on two OpenAI models to a 13.8x usage spike, and about a third of the new users stuck around after prices reverted to normal.
- New US Businesses Keep Spreading Out. AI Labs Refuse To. Stripe payments data shows startup formation scattering into smaller metros even as the labs building the tools that enable that scattering stay locked into San Francisco.
- Thinking Machines ditches scaffolding, trains a model past human SQL accuracy. A fine-tuned Kimi-K2.6 model called ReViSQL-K2.6 scored 92.97 percent on an expert-verified SQL benchmark, edging past the 92.96 percent human mark without any agentic pipeline wrapped around it.
- Codex gives persistent reasoning effort its own wire value. A new SDK variant means custom model providers can no longer count on the old pass-through behavior for that reasoning-effort setting, a small but breaking change for integrators.
- Halo Neuro Open-Sources a Voice Clone Small Enough for a Browser. Sopro V2 Turbo, a 120-million-parameter model, streams cloned speech from a laptop CPU and beats several rivals many times its size on Halo Neuro’s own tests.
- OpenAI Hires Away Meta’s India and Southeast Asia Chief. Sandhya Devanathan leaves Meta after a decade to run growth, partnerships, and regulatory engagement for OpenAI across Southeast Asia and Australia, working out of Singapore.