China’s open weight labs shipped bigger models and bigger business numbers today, while Apple and a Google DeepMind researcher escalated pressure on the West’s AI giants from opposite directions.
Beijing’s Open Weight Sprint: Bigger Models, Faster Releases
Alibaba and Moonshot both raced to out scale each other and Western labs this week, packing in more parameters and shipping developer tools faster than demand for either can be met.
- Alibaba’s Qwen3.8 Pushes Open Weights to 2.4 Trillion Parameters. Alibaba announced Qwen3.8, a 2.4 trillion parameter model headed for an open weight release, days after Moonshot’s 2.8 trillion parameter Kimi K3. A preview is already live through Alibaba’s Token Plan, Qoder, and QoderWork.
- Moonshot’s Kimi Code CLI Mirrors Claude Code’s Extension Model. Kimi Code is a single binary terminal coding agent that reads and edits code, runs shell commands, and calls MCP tools. Its subagents and hooks track Claude Code’s design closely enough to lower the switching cost for developers already using it.
- Moonshot Pauses Kimi K3 Signups as Demand Outstrips Its GPUs. Moonshot is capping new Kimi K3 subscriptions and redirecting compute to existing users, days before the model’s July 27 open weight release.
China’s AI Money Machine: Free Weights, Real Revenue
Giving models away is turning into a business model of its own, and Alibaba’s chip strategy now looks aimed at Nvidia’s software moat rather than just its silicon.
- Z.ai Nears $1 Billion in Sales While Giving Away Its Best Models. A large share of Z.ai’s revenue comes from on premises deployments for state owned enterprises and financial institutions, plus a fast growing cloud business, evidence that open weights can work as a sales funnel rather than a subsidy.
- Moonshot AI Targets a Hong Kong IPO at a $30 Billion Valuation. The Kimi K3 maker plans to list on the Hong Kong Stock Exchange within six months at a valuation near $30 billion.
- Alibaba Open Sources Zhenwu Chip Software to Loosen CUDA’s Grip. T-Head’s SAIL release, unveiled at WAIC in Shanghai, targets the software layer that locks developers into Nvidia rather than the chips themselves, an approach meant to make China’s AI ecosystem harder for any single government to shut down.
Legal and Governance Friction Builds in the West
Apple’s trade secret fight with OpenAI escalated and a Google DeepMind researcher walked away from a military AI deal, both signs that legal and ethical costs are catching up with the labs building frontier AI.
- Apple Widens Its OpenAI Trade Secret Fight With Preservation Letters. Apple sent legal letters to roughly 40 former employees now at OpenAI, instructing them to preserve documents related to a lawsuit accusing OpenAI of pulling hardware secrets and talent from Apple. OpenAI says it has seen no evidence the complaint has merit.
- Google DeepMind Researcher Quits Over Unrestricted Military AI Deal. Alex Turner resigned after failing to stop a Google contract that lets the Department of War use its AI models without limits on autonomous weapons.
Agent and Harness Engineering: What Actually Works in Production
Three separate teams found that more autonomy and more spend do not automatically produce better results, from a goal setting feature that can amplify bad decisions to a research pipeline that burned a subscription’s limit in half an hour.
- Claude Code and Codex Handle /goal Differently, Neither Handles It Well. A matched benchmark on an unpublished NP-hard problem found /goal winning most trials while making average results worse for both harnesses. Claude Code implements it as a session scoped Stop hook; Codex persists it as thread state and effectively grades its own work.
- Netflix Built Its Own LLM Inference Stack Instead of Buying APIs. Netflix deployed a vLLM and Triton inference platform inside its existing production infrastructure rather than pay for hosted API access, and production traffic exposed version pinning bugs and a GIL bottleneck along the way.
- Quesma’s Fix for AI Research Costs: Cheap Models First, Deep Research Last. One engineer burned a Claude Max plan’s limit in 30 minutes with zero output, then rebuilt the pipeline so cheap models search, pricier models only verify, and deep research runs last.
Model Access and Research Keep Widening
Anthropic, Google, and Sakana AI each moved on a different axis this week: who gets access to a model, where it runs, and how it learns in the first place.
- Claude Fable 5 Joins Max and Team Premium Plan Limits July 20. Anthropic is folding Fable 5 into subscription limits at 50 percent of allowance after admitting demand outran its own capacity forecasts. Pro and Team Standard users keep access through usage credits and a one time $100 credit.
- Sakana AI Trains Neural Nets Without Breaking a Core Brain Rule. Sakana’s Diffusing Blame method lets networks learn while obeying Dale’s principle and skipping backpropagation’s biologically implausible weight transport, using non-negative weight matrices to keep separate excitatory and inhibitory streams.
- Google Tests Gemini Live on Web With Portable Skills for All Chats. Pre-release builds show real time voice moving beyond mobile and Skills escaping the Spark subscription tier, though Google has not confirmed either change.
Today’s Quick Hits
- A $400 Million Loan Bets Non-Nvidia Chips Hold Their Value. Upper90 lent General Compute $400 million against inference specific chips instead of GPUs, testing whether non-Nvidia silicon holds resale value.