Four threads run through today’s edition. New models keep shipping with the vendor grading its own homework. xAI held pricing flat for Grok 4.7, Xiaomi open sourced a model it says matches Claude at a fraction of the price, and Alibaba showed off a chip it says triples its old one’s speed. None of it has independent confirmation.
New Models, Graded by the Companies That Built Them
xAI, Xiaomi and Alibaba all put out new models or chips this week, each with a benchmark table the company wrote itself. None of the numbers below have been checked by anyone outside the building that produced them.
- xAI ships Grok 4.7 at the same price, with a scorecard it didn’t win. xAI held pricing flat for its new coding model and released benchmarks, its own, showing Grok 4.7 losing three of seven comparisons to Fable 5.1. It is betting flat pricing and coding strength matter more than a clean sweep.
- Xiaomi open sources a model it says beats Claude and GPT-5.6 Sol on price. Xiaomi’s MiMo-V2.6 Pro claims to match Claude Opus 5 and GPT-5.6 Sol on agent benchmarks while costing a fraction of their API price, according to Xiaomi’s own release materials.
- Xiaomi says a smarter grader, not more attempts, fixed its coding model’s training. Xiaomi’s technical report says grading code quality instead of simple pass or fail stopped its agents from gaming their own tests, letting reinforcement learning keep improving the model for longer.
- Alibaba unveils a new AI chip and says it triples its old one’s speed. Alibaba’s stock rose after it showed the Zhenwu V900 chip, said it triples the prior chip’s speed, and laid out plans to grow its cloud data centers past 20 gigawatts by 2032.
Rumors, Theories, and the Race Nobody Can Fully See
None of this is announced. Claude users say they’ve spotted an unreleased model, and three independent writers each stake out a different read on who’s actually ahead and why the labs behave the way they do.
- Claude users say they’ve spotted an unreleased Anthropic model. Screenshots on X show Claude producing outputs users say look far more advanced than the current release. Talk of a model called Opus 5.5, including a rumored price, is still just talk, and Anthropic has not confirmed or denied any of it.
- China’s open AI models are pulling ahead of America’s open models. Researcher Nathan Lambert’s data shows Chinese open-weight models now taking over 80 percent of weekly token volume on OpenRouter and closing in on the closed frontier, while American open models fall further behind both.
- A writer says AI models are about to split into cheap and pricey parts. Jaya Gupta argues that as agents multiply, apps will route routine decisions to cheap models and save frontier AI, the kind Anthropic reportedly prices at high margins, for the problems nothing else can answer.
- Why AI labs are racing into ads, robots and cancer research. An essay argues OpenAI and Anthropic’s real technical lead over rivals is only one or two model generations, a gap narrow enough that both labs are chasing new revenue lines to protect it before it closes.
Agents Get More Rope, and Show Where It Runs Out
Meta’s Muse app proved people will install an AI agent, right up until a retailer decided it couldn’t be trusted at checkout. Elsewhere, new harnesses and benchmarks are testing what autonomy actually buys.
- Meta’s Muse agent app beat ChatGPT on downloads, then hit an Amazon wall. Muse pulled 730,000 iOS downloads in five days and passed Claude and Grok, but Amazon blocked it from its store, saying the app tried to check out without warning and captured customer credentials.
- Cognition brings its Devin coding agent’s cloud sessions into the terminal. Devin CLI can now hand a task to a cloud virtual machine and pull it back with a single command, and Cognition is giving away sessions with its SWE-2 model through October 8.
- Qwen’s new agent benchmark shows top models still fail most rebuilds. Alibaba’s RecreationWorld test asks agents to rebuild working software from scratch. Even the top scorer, OpenAI’s GPT-6 Astra, produced a flawless rebuild in under 3 percent of cases.
- AWS ships an agent harness that runs any AI model you pick. AWS’s new Strands harness bundles memory and tool use out of the box and says it runs 26 percent cheaper than rival setups, but leaves the choice of underlying model entirely up to the developer.
Who’s Actually in Control: Swarms, Referees, Money and a Spacecraft
A researcher’s math questions whether bigger agent swarms are worth their cost, mathematicians just won formal standing to publicly call out OpenAI’s own claims, investors are pricing a bet on self-improving AI, and one startup’s CEO wants to fly with no way to command his spacecraft from Earth.
- Adding more AI agents buys speed, not smarts, new math shows. Independent researcher Toby Ord’s analysis of OpenAI’s own published data finds that doubling a swarm’s agent count mostly finishes a task faster, without meaningfully raising what the system can actually solve.
- Mathematicians push back on OpenAI, so it built an outside advisory group. OpenAI says an internal model has cracked more than 100 open math problems, and has now given nine unpaid mathematicians standing to publicly criticize how the company communicates results like that one.
- Ex-Anthropic team’s startup Mirendil in talks at $5 billion, sources say. Mirendil is in talks for a new round that would value the company at five times its $1 billion seed price from three months ago, Bloomberg reported, with no deal yet closed.
- AstroForge will let its own AI model fly the next asteroid mission. After two prototype spacecraft failed to complete their missions, the startup’s CEO says he wants to fly a 2027 mission with no way to command it from Earth, handing real-time flight decisions to its in-house model, Solo.
Quick Hits
- Tomasz Tunguz says tiny AI models are quietly replacing old code rules. Venture capitalist Tomasz Tunguz swapped a quarter of his AI agent’s classification calls for narrow “decider” models and says accuracy jumped from 47 to over 80 percent while cost fell by roughly 99 percent in his own tests.
- StepFun’s Step 5 matches Kimi K3 at far lower cost, testing finds. Artificial Analysis testing puts StepFun’s new Step 5 Preview model at the same benchmark score as Moonshot’s Kimi K3, reaching it for roughly 2.8 times less cost per task, though it still trails on agentic tasks.
- Aikido shrinks a 1.5TB security model to run entirely on your own servers. Aikido compressed the open-weight GLM-5.3 model down to 328GB for its new Altar pentesting model, letting security teams run frontier-grade testing fully inside their own network, even air-gapped, while keeping most of the original’s flaw detection, the company says.
- Developer open sources small AI models that answer yes/no questions locally. Jared Palmer’s open source Kev models run locally on a laptop or single GPU and return a confidence score with each yes or no answer, copying a paid hosted service’s format at no cost, though Palmer’s own tests show it trailing slightly on accuracy.