TechTimes reported on 6 October that the gap between US and Chinese AI has fallen to 3 percent, citing a note by Robert Lea, a senior analyst at Bloomberg Intelligence. Read closely, the figure describes the distance between two specific models, on one benchmark’s overall score, on a single day. It does not describe two countries.
Here is what the article actually gives. On the 4 October snapshot of LiveBench, a public leaderboard, DeepSeek’s V4.1 Flash scored 81.1 overall. Anthropic’s Claude Fable 5.1, run at its maximum effort setting, scored 83.4. That is 2.3 points apart, which TechTimes says rounds to roughly 3 percent of Anthropic’s score. So the number is a relative difference between a best Chinese model and a best American model, not three points and not a share of capability.
The article leaves several things open. It does not say who runs LiveBench or whether that organization has any stake in the result; it only describes the design, with questions refreshed monthly from recent competitions, papers and news. It does not say whether Lea’s note states the percentage itself or whether TechTimes did the rounding. The earlier gaps it quotes, about 9 percent in May and roughly 15 percent before that, come with no stated models or method, and since LiveBench swaps its questions each month, those snapshots are not the same test.
The coding claim needs the same care. Agentic coding means a model is handed a software task and works through it largely alone: it writes code, runs it, reads the errors, and tries again until the thing works. On LiveBench’s agentic coding sub-score, TechTimes reports V4.1 Flash at 77.3 against 66.1 for that same Anthropic model. The article cites no OpenAI or Google result on this leaderboard, so the lead is over one named competitor on one sub-score.
A second figure comes from DeepSeek’s own technical report: 74.2 percent on DeepSWE v1.1, another agentic coding test, against 74.0 percent for Anthropic’s Opus 5 and 73.0 percent for OpenAI’s GPT-5.6 Sol. Those are the developer’s numbers, the margin over Opus 5 is 0.2 points, and TechTimes says no independent validation has been published.
A different yardstick gives a different answer. The article notes that the US government’s NIST, through its CAISI evaluation unit, tested DeepSeek’s earlier V4 Pro in May on nine benchmarks fixed in advance, and put the gap at about eight months of development. That is a measure of time, on a previous model, by a government body using a method set in advance. TechTimes itself says the two findings are not in conflict because they measure different things.
On cost, TechTimes lists peak prices of 30 cents for every million tokens fed in and $1.20 for every million produced. It also says the weights are open, so a team can host the model itself rather than send code to DeepSeek’s API, which the article flags as subject to Chinese data-access law.
Anyone who sees “3 percent” next week should ask for three labels: LiveBench composite, 4 October, one model against one model. Teams weighing V4.1 Flash for coding should run it on their own repository before a leaderboard decides for them.
Reported by Devin Culbertson for TechTimes, published 6 October 2026.