A Huawei-led consortium has finally attached hard numbers to a claim that spent six weeks circulating without them. In a July 22 technical report, the group says it ran full-parameter post-training on DeepSeek’s V4 family, including the 1.6-trillion-parameter V4-Pro, entirely on Huawei’s Ascend NPU SuperPOD hardware. The headline figure is 34.22% model FLOPs utilization (MFU), the share of a chip’s peak theoretical throughput actually converted into productive training work. The team says that is nearly three times more efficient than an open-source baseline recipe, a 2.93-fold gain by its own measurement.
Implicator.ai reported July 23 that the number matters mainly because of what preceded it. Tom’s Hardware had described the original June 6 announcement as arriving with no benchmarks, no run duration and no head-to-head Nvidia comparison, and grouped it with what the outlet called, in its own words, “a series of dubious claims.” DeepSeek did not comment at the time. Wednesday’s paper supplies the missing efficiency figure and baseline comparison. Both numbers still come from the Huawei-led team itself; no outside lab has replicated them.
The report does not resolve the argument over Nvidia’s role. It covers post-training only and says nothing about which chips trained V4 in the first place. Liu Zhiyuan, a Tsinghua University computer-science professor, told MIT Technology Review that DeepSeek “appears to have adapted only part of V4’s training process for Chinese chips” and may have relied on Nvidia hardware for most of the original run. People described only as sources with knowledge of the work added that Huawei’s chips currently perform better at inference than at training. Notably, the report never places its 34.22% MFU figure next to what a comparable Nvidia cluster achieves, so readers cannot judge whether the result is efficient or merely functional.
Some outside detail supports the scale of the effort. Separate reporting from June, citing figures from Shenzhen’s municipal government via the South China Morning Post, put the cluster at 1,000 or more Ascend 910C chips. In earlier DeepSeek testing, the 910C delivered about 60% of what an Nvidia H100 achieves in inference workloads, a comparison that says nothing about training efficiency.
Whether Chinese labs can train frontier models competitively on domestic silicon is the question underneath every U.S. export control aimed at Nvidia and AMD chips. A report that a Huawei-organized consortium wrote, reviewed and published itself is evidence toward answering that question. It is not independent verification of it, and its scope, post-training rather than the far more compute-intensive pre-training run, means it answers a narrower version of the question than the one Washington is actually asking.
The timing sits alongside a separate story: this week’s reporting that U.S. officials are investigating Chinese AI firms’ access to advanced processors and have signaled sanctions or blacklisting could follow. That inquiry concerns chip access and training practices broadly. It is not a response to this specific MFU figure, and treating the two as one story would overstate what either has actually shown.
Teams weighing Chinese open-weight models on cost or availability grounds should read the 34.22% figure as a data point on post-training economics on Ascend hardware, not as proof the pre-training compute gap with Nvidia has closed.
Reported by Implicator.ai on July 23, 2026.