DeepSeek has open-sourced a set of programming tools that let the code behind its AI training run on Huawei’s Ascend chips, the Chinese lab said in a post on its official WeChat account on 30 September. Geopolitechs, the newsletter that analyzed the announcement, calls the target the software gap that has held Chinese hardware back, not the silicon itself.
Nvidia’s lead in training large models rests on more than its processors. Its CUDA software gives developers mature ways to write programs, tune speed, and make thousands of chips cooperate. Huawei’s Ascend line has had credible hardware but a thinner toolbox, which makes it hard for engineers to extract what the chips can really do.
DeepSeek, the Hangzhou-based lab behind the V4 model family, says it has now filled part of that toolbox. The centerpiece is TileLang, a high-level language that lets a programmer describe what a chip should do without hand-writing the low-level instructions. TileLang, which ships with its own compiler toolchain, anchors the package. The rest is maths code and the networking layer that lets many chips work together. DeepSeek says TileLang already handled the majority of operators when it trained the V4 family, and that each TileLang operator in its training work now has a fast Ascend counterpart. Among the libraries are ones for matrix math, communication between chips, long-text attention, and data selection. DeepSeek says each piece mirrors one it previously released for Nvidia hardware.
The performance claim belongs to DeepSeek alone. The company says that in several key tests, these components come close to the ceiling of what the Ascend hardware allows. Neither the announcement nor Geopolitechs’ write-up cites independent measurements, so outside developers will have to confirm that on their own machines.
The strategic idea, in Geopolitechs’ reading, is a buffer layer between model and hardware. If DeepSeek’s team writes in TileLang, the same code could in principle target either Nvidia GPUs or Ascend chips. Geopolitechs says that, if it works, switching a workload from Nvidia to Chinese hardware would get cheaper and less painful. That is the analyst’s inference about DeepSeek’s intent, not a stated goal in the post, which speaks instead of building “independent and controllable” software ecosystems.
Huawei is a visible partner. DeepSeek thanks Huawei’s engineers for what it calls strong, wholehearted support, and says the two companies are jointly building a 128-card supernode on the Ascend 950 chip while tuning both computation and chip-to-chip traffic. Put plainly, a supernode is a tightly linked cluster designed to behave like one large machine.
The announcement stops short of one claim that would matter most. DeepSeek does not say, in the text Geopolitechs reproduces, that V4 was trained on Ascend hardware. It says the software now exists on Ascend, which is a necessary step but a smaller one. Training a frontier model across thousands of chips depends on reliability and supply as much as on code, and neither is addressed.
The competitive angle is plain. Geopolitechs frames the underlying worry as what Chinese labs do if they can no longer buy Nvidia’s most advanced GPUs, and CUDA is what makes those chips hard to replace. Every tool that lowers the cost of leaving CUDA weakens that leverage a little, and DeepSeek is publishing the tools for anyone to use.
For developers and buyers watching Chinese AI capacity, the number to follow is not the announcement but adoption: how many outside projects start building on the new Ascend repositories on GitHub over the next quarter.
Reported by Geopolitechs on 30 September 2026, analyzing DeepSeek’s announcement on its official WeChat account.