
DeepSeek has opened up a piece of its playbook for running artificial intelligence on Chinese-made silicon, releasing a free software package built to make Huawei’s chips work more like Nvidia’s without ever touching Nvidia’s code. The move, announced on September 30, 2026, marks one of the most concrete steps yet toward loosening China’s grip-and-grip-back relationship with Nvidia’s dominant chip programming standard. At its core, the DeepSeek Huawei AI software release is an attempt to give developers a real reason to build on Chinese hardware instead of defaulting to what everyone already knows.
Key takeaways
DeepSeek open-sourced a full programming toolkit for Huawei’s Ascend AI accelerators, announced via its WeChat account on September 30, 2026.
The package includes TileLang, a high-level language positioned as a domestic alternative to Nvidia’s CUDA, plus libraries such as DeepGEMM, FlashMLA, TileKernel, DeepSelect and DeepEP.
The tools are optimized for Huawei Ascend 950 chips and support a “supernode” setup linking up to 128 accelerators.
Huawei reportedly gave extensive support during development, according to Bloomberg’s Luz Ding.
Separately, DeepSeek has plans for an Inner Mongolia data center capable of housing no fewer than 160,000 Huawei Ascend accelerators.
DeepSeek releases open-source software for Huawei’s Ascend AI chips
DeepSeek‘s announcement, made through its WeChat account and reported by Bloomberg’s Luz Ding, confirmed that the Chinese AI startup has made the entire toolkit open-source and free to download. That detail matters: giving the software away removes one of the biggest arguments developers have for sticking with familiar, paid or license-bound alternatives.
At the center of the release sits TileLang, described as China’s answer to CUDA — the software platform widely treated as the global standard for programming AI chips. TileLang gives developers a high-level abstraction layer, meaning programmers who already know Nvidia’s workflow can, in theory, shift toward Huawei hardware without having to learn a completely unfamiliar system from scratch.
A stack built in layers, not a single tool
Beyond TileLang, DeepSeek packaged a set of compute and communication libraries — DeepGEMM, FlashMLA, TileKernel, DeepSelect and DeepEP — each covering a different piece of the AI development chain, from matrix multiplication to memory handling to communication between chips. The release also includes benchmarking tools meant to help developers with kernel development, giving them a standardized way to measure performance and fine-tune code for Huawei’s architecture. Put together, the package reads less like a single product and more like an attempt to replicate, layer by layer, the kind of ecosystem CUDA has spent years building around Nvidia hardware.
Why the Ascend 950 chip and its “supernode” setup matter
The toolkit is optimized specifically for Huawei’s Ascend 950 chips and supports what DeepSeek calls a “supernode” configuration, linking up to 128 of those accelerators. That scale signals an intent to support serious, large-batch AI workloads rather than small experimental projects — the kind of setup needed to train or run cutting-edge models at competitive speed.
This matters because software without adoption is just code sitting on a server. CUDA’s dominance isn’t really about its technical elegance; it’s about the sheer size of the ecosystem built around it over more than a decade. Rewriting code, retraining engineering teams, and rearchitecting entire workflows to move away from CUDA have always felt too costly for most developers. By making its toolkit free and modeling TileLang closely on familiar CUDA-style workflows, DeepSeek is explicitly trying to shrink that switching cost, giving Chinese developers — and potentially others — a lower-friction path toward Huawei’s chips.
Huawei’s hand in the project and the CUDA dependency problem
According to Bloomberg’s reporting, Huawei Technologies provided extensive support throughout the toolkit’s development, underscoring how closely the two companies are working together on technology positioned as an alternative to Nvidia. That partnership isn’t happening in a vacuum. It’s a direct response to a supply problem that has been building for years.
Nvidia’s CUDA framework isn’t just popular in AI circles — it’s close to a requirement. Because nearly every AI model, training pipeline, and inference system across the globe depends on it, that dependence turned into a real strategic vulnerability for China after the United States enacted export controls limiting access to Nvidia’s top-tier chips. Beijing could push companies like Huawei to design competing silicon, but hardware without a mature software layer to program it risks becoming, in practice, very expensive dead weight. This is exactly the gap the new toolkit is meant to close, giving Huawei’s Ascend chips a fighting chance at real-world usability outside of Nvidia’s ecosystem.
Building the infrastructure to match the software
Software alone doesn’t create an ecosystem, and DeepSeek appears to be moving on that front as well. With plans underway for a sizable Inner Mongolia facility built to accommodate a minimum of 160,000 Huawei Ascend accelerators, the company is set to give the new toolkit a genuine stress test rather than treating it merely as a hypothetical exercise.
The release also builds on groundwork DeepSeek laid earlier in 2026. That April, the firm modified its V4 AI model to operate on Huawei hardware, thereby demonstrating that a state-of-the-art model was capable of functioning on chips other than Nvidia’s. The toolkit released on September 30 generalizes that earlier effort, handing any developer the means to attempt the same migration with their own models rather than relying on DeepSeek to do it for them.
Taken together, the sequence — a working model on Huawei hardware in April, followed by a full open-source programming stack in September — suggests DeepSeek is trying to build proof of concept and infrastructure in parallel, rather than waiting for one to be finished before starting the other. Whether that translates into broad developer adoption outside DeepSeek’s own projects remains an open question, but the pieces being assembled point toward a more serious, sustained push to make Huawei’s Ascend chips a viable production environment rather than just a fallback option.
Article produced with the assistance of artificial intelligence and reviewed by the editorial team.
