DeepSeek Opens AI Software Stack for Huawei Chips in Fresh Challenge to Nvidia’s Ecosystem
Chinese AI company DeepSeek is expanding its work with Huawei by open-sourcing a broad set of software components designed for Huawei's Ascend AI processors, taking a deeper step toward building an AI computing stack that can operate beyond Nvidia's dominant ecosystem.
The release, announced on September 30, covers DeepSeek's TileLang programming framework, compute libraries and distributed communication tools for the Ascend platform.
DeepSeek says the new components correspond directly with software it had previously developed for Nvidia-based systems, creating a path for similar AI workloads to run across different accelerator architectures. Reuters reported that the move is part of a broader effort by Chinese technology companies to reduce reliance on Nvidia's software ecosystem.
The collaboration also extends to hardware infrastructure.
DeepSeek and Huawei are jointly developing a 128-chip supernode based on Huawei's Ascend 950 processors, with the companies working on both computing performance and communication between accelerators.
That matters because training increasingly large AI models requires more than powerful individual chips. Hundreds or thousands of processors have to work together efficiently, and communication between those processors can become a major bottleneck.
DeepSeek is opening more than model compatibility
The significance of the release goes beyond making a DeepSeek model compatible with Huawei hardware.
DeepSeek is opening software components closer to the hardware itself.
The September 30 release includes TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA and DeepSelect, covering programming, matrix computation, inter-device communication, memory and vector operations, attention mechanisms and data selection.
Together, these tools form a substantial portion of the software infrastructure required to optimise AI workloads.
That is important because one of Nvidia's biggest advantages is not simply its GPUs.
It is the software ecosystem surrounding them.
Nvidia's CUDA platform has become deeply embedded in AI development, giving developers tools and libraries that allow them to write and optimise applications for Nvidia hardware.
DeepSeek's latest move targets a similar layer of the technology stack.
TileLang is at the centre of the strategy
One of the most significant components in the release is TileLang, DeepSeek's open-source high-level programming language for AI computation.
The idea is to allow developers to express complex AI operations without having to work directly with every low-level hardware instruction.
DeepSeek says its Ascend implementation can provide a higher-level programming approach while still allowing developers to exploit the underlying capabilities of Huawei's chips. The company also says every TileLang operator used in its training has a corresponding high-performance implementation on Ascend.
That portability could become increasingly important as AI companies work with multiple accelerator architectures.
Instead of rebuilding an entire software stack whenever the underlying hardware changes, developers can potentially maintain a common programming approach while adapting the lower layers to different chips.
That is one of the reasons the release is more significant than a conventional hardware port.
Six pieces of the Ascend software stack
DeepSeek's release covers several different parts of the AI computing pipeline.
DeepGEMM accelerates general matrix multiplication, one of the fundamental operations behind modern AI models.
DeepEP focuses on communication between accelerator devices, particularly important for distributed AI workloads where different processors need to exchange data efficiently.
TileKernels provides commonly used vector-computation and memory-access operations.
FlashMLA provides specialised attention operations designed for DeepSeek's model architectures and workloads.
DeepSelect handles efficient data-selection operations.
TileLang provides the higher-level programming layer connecting developers to the underlying hardware.
The result is a broader software toolkit rather than a single compatibility layer. DeepSeek said that performance in several key tests approached the hardware limits of the Ascend platform. Those performance figures are company-reported rather than an independent benchmark.
DeepSeek and Huawei are also building a 128-chip supernode
The partnership is moving beyond software.
DeepSeek said the two companies are jointly advancing a 128-card supernode based on Ascend 950 processors, with optimisation work covering both computation and communication.
The distinction is important.
A large AI system is only as effective as its weakest bottleneck.
A processor can deliver high theoretical computing performance, but if data cannot move between chips quickly enough, the overall system may fail to achieve that potential.
As models become larger and training workloads become increasingly distributed, networking and memory movement can consume a significant portion of the system's performance budget.
DeepSeek's collaboration with Huawei therefore targets the entire computing system rather than simply the processor.
Why this matters for Huawei
For Huawei, the partnership provides something potentially more valuable than another hardware customer: real-world optimisation from a major AI model developer.
Huawei has been rapidly expanding its Ascend ecosystem as China seeks alternatives to Nvidia hardware.
The company unveiled the Atlas 960E SuperPoD earlier this month and said its systems are being designed to support training and inference for models with up to 10 trillion parameters. Huawei also said its upgraded TaiShan 950 SuperPoD can support systems containing as many as 4,096 NPUs.
But hardware is only one part of building a competitive AI platform.
Developers also need compilers, libraries, communication frameworks, debugging tools and optimised kernels.
DeepSeek's work provides Huawei with an important software ecosystem contribution while giving DeepSeek another hardware platform on which to develop and deploy AI workloads.
Why this matters for DeepSeek
The benefits run in the other direction too.
DeepSeek has become one of China's most closely watched AI developers, meaning its computing requirements are significant.
Supporting Huawei's Ascend platform gives the company greater hardware flexibility.
That does not mean DeepSeek has suddenly abandoned Nvidia.
The more accurate description is that DeepSeek is building software that can work across multiple accelerator environments.
That distinction matters.
The strength of Nvidia's ecosystem is partly its maturity and developer adoption. Replicating that advantage requires more than making a few libraries compatible with another chip.
It requires developers to actually use the tools, maintain them and build applications around them.
Open-source distribution could help with that process by allowing researchers and companies outside DeepSeek to inspect, adapt and contribute to the software.
China is building more of the AI stack at home
The broader significance is geopolitical as well as technical.
Restrictions on advanced semiconductor exports have pushed Chinese technology companies to develop greater domestic capability across AI hardware and software.
That creates pressure to solve multiple problems simultaneously.
China needs competitive AI processors.
It needs manufacturing capacity.
It needs high-speed networking.
It needs compilers and development frameworks.
And it needs AI models that can be efficiently trained and deployed on domestic hardware.
DeepSeek's latest release sits directly at the intersection of those efforts.
Instead of treating the chip and the model as separate technologies, the partnership is bringing AI models, software, processors and system architecture closer together.
That could eventually make China's AI ecosystem less dependent on any single foreign hardware or software platform.
But building a complete alternative is a much larger task than producing one successful AI model or processor.
The software battle may be as important as the chip battle
Nvidia's competitive position demonstrates why.
A powerful AI accelerator is difficult to sell at scale if developers cannot easily program it or move existing workloads onto it.
CUDA has become a foundational component of modern AI development precisely because it gives developers an extensive software environment around Nvidia hardware.
DeepSeek's decision to open-source tools for Huawei's Ascend platform therefore highlights a broader battle over who controls the software layer connecting AI models to AI chips.
If developers can use common tools across multiple hardware platforms, hardware choice becomes more flexible.
That could make it easier for Chinese AI companies to adopt domestic accelerators without having to completely rewrite their software for every new chip architecture.
It could also create opportunities for other chipmakers to compete without having to reproduce Nvidia's entire ecosystem independently.
The most important part of DeepSeek's Huawei partnership may not be the 128-chip supernode.
It is the decision to open-source the software required to make alternative AI hardware useful at scale.
The AI semiconductor race is often described as a battle between chips.
But chips alone do not run modern AI systems.
They need compilers, kernels, libraries, communication systems and developer tools.
That is where Nvidia's ecosystem has historically been powerful.
DeepSeek is now helping Huawei build out an alternative path by bringing its own AI software expertise to Ascend hardware.
The immediate result is not the disappearance of Nvidia from China's AI ecosystem.
It is something more gradual: a broader Chinese AI stack in which models, software and hardware are increasingly designed to work together without depending on a single foreign platform.
If that ecosystem attracts enough developers and achieves competitive performance at scale, the long-term impact could extend well beyond DeepSeek and Huawei.
The real test will be whether the software can turn alternative chips into platforms developers actually want to build on.
