DeepSeek Publishes AI Tools for Huawei Chips, but Hardware Access Limits Their Reach
An October 7 report details narrow hardware requirements and benchmarks run on a nonpublic test kit. Familiar programming interfaces offer a migration path, not proof of a performance advantage.
DeepSeek’s new Ascend libraries bring CUDA-familiar interfaces to Huawei’s AI stack: DeepGEMM-Ascend covers computation, while DeepEP-Ascend handles data routing and communication for mixture-of-experts workloads. The practical opening is easier porting, not demonstrated Nvidia-beating performance: both require Ascend 950 hardware, DeepEP also needs UBMEM, and DeepGEMM needs CANN 9.20. Published benchmarks used a proof-of-concept kit DeepSeek says was not publicly distributed, limiting reproducibility and near-term reach.
01
DeepGEMM-Ascend supports BF16, FP8 and FP4 operations and retains the Python package name and interface structure of its CUDA counterpart.
02
DeepEP-Ascend routes data among mixture-of-experts components and requires Ascend 950 chips with UBMEM connectivity.
03
GitHub snapshots showed 515 stars for DeepGEMM-Ascend and 224 for DeepEP-Ascend; those counts do not establish deployments or hardware switching.
Developers have a more familiar route to running DeepSeek workloads on Huawei chips—but access to the right hardware remains a constraint. DeepSeek released two open-source libraries with Huawei’s support. TechRadar’s October 7 coverage details the limits: both target Ascend 950 hardware, and published benchmarks used a test kit DeepSeek said was not publicly distributed.
Familiar interfaces for a different chip
The release addresses two linked jobs in large AI systems: doing calculations on individual processors and moving data between them. Tom’s Hardware described the Huawei-supported release on October 1, following a September 30 Reuters report. The companies also optimized computation and communication on a system built around 128 Ascend 950 chips.
DeepGEMM-Ascend handles matrix multiplication and other calculations used in DeepSeek models. It supports BF16, FP8 and FP4 operations—different numerical formats—and uses the same programming interfaces as the existing DeepGEMM library.
DeepEP-Ascend handles communication during training and inference, the processes of building and running a model. For mixture-of-experts models, it routes data to the different expert components and combines their outputs.
Nvidia’s software advantage extends beyond its processors. CUDA, its platform for programming GPUs, supplies a mature development environment and libraries optimized for that hardware. DeepSeek’s tools address the corresponding software work on Ascend: giving developers ways to optimize calculations and communication, rather than simply providing another chip on which to run a model.
The interface choices connect the new tools to their Nvidia counterparts. DeepGEMM-Ascend retains its CUDA sibling’s Python package name and interface structure. DeepEP-Ascend aligns its public buffer interfaces with the Nvidia version. TechRadar characterizes this approach as reducing switching costs, rather than establishing a performance advantage.
The installation boundary
The requirements are specific. DeepGEMM-Ascend needs Ascend 950-series hardware and version 9.20 of CANN, Huawei’s underlying software platform for AI workloads. DeepEP-Ascend requires Ascend 950 chips with UBMEM connectivity. Both libraries were developed and tested on Ascend 950, rather than presented as general-purpose tools for every Huawei accelerator.
DeepSeek says Huawei provided full support during development. The company described the tools as intended to simplify programming while helping developers use the hardware’s performance. Those are development goals, not independently measured findings that the release outperforms Nvidia’s software or hardware.
Early public interest is much smaller than for the Nvidia-oriented versions. TechRadar recorded 515 GitHub stars and 39 forks for DeepGEMM-Ascend, against 8,522 stars and 1,359 forks for DeepGEMM. DeepEP-Ascend had 224 stars, while DeepEP had more than 10,200. These are snapshots of repository activity, not counts of working deployments or evidence that developers have switched hardware.
Two new repositories, several existing projects
The release also needs separating from the older software around it. DeepGEMM-Ascend and DeepEP-Ascend are new repositories. TileKernels, DeepSelect and FlashMLA are existing projects that received Ascend-targeted updates. Counting every named component as a newly created tool would overstate what changed.
TileLang, the higher-level language for writing optimized computational routines, is older too. It was open-sourced in January 2025, and an Ascend adapter repository appeared on September 29, 2025. The September 30, 2026 update added native Ascend 950 support, including code generation, automatic scheduling and synchronization.
DeepSeek describes TileLang’s programming model as simpler than CUDA, with the aim of improving development efficiency and simplifying code. But TileLang already supported Nvidia and other hardware. Its Ascend work does not make it an exclusively Huawei language: developers can use the higher-level programming layer for Nvidia GPUs as well.
Sources
tomshardware.comDeepSeek and Huawei release open-source Ascend AI programming tools to reduce reliance on Nvidia CUDA ecosystem — tools include compute and communication libraries, as well as Ascend support for TileLang
techradar.comHuawei and DeepSeek release open source AI tools in a bid to lower Nvidia exposure — but will programmers make the switch?
Reader comments
Newest comments first. Replies stay oldest first.