DeepSeek has published Huawei Ascend versions of two of the low-level libraries its own models are trained and served with, putting code that previously assumed NVIDIA hardware into the hands of anyone running Chinese accelerators.
The repositories appeared on DeepSeek’s GitHub organisation on 29 and 30 September. DeepGEMM-Ascend is a port of DeepGEMM (the matrix-multiplication kernel library behind DeepSeek’s models) to Huawei Ascend NPUs, covering BF16, FP8 and FP4 arithmetic. A companion repository, DeepEP-Ascend, does the same for DeepEP, the communication library that shuttles tokens between experts in a mixture-of-experts model. Both are MIT-licensed and validated on the Ascend 950 series with Huawei’s CANN 9.20 toolkit.
The performance claims are DeepSeek’s own and have not been independently reproduced. The README reports dense matrix multiplication reaching up to 99.8% of the hardware’s theoretical limit, with comparable utilisation for the mixture-of-experts kernels. Different renderings of the page give slightly different TFLOPS totals, so the utilisation percentages are the sturdier number.
DeepSeek thanks Huawei for “technical support and engineering expertise” during development. The significance is less the speed than the software: the practical barrier to training and serving large models on domestic Chinese chips has been tooling, not silicon.
