DeepSeek ports high-performance GEMM library to Huawei Ascend NPUs
DeepGEMM-Ascend brings API-compatible, near-peak matrix multiplication performance to Huawei Ascend 950 hardware, supporting FP4, FP8, and BF16 workloads.
DeepSeek AI has released DeepGEMM-Ascend, a specialized library for matrix multiplication on Huawei Ascend neural processing units. Announced on September 30, 2026, this open-source project ports the original DeepGEMM framework to the Ascend ecosystem, targeting the Ascend 950 series hardware.
The release provides engineers with a tool to achieve near-peak hardware utilization for large language model inference and training tasks without rewriting low-level kernel code. By abstracting complex hardware constraints, it allows developers to use familiar APIs while leveraging Ascend-specific optimizations.
What happened
DeepGEMM-Ascend is a direct port of the DeepGEMM library, designed specifically for the Huawei Ascend platform. The initial release supports the Ascend 950 devices and requires the CANN 9.20 toolkit. It maintains full API compatibility with the upstream DeepGEMM project, meaning users can install the package and continue using the same development workflow they would on other supported platforms.
The library supports several critical data types and operations used in modern AI models, including BF16, FP8, and FP4 general matrix multiplication (GEMM). It also includes specialized support for MQA logits and MegaMoE operators, which are essential for efficient mixture-of-experts architectures. This broad support ensures that developers working with cutting-edge model structures can deploy them effectively on Ascend hardware.
A key feature of this release is its lightweight abstraction layer over the Ascend MAD (matrix multiply-add) primitives. This layer hides intricate details such as fractal matrix layouts, alignment constraints, address calculations, and verbose low-level parameters. By managing these complexities internally, the library enables GEMM kernels to remain concise in code while maintaining high efficiency in execution.
How it works
DeepGEMM-Ascend achieves its performance by employing Ascend-specific optimization techniques that push the hardware close to its theoretical limits. The library utilizes sparse data loading and coroutine-based pipelining to minimize idle time and maximize throughput. These methods allow the kernels to handle data movement and computation concurrently, reducing bottlenecks that often plague high-performance computing tasks.
The implementation serves as a reference for extreme performance optimization on the Ascend platform. Despite having a lightweight codebase, the library demonstrates the ability to reach peak hardware performance across a wide variety of matrix shapes. This is particularly important for dynamic workloads where tensor dimensions change frequently during inference or training phases.
For developers integrating this library, the scaling factor format differs slightly from NVIDIA implementations. On Ascend, each pair of UE8M0 scaling factors along the K dimension is packed into an int16, with values stored in MN-major order for optimal hardware efficiency. The library provides utility functions to transform scaling factors into this required layout, ensuring seamless interoperability with existing data pipelines.
Key details
- Hardware Support: Developed and validated on the Huawei Ascend 950 series NPUs.
- Software Requirements: Requires CANN 9.20 toolkit, torch_npu, Python 3.10+, and C++20 compiler support.
- Data Types: Supports BF16, FP8, and FP4 GEMM operations, along with MQA logits and MegaMoE operators.
- Performance: Dense GEMM operations reach up to 99.8% of the hardware limit for BF16 and 99.5% for FP8 formats.
- Optimization Techniques: Uses sparse data loading and coroutine-based pipelining to approach hardware performance limits.
- License: Released under the MIT License, allowing for broad adoption and modification.
Why it matters
For software engineers building AI infrastructure, this release significantly lowers the barrier to entry for using Huawei Ascend hardware. Historically, optimizing kernels for non-NVIDIA accelerators has required deep expertise in specific hardware architectures and manual management of memory layouts. DeepGEMM-Ascend abstracts these difficulties, allowing teams to focus on model architecture rather than low-level device tuning.
The high utilization rates reported in the benchmarks suggest that organizations can achieve cost-effective scaling on Ascend clusters without sacrificing computational efficiency. With dense GEMM operations reaching nearly 100% of the hardware limit, the library ensures that expensive NPU resources are not wasted due to software inefficiencies. This is crucial for large-scale deployments where even small percentage gains in utilization translate to significant savings in energy and time.
Furthermore, the support for advanced features like MegaMoE and MQA logits indicates that the library is ready for next-generation model architectures. As AI models grow larger and more complex, the ability to efficiently handle mixture-of-experts routing and multi-query attention becomes a competitive advantage. This library provides a robust foundation for implementing these features on Ascend hardware.
What you can do
- Install the library using pip with the no-build-isolation flag to ensure proper compilation of C++ extensions.
- Verify your environment meets the requirements, specifically checking for CANN 9.20 and Python 3.10 or higher.
- Use the provided utility functions to transform scaling factors into the required MN-major order for FP4 and FP8 operations.
- Benchmark your specific workloads using the included test suite to compare performance against theoretical hardware limits.
- Explore the source code as a reference for implementing custom optimizations on the Ascend platform.
- Configure environment variables like DG_JIT_CACHE_DIR to manage compiled kernel caches and improve startup times.



