deepseek-ai/DeepGEMM-Ascend

DeepGEMM-Ascend: clean and efficient matrix multiplication kernel library for Huawei Ascend NPUs

Open original source ↗

From the publisher

README & documentation

Source preview

Read the project’s overview, installation instructions and usage examples. The original README is the source of truth.

DeepGEMM Ascend is a port of DeepGEMM to the HUAWEI Ascend platform. It is fully API-compatible with DeepGEMM and supports BF16, FP8, FP4 GEMM, MQA logits, and MegaMoE. On Ascend platforms, users can simply install the package and use the same APIs and development workflow as DeepGEMM on other supported platforms. DeepGEMM Ascend provides a lightweight abstraction over the Ascend MAD (matrix multiply-add) primitives, hiding much of the complexity associated with fractal layouts, alignment constraints, address calculations, and verbose low-level parameters. This enables GEMM kernels to remain both concise and efficient. DeepGEMM Ascend makes extensive use of Ascend-specific optimization techniques, such as sparse data loading and coroutine-based pipelining, to approach the performance limits of the Ascend hardware. These implementations can also serve as references for extreme performance optimization on the Ascend platform. Despite its lightweight codebase, DeepGEMM Ascend can achieve peak hardware performance across a wide range of matrix shapes.…

deepseek-ai/DeepGEMM-Ascend on GitHub A short preview, not the full document.

Read the full README ↗ · Preview checked 2026-09-30T12:11:15.286Z

Inside the original README — Document outline
  1. DeepGEMM Ascend
  2. News
  3. Quick Start
  4. Requirements
  5. Development
  6. Installation
  7. Interfaces
  8. Kernel Interface
  9. Utilities
  10. Environment Variables
  11. Performance
  12. Dense GEMM

Headings are captured from the source. Links open the publisher’s document, not a locally hosted copy.

Links from the README

References supplied by the publisher, not independently verified endorsements. Check the destination before downloading files or entering credentials.

Documentation belongs to its respective authors. Reported project/model license: MIT. A listing is not a grant of reuse or training rights. Confirm the document’s own terms at the source.

What this repository does

DeepGEMM-Ascend: clean and efficient matrix multiplication kernel library for Huawei Ascend NPUs

Repository facts

Owner
deepseek-ai
Primary language
C++
Stars
295
Forks
12
Open issues + pull requests
2
License
MIT
Archived
No
Default branch
main
Created
2026-09-29T15:49:55.000Z
Last push
2026-09-30T01:11:39.000Z

Topics and intended use

No topics were included in the latest source metadata.

Review the README for scope, installation, examples and limitations. We do not run repository code or certify it.

Evaluate before installing

Review licensing and dependencies, inspect recent commits and unresolved issues, and test in an isolated environment before production use. Stars and forks alone cannot answer these questions.

README and project files

Issues and maintenance discussion

Releases and changelog

Source and freshness

Source: GitHub. Metadata observed 2026-09-30T12:05:16.119Z. Daily imports are snapshots, not real-time monitoring.

Popularity and source listings do not establish security, suitability, licensing rights or benchmark performance.

Open original source

Related repositories

tensorflow/tensorflow

ggml-org/llama.cpp

react/react-native

electron/electron

godotengine/godot

opencv/opencv