deepseek-ai / DeepGEMM

deepseek-ai / DeepGEMM

DeepGEMM is a unified, high-performance tensor core kernel library that brings together the key computation primitives of modern large language models — GEMMs (FP8, FP4, BF16), fused MoE with overlapped communication (Mega MoE), MQA scoring for the lightning indexer, HyperConnection (HC), and more — into a single, cohesive CUDA codebase. DeepGEMM 是一个统一的高性能 Tensor Core 内核库,它将现代大语言模型的关键计算原语(如 GEMM (FP8, FP4, BF16)、带重叠通信的融合 MoE (Mega MoE)、用于闪电索引器的 MQA 评分、HyperConnection (HC) 等)整合到一个统一且紧凑的 CUDA 代码库中。

All kernels are compiled at runtime through DeepJIT, requiring no CUDA compilation during installation. DeepGEMM leverages some concepts from CUTLASS and CuTe, but avoids heavy reliance on their templates or algebras. The library is designed for simplicity, with only a limited number of core kernel functions, making it a clean and accessible resource for learning NVIDIA GPU kernel optimization techniques. Despite its lightweight design, DeepGEMM’s performance matches or exceeds expert-tuned libraries across various matrix shapes. 所有内核均通过 DeepJIT 在运行时编译,安装时无需进行 CUDA 编译。DeepGEMM 借鉴了 CUTLASS 和 CuTe 的部分概念,但避免了对它们模板或代数系统的过度依赖。该库设计简洁,核心内核函数数量有限,是学习 NVIDIA GPU 内核优化技术的清晰且易于上手的资源。尽管设计轻量,DeepGEMM 在各种矩阵形状下的性能均能达到甚至超过专家调优的库。

News

新闻

  • 2026.09.30: DeepGEMM Ascend is available! Check DeepGEMM-Ascend for more details. Add more optimizations, including locality domain features, check #46 for more details. 2026.09.30: DeepGEMM Ascend 已发布!详情请查看 DeepGEMM-Ascend。增加了更多优化,包括局部域(locality domain)特性,详情请查看 #46。
  • 2026.09.10: Sparse Indexer, Mega Gate, Mega mHC, DeepJIT, MoE and Indexer optimizations and more. Please see #432 for more details. 2026.09.10: 稀疏索引器、Mega Gate、Mega mHC、DeepJIT、MoE 和索引器优化等。详情请查看 #432。
  • 2026.04.16: Mega MoE, FP8xFP4 GEMM, FP4 Indexer, PDL, faster JIT compilation and more. Please see #304 for more details. For Mega MoE benchmarks, refer to #316. 2026.04.16: Mega MoE、FP8xFP4 GEMM、FP4 索引器、PDL、更快的 JIT 编译等。详情请查看 #304。关于 Mega MoE 的基准测试,请参考 #316。
  • 2025.09.28: DeepGEMM now supports scoring kernels (weighted ReLU MQA logits) for the lightning indexer for DeepSeek v3.2. Please see #200 for more details. 2025.09.28: DeepGEMM 现已支持 DeepSeek v3.2 闪电索引器的评分内核(加权 ReLU MQA logits)。详情请查看 #200。
  • 2025.07.20: DeepGEMM now supports both SM90/SM100, and has a full refactor with a low-CPU-overhead JIT CPP module. As NVCC 12.9 will automatically do the FFMA interleaving, all post optimizations will be no longer supported. Please see #112 for more details. 2025.07.20: DeepGEMM 现已同时支持 SM90/SM100,并进行了全面重构,引入了低 CPU 开销的 JIT CPP 模块。由于 NVCC 12.9 将自动执行 FFMA 交错,后续优化将不再支持。详情请查看 #112。
  • 2025.05.14: DeepGEMM now offers weight gradient kernels for dense and MoE backward! See #95 for details. 2025.05.14: DeepGEMM 现已提供用于 Dense 和 MoE 反向传播的权重梯度内核!详情请查看 #95。
  • 2025.04.18: DeepGEMM now achieves up to 1550 TFLOPS on H800! See #74, #78, #81, #86 and 340d988 for details. 2025.04.18: DeepGEMM 在 H800 上现可达到 1550 TFLOPS!详情请查看 #74, #78, #81, #86 以及 340d988。

Quick start

快速开始

Requirements

环境要求

  • NVIDIA SM90 or SM100 architecture GPU
  • NVIDIA SM90 或 SM100 架构 GPU
  • Python 3.8 or higher
  • Python 3.8 或更高版本
  • Compilers and standard libraries with C++20 <format> support
  • 支持 C++20 <format> 的编译器和标准库
  • CUDA Toolkit 12.9 or higher
  • CUDA Toolkit 12.9 或更高版本
  • PyTorch 2.3 or higher
  • PyTorch 2.3 或更高版本
  • CUTLASS 4.0 or higher (could be cloned by Git submodule)
  • CUTLASS 4.0 或更高版本(可通过 Git 子模块克隆)

Development

开发

# Submodule must be cloned
git clone --recursive git@github.com:deepseek-ai/DeepGEMM.git
cd DeepGEMM
# Link some essential includes and build the C++ extension
cat develop.sh
./develop.sh

Installation

安装

cat install.sh
./install.sh

Then, import deep_gemm in your Python project, and enjoy! 之后,在你的 Python 项目中导入 deep_gemm 即可使用!


Interfaces

接口

Notices

注意事项

This library provides optimized GEMM kernels for NVIDIA GPUs with a naming convention: D = C + A @ B. The input shape layout is NT (non-transposed A, transposed B). While the SM90 implementation supports only the NT memory layout (row-major, col-major), the SM100 implementation supports all memory layouts (NT, TN, NN, TT). For example, fp8_gemm_nt will do a D = C + A @ B.T. 本库为 NVIDIA GPU 提供优化的 GEMM 内核,命名约定为 D = C + A @ B。输入形状布局为 NT(A 不转置,B 转置)。SM90 实现仅支持 NT 内存布局(行主序,列主序),而 SM100 实现支持所有内存布局(NT, TN, NN, TT)。例如,fp8_gemm_nt 将执行 D = C + A @ B.T。

For both architectures, the LHS scaling factor is required to have a TMA-aligned and transposed layout. And the data format for the scaling factor of SM90 and SM100 is different: 对于两种架构,左侧(LHS)缩放因子均要求具有 TMA 对齐且转置的布局。此外,SM90 和 SM100 的缩放因子数据格式不同:

  • SM90 requires scaling factors in FP32 format.
  • SM90 要求缩放因子为 FP32 格式。
  • SM100 requires scaling factors in packed UE8M0 format, which packs 4 UE8M0 into a single torch.int.
  • SM100 要求缩放因子为打包的 UE8M0 格式,即将 4 个 UE8M0 打包进一个 torch.int 中。

Please note that operations like input transposition or FP8 casting must be handled separately by the user, please implement or fuse them into prior kernels independently. While the library provides some simple PyTorch utility functions, these may result in slower performance, but our primary focus is on optimizing the GEMM kernels themselves. 请注意,输入转置或 FP8 类型转换等操作必须由用户单独处理,请独立实现或将其融合到之前的内核中。虽然本库提供了一些简单的 PyTorch 工具函数,但它们可能会导致性能下降,我们的主要重点是优化 GEMM 内核本身。

Normal dense GEMMs (non-grouped)

普通稠密 GEMM(非分组)

To perform a basic non-grouped FP8 GEMM, call the fp8_gemm_{nt, nn, tn, tt} function. For more details, please refer to the function documentation. 要执行基本的非分组 FP8 GEMM,请调用 fp8_gemm_{nt, nn, tn, tt} 函数。更多详情,请参考函数文档。

Grouped GEMMs (contiguous layout)

分组 GEMM(连续布局)

Unlike traditional grouped GEMMs in CUTLASS, DeepGEMM groups only the M-axis, while N and K must remain fixed. This design is tailored for scenarios where experts in an MoE model share the same shape. For training forward passes or inference prefilling, where each expert may process a varying number of tokens, we concatenate these tokens into a single tensor, referred to as the “contiguous” layout. 与 CUTLASS 中传统的分组 GEMM 不同,DeepGEMM 仅对 M 轴进行分组,而 N 和 K 必须保持固定。这种设计专为 MoE 模型中专家共享相同形状的场景而定制。对于训练前向传播或推理预填充(prefilling),由于每个专家处理的 token 数量可能不同,我们将这些 token 连接成一个张量,称为“连续(contiguous)”布局。

Note that each expert segment must be aligned to the GEMM M block size (get_mk_alignment_for_contiguous_layout()). For more information, please refer to the m_grouped_fp8_gemm_{nt, nn}_contiguous function documentation. We also provide a K-axis-grouped API for MoE weight backward (with M and N must remain fixed), please refer to k_grouped_fp8_gemm_tn_contiguous for more information. 请注意,每个专家段必须与 GEMM M 块大小对齐(使用 get_mk_alignment_for_contiguous_layout())。更多信息,请参考 m_grouped_fp8_gemm_{nt, nn}_contiguous 函数文档。我们还为 MoE 权重反向传播提供了 K 轴分组 API(此时 M 和 N 必须保持固定),详情请参考 k_grouped_fp8_gemm_tn_contiguous。

Grouped GEMMs (masked layout)

分组 GEMM(掩码布局)

During the inference decoding phase, when CUDA graph is enabled and the CPU is unaware of the number of tokens each expert receives, we support masked grouped GEMMs. By providing a mask tensor, the kernel computes only the valid portions. Use m_grouped_fp8_gemm_nt_masked for this purpose and consult the relevant documentation. An example usage is to use the output of low-latency kernels from DeepEP as input. 在推理解码阶段,当启用 CUDA Graph 且 CPU 不知道每个专家接收的 token 数量时,我们支持掩码分组 GEMM。通过提供掩码张量,内核仅计算有效部分。请使用 m_grouped_fp8_gemm_nt_masked 并查阅相关文档。一个示例用法是将 DeepEP 的低延迟内核输出作为输入。

V3.2 MQA kernels for the indexer

用于索引器的 V3.2 MQA 内核

The kernel family has two versions, non-paged (for prefilling) and paged (for decoding). Take the non-paged version fp8_fp4_mqa_logits as an example. 该内核系列有两个版本:非分页(用于预填充)和分页(用于解码)。以非分页版本 fp8_fp4_mqa_logits 为例:

Its main inputs are: 其主要输入为:

  • q, a (q_data, q_sf) tuple; SM100 accepts MXFP4/MXFP8 data with packed UE8M0 scales.
  • q,一个 (q_data, q_sf) 元组;SM100 接受带有打包 UE8M0 缩放因子的 MXFP4/MXFP8 数据。
  • kv, a (kv_data, kv_sf) tuple with shape [seq_len_kv, head_dim] logically.
  • kv,一个逻辑形状为 [seq_len_kv, head_dim] 的 (kv_data, kv_sf) 元组。
  • weights, tensor with shape [seq_len, num_heads] (BF16 on SM100).
  • weights,形状为 [seq_len, num_heads] 的张量(在 SM100 上为 BF16)。
  • cu_seq_len_k_start and cu_seq_len_k_end, int tensor with shape [seq_len].
  • cu_seq_len_k_start 和 cu_seq_len_k_end,形状为 [seq_len] 的整数张量。
  • max_seqlen_k, the maximum valid KV span of any query row.
  • max_seqlen_k,任意查询行的最大有效 KV 跨度。

The output is compressed to [seq_len, max_seqlen_k]; row i stores its valid KV span starting at column zero. For each token i in q, it will iterate all tokens j from [cu_seq_len_k_start[i], cu_seq_len_k_end[i]), and calculate the corresponding compressed logit as: 输出被压缩为 [seq_len, max_seqlen_k];第 i 行存储其从第 0 列开始的有效 KV 跨度。对于 q 中的每个 token i,它将遍历 [cu_seq_len_k_start[i], cu_seq_len_k_end[i]) 范围内的所有 token j,并计算相应的压缩 logit:

kv_j = kv[0][j, :] * kv[1][j].unsqueeze(1) # [head_dim]
out_ij = q[i, :, :] @ kv_j # [num_heads]
out_ij = out_ij.relu() * weights[i, :] # [num_heads]