Open access
Jul 2026
Realizable N:M Sparse Transformer Inference via Search-Kernel Co-design
This work designs MD-SpMM, an N:M sparse CUDA kernel that reorganizes sparse GEMM into micro-dense, Tensor-Core-aligned dataflow and uses inference-aware adaptive parallelism to sustain utilization and performs layer-wise sparsity search under explicit end-to-end latency budgets.
Yi-Ming Liu, Wenqi Lou, Zhiguang Wang et al.
· European Conference on Paral... · 0 citations