Skip to content
Book Open access

AFH-SpMM: Auto-Fit Heterogeneous Block Sparse-Dense Matrix Multiplication on Tensor Core GPUs

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 25 references

TL;DR

AFH-SpMM is a novel Auto-Fit Heterogeneous SpMM framework designed for adaptively parallelizing sparse-dense matrix multiplication on Tensor Core-equipped GPUs that achieves average speedups of 1.33 ×, and often leads cuSPARSE, ASpT, Sputnik, RoDe, Acc-SpMM, and MP-SpMM, with especially strong gains on medium and large matrices.

Abstract

Sparse-dense matrix multiplication (SpMM), a fundamental computational kernel in graph analytics and scientific computing, can be substantially accelerated on modern GPUs by leveraging both dense and sparse Tensor Core units, thereby enabling high-throughput computation. However, existing methods often fail to exploit these hardware units effectively when applied to real-world sparse matrices that exhibit strong local structural heterogeneity. In particular, methods that either (i) reorganize all regions into dense-like tiles or (ii) aggressively convert them into strict 2:4 structured sparsity blocks typically incur low effective block density, substantial padding overhead, and nontrivial preprocessing costs. To address these challenges, we propose AFH-SpMM, a novel Auto-Fit Heterogeneous SpMM framework designed for adaptively parallelizing sparse-dense matrix multiplication on Tensor Core-equipped GPUs. Using a 16-row window as the basic processing granularity, AFH-SpMM adaptively maps local regions to two hardware-efficient computation paths: 16 × 16 dense tiles targeted to Dense Tensor Cores and 16 × 8 row-wise 2:4 structured-sparse tiles targeted to Sparse Tensor Cores. For the sparse computation path, AFH-SpMM further exploits local column proximity to mitigate subsequent memory-access and address-generation overheads. At runtime, the two block types are executed within a single fused kernel, while preserving distinct operand layouts and specialized MMA pipelines for each path. Experiments on 600 SuiteSparse matrices across NVIDIA RTX PRO 6000, H100, and A800 show that AFH-SpMM achieves average speedups of 1.33 × (up to 5.72 ×), and often leads cuSPARSE, ASpT, Sputnik, RoDe, Acc-SpMM, and MP-SpMM, with especially strong gains on medium and large matrices.

Read PDF

Similar papers

Book Open access Sep 2026

CoTC-SpMM: A Cooperative Tensor–CUDA Cores Scheme for Efficient Sparse Matrix Multiplication

Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in scientific computing and artificial intelligence. However, the inherent sparsity and irregularity of real-world datasets present significant challenges for developing high-performance SpMM kernels on modern GPUs. Existing approaches typically focu...

Qi Du, Sheng-Le Lin, Yue-Dan Chen et al. · 0 citations
Book Open access Sep 2026

TileSpMM: A Variable-Size Tiled Algorithm for Sparse Matrix-Matrix Multiplication on Tensor Cores

This work proposes TileSpMM, which breaks the static-granularity bottleneck through a variable-size tiling algorithm that dynamically adapts to local sparsity patterns, and is equipped with an adaptive load-balancing strategy and customized granularity-specific kernels to improve hardware utilization and mitigate compu...

Hongwei Zeng, Shu-Qin Feng, Hao-Cheng Lian et al. · 0 citations
Book Open access Aug 2026

DB-SpMSpV: Dual-View Blocked Sparse Matrix-Sparse Vector Multiplication for Dynamic GPU Workloads

DB-SpMSpV is presented, a dual-view blocked SpMSpV framework for dynamic GPU workloads that uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance and is integrated into DB-BFS and DB-Decoding.

Xing Cong, Chen-Hao Xie, Rui Wang et al. · 0 citations
Open access

CB-Sparse:A Cache-Friendly Data Aggregating Algorithm for Block-Based Sparse Matrix Multiplication on GPUs

Sparse matrix multiplications—including SpMV, SpMM, and SpGEMM—are fundamental to scientific computing, graph analytics, and machine learning. Despite extensive GPU-focused optimizations such as custom sparse formats and load balance, CSR-style and block-based methods can still underexploit fine-grained cache locality...

Xing Cong, Fu-Kai Sun, Yi-Ding Liu et al. · 0 citations
Preprint Sep 2026

Scaling Fourier-Based Sparse Matrix Analysis on GPUs

Sparse computations are important workloads in applications such as scientific computing, graph neural networks (GNNs), and machine learning. While many sparse operations can benefit from modern GPUs, the sparsity pattern remains important to performance because it affects memory coalescing, block organization, and loa...

Rui-Feng Zhang, Sai Akhil Varma Manthena, Jia-Jia Li et al. · 0 citations
Open access Sep 2026

ADEM: Accelerating Sparse Matrix Multiplication with Adaptive Dataflow and Efficient Merging

This work proposes the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support the dataflow paradigm while enhancing data reuse, and decouples the multiplication and merging phases.

Sheng-Bai Luo, Sheng Ma, Bo Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.