It is found that no single kernel dominates across all graph structures and that effective Tensor Core utilization reaches only 5–20% on irregular GNN matrices, while graph reordering is broadly beneficial, yielding gains of up to $43\times $ when it enables Sparse Tensor Core execution.
Abstract
Sparse Matrix-Dense Matrix Multiplication (SpMM) is a dominant computational bottleneck in Graph Neural Network (GNN) inference and training. Representative studies report that SpMM consumes roughly 30% of the execution time in some Graph Convolutional Network (GCN) settings and over 80% in full-batch GraphSAGE training. Despite the rapid growth of GPU SpMM optimization techniques, spanning CUDA core kernels, Tensor Core acceleration, adaptive hybrid execution, autotuning, graph reordering, and framework integration, no dedicated survey has focused on this subfield. This paper presents the first such survey, covering 52 GPU-accelerated SpMM methods for GNN workloads published between 2019 and 2026. We constructed the corpus from IEEE Xplore, the ACM Digital Library, USENIX, arXiv, and Google Scholar, screening the studies first by title and abstract and then by full text. We included GPU-based SpMM kernels and GNN aggregation systems and excluded CPU-only, non-SpMM, abstract-only, and duplicate-version papers. We organize the literature into six technique categories and compare the methods using a ten-dimensional framework. Representative dimensions include sparse format, hardware target, parallelism strategy, load balancing, preprocessing cost, and open-source availability. We find that no single kernel dominates across all graph structures and that effective Tensor Core utilization reaches only 5–20% on irregular GNN matrices. Graph reordering is broadly beneficial, yielding gains of up to $43\times $ when it enables Sparse Tensor Core execution. Because the surveyed literature is overwhelmingly based on CUDA and Tensor Cores, our analysis is NVIDIA-centered. Nevertheless, we distinguish architecture-level insights that generalize to AMD and Intel accelerators from vendor-specific implementation details. We conclude with eight open challenges, including the no-single-winner problem, the Tensor Core utilization gap, standardized benchmarking, and graph-to-kernel compilation.
Sparse-dense matrix multiplication (SpMM), a fundamental computational kernel in graph analytics and scientific computing, can be substantially accelerated on modern GPUs by leveraging both dense and sparse Tensor Core units, thereby enabling high-throughput computation. However, existing methods often fail to exploit...
Zhi-Rui Chen, Heng Zhang, Kai-Fan Jia· Proceedings of the Internati...· 0 citations
Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in scientific computing and artificial intelligence. However, the inherent sparsity and irregularity of real-world datasets present significant challenges for developing high-performance SpMM kernels on modern GPUs. Existing approaches typically focu...
Qi Du, Shengle Lin, Yuedan Chen et al.· Proceedings of the Internati...· 0 citations
Matrix multiplication is a fundamental computation kernel in many parallel and sequential scientific applications. We target FP32 matrix multiplication on GPUs, a setting required by numerous HPC and scientific workloads. Alternative Basis Matrix Multiplication (ABMM) is a practical Strassen-like algorithm that reduces...
Yao Liu, Ye-Wen Li, Zhonghai Zhang et al.· Proceedings of the Internati...· 0 citations
Sparse GEneral Matrix Multiplication (SpGEMM) is one of the most vital kernels in massive research domains, including bioinformatics, graph analytics, and machine learning. Moreover, with the prosperity of the Big Data era, nonzero elements in sparse matrices of SpGEMM boost rapidly into the magnitude of billions. Thus...
Ming Dun, Cheng Zhang, Shuhan Song et al.· IEEE Transactions on Paralle...· 0 citations
The proposed DistSpMM proposes DistSpMM, a co-design framework integrating data layout, pipelining, and communication strategies, which introduces HSDMA, a lightweight algorithm that reduces communication by optimizing the dense matrix allocation.
Junyu Gu, Jue Wang, Zhikuang Xin et al.· ACM Transactions on Architec...· 0 citations
DB-SpMSpV is presented, a dual-view blocked SpMSpV framework for dynamic GPU workloads that uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance and is integrated into DB-BFS and DB-Decoding.
Xing Cong, Chen-Hao Xie, Rui Wang et al.· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.