Aug 2026· ACM Transactions on Architecture and Code Optimization (TACO)· 0 citations· 37 references
TL;DR
The proposed DistSpMM proposes DistSpMM, a co-design framework integrating data layout, pipelining, and communication strategies, which introduces HSDMA, a lightweight algorithm that reduces communication by optimizing the dense matrix allocation.
Abstract
Sparse matrix-dense matrix multiplication (SpMM) is a core operation in scientific computing and deep learning. On multi-GPU platforms, its scalability is limited by communication bottlenecks. To address this, we propose DistSpMM, a co-design framework integrating data layout, pipelining, and communication strategies. DistSpMM introduces HSDMA, a lightweight algorithm that reduces communication by optimizing the dense matrix allocation. DistSpMM features a topology-aware two-stage pipeline that manages the IB/NVLink bandwidth disparity to maximize the overlap of computation and communication. Finally, DistSpMM employs an adaptive selector that uses a performance model to dynamically choose the optimal communication granularity (coarse vs. fine-grained) based on data sparsity and network tier. Experiments on diverse real-world datasets demonstrate the superior performance of our method. It achieves average speedups of 1.6 × to 2.6 × in single-node multi-GPU environments and 4.0 × to 5.1 × in multi-node multi-GPU environments over the baseline. Compared to the best state-of-the-art implementations, our method delivers up to 2.0 × speedup.
This work proposes TileSpMM, which breaks the static-granularity bottleneck through a variable-size tiling algorithm that dynamically adapts to local sparsity patterns, and is equipped with an adaptive load-balancing strategy and customized granularity-specific kernels to improve hardware utilization and mitigate compu...
Hongwei Zeng, Shu-Qin Feng, Hao-Cheng Lian et al.· Proceedings of the Internati...· 0 citations
Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in scientific computing and artificial intelligence. However, the inherent sparsity and irregularity of real-world datasets present significant challenges for developing high-performance SpMM kernels on modern GPUs. Existing approaches typically focu...
Qi Du, Sheng-Le Lin, Yue-Dan Chen et al.· Proceedings of the Internati...· 0 citations
Sparse matrix multiplications—including SpMV, SpMM, and SpGEMM—are fundamental to scientific computing, graph analytics, and machine learning. Despite extensive GPU-focused optimizations such as custom sparse formats and load balance, CSR-style and block-based methods can still underexploit fine-grained cache locality...
Xing Cong, Fu-Kai Sun, Yi-Ding Liu et al.· 0 citations
This work proposes the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support the dataflow paradigm while enhancing data reuse, and decouples the multiplication and merging phases.
Sheng-Bai Luo, Sheng Ma, Bo Wang et al.· ACM Transactions on Architec...· 0 citations
RODIS, a two-level acceleration scheme that utilizes row-orchestration at the data loading level to achieve inter-block load balancing and employs dynamic instruction scheduling at the underlying computation level to enhance MAC utilization, is proposed.
Jian-Feng Cui, Bo Yuan, Ze-Kun Jiang et al.· Proceedings of the Internati...· 0 citations
BAG (Basis Alternative Matrix Multiplication on GPUs), a GPU-oriented implementation of ABMM for the NVIDIA Ampere architecture is presented, designed to shrink workspace and eliminate redundant global-memory traffic, and introduce a cost-model-based recursion policy together with a Roofline-guided blocking strategy to...
Yao Liu, Ye-Wen Li, Zhong-Hai Zhang et al.· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.