Aug 2026· Proceedings of the International Conference on Parallel Processing· 0 citations· 43 references
Computer Science
TL;DR
DB-SpMSpV is presented, a dual-view blocked SpMSpV framework for dynamic GPU workloads that uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance and is integrated into DB-BFS and DB-Decoding.
Abstract
Sparse Matrix-Sparse Vector Multiplication (SpMSpV) is a core primitive in graph traversal, sparse linear algebra, and sparse model inference. Its input vector is often dynamically sparse, so the best GPU execution path depends on both global sparsity and the local vector-block distribution. Existing GPU SpMSpV methods often bind storage layouts, push/pull traversal, and kernels together, making fine-grained adaptation difficult without extra storage or scheduling overhead. This paper presents DB-SpMSpV, a dual-view blocked SpMSpV framework for dynamic GPU workloads. DB-SpMSpV partitions the matrix into fixed-size 2D blocks, maintains block-level CSR/CSC views at the high level, and reuses a single low-level block payload to support both row-driven pull and column-driven push. At runtime, it selects the global traversal path based on input block sparsity, chooses block microkernels from the local matrix/vector block structure, and uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance. We further integrate the framework into DB-BFS and DB-Decoding. We evaluate DB-SpMSpV on NVIDIA A100 and RTX 4090 using SuiteSparse matrices, symmetric graphs, and three open-source LLMs. Across input sparsities, DB-SpMSpV achieves average speedups of 5.48 × –64.34 × over cuSPARSE and 2.36 × –14.01 × over TileSpMSpV on A100, with similar gains on RTX 4090. DB-BFS further improves end-to-end graph traversal by 2.66 × over TileBFS on A100 and 3.60 × on RTX 4090 on average, while DB-Decoding accelerates single-token linear layers by up to 4.50 ×.
Celty, a co-designed sparse format, GPU kernel, and SIMT microarchitecture for efficient spMspV in LLM inference is proposed, which introduces a Run-Length Compressed CSC format that enables vectorized loading of compressed weight columns and exploits both sparsity sources to skip unnecessary memory accesses.
Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in scientific computing and artificial intelligence. However, the inherent sparsity and irregularity of real-world datasets present significant challenges for developing high-performance SpMM kernels on modern GPUs. Existing approaches typically focu...
Qi Du, Sheng-Le Lin, Yue-Dan Chen et al.· Proceedings of the Internati...· 0 citations
AFH-SpMM is a novel Auto-Fit Heterogeneous SpMM framework designed for adaptively parallelizing sparse-dense matrix multiplication on Tensor Core-equipped GPUs that achieves average speedups of 1.33 ×, and often leads cuSPARSE, ASpT, Sputnik, RoDe, Acc-SpMM, and MP-SpMM, with especially strong gains on medium and large...
Zhi-Rui Chen, Heng Zhang, Kai-Fan Jia· Proceedings of the Internati...· 0 citations
This work proposes the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support the dataflow paradigm while enhancing data reuse, and decouples the multiplication and merging phases.
Sheng-Bai Luo, Sheng Ma, Bo Wang et al.· ACM Transactions on Architec...· 0 citations
RODIS, a two-level acceleration scheme that utilizes row-orchestration at the data loading level to achieve inter-block load balancing and employs dynamic instruction scheduling at the underlying computation level to enhance MAC utilization, is proposed.
Jian-Feng Cui, Bo Yuan, Ze-Kun Jiang et al.· Proceedings of the Internati...· 0 citations
The proposed DistSpMM proposes DistSpMM, a co-design framework integrating data layout, pipelining, and communication strategies, which introduces HSDMA, a lightweight algorithm that reduces communication by optimizing the dense matrix allocation.
Junyu Gu, Jue Wang, Zhikuang Xin et al.· ACM Transactions on Architec...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.