Skip to content
Book Open access

DB-SpMSpV: Dual-View Blocked Sparse Matrix-Sparse Vector Multiplication for Dynamic GPU Workloads

Aug 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 43 references
Computer Science

TL;DR

DB-SpMSpV is presented, a dual-view blocked SpMSpV framework for dynamic GPU workloads that uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance and is integrated into DB-BFS and DB-Decoding.

Abstract

Sparse Matrix-Sparse Vector Multiplication (SpMSpV) is a core primitive in graph traversal, sparse linear algebra, and sparse model inference. Its input vector is often dynamically sparse, so the best GPU execution path depends on both global sparsity and the local vector-block distribution. Existing GPU SpMSpV methods often bind storage layouts, push/pull traversal, and kernels together, making fine-grained adaptation difficult without extra storage or scheduling overhead. This paper presents DB-SpMSpV, a dual-view blocked SpMSpV framework for dynamic GPU workloads. DB-SpMSpV partitions the matrix into fixed-size 2D blocks, maintains block-level CSR/CSC views at the high level, and reuses a single low-level block payload to support both row-driven pull and column-driven push. At runtime, it selects the global traversal path based on input block sparsity, chooses block microkernels from the local matrix/vector block structure, and uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance. We further integrate the framework into DB-BFS and DB-Decoding. We evaluate DB-SpMSpV on NVIDIA A100 and RTX 4090 using SuiteSparse matrices, symmetric graphs, and three open-source LLMs. Across input sparsities, DB-SpMSpV achieves average speedups of 5.48 × –64.34 × over cuSPARSE and 2.36 × –14.01 × over TileSpMSpV on A100, with similar gains on RTX 4090. DB-BFS further improves end-to-end graph traversal by 2.66 × over TileBFS on A100 and 3.60 × on RTX 4090 on average, while DB-Decoding accelerates single-token linear layers by up to 4.50 ×.

Read PDF

Similar papers

Preprint Aug 2026

Celty: SpMSpV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference

Celty, a co-designed sparse format, GPU kernel, and SIMT microarchitecture for efficient spMspV in LLM inference is proposed, which introduces a Run-Length Compressed CSC format that enables vectorized loading of compressed weight columns and exploits both sparsity sources to skip unnecessary memory accesses.

Ruokai Yin, Priyadarshini Panda · 0 citations
Book Open access Sep 2026

CoTC-SpMM: A Cooperative Tensor–CUDA Cores Scheme for Efficient Sparse Matrix Multiplication

Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in scientific computing and artificial intelligence. However, the inherent sparsity and irregularity of real-world datasets present significant challenges for developing high-performance SpMM kernels on modern GPUs. Existing approaches typically focu...

Qi Du, Sheng-Le Lin, Yue-Dan Chen et al. · 0 citations
Book Open access Sep 2026

AFH-SpMM: Auto-Fit Heterogeneous Block Sparse-Dense Matrix Multiplication on Tensor Core GPUs

AFH-SpMM is a novel Auto-Fit Heterogeneous SpMM framework designed for adaptively parallelizing sparse-dense matrix multiplication on Tensor Core-equipped GPUs that achieves average speedups of 1.33 ×, and often leads cuSPARSE, ASpT, Sputnik, RoDe, Acc-SpMM, and MP-SpMM, with especially strong gains on medium and large...

Zhi-Rui Chen, Heng Zhang, Kai-Fan Jia · 0 citations
Open access Sep 2026

ADEM: Accelerating Sparse Matrix Multiplication with Adaptive Dataflow and Efficient Merging

This work proposes the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support the dataflow paradigm while enhancing data reuse, and decouples the multiplication and merging phases.

Sheng-Bai Luo, Sheng Ma, Bo Wang et al. · 0 citations
Book Open access Sep 2026

RODIS: Accelerating Sparse Matrix Multiplication on GPU via Row-Orchestration and Dynamic Instruction Scheduling

RODIS, a two-level acceleration scheme that utilizes row-orchestration at the data loading level to achieve inter-block load balancing and employs dynamic instruction scheduling at the underlying computation level to enhance MAC utilization, is proposed.

Jian-Feng Cui, Bo Yuan, Ze-Kun Jiang et al. · 0 citations
Open access Aug 2026

DistSpMM: Accelerating Sparse Matrix Dense Matrix Multiplication on GPUs

The proposed DistSpMM proposes DistSpMM, a co-design framework integrating data layout, pipelining, and communication strategies, which introduces HSDMA, a lightweight algorithm that reduces communication by optimizing the dense matrix allocation.

Junyu Gu, Jue Wang, Zhikuang Xin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.