Skip to content
Book Open access

RODIS: Accelerating Sparse Matrix Multiplication on GPU via Row-Orchestration and Dynamic Instruction Scheduling

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 27 references

TL;DR

RODIS, a two-level acceleration scheme that utilizes row-orchestration at the data loading level to achieve inter-block load balancing and employs dynamic instruction scheduling at the underlying computation level to enhance MAC utilization, is proposed.

Abstract

Following the Scaling Law, Deep Neural Networks face immense parameter scales and costs, making high-compression unstructured pruning attractive. Additionally, rising edge-side deployment demands have revived interest in activation sparsity. Consequently, unstructured sparse-dense and sparse-sparse matrix multiplications are becoming prevalent in LLM training and inference. However, as the most common general-purpose AI accelerators, GPUs do not provide efficient support for unstructured sparse matrices. Although existing works have proposed GPU architectural enhancements for unstructured sparsity, they are often limited to optimizing Tensor Core computation patterns through dataflows such as outer-product and row-by-row. These methods fail to adequately resolve the poor load balancing and low Multiply-Accumulate utilization in unstructured sparse workloads, especially lacking targeted optimizations for the latest GPU architectural features. To address these issues, we propose RODIS, a two-level acceleration scheme. It utilizes row-orchestration at the data loading level to achieve inter-block load balancing and employs dynamic instruction scheduling at the underlying computation level to enhance MAC utilization. Experimental results show that with minimal area and power overhead, RODIS achieves an average 1.35 × performance improvement compared to other state-of-the-art sparse Tensor Core schemes.

Read PDF

Similar papers

Open access Sep 2026

ADEM: Accelerating Sparse Matrix Multiplication with Adaptive Dataflow and Efficient Merging

This work proposes the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support the dataflow paradigm while enhancing data reuse, and decouples the multiplication and merging phases.

Sheng-Bai Luo, Sheng Ma, Bo Wang et al. · 0 citations
Oct 2026

SpMM-GO: Sparse Matrix-Matrix Multiplication Acceleration via Hybrid Gustavson and Outer-Product Dataflows

Sparse matrix-matrix multiplication (SpMM) is critical for graph analytics and learning tasks, yet its irregular sparsity poses challenges for hardware acceleration. Traditional static dataflows fail to adapt to local sparsity variations, causing load imbalance. We introduce SpMM-GO, a hybrid FPGA accelerator integrati...

Shang-Shang Yao, Yan Yan, Zuo-Ning Chen · 0 citations
Open access

CB-Sparse:A Cache-Friendly Data Aggregating Algorithm for Block-Based Sparse Matrix Multiplication on GPUs

Sparse matrix multiplications—including SpMV, SpMM, and SpGEMM—are fundamental to scientific computing, graph analytics, and machine learning. Despite extensive GPU-focused optimizations such as custom sparse formats and load balance, CSR-style and block-based methods can still underexploit fine-grained cache locality...

Xing Cong, Fu-Kai Sun, Yi-Ding Liu et al. · 0 citations
Open access Aug 2026

DistSpMM: Accelerating Sparse Matrix Dense Matrix Multiplication on GPUs

The proposed DistSpMM proposes DistSpMM, a co-design framework integrating data layout, pipelining, and communication strategies, which introduces HSDMA, a lightweight algorithm that reduces communication by optimizing the dense matrix allocation.

Junyu Gu, Jue Wang, Zhikuang Xin et al. · 0 citations
Book Open access Sep 2026

CoTC-SpMM: A Cooperative Tensor–CUDA Cores Scheme for Efficient Sparse Matrix Multiplication

Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in scientific computing and artificial intelligence. However, the inherent sparsity and irregularity of real-world datasets present significant challenges for developing high-performance SpMM kernels on modern GPUs. Existing approaches typically focu...

Qi Du, Sheng-Le Lin, Yue-Dan Chen et al. · 0 citations
Book Open access Aug 2026

DB-SpMSpV: Dual-View Blocked Sparse Matrix-Sparse Vector Multiplication for Dynamic GPU Workloads

DB-SpMSpV is presented, a dual-view blocked SpMSpV framework for dynamic GPU workloads that uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance and is integrated into DB-BFS and DB-Decoding.

Xing Cong, Chen-Hao Xie, Rui Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.