Sep 2026· Proceedings of the International Conference on Parallel Processing· 0 citations· 27 references
TL;DR
RODIS, a two-level acceleration scheme that utilizes row-orchestration at the data loading level to achieve inter-block load balancing and employs dynamic instruction scheduling at the underlying computation level to enhance MAC utilization, is proposed.
Abstract
Following the Scaling Law, Deep Neural Networks face immense parameter scales and costs, making high-compression unstructured pruning attractive. Additionally, rising edge-side deployment demands have revived interest in activation sparsity. Consequently, unstructured sparse-dense and sparse-sparse matrix multiplications are becoming prevalent in LLM training and inference. However, as the most common general-purpose AI accelerators, GPUs do not provide efficient support for unstructured sparse matrices. Although existing works have proposed GPU architectural enhancements for unstructured sparsity, they are often limited to optimizing Tensor Core computation patterns through dataflows such as outer-product and row-by-row. These methods fail to adequately resolve the poor load balancing and low Multiply-Accumulate utilization in unstructured sparse workloads, especially lacking targeted optimizations for the latest GPU architectural features. To address these issues, we propose RODIS, a two-level acceleration scheme. It utilizes row-orchestration at the data loading level to achieve inter-block load balancing and employs dynamic instruction scheduling at the underlying computation level to enhance MAC utilization. Experimental results show that with minimal area and power overhead, RODIS achieves an average 1.35 × performance improvement compared to other state-of-the-art sparse Tensor Core schemes.
This work proposes the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support the dataflow paradigm while enhancing data reuse, and decouples the multiplication and merging phases.
Sheng-Bai Luo, Sheng Ma, Bo Wang et al.· ACM Transactions on Architec...· 0 citations
Sparse matrix-matrix multiplication (SpMM) is critical for graph analytics and learning tasks, yet its irregular sparsity poses challenges for hardware acceleration. Traditional static dataflows fail to adapt to local sparsity variations, causing load imbalance. We introduce SpMM-GO, a hybrid FPGA accelerator integrati...
Shang-Shang Yao, Yan Yan, Zuo-Ning Chen· IEEE Transactions on Very La...· 0 citations
Sparse matrix multiplications—including SpMV, SpMM, and SpGEMM—are fundamental to scientific computing, graph analytics, and machine learning. Despite extensive GPU-focused optimizations such as custom sparse formats and load balance, CSR-style and block-based methods can still underexploit fine-grained cache locality...
Xing Cong, Fu-Kai Sun, Yi-Ding Liu et al.· 0 citations
The proposed DistSpMM proposes DistSpMM, a co-design framework integrating data layout, pipelining, and communication strategies, which introduces HSDMA, a lightweight algorithm that reduces communication by optimizing the dense matrix allocation.
Junyu Gu, Jue Wang, Zhikuang Xin et al.· ACM Transactions on Architec...· 0 citations
Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in scientific computing and artificial intelligence. However, the inherent sparsity and irregularity of real-world datasets present significant challenges for developing high-performance SpMM kernels on modern GPUs. Existing approaches typically focu...
Qi Du, Sheng-Le Lin, Yue-Dan Chen et al.· Proceedings of the Internati...· 0 citations
DB-SpMSpV is presented, a dual-view blocked SpMSpV framework for dynamic GPU workloads that uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance and is integrated into DB-BFS and DB-Decoding.
Xing Cong, Chen-Hao Xie, Rui Wang et al.· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.