Skip to content
Open access

ADEM: Accelerating Sparse Matrix Multiplication with Adaptive Dataflow and Efficient Merging

Sep 2026 · ACM Transactions on Architecture and Code Optimization (TACO) · 0 citations · 23 references

TL;DR

This work proposes the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support the dataflow paradigm while enhancing data reuse, and decouples the multiplication and merging phases.

Abstract

Sparse Matrix-Sparse Matrix Multiplication (SpMSpM) is a crucial computational kernel widely used in scientific computing and machine learning. The varying sparse patterns across different matrices pose significant challenges for conventional accelerators with fixed dataflow architectures. Although recent studies have explored dynamic dataflow approaches to better capture memory access patterns under diverse sparsity conditions, these solutions still struggle to simultaneously improve data reuse, load balance, and merging efficiency. To address these limitations, we present an SpMSpM accelerator based on adaptive dataflow and efficient merging (ADEM). ADEM is carefully designed from four key aspects. First, we propose the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support our dataflow paradigm while enhancing data reuse. We then present a cache-aware dataflow to mitigate memory overflow issues. Furthermore, ADEM decouples the multiplication and merging phases, utilizing the SFT structure to enable fine-grained task scheduling for improved load balancing. Finally, the accelerator incorporates heterogeneous merging units specifically optimized for handling two distinct types of merging operations, thereby significantly improving merger utilization. Compared with the state-of-the-art baseline system, ADEM achieves average speedups of 1.25 ×, 1.73 ×, and 1.16 × on VGG-16, ResNet-50, and SuiteSparse workloads, respectively.

Read PDF

Similar papers

Book Open access Sep 2026

CoTC-SpMM: A Cooperative Tensor–CUDA Cores Scheme for Efficient Sparse Matrix Multiplication

Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in scientific computing and artificial intelligence. However, the inherent sparsity and irregularity of real-world datasets present significant challenges for developing high-performance SpMM kernels on modern GPUs. Existing approaches typically focu...

Qi Du, Sheng-Le Lin, Yue-Dan Chen et al. · 0 citations
Book Open access Sep 2026

RODIS: Accelerating Sparse Matrix Multiplication on GPU via Row-Orchestration and Dynamic Instruction Scheduling

RODIS, a two-level acceleration scheme that utilizes row-orchestration at the data loading level to achieve inter-block load balancing and employs dynamic instruction scheduling at the underlying computation level to enhance MAC utilization, is proposed.

Jian-Feng Cui, Bo Yuan, Ze-Kun Jiang et al. · 0 citations
Book Open access Sep 2026

TileGEMM: Boosting the Performance of GEMM on AMX-Powered CPUs by Exploiting Data Reuse

General Matrix Multiplication (GEMM) is the cornerstone of high-performance computing and deep learning. Its efficiency significantly influences the performance of applications ranging from large language models to scientific simulations. Intel Advanced Matrix Extensions (AMX) significantly boost matrix operations thro...

Kang-Kang Chen, Hua-You Su, Meng-Han Jia et al. · 0 citations
Open access Sep 2026

DCSR-GCN: A High-Performance GCN Accelerator Based on Dynamic Compression and Sparsity Reordering

Graph Convolutional Networks (GCNs) are widely used in tasks involving irregular graph data, such as recommendation. The hybrid execution pattern of sparse aggregation and dense combination during inference limits the efficiency of general processors like CPU and GPU. Therefore, designing dedicated accelerators for GCN...

Jun-Sheng Chang, Yi-Min Zhao, Yu-Xin Huang et al. · 0 citations
Open access Aug 2026

DistSpMM: Accelerating Sparse Matrix Dense Matrix Multiplication on GPUs

The proposed DistSpMM proposes DistSpMM, a co-design framework integrating data layout, pipelining, and communication strategies, which introduces HSDMA, a lightweight algorithm that reduces communication by optimizing the dense matrix allocation.

Junyu Gu, Jue Wang, Zhikuang Xin et al. · 0 citations
Book Open access Aug 2026

DB-SpMSpV: Dual-View Blocked Sparse Matrix-Sparse Vector Multiplication for Dynamic GPU Workloads

DB-SpMSpV is presented, a dual-view blocked SpMSpV framework for dynamic GPU workloads that uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance and is integrated into DB-BFS and DB-Decoding.

Xing Cong, Chen-Hao Xie, Rui Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.