Sep 2026· ACM Transactions on Architecture and Code Optimization (TACO)· 0 citations· 23 references
TL;DR
This work proposes the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support the dataflow paradigm while enhancing data reuse, and decouples the multiplication and merging phases.
Abstract
Sparse Matrix-Sparse Matrix Multiplication (SpMSpM) is a crucial computational kernel widely used in scientific computing and machine learning. The varying sparse patterns across different matrices pose significant challenges for conventional accelerators with fixed dataflow architectures. Although recent studies have explored dynamic dataflow approaches to better capture memory access patterns under diverse sparsity conditions, these solutions still struggle to simultaneously improve data reuse, load balance, and merging efficiency. To address these limitations, we present an SpMSpM accelerator based on adaptive dataflow and efficient merging (ADEM). ADEM is carefully designed from four key aspects. First, we propose the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support our dataflow paradigm while enhancing data reuse. We then present a cache-aware dataflow to mitigate memory overflow issues. Furthermore, ADEM decouples the multiplication and merging phases, utilizing the SFT structure to enable fine-grained task scheduling for improved load balancing. Finally, the accelerator incorporates heterogeneous merging units specifically optimized for handling two distinct types of merging operations, thereby significantly improving merger utilization. Compared with the state-of-the-art baseline system, ADEM achieves average speedups of 1.25 ×, 1.73 ×, and 1.16 × on VGG-16, ResNet-50, and SuiteSparse workloads, respectively.
Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in scientific computing and artificial intelligence. However, the inherent sparsity and irregularity of real-world datasets present significant challenges for developing high-performance SpMM kernels on modern GPUs. Existing approaches typically focu...
Qi Du, Sheng-Le Lin, Yue-Dan Chen et al.· Proceedings of the Internati...· 0 citations
RODIS, a two-level acceleration scheme that utilizes row-orchestration at the data loading level to achieve inter-block load balancing and employs dynamic instruction scheduling at the underlying computation level to enhance MAC utilization, is proposed.
Jian-Feng Cui, Bo Yuan, Ze-Kun Jiang et al.· Proceedings of the Internati...· 0 citations
General Matrix Multiplication (GEMM) is the cornerstone of high-performance computing and deep learning. Its efficiency significantly influences the performance of applications ranging from large language models to scientific simulations. Intel Advanced Matrix Extensions (AMX) significantly boost matrix operations thro...
Kang-Kang Chen, Hua-You Su, Meng-Han Jia et al.· Proceedings of the Internati...· 0 citations
Graph Convolutional Networks (GCNs) are widely used in tasks involving irregular graph data, such as recommendation. The hybrid execution pattern of sparse aggregation and dense combination during inference limits the efficiency of general processors like CPU and GPU. Therefore, designing dedicated accelerators for GCN...
Jun-Sheng Chang, Yi-Min Zhao, Yu-Xin Huang et al.· ACM Transactions on Design A...· 0 citations
The proposed DistSpMM proposes DistSpMM, a co-design framework integrating data layout, pipelining, and communication strategies, which introduces HSDMA, a lightweight algorithm that reduces communication by optimizing the dense matrix allocation.
Junyu Gu, Jue Wang, Zhikuang Xin et al.· ACM Transactions on Architec...· 0 citations
DB-SpMSpV is presented, a dual-view blocked SpMSpV framework for dynamic GPU workloads that uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance and is integrated into DB-BFS and DB-Decoding.
Xing Cong, Chen-Hao Xie, Rui Wang et al.· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.