SpMM-GO: Sparse Matrix-Matrix Multiplication Acceleration via Hybrid Gustavson and Outer-Product Dataflows
Abstract
Sparse matrix-matrix multiplication (SpMM) is critical for graph analytics and learning tasks, yet its irregular sparsity poses challenges for hardware acceleration. Traditional static dataflows fail to adapt to local sparsity variations, causing load imbalance. We introduce SpMM-GO, a hybrid FPGA accelerator integrating Gustavson’s algorithm and Outer-Product dataflows. A Dynamic Density-Aware Dispatcher assigns tiles at runtime: sparse tiles go to Gustavson dataflow PEs to minimize reduction overhead, while dense tiles go to Outer-Product dataflow PEs to maximize computational density. To ensure load balancing, we employ a feedback-driven scheduling algorithm that adjusts dispatch thresholds based on PE workload. Experiments show SpMM-GO achieves a geometric-mean speedup of <inline-formula> <tex-math notation="LaTeX">$10.5\times $ </tex-math></inline-formula> (up to <inline-formula> <tex-math notation="LaTeX">$25.4\times $ </tex-math></inline-formula>) over a CPU baseline at the reference 2S+2D configuration, demonstrating the high efficiency and adaptability of our dynamic hybrid dataflow architecture.