Sep 2026· IEEE Transactions on Very Large Scale Integration (vlsi) Systems· Vol 34, pp. 2832-2845· 0 citations· 39 references
Abstract
Transformers outperform traditional neural networks but face high computational and memory costs, limiting edge device deployment. Although many hardware accelerators aim to address this, the original Transformer structure still restricts the optimization effect. A recent breakthrough, mixture-of-depths (MoDs), employs conditional computation and effectively reduces the computational complexity of large language models, providing a valuable opportunity for designing an efficient hardware accelerator. However, when applied to vision transformers, MoD suffers from accuracy degradation and excessive external memory access (EMA). Therefore, this article presents ME-MoD, the first memory-efficient MoD-based vision transformer inference accelerator, leveraging the idea of reordering and algorithm-hardware-dataflow codesign. Algorithmically, distribution adjustment forward (DAF) and routing decision forward (RDF) techniques restore accuracy and alleviate memory access costs through token reordering. Architecturally, a LayerNorm-Routing (L-R) fusion module and a token reordering and sequential recording module enhance computational efficiency while minimizing memory overhead. In addition, a token-stationary layer fusion dataflow and an on-chip dynamic memory module are designed, which further optimizes the EMA caused by the intermediate results of interlayer computation of valid tokens routed by MoD. With negligible accuracy loss, our ME-MoD accelerator achieves $1.62\times $ inference speed up, eliminates 46.5% of the external memory bandwidth requirement and 45.2% of energy consumption compared with standard MoD. It achieves 23.6 TOPS/W energy efficiency, which is $4.02\times $ improvements compared with state-of-the-art (SOTA) designs.
The efficient implementation of the softmax is critical for optimizing transformer hardware accelerators. Unlike its role as a static, one-time classifier in convolutional neural networks (CNNs), softmax in transformers is core to achieving dynamic contextual awareness, generating attention weights that enable the mode...
Bangzheng He, Bang-Xin Qin, Han Wang et al.· IEEE Transactions on Very La...· 0 citations
A systematic review of FPGA-based DL deployment from a cross-layer perspective spanning model, compiler, architecture, runtime, and electronic design automation (EDA) is presented, highlighting that reliable cross-study comparison requires careful consideration of model configuration, precision, execution phase, batch...
Shuo Wang, Lei Chen, Chunsheng Tian et al.· Electronics· 0 citations
A memory-efficient hardware accelerator for depthwise separable convolution that minimizes off-chip memory traffic and parameter storage and performs the depthwise separable convolution with a small number of logic resources and on-chip memory, confirming its feasibility for resource-constrained edge devices.
Jaeseong Kim, Taehong Min, Chaebin Lee et al.· Electronics· 0 citations
This paper presents a low-power sparse convolution accelerator for edge devices, fabricated and validated in a 16 nm process that adopts a bitmap-based format for compression in both data transmission and computation, effectively reducing memory and bandwidth overhead.
Jingyue Zhuge, Johannes Partzsch, Christian Mayr· arXiv.org· 0 citations
FluxBin is proposed, an algorithm-kernel co-design that synergizes post-training quantization with a highly optimized CUDA kernel and introduces Decoupled Row-Column Binary Decomposition to enhance representational capacity while maintaining hardware efficiency.
Qingyao Yang, Run-Ming Yang, He Xiao et al.· 0 citations