An Efficient HLS-Based Hardware Accelerator with Resource Optimization for Transformer Models
Abstract
Deploying Transformer models on FPGA and System-on-Chip (SoC) platforms remains challenging due to their substantial computational complexity, large memory footprint, and high hardware resource requirements, particularly in multi-head attention and stacked encoder-decoder layers. This paper proposes a hardware-efficient Transformer acceleration framework based on high-level synthesis (HLS) for resource-constrained FPGA platforms. The proposed framework introduces a reuse-oriented architecture that minimizes redundant hardware instantiations across Transformer components with identical computational behaviors. Instead of implementing independent processing units for each attention head, a shared attention engine is iteratively reused across multiple heads to reduce arithmetic and memory overhead. In addition, shared hardware blocks are employed for Add & Norm operations as well as encoder and decoder layers through iterative execution, significantly reducing logic utilization and on-chip buffering requirements while preserving the original Transformer computation flow. To further improve hardware efficiency, low-precision quantization techniques, including 16-bit Fixed point and 8-bit Integer, are integrated to reduce hardware computational cost, memory bandwidth demand, and power consumption. The proposed framework is evaluated using BERT (MRPC benchmark) and Transformer-based models on FPGA platforms. Experimental results show that the proposed design reduces BRAM utilization by up to 87.5%, LUT utilization by 80.4%, and FF utilization by 64.2% compared with conventional 32-bit Floating point implementations, while 8-bit Integer quantization reduces power consumption from 1.83 W to 1.01 W and increases the maximum operating frequency from 124.08 MHz to 262.12 MHz.