Skip to content
Conference

An FPGA-Based Unified Processing Element for INT8/Binary Quantization and its Scalable Array Design

Aug 2026 · 2026 2nd International Conference on Electronic Information, Computer and Aerospace Remote Sensing (EICARS) · pp. 415-419 · 0 citations · 13 references

Abstract

The widespread deployment of deep neural networks on edge devices faces a severe imbalance between computational demand and available power, while devices frequently switch between low-power standby and highperformance detection modes. Existing general-purpose processors, graphics processing units, and fixed-precision application-specific integrated circuits struggle to reconcile energy efficiency with architectural flexibility. This paper presents an FPGA-based INT8/Binary quantized Unified Processing Element (UPE) and its scalable array. The design decomposes INT8 two's-complement multiplication into eight signed partial products, maps Binary multiply-accumulate operations to eight masked-XNOR terms, and uses input-side mode selection to share one three-level balanced adder tree and one 32-bit accumulator, thereby supporting high-precision and bit-level-parallel computation in a single core. Implementationmatched extension experiments further quantify the LUT, FF, and CARRY4 changes of high-width C3 relative to C0, exposing the implementation cost of a wider reduction tree and accumulated state. Compared with a Separate Dual PE composed of independent INT8 and Binary PEs, the proposed UPE reduces LUTs, FFs, and CARRY4s by 8.90%, 50.00%, and 28.57%, respectively. It reaches 131.579 MHz with a 1.32% frequency cost; the parameterized $4 \times 4$ and $8 \times 8$ arrays reach 119.048 MHz and 111.111 MHz. At the same array scale, Binary mode reduces total power by 31.17% and 55.06% relative to INT8 mode on the $\mathbf{4} \times \mathbf{4}$ and $\mathbf{8} \times \mathbf{8}$ arrays. On LeNet-5/MNIST, the FP32, INT8-quantized, and Binary-quantized models achieve test accuracies of 99.54%, 99.52%, and 96.30%. These results show that the unified structure removes duplicated dual-precision resources at a limited timing cost and offers a scalable FPGA core for flexible switching between low-power and high-performance modes on edge devices.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.