Jul 2026· International Journal of AI Electronics and Nexus Energy· Vol 2, pp. 124-127· 0 citations· 1 references
TL;DR
A combined hardware-andalgorithm design that pairs a carefully regularized binarized neural network (BNN) with a compact, variable-precision FPGA processor to enable accurate, real-time object detection without relying on off-chip memory is proposed.
Abstract
Convolutional neural networks (CNNs) achieve strong accuracy in object detection, but their heavy computation and memory demands make them difficult to run on embedded and mobile hardware. This work proposes a combined hardware-andalgorithm design that pairs a carefully regularized binarized neural network (BNN) with a compact, variable-precision FPGA processor. A new building block called DenseToRes is introduced to reduce the accuracy loss usually caused by aggressive 1-bit quantization. The supporting processor stores the entire trained network in on-chip memory instead of external DRAM and performs its multiplyaccumulate (MAC) operations using an XNOR-based processing element that can flexibly handle 1-, 2-, 4- and 8-bit operand widths from one shared gate array. Implemented on a Xilinx FPGA, the design achieves real-time performance of 64.51 frames per second with 64.92% mean average precision (mAP) on the PASCAL VOC dataset, while consuming only 6.58 W. The results show that jointly optimizing network structure and hardware precision can enable accurate, real-time object detection without relying on off-chip memory.
This paper presents a hardware-efficient object detection accelerator based on XNOR-driven variable-precision computation for real-time edge artificial intelligence. The proposed network combines DenseToRes and transition layers to preserve feature information under aggressive quantization. Binary convolution is execut...
Javeed Md, Srinivasa Reddy Dumpa, K. Saisri et al.· Adolescência e Saúde· 0 citations
Field-programmable gate arrays (FPGAS) have emerged as a powerful platform for real-time image processing due to their inherent parallelism and configurability. This paper presents an optimized hardware implementation of fundamental image processing algorithms including Sobel edge detection, Thresholding contrast stret...
Pramod Moud, P. Sharma· International Journal of Lat...· 0 citations
Object detection at the edge requires a difficult balance among detection accuracy, deterministic latency, memory bandwidth, and energy consumption. Existing binarized accelerators replace multipliers with XNOR and population-count logic, but many designs use a fixed binary datapath or select precision only at the laye...
Budidha Sriman and D. Sateesh· International Journal of Adv...· 0 citations
In an effort to mitigate processing delays and latency in the traditional edge detection in the vision based systems such as robotics and surveillance platforms, this paper attempts to introduce an FPGA-based 3×3 convolution accelerator. The proposed architecture employs a multiply-accumulate (MAC) unit and fixed-point...
Aysha Pareekutty, R. Megalingam· International Conference on...· 0 citations
The widespread deployment of deep neural networks on edge devices faces a severe imbalance between computational demand and available power, while devices frequently switch between low-power standby and highperformance detection modes. Existing general-purpose processors, graphics processing units, and fixed-precision...
Zi-Qing Mai, Zhan-Peng Jiang· 2026 2nd International Confe...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.