Aug 2026· Bulletin of Electrical Engineering and Informatics· 0 citations· 26 references
TL;DR
A hardware-level data layout technique using a memory-centric accelerator architecture that improves memory performance in FPGA-based CNN accelerators in a scalable and hardware-efficient manner without requiring a large computational burden.
Abstract
Memory row conflicts (MRCs) continue to be a major bottleneck that results in higher latency, ineffective double data rate (DDR) usage, and decreased effective bandwidth in field-programmable gate array (FPGA-based) convolutional neural network (CNN) accelerators. The majority of current effort focuses on computational optimization, frequently ignoring inefficient memory access. In order to reduce memory reference codes (MRCs), this research suggests a hardware-level data layout technique using a memory-centric accelerator architecture. In order to improve hit rates and row buffer locality, the architecture incorporates a dynamic cursor-based address mapping method that adjusts to different feature map sizes across CNN layers and a dual-DDR setup for concurrent data access. The experimental results on VGG16, YOLOv2, and AlexNet show an 18% reduction in MRCs, a 40% increase in throughput, and a 23% decrease in latency compared to state-of-the-art techniques. The design uses a Xilinx Kintex-7 FPGA with low power usage of 1.52 W. The suggested method improves memory performance in FPGA-based CNN accelerators in a scalable and hardware-efficient manner without requiring a large computational burden.
A memory-efficient hardware accelerator for depthwise separable convolution that minimizes off-chip memory traffic and parameter storage and performs the depthwise separable convolution with a small number of logic resources and on-chip memory, confirming its feasibility for resource-constrained edge devices.
Jaeseong Kim, Taehong Min, Chaebin Lee et al.· Electronics· 0 citations
To mitigate interconnect scaling bottlenecks $\left(O\left(N^{2}\right)\right)$ and Non-Uniform Memory Access (NUMA) congestion in Programmable Multi-Core Accelerators (PMCAs), this paper introduces a multi-cluster architecture that replaces inter-cluster communication with localized data replication within ScratchPad...
Chanon Khongprasongsiri, P. Tanguy, Kevin J. M. Martin et al.· IEEE International Conferenc...· 0 citations
Graph Convolutional Networks (GCNs) are widely used in tasks involving irregular graph data, such as recommendation. The hybrid execution pattern of sparse aggregation and dense combination during inference limits the efficiency of general processors like CPU and GPU. Therefore, designing dedicated accelerators for GCN...
Jun-Sheng Chang, Yi-Min Zhao, Yu-Xin Huang et al.· ACM Transactions on Design A...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.