A high-speed CNN classification model with full layer-wise control is proposed, enabling seamless streaming of pixel data across all CNN layers via line buffers, resulting in high classification speed, low-power, and low energy consumption per classification.
Abstract
Convolutional Neural Networks (CNNs) are extensively used in advanced image processing applications. However, their computational complexity makes real-time deployment challenging. In this research, a high-speed CNN classification model with full layer-wise control is proposed, enabling seamless streaming of pixel data across all CNN layers via line buffers. The proposed architecture employs a novel, optimized, pipelined, layer-wise streaming control mechanism that enables concurrent execution across multiple layers and achieves ultra-high processing speed suitable for real-time applications. The learnable parameters (weights and biases) obtained after training the Optimized CNN (O-CNN) model are used to design the O-CNN model’s hardware architecture. These parameters are stored on-chip to eliminate frequent fetching of data from memory and thereby improve speed while reducing power consumption. Furthermore, line-buffer-based dataflow and on-chip storage of fixed hardware parameters minimize memory access overhead, resulting in high classification speed, low-power, and low energy consumption per classification. The operations of different layers, clock cycles requirements, and timing behaviour are verified through HDL simulation. The proposed implementation uses a signed 32-bit, Q18.13 fixed-point number format and is implemented on the Xilinx ZCU104 FPGA board. With 796 cycles at a 15 ns clock (66.67 MHz) and a positive worst slack of 1.439 ns, the design achieves a processing speed of $11.94~\mu $ s per image while consuming only 1.745 W of power.
Convolutional Neural Networks (CNNs) are the most common deep learning architecture used for video pro-cessing enhancement. Particularly, the Multi-Deep Convolutional Neural Network (MD-CNN) model, embedded into the 3D-HEVC encoder, was able to extract the optimal CTU partition structure efficiently in the depth map an...
Nacir Omran, Amna Maraoui, I. Werda et al.· International Journal of Adv...· 0 citations
CNN inference on edge hardware is constrained by memory bandwidth, redundant logic, and the area overhead of fully parallel multiply-accumulate (MAC) units. This paper presents a hardware-efficient CNN accelerator with an SRAM-based architecture optimised for classifying 28×28 grayscale images. Input images are ingeste...
Avinash Krishna Pk, R. S· International Conference Inn...· 0 citations
Real-time image segmentation is essential for edge-based intelligent systems. Deep learning models are efficient for image segmentation compared to conventional techniques. Among other deep neural networks (DNN), MobileNet convolution neural network (CNN) model utilizes depthwise separable convolutions to reduce comput...
Shreyas V, Sinchana P. Shetty, Sripriya B. S. et al.· International Conference on...· 0 citations
A flexible and scalable hardware/software framework for training CNNs on Ethernet-connected FPGA clusters using tightly pipelined layer parallelism, highlighting the potential of FPGA-based, layer-parallel training for separable-convolution-dominated CNNs.
Philipp Kreowsky, Justin Knapheide, B. Stabernack· ACM Transactions on Reconfig...· 1 citation
The increasing adoption of modern embedded platforms, edge devices, and AI driven systems has led to higher computational demands. To facilitate that, there should be hardware acceleration techniques capable of delivering higher throughput with minimal latency. Most of the traditional hardware accelerator architectures...
K. O. Y. N. Karunanayake, H. D. I. J. A. Deshapriya, A. T. Saiamirthan et al.· Moratuwa Engineering Researc...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.