Skip to content
Open access

A Memory-Efficient Depthwise Separable Convolution Accelerator Using Run-Length Coding

Aug 2026 · Electronics · 0 citations · 14 references

TL;DR

A memory-efficient hardware accelerator for depthwise separable convolution that minimizes off-chip memory traffic and parameter storage and performs the depthwise separable convolution with a small number of logic resources and on-chip memory, confirming its feasibility for resource-constrained edge devices.

Abstract

On edge devices, convolutional neural network (CNN) inference is bottlenecked mainly by memory bandwidth, owing to the frequent memory accesses to feature maps and parameters. To address this challenge, we propose a memory-efficient hardware accelerator for depthwise separable convolution that minimizes off-chip memory traffic and parameter storage. The proposed architecture employs three key techniques: (1) a 64-bit run-length coding (RLC) packet compression that exploits feature-map sparsity after ReLU, (2) a mixed-precision scheme that represents feature maps and weights at different precisions, and (3) separated depthwise and pointwise convolution units. In particular, feature maps are transferred in RLC-compressed form, which reduces the amount of data exchanged with the host. The compressed data are decoded row by row, so the on-chip buffers hold only the rows required for computation rather than a complete feature map. In software simulation on ImageNet, the mixed-precision scheme reduced the parameter storage by 49.22% at a cost of 6.24 percentage points (pp) in Top-1 accuracy, and the RLC reduced the data by up to 56.07% in the deeper layers. Implemented on a Xilinx ZCU-104 FPGA, the proposed accelerator performs the depthwise separable convolution with a small number of logic resources and on-chip memory, confirming its feasibility for resource-constrained edge devices.

Read PDF

Similar papers

Open access Aug 2026

Flash-DWC: Making Depthwise Convolution Compute-Efficient on GPUs

By making DWC compute-efficient, Flash-DWC promotes the use of larger and more expressive filters and extends forward and backward propagation to both forward and backward propagation for efficient end-to-end training.

Zhiyi Zhang, Yang Zhao, Jing-Wei Sun et al. · 0 citations
Open access Aug 2026

Hardware-level data layout approach to mitigate the memory row conflicts on FPGA-based CNN accelerators

A hardware-level data layout technique using a memory-centric accelerator architecture that improves memory performance in FPGA-based CNN accelerators in a scalable and hardware-efficient manner without requiring a large computational burden.

S. Prasad, Suman Jayakumar, Bellary Kursheed et al. · 0 citations
Preprint Aug 2026

Enabling Memory-efficient Im2win Convolution with Multi-precision Support on GPU CUDA and Tensor Cores

This work extends the im2win paradigm -- a universal, memory-efficient convolution method with contiguous memory access for all kernel sizes -- to run efficiently in full precision on CUDA cores and half precision on tensor cores, establishing im2win as a unified, high-performance convolution framework for modern GPU a...

Xiang Fu, Jixiang Ma, Xinpeng Zhang et al. · 1 citation
Open access Aug 2026

On-chip non-volatile all-optical residual neural network accelerator

Optical neural networks (ONNs) hold substantial potential in artificial intelligence, promising faster processing speed and reduced energy consumption compared to traditional electronic neural networks, by implementing matrix operations with optical computations. Current ONN architectures predominantly rely on single-...

Zhi-Qiang Quan, Bing Han, Xiaoxiao Ma et al. · 0 citations
Conference Open access 2026

Key Algorithms of Convolutional Neural Networks and Hardware Implementation of Image Processing

Edge computing and artificial intelligence have made the efficient deployment of machine vision algorithms on low-power hardware a critical challenge for integrated circuit design. Given data-intensive image pixels and deep neural network tensors, traditional von Neumann architectures inevitably encounter severe memory...

Linenxu Zhang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.