Skip to content

A Low-Power Sparse Convolution Accelerator with Idle-First-Task-Assignment for Edge Vision

Jul 2026 · arXiv.org · Vol abs/2607.26835 · 0 citations · 25 references
Computer Science

TL;DR

This paper presents a low-power sparse convolution accelerator for edge devices, fabricated and validated in a 16 nm process that adopts a bitmap-based format for compression in both data transmission and computation, effectively reducing memory and bandwidth overhead.

Abstract

In recent years, edge-vision monitoring systems for applications such as smart animal husbandry have faced strict tripartite constraints: maintaining input resolution under extremely limited transmission bandwidth and strict power budgets. Conventional dense convolutional neural networks (CNNs) cannot satisfy the resource limits of such constrained IoT nodes. To address this challenge, this paper presents a low-power sparse convolution accelerator for edge devices, fabricated and validated in a 16 nm process. First, the accelerator adopts a bitmap-based format for compression in both data transmission and computation, effectively reducing memory and bandwidth overhead. Second, to mitigate load imbalance in sparse computation, an Idle-First-Task-Assignment (IFTA) dynamic scheduling strategy is proposed, significantly reducing processing-element (PE) idle time and improving multiplier utilization. In addition, a dedicated dataflow is designed to support and accelerate depthwise separable convolution (DWConv), which is widely used in lightweight networks. Experimental results show that the chip occupies only 0.5~mm$^2$ core area and consumes as little as 12--16~mW. On ImageNet, for sparse VGG16 and MobileNetV2, the proposed accelerator achieves 6.5$\times$ and 2.8$\times$ speedups, respectively, over traditional dense accelerators, and also delivers significant performance gains over the existing sparse accelerator.

View source

Similar papers

Open access Aug 2026

A Memory-Efficient Depthwise Separable Convolution Accelerator Using Run-Length Coding

A memory-efficient hardware accelerator for depthwise separable convolution that minimizes off-chip memory traffic and parameter storage and performs the depthwise separable convolution with a small number of logic resources and on-chip memory, confirming its feasibility for resource-constrained edge devices.

Jaeseong Kim, Taehong Min, Chaebin Lee et al. · 0 citations
Conference Open access 2026

Key Algorithms of Convolutional Neural Networks and Hardware Implementation of Image Processing

Edge computing and artificial intelligence have made the efficient deployment of machine vision algorithms on low-power hardware a critical challenge for integrated circuit design. Given data-intensive image pixels and deep neural network tensors, traditional von Neumann architectures inevitably encounter severe memory...

Linenxu Zhang · 0 citations
Aug 2026

Hardware-Aware Neural Network Deployment on Multi-Core in-Memory Computing Systems: A Compiler Perspective

Conventional compute-in-memory (CIM) deployment flows usually assume a fixed crossbar geometry, although DNN layers often have different channel counts and matrix shapes. The resulting shape mismatch leaves part of the array capacity unused and can increase the number of split-and-transfer operations during inference....

Kaiwen Deng, Sifan Sun, Hanjie Liu et al. · 0 citations
Book Open access Aug 2026

A Differentiable Simulator for Optimizing Time-Domain Analog CNN Accelerators

Edge intelligence promises responsive, private, and energy-efficient sensing without continual dependence on remote compute. This demands convolutional neural network (CNN) accelerators that deliver substantially higher throughput and energy efficiency than conventional digital pipelines while preserving high accuracy....

Mark Horton, Changwoo Park, Tergel Molom-Ochir et al. · 0 citations
Conference Aug 2026

Designing and Building an FPGA Accelerator That Uses Less Energy for DNN Inference

Deep Neural Networks (DNNs) are critical to modern AI applications, yet their deployment on standard CPUs and GPUs is constrained by high power consumption and computational latency, particularly in resource-constrained edge environments. To address these limitations, this paper presents the design and implementation o...

P. V. G. K. Rao, Dudekula Raziya · 0 citations
Sep 2026

A High-Performance Neural Rendering Accelerator With Dual-Lane Micro-MLPs and Hierarchical Latency-Hiding Scheduling

While Neural Radiance Fields (NeRF) have transformed 3D vision, their prohibitive computational and memory demands restrict real-time deployment on power-constrained edge devices. To bridge this gap, we propose a novel neural rendering accelerator that orchestrates three architectural innovations to maximize throughput...

Cheng Zhang, Yuefeng Zhang, Wenkai Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.