Channel Group-Shared (CGS) low-rank approximation, a novel Singular Value Decomposition (SVD)-based parameter-sharing strategy, enables the feasible deployment of pre-trained large-kernel CNN models on edge devices, thereby bridging the gap between high-performance vision models and practical edge deployment.
Abstract
Large-kernel Convolutional Neural Networks (CNNs) deliver remarkable performance in vision tasks by significantly expanding receptive fields, yet their quadratic parameter growth critically impedes storage-efficient edge deployment. While existing efficient architectures adopt parameter-efficient depthwise separable convolution backbones that leverage techniques like low-rank approximation and weight sharing to compress depthwise convolutions, we identify a critical oversight: pointwise convolutions dominate parameter volume (>87% in models like RepLKNet-31B) and constitute the primary deployment bottleneck on resource-constrained edge devices. This results in prohibitive storage costs and severe memory-loading constraints on resource-limited devices (e.g., smartphones with 4-12 GB Random Access Memory (RAM)). To overcome this, we propose Channel Group-Shared (CGS) low-rank approximation, a novel Singular Value Decomposition (SVD)-based parameter-sharing strategy. CGS constructs a structured low-rank paradigm isomorphic to SVD decomposition, comprising shared (high-parameter-cost) down/up-projection matrices across channel groups within a layer and channel-group-specific (low-parameter-cost) scalable diagonal matrices. This group-sharing design achieves significant parameter reduction. Extensive experiments demonstrate that large-kernel CNNs (RepLKNet, ConvNeXt, SLaK) enhanced with CGS strike an empirically favorable balance between competitive performance and substantially reduced storage costs. Crucially, by alleviating storage constraints, reducing memory bandwidth pressure during loading, and minimizing model loading latency, CGS enables the feasible deployment of pre-trained large-kernel CNN models on edge devices, thereby bridging the gap between high-performance vision models and practical edge deployment.
VC-Attention is proposed, a training-free low-bit attention framework that addresses diffusion Transformers and quantization scale by pairing Value smoothing with a fused probability Cast, and improves fidelity over low-bit baselines.
By making DWC compute-efficient, Flash-DWC promotes the use of larger and more expressive filters and extends forward and backward propagation to both forward and backward propagation for efficient end-to-end training.
Zhiyi Zhang, Yang Zhao, Jing-Wei Sun et al.· ACM Transactions on Architec...· 0 citations
The reformulation preserves the benefits of WTConv while substantially reducing its execution time and memory footprint, removing the systems overhead that previously limited its practical efficiency.
Amit Aflalo, Shahaf E. Finder, Roy Amoyal et al.· 0 citations
This work proposes a unified compression framework, SaP, that jointly leverages basis sharing and coefficient pruning to reduce redundancies in both convolutional filters and their resulting feature maps and highlights the potential of combining complementary compression strategies to create highly efficient, deployabl...
This paper proposes a lightweight high-resolution pose estimation network, UG-HRNet, which significantly reduces model complexity while maintaining competitive performance, serving as a practical alternative to existing popular lightweight networks.
Yong-Feng Qi, Jun-Teng Zhang· International Journal of Mac...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.