Skip to content
Preprint

Enabling Memory-efficient Im2win Convolution with Multi-precision Support on GPU CUDA and Tensor Cores

Aug 2026 · 1 citation · 16 references
Computer Science

TL;DR

This work extends the im2win paradigm -- a universal, memory-efficient convolution method with contiguous memory access for all kernel sizes -- to run efficiently in full precision on CUDA cores and half precision on tensor cores, establishing im2win as a unified, high-performance convolution framework for modern GPU architectures.

Abstract

Convolution is a principal computational bottleneck in deep neural networks, and its efficiency depends on tight integration between algorithms and GPU hardware. Existing GPU convolution methods suffer from large memory overhead, poor cache utilization, limited effectiveness across kernel sizes, or numerical instability. This work extends the im2win paradigm -- a universal, memory-efficient convolution method with contiguous memory access for all kernel sizes -- to run efficiently in full precision on CUDA cores and half precision on tensor cores. By introducing new kernel designs and optimizations such as zig-zag memory access and asynchronous data movement, im2win efficiently exploits hardware-accelerated half-precision matrix multiply-accumulate operations. Across twelve CNN benchmarks, im2win achieves up to 2.8x higher TFLOPS than its CUDA core implementation, 1.4x higher than cuDNN, and 6.4x higher than GEMM-based convolution with cuBLAS, while using as little as 53% and 35% of their memory, respectively. These results establish im2win as a unified, high-performance convolution framework for modern GPU architectures.

View source

Similar papers

Open access Aug 2026

Flash-DWC: Making Depthwise Convolution Compute-Efficient on GPUs

By making DWC compute-efficient, Flash-DWC promotes the use of larger and more expressive filters and extends forward and backward propagation to both forward and backward propagation for efficient end-to-end training.

Zhiyi Zhang, Yang Zhao, Jing-Wei Sun et al. · 0 citations
Book Open access Sep 2026

CoTC-SpMM: A Cooperative Tensor–CUDA Cores Scheme for Efficient Sparse Matrix Multiplication

Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in scientific computing and artificial intelligence. However, the inherent sparsity and irregularity of real-world datasets present significant challenges for developing high-performance SpMM kernels on modern GPUs. Existing approaches typically focu...

Qi Du, Shengle Lin, Yuedan Chen et al. · 0 citations
Open access Aug 2026

A Memory-Efficient Depthwise Separable Convolution Accelerator Using Run-Length Coding

A memory-efficient hardware accelerator for depthwise separable convolution that minimizes off-chip memory traffic and parameter storage and performs the depthwise separable convolution with a small number of logic resources and on-chip memory, confirming its feasibility for resource-constrained edge devices.

Jaeseong Kim, Taehong Min, Chaebin Lee et al. · 0 citations
Book Open access Aug 2026

LayUp: Layer-wise Parallelization for Energy-Efficient Edge LLM Training Exploiting Unified Memory Characteristics

This paper proposes LayUp, a layer-wise training optimization framework for edge devices with unified memory architectures that achieves speedup and energy reduction compared to baseline GPU-only training for GPT-2 models on the NVIDIA Jetson Orin NX, while conventional offloading increases latency.

Bang-San Lee, Young-Ho Gong · 0 citations
Open access Aug 2026

GPU and CPU Memory Co-Optimization in Heterogeneous Pipeline Parallelism for Efficient Large Language Model Fine-Tuning on Commodity Servers

Tiny-Pipe comprises a holistic layer packing method that simultaneously reduces GPU memory footprint and improves training performance, an active CPU memory management that alleviates CPU memory pressure by eliminating redundant parameters, and a layer-wise runtime swapping strategy that further enhances overall perfor...

Yu-Quan Ding, Jie Shao · 0 citations
Book Open access Sep 2026

AFH-SpMM: Auto-Fit Heterogeneous Block Sparse-Dense Matrix Multiplication on Tensor Core GPUs

Sparse-dense matrix multiplication (SpMM), a fundamental computational kernel in graph analytics and scientific computing, can be substantially accelerated on modern GPUs by leveraging both dense and sparse Tensor Core units, thereby enabling high-throughput computation. However, existing methods often fail to exploit...

Zhi-Rui Chen, Heng Zhang, Kai-Fan Jia · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.