This work extends the im2win paradigm -- a universal, memory-efficient convolution method with contiguous memory access for all kernel sizes -- to run efficiently in full precision on CUDA cores and half precision on tensor cores, establishing im2win as a unified, high-performance convolution framework for modern GPU architectures.
Abstract
Convolution is a principal computational bottleneck in deep neural networks, and its efficiency depends on tight integration between algorithms and GPU hardware. Existing GPU convolution methods suffer from large memory overhead, poor cache utilization, limited effectiveness across kernel sizes, or numerical instability. This work extends the im2win paradigm -- a universal, memory-efficient convolution method with contiguous memory access for all kernel sizes -- to run efficiently in full precision on CUDA cores and half precision on tensor cores. By introducing new kernel designs and optimizations such as zig-zag memory access and asynchronous data movement, im2win efficiently exploits hardware-accelerated half-precision matrix multiply-accumulate operations. Across twelve CNN benchmarks, im2win achieves up to 2.8x higher TFLOPS than its CUDA core implementation, 1.4x higher than cuDNN, and 6.4x higher than GEMM-based convolution with cuBLAS, while using as little as 53% and 35% of their memory, respectively. These results establish im2win as a unified, high-performance convolution framework for modern GPU architectures.
By making DWC compute-efficient, Flash-DWC promotes the use of larger and more expressive filters and extends forward and backward propagation to both forward and backward propagation for efficient end-to-end training.
Zhiyi Zhang, Yang Zhao, Jing-Wei Sun et al.· ACM Transactions on Architec...· 0 citations
Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in scientific computing and artificial intelligence. However, the inherent sparsity and irregularity of real-world datasets present significant challenges for developing high-performance SpMM kernels on modern GPUs. Existing approaches typically focu...
Qi Du, Shengle Lin, Yuedan Chen et al.· Proceedings of the Internati...· 0 citations
A memory-efficient hardware accelerator for depthwise separable convolution that minimizes off-chip memory traffic and parameter storage and performs the depthwise separable convolution with a small number of logic resources and on-chip memory, confirming its feasibility for resource-constrained edge devices.
Jaeseong Kim, Taehong Min, Chaebin Lee et al.· Electronics· 0 citations
This paper proposes LayUp, a layer-wise training optimization framework for edge devices with unified memory architectures that achieves speedup and energy reduction compared to baseline GPU-only training for GPT-2 models on the NVIDIA Jetson Orin NX, while conventional offloading increases latency.
Bang-San Lee, Young-Ho Gong· International Symposium on L...· 0 citations
Tiny-Pipe comprises a holistic layer packing method that simultaneously reduces GPU memory footprint and improves training performance, an active CPU memory management that alleviates CPU memory pressure by eliminating redundant parameters, and a layer-wise runtime swapping strategy that further enhances overall perfor...
Sparse-dense matrix multiplication (SpMM), a fundamental computational kernel in graph analytics and scientific computing, can be substantially accelerated on modern GPUs by leveraging both dense and sparse Tensor Core units, thereby enabling high-throughput computation. However, existing methods often fail to exploit...
Zhi-Rui Chen, Heng Zhang, Kai-Fan Jia· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.