Skip to content
Preprint

Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

The reformulation preserves the benefits of WTConv while substantially reducing its execution time and memory footprint, removing the systems overhead that previously limited its practical efficiency.

Abstract

Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive field exponentially with the number of decomposition levels while keeping the parameter count linear. However, its reference implementation is severely memory-bound due to excessive data movement through high-bandwidth memory (HBM). We develop an I/O model of WTConv to characterize this bottleneck and use it to guide three algebraic reformulations: (1) recomputing the inexpensive Haar analysis butterfly on chip, (2) collapsing the multi-level synthesis cascade into a single closed-form pass indexed by output-coordinate bits, and (3) folding learned per-channel scales into the convolution weights. Together, these reformulations enable an I/O-aware fused implementation that substantially reduces HBM traffic. We evaluate the WTConvNeXt configuration across decomposition levels and a broad range of tensor shapes. Despite performing comparable arithmetic, the reference WTConv is substantially slower than the depthwise convolution it replaces. Our reformulation reduces modeled HBM traffic by approximately $2.55\times$, yielding up to a $4.35\times$ training speedup over the reference while roughly halving peak memory usage. Thus, our reformulation preserves the benefits of WTConv while substantially reducing its execution time and memory footprint, removing the systems overhead that previously limited its practical efficiency.

View source

Similar papers

Preprint Aug 2026

BaKron: Efficient Quantization with Kronecker-Factored Hessians

BaKron is an efficient solver that combines anti-diagonal parallelism with a recursive divide-and-conquer construction that matches the cubic scaling of GPTQ while exploiting richer curvature information.

Johann Birnick, R. Saab · 2 citations
Open access Aug 2026

Flash-DWC: Making Depthwise Convolution Compute-Efficient on GPUs

By making DWC compute-efficient, Flash-DWC promotes the use of larger and more expressive filters and extends forward and backward propagation to both forward and backward propagation for efficient end-to-end training.

Zhiyi Zhang, Yang Zhao, Jing-Wei Sun et al. · 0 citations
Preprint Aug 2026

Is Haar Enough? Exploring Symlets and Coiflets for Wavelet Convolution Layers

Wavelet convolution layers have recently emerged as an efficient mechanism for enlarging receptive fields through multiresolution analysis, but prior work has fixed the wavelet basis to Haar or Daubechies at a chosen decomposition depth, leaving open whether a different basis can shift the underlying efficiency frontie...

Md Rifat Ur Rahman · 0 citations
Preprint Aug 2026

MonoMoE: An Efficient Fused Mega-kernel for Quantized MoE Decoding

Mixture-of-Experts (MoE) layers increase model capacity without proportionally increasing arithmetic, but their sparse expert computation is difficult to execute efficiently during autoregressive decode. Existing grouped and batched GEMMs are token-major: they construct expert-local token tiles and obtain parallelism f...

Yu Gong, Kailash Budhathoki, Taeho Kim et al. · 0 citations
#large language models Book Open access Sep 2026

SSQT: A Hardware-Friendly Fusion Compression Framework of Structured Sparsification and Sensitivity-Driven Quantization for Large-Scale Language Models

Large language models (LLMs) are often memory-bandwidth bound during autoregressive decoding, so reducing weight storage does not automatically produce parallel speedup when the compressed representation is irregular. We present SSQT, a post-training framework that jointly applies hardware-aligned structured sparsifica...

Qian-Sheng Song, Guolin Tang · 0 citations
Review Aug 2026

Transforms for LLM Quantization: The Great Inversion and Format Co-Design

This work identifies and formalizes the principle that organizes the Great Inversion, the Great Inversion: allocation-flexible coding rewards energy concentration, whereas the grouped shared-scale quantization a deployed matrix instruction performs rewards within-group flattening.

Ehsan Jokar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.