Skip to content

AVQ-Attention: Adaptive Vector-Quantized Attention

Jul 2026 · arXiv.org · Vol abs/2607.12789 · 0 citations · 41 references
Computer Science

TL;DR

This work proposes Adaptive Vector-Quantized (AVQ) Attention, which adaptively allocates codebook capacity based on attention importance and develops an implementation using custom Triton kernels that enables the full adaptive refinement process, including importance scoring, child codeword insertion, and parent contribution replacement, to be carried out within the tiled computation paradigm of Flash Attention with minimal overhead.

Abstract

The $\mathcal{O}(N^2)$ complexity of attention over $N$ tokens remains a computational bottleneck in transformer models. Vector-Quantized (VQ) attention reduces this to $\mathcal{O}(MN)$ by representing keys with $M$ codewords, but applies uniform codebook capacity regardless of where attention mass concentrates: high-attention regions of key space may be coarsely approximated while low-attention regions waste representational capacity. We propose Adaptive Vector-Quantized (AVQ) Attention, which adaptively allocates codebook capacity based on attention importance. Starting from a small set of codewords, our method identifies the most important codes during the forward pass and refines them with pre-learned child codewords, achieving fine-grained quantization where it matters most while maintaining coarse quantization elsewhere. We develop an implementation using custom Triton kernels that enables the full adaptive refinement process, including importance scoring, child codeword insertion, and parent contribution replacement, to be carried out within the tiled computation paradigm of Flash Attention with minimal overhead. Our approach maintains $\mathcal{O}(MN)$ complexity while achieving improved accuracy-efficiency trade-offs compared to fixed-codebook VQ-attention.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

HyQuant: Hybrid-Precision Quantization for LLM Attention

Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention.

Jiarui Ding, Bin Xing, Yu Zhang et al. · 0 citations
Preprint Aug 2026

SQuad: Sub-Quadratic Attention Distillation for Efficient Video Generation

This work proposes SQuad, a Sub-Quadratic Attention Distillation framework that achieves a complexity of $\mathcal{O}(n\sqrt{n})$ in the resulting distilled Attention, naturally balancing the efficiency v/s expressivity trade-off.

Animesh Karnewar, Denis Korzhenkov, A. Habibian et al. · 0 citations
Preprint Aug 2026

SoftWater: Class-Aware Rate Allocation for Softmax Quantization

Across five models from 1B to 32B, SoftWater outperforms the released WaterSIC quantizer at matched head rates on 59 of 60 test points, using none of that pipeline's refinements and cutting head-induced KL by $6.5\times$--$8.3\times at 2 bits.

Joao V. Cavalcanti, Ashia C. Wilson · 2 citations
Jul 2026

ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers

ELSAA is proposed, an efficient low-rank and sparse approximation of attention that gives a practical framework for constructing low-rank and sparse attention outputs without materializing the full quadratic score matrix, aiming to enable longer-context training while preserving both sharp token-level interactions and...

Mahdi Heidari, Mohammad Mahdi Rahimi, Jaekyun Moon · 0 citations
#machine learning Preprint Sep 2026

EFQ-Softmax: Exp-Free Quantization for Softmax

EFQ-Softmax (Exp-Free Quantization for Softmax), a low-bit probability-generation method that directly maps shifted attention scores to block-scaled E2M1 operands, is proposed and shown to replace the conventional exp-then-quantize path while preserving end-to-end model quality.

Hao-Hui Han, Yu-Ming Wan, Hong-Ning Wang et al. · 0 citations
Preprint Aug 2026

APT: Accelerating Diffusion Transformers via Attention Probability-Guided Pruning and Quantization

APT, a software-hardware co-designed accelerator for high-resolution DiTs that leverages attention probabilities as a unified importance metric to jointly optimize computation through fine-grained pruning and adaptive precision scaling, and is evaluated on SOTA DiT models including PixArt, Stable Diffusion 3, and FLUX.

Sungyeob Yoo, Seeyeon Kim, Joonyong Park et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.