Skip to content

Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework

Jul 2026 · arXiv.org · Vol abs/2607.15933 · 1 citation · 60 references
Computer Science

TL;DR

This work introduces principled criteria for desirable VQ behavior and demonstrates that aligning feature and code vector distributions provides a unifying mechanism for mitigating training instability and codebook collapse, and instantiate this framework using a Wasserstein-based objective with an efficient closed-form under a mild Gaussian approximation.

Abstract

The effectiveness of modern visual representation learning and autoregressive models critically depends on vector quantization (VQ), which discretizes continuous feature representations using a learnable codebook. Despite its widespread use, existing VQ methods often suffer from training instability and codebook collapse, arising from gradient mismatch induced by the straight-through estimator and the under-utilization of code vectors. In this work, we show that both issues can be traced to a fundamental mismatch between the distributions of feature vectors and code vectors, leading to inefficient representation and information loss. Building on this observation, we propose a distributional matching framework for vector quantization. We introduce principled criteria for desirable VQ behavior and demonstrate through theoretical analysis and empirical evaluation that aligning feature and code vector distributions provides a unifying mechanism for mitigating training instability and codebook collapse. We instantiate this framework using a Wasserstein-based objective with an efficient closed-form under a mild Gaussian approximation, and further show that a nonparametric alternative based on maximum mean discrepancy yields comparable performance. Extensive experiments on visual tokenization benchmarks support the effectiveness and robustness of the proposed approach.

View source

Similar papers

#machine learning Preprint Sep 2026

A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization

Discrete visual tokenization, predominantly driven by vector, scalar, and product quantization, lacks a unified conceptual framework for understanding quantization tradeoffs. In this paper, we propose a unified rate--distortion perspective on modern discrete visual tokenization. By viewing quantization as lossy compres...

Xianghong Fang, Wenlong Mou, Yuan Yuan et al. · 0 citations
Preprint Aug 2026

Gromov-Wasserstein Quantization and Clustering: Structure, Rates, and Algorithms

Numerical experiments show that GW quantization opens up many modeling possibilities beyond normal clustering methods and that the introduced algorithm leads to useful numerical solutions with approximation quality often in line with theoretically optimal rates.

F. Beier, S. Eckstein · 0 citations
Jul 2026

dRAE: Representation Autoencoder with Hyper-Spherical Codes

This work proposes Hyper-Spherical Quantization (HSQ), which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning.

Tianren Ma, Lin Long, Chu-Yan Chen et al. · 0 citations
#machine learning Preprint Sep 2026

When Does Low-Bit Quantization Preserve the Decisions of Vector Search?

Low-bit quantization can achieve high recall on some vector representations and fail sharply on others, while average distortion and global rank correlation do not explain the difference. We study quantized vector search at the level of the comparisons consumed by ranking and graph-pruning algorithms. Our first result...

Wen-Xuan Xiao, Xu Cao · 0 citations
Preprint Aug 2026

Hadamard-Domain Model Quantization for Learned Image Coding

Hadamard-Transform-domain Quantization (HaTQ), which uses orthogonal Hadamard reparameterization before quantization to redistribute weight and activation responses in the original domain across channels, and is compatible with integer-only execution.

Jun-Qi Shi, Chongzhi Wang, Yiwen He et al. · 0 citations
Preprint Aug 2026

BaKron: Efficient Quantization with Kronecker-Factored Hessians

BaKron is an efficient solver that combines anti-diagonal parallelism with a recursive divide-and-conquer construction that matches the cubic scaling of GPTQ while exploiting richer curvature information.

Johann Birnick, R. Saab · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.