Discrete visual tokenization, predominantly driven by vector, scalar, and product quantization, lacks a unified conceptual framework for understanding quantization tradeoffs. In this paper, we propose a unified rate--distortion perspective on modern discrete visual tokenization. By viewing quantization as lossy compres...
Xianghong Fang, Wenlong Mou, Yuan Yuan et al.· 0 citations
VQ-Transplant democratizes quantization research by enabling resource-efficient integration of novel VQ techniques while matching industry-level reconstruction performance.
Xianghong Fang, Yuan Yuan, Dehan Kong et al.· arXiv.org· 1 citation
This work introduces principled criteria for desirable VQ behavior and demonstrates that aligning feature and code vector distributions provides a unifying mechanism for mitigating training instability and codebook collapse, and instantiate this framework using a Wasserstein-based objective with an efficient closed-for...