VQ-Transplant democratizes quantization research by enabling resource-efficient integration of novel VQ techniques while matching industry-level reconstruction performance.
Abstract
Vector Quantization (VQ) underpins modern discrete visual tokenization. However, training quantization modules for state-of-the-art VQ-based models requires significant computational resources which, in practice, all but prevents the development of novel, cutting-edge VQ techniques under resource constraints. To address this limitation, we propose {\bf VQ-Transplant}, a simple framework that enables plug-and-play integration of new VQ modules into frozen, pre-trained tokenizers by replacing their native VQ modules. Crucially, the proposed transplantation process preserves all encoder-decoder parameters, obviating the need for costly end-to-end retraining when modifying the quantization method. To mitigate decoder-quantization mismatch, we introduce a lightweight decoder adaptation strategy (trained for only 5 epochs on ImageNet-1k) to align feature priors with the new quantization space. In our empirical evaluation, we find that VQ-Transplant allows obtaining near state-of-the-art reconstruction fidelity for industry-level models like VAR while reducing the training cost by 95\%. VQ-Transplant democratizes quantization research by enabling resource-efficient integration of novel VQ techniques while matching industry-level reconstruction performance.
The Channel-wise Token-balanced Output-Aware Clipping (CTOAC) method, which learns per-input-channel clipping bounds by minimizing a token-balanced reconstruction loss on the corresponding linear outputs, is proposed.
The “Prompt for Quantization” (P4Q) is proposed, by integrating PTQ with Parameter-Efficient Fine-Tuning (PEFT) techniques, and demonstrates that P4Q significantly enhances the performance of low-bit CLIP while reducing deployment costs.
H. Sun, Runqi Wang, Yanjing Li et al.· ACM Transactions on Multimed...· 0 citations
Extensive experiments demonstrate that the on-device latency-informed design combined with the tailored training strategy establishes a new state-of-the-art for efficient LVLM encoding, significantly outperforming existing encoder-centric baselines while operating on-device at nearly 1.7xthe speed.
Ioannis Maniadis Metaxas, Adrian Bulat, Alberto Baldrati et al.· arXiv.org· 0 citations
Large-scale vision-language models (VLM) such as CLIP enable strong open-vocabulary reasoning, yet deploying these capabilities on resource-constrained edge devices remains challenging. EdgeVL addresses this problem by distilling CLIP representations into lightweight multi-modal encoders and applying quantization-aware...
Jinwoo Jeon, GyuYeop Do, Yunkyu Lim et al.· 0 citations
This work systematically identifies the computational bottlenecks and proposes GVC-RT, which redesigns the generative latent coding framework to realize real-time video coding without sacrificing compression performance, and introduces a lightweight de-tokenizer architecture to resolve the final latency bottleneck duri...
Tian-Jian Dang, Si-Xian Wang, Lei Luo et al.· 1 citation· ⚡1
Efficient Point Masked Autoencoders (EP-MAE), a new framework designed to significantly reduce the training cost of 3D self-supervised pre-training while maintaining strong representation quality, and provides a scalable and effective foundation for future 3D neural network models is presented.
Jian Zhu, Jiale Zhao, Cheng Lin et al.· Neural Networks· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.