Skip to content

VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers

Jul 2026 · arXiv.org · Vol abs/2607.19575 · 1 citation · 58 references
Computer Science

TL;DR

VQ-Transplant democratizes quantization research by enabling resource-efficient integration of novel VQ techniques while matching industry-level reconstruction performance.

Abstract

Vector Quantization (VQ) underpins modern discrete visual tokenization. However, training quantization modules for state-of-the-art VQ-based models requires significant computational resources which, in practice, all but prevents the development of novel, cutting-edge VQ techniques under resource constraints. To address this limitation, we propose {\bf VQ-Transplant}, a simple framework that enables plug-and-play integration of new VQ modules into frozen, pre-trained tokenizers by replacing their native VQ modules. Crucially, the proposed transplantation process preserves all encoder-decoder parameters, obviating the need for costly end-to-end retraining when modifying the quantization method. To mitigate decoder-quantization mismatch, we introduce a lightweight decoder adaptation strategy (trained for only 5 epochs on ImageNet-1k) to align feature priors with the new quantization space. In our empirical evaluation, we find that VQ-Transplant allows obtaining near state-of-the-art reconstruction fidelity for industry-level models like VAR while reducing the training cost by 95\%. VQ-Transplant democratizes quantization research by enabling resource-efficient integration of novel VQ techniques while matching industry-level reconstruction performance.

View source

Similar papers

Jul 2026

P4Q: Learning to Prompt for Quantization in Low-Bit CLIP

The “Prompt for Quantization” (P4Q) is proposed, by integrating PTQ with Parameter-Efficient Fine-Tuning (PEFT) techniques, and demonstrates that P4Q significantly enhances the performance of low-bit CLIP while reducing deployment costs.

H. Sun, Runqi Wang, Yanjing Li et al. · 0 citations
Jul 2026

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

Extensive experiments demonstrate that the on-device latency-informed design combined with the tailored training strategy establishes a new state-of-the-art for efficient LVLM encoding, significantly outperforming existing encoder-centric baselines while operating on-device at nearly 1.7xthe speed.

Ioannis Maniadis Metaxas, Adrian Bulat, Alberto Baldrati et al. · 0 citations
Preprint Sep 2026

Efficient Quantization-Aware Distillation with Cross-Modal Alignment for Edge Vision-Language Models

Large-scale vision-language models (VLM) such as CLIP enable strong open-vocabulary reasoning, yet deploying these capabilities on resource-constrained edge devices remains challenging. EdgeVL addresses this problem by distilling CLIP representations into lightweight multi-modal encoders and applying quantization-aware...

Jinwoo Jeon, GyuYeop Do, Yunkyu Lim et al. · 0 citations
Preprint Aug 2026

GVC-RT: Towards Real-Time Generative Video Compression at Ultra-Low Bitrates

This work systematically identifies the computational bottlenecks and proposes GVC-RT, which redesigns the generative latent coding framework to realize real-time video coding without sacrificing compression performance, and introduces a lightweight de-tokenizer architecture to resolve the final latency bottleneck duri...

Tian-Jian Dang, Si-Xian Wang, Lei Luo et al. · 1 citation · ⚡1
Aug 2026

EP-MAE: A resource-efficient masked autoencoding framework for 3D neural representation learning.

Efficient Point Masked Autoencoders (EP-MAE), a new framework designed to significantly reduce the training cost of 3D self-supervised pre-training while maintaining strong representation quality, and provides a scalable and effective foundation for future 3D neural network models is presented.

Jian Zhu, Jiale Zhao, Cheng Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.