Jul 2026· IEEE Transactions on Image Processing· Vol PP, pp. 1-1· 0 citations
Medicine
TL;DR
This work proposes a novel scalable product quantization framework within an end-to-end network, termed Tensor-based Codebook Product Quantization (TCPQ), an innovative attempt to integrate tensor theory with product quantization methods.
Abstract
Scalable product quantization has recently attracted considerable attention in large-scale image retrieval, as it avoids the need to train multiple models for generating quantization codes of varying lengths. However, most existing approaches primarily concentrate on approximating the ground-truth similarity between image pairs, while overlooking the correlations among sub-codebooks and among codewords. In addition, limited work has addressed the challenge that increasing the number of subspaces or codewords substantially raises memory consumption. To address these limitations, we propose a novel scalable product quantization framework within an end-to-end network, termed Tensor-based Codebook Product Quantization (TCPQ). This work represents an innovative attempt to integrate tensor theory with product quantization methods. The framework leverages tensor-based methods to capture spatial correlations among sub-codebooks and among codewords, and adopts lightweight codebooks for efficiency. For optimization, a subspace-wise unbiased supervised contrastive loss is proposed to bring embeddings of the same class closer together, push embeddings of different classes farther apart within quantization subspaces, and precisely regulate the minimum distance between positive and negative samples. In addition, the ArcFace loss is incorporated to enhance the discriminative power of the learned features, and an orthogonality constraint is imposed on the factor matrix along the codebook dimension to avoid overfitting. Extensive experiments on three large-scale real-world benchmarks demonstrate that TCPQ achieves state-of-the-art retrieval performance.
Hadamard-Transform-domain Quantization (HaTQ), which uses orthogonal Hadamard reparameterization before quantization to redistribute weight and activation responses in the original domain across channels, and is compatible with integer-only execution.
Jun-Qi Shi, Chongzhi Wang, Yiwen He et al.· 0 citations
This work proposes Hyper-Spherical Quantization (HSQ), which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning.
Tianren Ma, Lin Long, Chu-Yan Chen et al.· arXiv.org· 0 citations
We revisit Kashin-decomposition-based weight quantization for large language models and propose an improved algorithm with stronger convergence properties and structured, efficient orthogonal transforms. The method retains the core factorization of each weight into two components -- one with bounded infinity norm and t...
Daria Cherniuk, A. Rudikov, B. Kashin et al.· 0 citations
A training-free framework, BAL-ANCER, which achieves global budget allocation for mixed-precision quantization and low-rank correction through information-guided subspace matrices, allowing a principled greedy allocator to distribute compression bits and ranks across the entire model.
Si-Nuo Fan, Ying-Jie Lao· Proceedings of the Thirty-Fi...· 0 citations
Large-scale image retrieval requires compact representations without substantially sacrificing retrieval accuracy. However, Vision Transformer Hashing (VTS) concatenates all output tokens before hash projection, resulting in a high-dimensional hashing head with considerable model and memory overhead. We replace this to...
Discrete visual tokenization, predominantly driven by vector, scalar, and product quantization, lacks a unified conceptual framework for understanding quantization tradeoffs. In this paper, we propose a unified rate--distortion perspective on modern discrete visual tokenization. By viewing quantization as lossy compres...
Xianghong Fang, Wenlong Mou, Yuan Yuan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.