The Channel-wise Token-balanced Output-Aware Clipping (CTOAC) method, which learns per-input-channel clipping bounds by minimizing a token-balanced reconstruction loss on the corresponding linear outputs, is proposed.
Abstract
State space models (SSMs), particularly Mamba, have emerged as efficient alternatives to attention-based architectures and have been extended to vision through ViM, VMamba, and Visual State Space Duality (VSSD). Yet the low-bit post-training quantization (PTQ) behavior of VSSD remains insufficiently understood. A weight-activation split on VSSD-Tiny identifies activation quantization as the dominant low-bit bottleneck, while representative inputs to selected VSSD-backbone linear layers exhibit strong channel-wise magnitude variation and token-localized extremes. We propose the Channel-wise Token-balanced Output-Aware Clipping (CTOAC) method, which learns per-input-channel clipping bounds by minimizing a token-balanced reconstruction loss on the corresponding linear outputs. Only the selected linear layers and their input activations are quantized; other backbone operations retain their original precision. Across VSSD-Tiny, VSSD-Small, and VSSD-Base, the proposed CTOAC method retains ImageNet-1K accuracy and remains substantially more robust than the evaluated baselines at more aggressive precision settings. Applying the same quantization scope to VSSD backbones on COCO and ADE20K preserves strong object detection, instance segmentation, and semantic segmentation performance. An optimized RTX 4090 deployment configuration achieves up to 1.42x end-to-end speedup over FP32.
A cross-modal structural sensitivity asymmetry in VLMs is revealed and SeGO is proposed, a unified structural sensitivity-aware sparse optimization framework that achieves the balance among model parameter amount, quantization accuracy and scaling factors’ search efficiency on InternVL2 and LLaVA series.
Tian-Qi Zhao, Xin-Rui Cheng, Yang Su et al.· Proceedings of the Thirty-Fi...· 0 citations
Many post-training quantization (PTQ) methods use layer-wise reconstruction, second-order proxy objectives, or activation-aware transformations to reduce quantization-induced error. Whether that error signal predicts the downstream functional impact of quantizing an individual attention projection has not been directly...
Kasun Dewage, Marianna Pensky, Suranadi De Silva· 0 citations
VC-Attention is proposed, a training-free low-bit attention framework that addresses diffusion Transformers and quantization scale by pairing Value smoothing with a fused probability Cast, and improves fidelity over low-bit baselines.
Real-Q (Real-time E2E-loss Aligned LLM Quantization), a novel PTQ paradigm that breaks this compromise: instead of diluting the objective for the sake of analytic tractability, REAL-Q targets an end-to-end-aligned surrogate of the global loss and refines it via fine-grained, dynamic Block-wise Gradient Descent applied...
Qian Zhang, Yao-Ming Li, Zheng Tan et al.· 0 citations
Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-textual token ratios...
Zhen-Hao Shang, Hai-Zhao Jing, Hao-Kui Zhang et al.· 0 citations
Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based procedures such as SpinQuant and computationally friendlier gradient-free approaches such as D...
Mehdi Makni, Ryan Lucas, Rahul Mazumder· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.