Skip to content
Preprint

Channel-Wise and Token-Aware Post-Training Quantization for Visual State Space Duality

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

The Channel-wise Token-balanced Output-Aware Clipping (CTOAC) method, which learns per-input-channel clipping bounds by minimizing a token-balanced reconstruction loss on the corresponding linear outputs, is proposed.

Abstract

State space models (SSMs), particularly Mamba, have emerged as efficient alternatives to attention-based architectures and have been extended to vision through ViM, VMamba, and Visual State Space Duality (VSSD). Yet the low-bit post-training quantization (PTQ) behavior of VSSD remains insufficiently understood. A weight-activation split on VSSD-Tiny identifies activation quantization as the dominant low-bit bottleneck, while representative inputs to selected VSSD-backbone linear layers exhibit strong channel-wise magnitude variation and token-localized extremes. We propose the Channel-wise Token-balanced Output-Aware Clipping (CTOAC) method, which learns per-input-channel clipping bounds by minimizing a token-balanced reconstruction loss on the corresponding linear outputs. Only the selected linear layers and their input activations are quantized; other backbone operations retain their original precision. Across VSSD-Tiny, VSSD-Small, and VSSD-Base, the proposed CTOAC method retains ImageNet-1K accuracy and remains substantially more robust than the evaluated baselines at more aggressive precision settings. Applying the same quantization scope to VSSD backbones on COCO and ADE20K preserves strong object detection, instance segmentation, and semantic segmentation performance. An optimized RTX 4090 deployment configuration achieves up to 1.42x end-to-end speedup over FP32.

View source

Similar papers

Conference Open access Sep 2026

SeGO: Sensitivity-Aware Golden Optimization for Large-Scale VLM Quantization

A cross-modal structural sensitivity asymmetry in VLMs is revealed and SeGO is proposed, a unified structural sensitivity-aware sparse optimization framework that achieves the balance among model parameter amount, quantization accuracy and scaling factors’ search efficiency on InternVL2 and LLaVA series.

Tian-Qi Zhao, Xin-Rui Cheng, Yang Su et al. · 0 citations
#machine learning Preprint Aug 2026

Component Type, Not Reconstruction Error, Predicts Attention Quantization Sensitivity

Many post-training quantization (PTQ) methods use layer-wise reconstruction, second-order proxy objectives, or activation-aware transformations to reduce quantization-induced error. Whether that error signal predicts the downstream functional impact of quantizing an individual attention projection has not been directly...

Kasun Dewage, Marianna Pensky, Suranadi De Silva · 0 citations
Preprint Sep 2026

VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

VC-Attention is proposed, a training-free low-bit attention framework that addresses diffusion Transformers and quantization scale by pairing Value smoothing with a fused probability Cast, and improves fidelity over low-bit baselines.

Xing-Yang Li, Dong-Yun Zou, Shi-Ning Zhang et al. · 1 citation · ⚡1
#artificial intelligence Preprint Aug 2026

REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

Real-Q (Real-time E2E-loss Aligned LLM Quantization), a novel PTQ paradigm that breaks this compromise: instead of diluting the objective for the sake of analytic tractability, REAL-Q targets an end-to-end-aligned surrogate of the global loss and refines it via fine-grained, dynamic Block-wise Gradient Descent applied...

Qian Zhang, Yao-Ming Li, Zheng Tan et al. · 0 citations
Preprint Sep 2026

SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization

Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-textual token ratios...

Zhen-Hao Shang, Hai-Zhao Jing, Hao-Kui Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs

Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based procedures such as SpinQuant and computationally friendlier gradient-free approaches such as D...

Mehdi Makni, Ryan Lucas, Rahul Mazumder · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.