Skip to content
Conference

Sensitivity-Guided Mixed-Precision Post-Training Quantization for MambaVision

Aug 2026 · International Conference on Advanced Mechatronic Systems · pp. 109-114 · 0 citations · 26 references

Abstract

Hybrid vision backbones such as MambaVision combine convolutional layers, Mamba blocks (based on selective state space models), and self-attention within a single architecture. However, the heterogeneous operator composition of such models poses new challenges for post-training quantization (PTQ): uniform bit-width assignment causes catastrophic accuracy collapse due to widely varying per-block quantization sensitivity. In this work, we present a systematic per-block sensitivity analysis of MambaVision, revealing that mixer blocks (Mamba and attention) are remarkably robust to quantization down to 4-bit weights and 4-bit activations (average accuracy drop of only $\text{0. 0 6 \%})$, while a small number of bottleneck blocks (stem and downsampling layers) are extremely fragile under the same conditions. Based on these findings, we propose a sensitivity-guided mixed-precision assignment strategy that keeps critical blocks at 16-bit floating point, quantizes robust mixer blocks to 4-bit weights and 4-bit activations, and applies 8-bit weights and 8-bit activations to moderately sensitive blocks. Our best configuration achieves 80.66% Top-1 accuracy on ImageNet-1K with GPTQ (-3.29% from 16-bit floating point), recovering +4.68% over uniform 8-bit quantization at a comparable average bit-width of 8.14 bits, while reducing model size by 53.3% and bit operations by 70.8%. We further provide a Hessian mismatch analysis explaining the counterintuitive finding that GPTQ underperforms round-to-nearest under uniform quantization but recovers its advantage under mixed-precision, where 16-bit floating point bottleneck blocks act as error propagation firewalls.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

HyQuant: Hybrid-Precision Quantization for LLM Attention

Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention.

Jiarui Ding, Bin Xing, Yu Zhang et al. · 0 citations
Open access 2026

Threshold-Tuning Coordinate Attention for Binarized Neural Networks

—Binary Neural Networks (BNNs) are highly attractive for mobile and embedded vision due to their extremely low memory footprint and efficient bit-level convolutions. However, binarization often causes severe information loss and weakens the effectiveness of conventional attention modules designed for full-precision net...

Shao-Qing Wu, Hiroyuki Yamauchi · 0 citations
Conference Open access Sep 2026

SeGO: Sensitivity-Aware Golden Optimization for Large-Scale VLM Quantization

A cross-modal structural sensitivity asymmetry in VLMs is revealed and SeGO is proposed, a unified structural sensitivity-aware sparse optimization framework that achieves the balance among model parameter amount, quantization accuracy and scaling factors’ search efficiency on InternVL2 and LLaVA series.

Tian-Qi Zhao, Xin-Rui Cheng, Yang Su et al. · 0 citations
#machine learning Preprint Sep 2026

RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models

K-Means clustering is addressed with K-Means clustering, achieving near-lossless accuracy and a mean speed-up over the full-precision model, and it is revealed that excluding from quantization the layers whose speed-up is negligible, regardless of their sensitivity, can be counterproductive.

D. Población-Criado, D. García-Gasulla, Eduardo Quiñones · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.