Preprint
Sep 2026
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention
VC-Attention is proposed, a training-free low-bit attention framework that addresses diffusion Transformers and quantization scale by pairing Value smoothing with a fused probability Cast, and improves fidelity over low-bit baselines.
Xing-Yang Li, Dong-Yun Zou, Shi-Ning Zhang et al.
· 1 citation
· ⚡1