Skip to content
Preprint

VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

Sep 2026 · 1 citation · ⚡ 1 influential · 64 references
Computer Science

TL;DR

VC-Attention is proposed, a training-free low-bit attention framework that addresses diffusion Transformers and quantization scale by pairing Value smoothing with a fused probability Cast, and improves fidelity over low-bit baselines.

Abstract

Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost, and a deployable low-bit kernel must be accurate and fast. Accuracy is limited by outliers: a block's quantization scale is set by its largest entries, leaving typical entries confined to a narrow range of representable values. Prior work smooths queries and keys, but value outliers follow no fixed channel or spatiotemporal structure and remain the dominant source of output error. Speed is limited by softmax: low-bit Tensor Cores accelerate only the two matrix multiplications, so the high-precision exponential between them becomes the longest pipeline stage on datacenter GPUs. We propose VC-Attention, a training-free low-bit attention framework that addresses both by pairing Value smoothing with a fused probability Cast. V-Smooth reorders value tokens by lightweight online clustering, so the tokens in a hardware block quantize well together. It quantizes only the residual after subtracting the block mean, and restores that mean from the row sum the online softmax already maintains. ExpCast-FP8 maps log-domain scores directly to E4M3 probability codes with one fused multiply-add, eliminating the FP32 exponential and the format conversion. We implement VC-Attention for B200, B300, H200, RTX PRO 6000, and RTX 5090. Across Wan2.2, LongCat-Video, HunyuanVideo-1.5, and MiniMax-H3, VC-Attention improves fidelity over low-bit baselines, speeds up the attention kernel over BF16 FlashAttention-4 by 1.46-1.59x on datacenter Blackwell and Hopper and by 2.3-3.6x on workstation cards, and generates a clip 1.13-1.19x and 1.36-1.70x faster end to end.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers

Most post-training quantization pipelines fit each weight matrix to its pretrained counterpart, one matrix at a time. Whether that proxy tracks what an attention block actually computes, or how errors in the Q, K and V projections compound inside the softmax, is rarely checked. We write the objective on the attention o...

Nafiseh HosseinpourFardi, Negar Alihadi, Mahmoudreza Babaei et al. · 0 citations
Preprint Aug 2026

APT: Accelerating Diffusion Transformers via Attention Probability-Guided Pruning and Quantization

APT, a software-hardware co-designed accelerator for high-resolution DiTs that leverages attention probabilities as a unified importance metric to jointly optimize computation through fine-grained pruning and adaptive precision scaling, and is evaluated on SOTA DiT models including PixArt, Stable Diffusion 3, and FLUX.

Sungyeob Yoo, Seeyeon Kim, Joonyong Park et al. · 0 citations
Preprint Aug 2026

LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration

LoSA is a training-free sparse-attention method that fixes a retained-mass threshold of 99% rather than a sparsity ratio: it measures exact block attention masses at one early dense step, keeps, for each head and query block, the smallest key/value block set meeting the threshold, and reuses the frozen block indices fo...

En-Huai Liu, Yun-Ke Wang, Yu-Tong Wang et al. · 0 citations
Preprint Aug 2026

SoftWater: Class-Aware Rate Allocation for Softmax Quantization

Across five models from 1B to 32B, SoftWater outperforms the released WaterSIC quantizer at matched head rates on 59 of 60 test points, using none of that pipeline's refinements and cutting head-induced KL by $6.5\times$--$8.3\times at 2 bits.

Joao V. Cavalcanti, Ashia C. Wilson · 2 citations
#artificial intelligence Preprint Sep 2026

DraftAttention2: Fast Video Diffusion with Low-Resolution-Guided Mixed-Precision Attention

Video generation has broad applications in content creation and entertainment. Diffusion transformers have advanced the quality of generated videos, but attention over spatiotemporal tokens becomes increasingly expensive as video resolution and duration increase. We present DraftAttention2, a training-free framework th...

Rui Ding, Haopeng Li, Wei-Ze Ma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.