Preprint
Jul 2026
LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models
LaCache is proposed, a training-free acceleration framework that alleviates operator-level redundancy through lossless caching and mixed precision, and inegrates a per-group FP8 quantization strategy for FFN layers, tailored to step-dependent activation distributions across the diffusion process.
Xingru Chen, Zelang Liang, Yongjia Ma et al.
· 0 citations