Skip to content

OnlineCache: Learning Dynamic Caching Policies with Error Correction for Efficient Diffusion Inference

Jul 2026 · arXiv.org · Vol abs/2607.29398 · 0 citations · 63 references
Computer Science

TL;DR

This paper proposes OnlineCache, a dynamic caching framework that jointly learns when to cache and how to correct approximation errors, and leverages policy gradient to train a lightweight network for adaptive speed-quality trade-offs, and incorporate a learnable corrector to mitigate caching-induced errors.

Abstract

Diffusion models have revolutionized generative tasks but incur high latency due to iterative denoising. While cache-based strategies accelerate inference by reusing intermediate features, they largely rely on static, sample-agnostic schedules. We argue that this rigidity overlooks two facts empirically validated in this paper: (i) generation difficulty varies across prompts, requiring adaptive resource allocation--complex inputs demand more computation while simpler ones require less; (ii) error sensitivity fluctuates across timesteps, where static policies may cache high-error steps or waste computation on low-error ones. We therefore propose OnlineCache, a dynamic caching framework that jointly learns when to cache and how to correct approximation errors. We leverage policy gradient to train a lightweight network for adaptive speed-quality trade-offs, and incorporate a learnable corrector to mitigate caching-induced errors. Both modules are jointly optimized under a bilevel optimization framework, with the policy targeting global generation quality and the corrector minimizing local errors. Our method automatically allocates computational resources across both samples and timesteps, improving overall generation quality. Extensive experiments demonstrate clear superiority. On FLUX.1-dev model, OnlineCache achieves nearly 3 speedup while preserving generation fidelity. On DiT and CogVideoX, it similarly delivers competitive acceleration without compromising quality; across all scenarios, it consistently outperforms existing cache-based acceleration baselines.

View source

Similar papers

Preprint Aug 2026

BAG: Budget-Aware Gating for Diffusion Caching

Diffusion caching is a lightweight strategy that accelerates Diffusion Transformers (DiTs) by reusing intermediate features across denoising steps, but existing paradigms face a fundamental trade-off: online heuristics lack global budget awareness, whereas static schedules lack instance adaptivity and fail to flexibly...

Tong Zhao, Ming Lei, Yucheng Han et al. · 0 citations
Jul 2026

Evolving Cache Schedules for Fast Diffusion Policy Inference

Diffusion policies achieve strong visuomotor control by iteratively denoising action chunks, but repeated denoising makes real-time deployment computationally demanding. Cache-based methods reduce inference cost by reusing intermediate activations, but existing training-free schedules typically allocate computation uni...

Si Wang, Kangye Ji, Dixiang Wang et al. · 0 citations
Preprint Aug 2026

From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which we identify as being significantly misaligned with final generation qu...

Xichen Ye, Yifan Wu, Zhikang Xie et al. · 0 citations
#artificial intelligence Preprint Aug 2026

EpaCache: Error-Propagation-Aware Caching for Accelerating Diffusion-Based Visual Generation

This work introduces Error-Propagation-Aware Cache (EpaCache), a training-free caching policy that adaptively allocates the reuse budget on timesteps with lower downstream impact and consistently improves the latency--fidelity trade-off over existing caching methods.

Yu-Han Liu, Zong-Wei Hong, Jinglun Li et al. · 0 citations
Jul 2026

LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models

LaCache is proposed, a training-free acceleration framework that alleviates operator-level redundancy through lossless caching and mixed precision, and inegrates a per-group FP8 quantization strategy for FFN layers, tailored to step-dependent activation distributions across the diffusion process.

Xingru Chen, Zelang Liang, Yongjia Ma et al. · 0 citations
Preprint Aug 2026

Archer: Adaptive Reuse of Cached Hidden States for Efficient Rollback in Diffusion Language Models

Adaptive Reuse of Cached Hidden States for Efficient Rollback (Archer) is introduced, a training-free KV caching method for rollback-capable DLMs that characterizes prompt reuse as a reversibility-aligned cache boundary, bounds its state-dependent approximation error, and gives a decoder-margin condition for preserving...

Xuning He, Zinan Sheng, Yong-Ding Tao et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.