Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning

ReCo (Reward-Coordinated Compression), a step-wise framework in which a lightweight process-reward estimator scores each completed step and drives three components: reward-adaptive KV-cache compression that shrinks the retained cache harder at high-reward steps and less at low-reward ones, and a confidence-based early stopping that triggers when the reasoning is reliable.

Qiyuan Zhu, Dezhi Li, Pengyu Cheng et al. · 0 citations