Large language models have made strong reasoning gains through supervised fine-tuning, reinforcement learning, and on-policy distillation, yet these post-training methods are usually evaluated only by final-answer accuracy. We study how they reshape confidence during reasoning. We introduce a three-stage calibration fr...
Shuhao Li, Guodong Du, Anhao Zhao et al.· arXiv.org· 0 citations
This work recast CoT compression along three dimensions: importance criterion, restructuring level, and compression budget, and yields condition-aware guidelines for matching compression to deployment context.
Siyang Lyu, Zhijing Sun, Xing-Hao Chen et al.· arXiv.org· 0 citations
Results show that separating reusable schema encoding from selective resource access substantially reduces agentic inference costs with limited effectiveness loss.
Yichu Fang, Si-Tong Wei, Haozhe Hu et al.· 3 citations
A three-stage calibration framework that evaluates confidence before, during, and after chain-of-thought generation, corresponding to difficulty estimation, early termination, and answer aggregation finds that OPD provides the most useful pre-reasoning confidence, SFT gives the strongest online signal for early stoppin...
Shuhao Li, Guodong Du, Anhao Zhao et al.· 0 citations
WIDE is presented, the first end-to-end differentiable token-level dynamic width pruning framework designed for both prefill and decode scenarios, and a pruning--kernel co-design framework that decomposes dynamic sparsity acceleration into mask reordering, hardware-agnostic block-level skipping, and hardware-dependent...