Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such a...
Hexuan Deng, Zi-Hao Yan, Xue-Bo Liu et al.· 0 citations
This paper proposes CRISP, a framework for training efficient deep search agents through critical step perception that distinguishes interactions that gather necessary evidence from redundant ones and shapes the training reward to preserve the former while pruning the latter, improving efficiency without sacrificing th...
Haosi Mo, Zihao Yan, Ruiqing Zhang et al.· 0 citations
MAP-Graph is introduced, a provenance-aware memory layer that represents agents, sources, memories, claims, and actions in a typed execution graph and supports provenance as an operational control signal, rather than only post-hoc audit metadata, within the evaluated setting.