Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such a...
Hexuan Deng, Zi-Hao Yan, Xue-Bo Liu et al.· 0 citations
Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repe...
Hexuan Deng, Yue Wang, Wen-Yu Jiang et al.· 0 citations
This work forms this challenge as Narrative Commitment Preservation (NCP), and introduces NCP-Bench, a benchmark of 100 narrative environments derived from movie synopses that each environment includes a structured narrative specification that can automatically check throughout the interaction between the player agent...
Yingpeng Ma, Jianhao Yan, Bei-Ning Shi et al.· 1 citation
FPCO-Dialog is introduced, a benchmark for evaluating correction and cooperation behavior in VLMs under repeated false premises, and reveals substantial and persistent cross-model differences in aggregate correction tendency, model-specific turn-wise dynamics, and systematic variation across false-premise types under t...
Jiayuan Ma, Yu-Qi Lu, Wei-Yang Guo et al.· 0 citations
This work introduces MemoNoveltyAgent, a multi-agent system designed to generate comprehensive and faithful novelty reports, and proposes a RAG-augmented checklist evaluation method that enables reliable and evidence-grounded assessments.
Jiajun Hou, Hexuan Deng, Wenxiang Jiao et al.· 0 citations
This paper proposes CRISP, a framework for training efficient deep search agents through critical step perception that distinguishes interactions that gather necessary evidence from redundant ones and shapes the training reward to preserve the former while pruning the latter, improving efficiency without sacrificing th...
Haosi Mo, Zihao Yan, Ruiqing Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.