A set of unsupervised signals and a Retrieval-Aware Agent Controller (RAAC) are introduced, which assists the agent in selecting optimal actions at each stage of the research process, resulting in more effective reasoning trajectories that improve overall performance while reducing unnecessary iterations, and consequently cost and latency.
Abstract
In this paper, we analyze the reasoning trajectories of a variety of DRAs and show that existing agents often suffer from reasoning stagnation: the majority of iterations contribute little or no improvement to final performance, while agents lack awareness of their trajectories and are therefore ineffective at adapting their search strategies or determining when to terminate. To address this issue, we introduce a set of unsupervised signals and a Retrieval-Aware Agent Controller (RAAC), which assists the agent in selecting optimal actions at each stage of the research process. RAAC incorporates key information retrieval principles, namely search novelty and information coverage, resulting in more effective reasoning trajectories that improve overall performance while reducing unnecessary iterations, and consequently cost and latency. Specifically on BrowseComp-Plus and across a large set of DRAs, adding RAAC reduces the number of search calls by an average of 14, significantly improves the best-performing DRA on recall and accuracy, and achieves an accuracy gain of up to 10% (3% on average).
ITER, an agent interaction-aware dense retriever trained using agent trajectory learning signals, is introduced andlations show that structured interaction history and pre-search reasoning provide complementary retrieval context, while previously visited and useful documents provide the strongest trajectory-relative su...
Hao-Dong Chen, Shuai Wang, Yu Yin et al.· 0 citations
This paper proposes TAPE (Trajectory Alignment and exPerience pool Evolution), a novel self-evolving fine-tuning framework designed to enhance the generalization and robustness of web search agents without requiring large-scale human annotation.
Minjae Rhee, Jitong Zou, Tianjun Mo et al.· IEEE Access· 3 citations
A framework that gives LLM agents two mechanisms for searching large solution spaces autonomously: a leaderboard scored on held-out data acts as a reward signal that drives each agent to refine its solutions over repeated submissions, and a loop that operates even with a single agent.
Dulmini Hettiarachchi, Andre Rusli, J. Young et al.· 0 citations
AgentIdeaBench is introduced, a multidisciplinary benchmark that evaluates scientific ideation under two matched settings, static observation and active exploration, and Scientific World Modeling is explored, a generation-time loop that refines a draft hypothesis through structured thought experiments.
Yunxiang Mo, Tianshi ZHENG, Yi-Sen Gao et al.· 0 citations
Even though AI is getting better quickly, Large Language Models (LLMs) are still very good at multi-step reasoning and structured problem solving. However, their performance during inference often relies heavily on initial prompts and set strategies. When outputs are not ideal, fixing them usually requires outside help...
Shreyas Kushwaha, Bhavya Shah, Y. Chauhan et al.· 2026 7th International Confe...· 0 citations
NIS-Agent is proposed, which applies context isolation at the two decision points most vulnerable to inertia bias: webpage triage and final-answer validation, and trains an 8B model to be intrinsically more resistant to inertia bias.
Xiang-Xin Zhang, Zhanwei Zhang, Zhihang Fu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.