Experiments show that HypoForge consistently outperforms existing AI scientist frameworks and skill-level variants, and demonstrates the effectiveness of the proposed stage-specific skill learning paradigms.
Abstract
Large language models (LLMs) have enabled AI scientist systems to automate scientific discovery, yet existing approaches most rely on static prompting or fixed workflows and fail to accumulate experience for continual improvement. We propose HypoForge, an experience-guided multi-agent framework that learns reusable scientific skills for automated hypothesis generation and hypothesis testing. HypoForge is built on the observation that these two stages involve different supervision signals. For hypothesis generation, where explicit feedback is unavailable, HypoForge adopts an adversarial generator--discriminator mechanism to improve reasoning through comparative critique. For hypothesis testing, where empirical feedback is available, HypoForge learns testing skills from execution outcomes and ground-truth results. By matching skill learning strategies with stage-specific supervision, HypoForge enables continual improvement without fine-tuning foundation models. Experiments on hypothesis generation and testing benchmarks show that HypoForge consistently outperforms existing AI scientist frameworks and skill-level variants. Further analysis demonstrates the effectiveness of the proposed stage-specific skill learning paradigms.
SOLID is proposed, a novel framework for self-improving OR language models without verified answers or external evaluators that improves solution accuracy for both general-purpose and OR-tuned models over outcome-only group-relative training.
Rui-Chen Zhu, Ming-Long Cao, Chen-Yu Zhou et al.· 0 citations
P-Bench is built, a benchmark comprising 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine and introduces Fisher-R1, an open-weight LLM agent trained for rigorous hypothesis testing using synthetic tasks and reinforcement learning.
Jia-Cheng Miao, Jin Mu, Guan-Hua Chen et al.· 0 citations
Agent self-evolution has primarily focused on learning how to act, while overlooking an equally important capability: learning to discover what an agent does not know. Existing approaches typically assume that failure discovery is given, focusing on how to repair failures once they are identified. We ask whether blind-...
Xiaoyi Bao, Yuanzhen Xie, Yunzhi Tan et al.· arXiv.org· 0 citations
CAST, a critique-aware training framework that converts sparse task outcomes into action-level supervision for critique learning and policy optimization, is presented, demonstrating that critique-aware training improves the robustness of LLM agents in realistic dynamic environments.
Amir Saeidi, Zeng Zhang, Rishi Singh et al.· 1 citation
SkillBoost is proposed, a three-stage framework that mitigates both risks: structured exploitation localizes observed failures to editable skill components, prior-guided exploration draws on prior knowledge in the LLM to generate diverse repair candidates, and verified acceptance commits a candidate only when it improv...
Hong-Qiang Lin, Chao Liu, Xiaofan Bai et al.· arXiv.org· 1 citation
This work proposes GSE, a globalized skill evolution framework that jointly optimizes skill compatibility and skill generalization, and maintains a Skill Relation Graph (SRG) that explicitly models and co-evolves inter-skill relationships.
Chen Yang, Jiashuo Tian, Zi-Qi Wang et al.· 3 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.