This work proposes DeepInstructor, an agentic framework that formulates idea evaluation as reasoning over structured scholarly experience, and introduces DeepInstruct, a dataset with controlled pairwise comparisons across novelty, significance, and feasibility.
Abstract
As automated scientific discovery advances, Large Language Models (LLMs) can now generate research ideas at an unprecedented scale, shifting the bottleneck from idea generation to idea evaluation. Existing evaluators mainly rely on parametric LLM knowledge or unstructured retrieval, producing judgments that lack the experience-grounded reasoning used by human instructors. To address this, we propose DeepInstructor, an agentic framework that formulates idea evaluation as reasoning over structured scholarly experience. DeepInstructor constructs an Experience Graph from 58,607 peer reviews and employs a ReAct-based agent to retrieve dimension-specific evidence for traceable evaluation. We further introduce DeepInstruct, a dataset with controlled pairwise comparisons across novelty, significance, and feasibility. Experiments show that DeepInstructor substantially outperforms existing baselines, improving Hit@1 and Hit@2 alignment with human judgments by 24.4% and 29.7%, respectively. Our findings suggest that scientific idea evaluation can be grounded in explicit reasoning over structured scholarly experience
Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely assess it by asking models to generate ideas from a static, curated set of reference papers. That passive setup departs from the retrieval-and-...
Yunxiang Mo, Tianshi ZHENG, Yi-Sen Gao et al.· 0 citations
The bottleneck is scientific judgment rather than coding, and genuine discovery remains out of reach, so TruthInsightBench makes this gap a measurable target.
Zhi-Bo Yang, Chen Zhang, Yue-Wei Zhang et al.· 0 citations
It is revealed that current agents can explore novel regions of the solution space but lack the capacity to convert this novelty into improved task performance, and that LLMs exhibit greater H-Creativity than medal-winning humans, yet achieve lower performance.
Shitanshu Bhushan, Yun-Xiang Zhang, Lu Wang· 1 citation· ⚡1
This work presents AgentPanel, a multi-agent forum for human--AI collaboration in scientific exploration, a multi-agent forum for human--AI collaboration in scientific exploration that outperforms a centralized multi-agent debate baseline and shows that users value AgentPanel for perspective diversity and exploration s...
Zhi-Yao Cui, Qianyi Wang, Hao Yan et al.· 1 citation
Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a fra...
Lorenzo Molfetta, Wai-Chung Kwan, Giacomo Frisoni et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.