This work presents VERA-RL, a reinforcement-learning formulation for scientific error detection over academic papers, and constructs VERA-13K, a 12,900-sample dataset organized into 4,300 matched chains, covering 6 scientific-error categories across the research workflow and broad natural-science domains.
Abstract
Multimodal large language models (MLLMs) are increasingly capable scientific assistants, yet they remain far from fully autonomous research. This transition requires models to actively inspect academic papers, build global evidence views, and make traceable judgments without prespecified issues or evidence. However, existing work provides limited task paradigms or training studies for such issue- and evidence-absent verification. We study this challenge through scientific error detection, where models must determine whether errors exist and justify them with evidence-based reasoning. To fill this gap, we present VERA-RL, a reinforcement-learning formulation for scientific error detection over academic papers. Following a Reason--Verify--Scan progression, we construct VERA-13K, a 12,900-sample dataset organized into 4,300 matched chains, covering 6 scientific-error categories across the research workflow and broad natural-science domains. We further introduce fine-grained rewards for reasoning completeness, evidence alignment, and error precision. Training Qwen3-VL-8B with VERA-RL substantially improves verifiable reasoning, approaching flagship MLLMs such as Gemini 3 Pro and Qwen3-VL-235B-A22B on Scan.
COGTRL is proposed, a trajectory-level reinforcement learning framework that trains LLMs to emulate cognitively grounded reasoning by jointly optimizing cognitive traces and the scientific steps produced in an interleaved manner.
Shrinidhi Kumbhar, Santosh Mashetty, Divij Handa et al.· 1 citation
Results provide evidence that access to a scalable, objective training signal can be a more important constraint than model scale alone, and can make otherwise ambiguous diagnostic reasoning tasks amenable to scalable reinforcement learning.
SOLID is proposed, a novel framework for self-improving OR language models without verified answers or external evaluators that improves solution accuracy for both general-purpose and OR-tuned models over outcome-only group-relative training.
Rui-Chen Zhu, Ming-Long Cao, Chen-Yu Zhou et al.· 0 citations
P-Bench is built, a benchmark comprising 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine and introduces Fisher-R1, an open-weight LLM agent trained for rigorous hypothesis testing using synthetic tasks and reinforcement learning.
Jia-Cheng Miao, Jin Mu, Guan-Hua Chen et al.· 0 citations
Self-Verification via Reinforcement Learning (SVRL), an RL-only finetuning framework that trains multimodal agents to verify and filter retrieved evidence within their own reasoning traces, reducing reliance on external verifiers at inference time is presented.
Vishwas Sathish, Viresh Ranjan, Xin-Liang Zhu et al.· 1 citation
Autonomous coding agents can remember an experiment yet carry forward a conclusion it does not justify. We reconstruct how evidence is reused in a 400-task NeuroGolf campaign, with selected wellbore-prediction records from the same operator as cross-domain comparisons. A numerical counterexample exposes an overbroad ex...
Bo-Da Cheng· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026