SPARK (Scientific Paper Abstracted Reasoning sKeleton), a paper-oriented synthesis framework built on Sci-Base, is introduced, a scientific reasoning dataset with substantially higher difficulty and diversity than existing resources and consistently outperforms existing scientific reasoning datasets.
This work presents a comprehensive framework for the development and evaluation of specialized scholarly agents, InternReviewer and InternAdvocate, by implementing an agentic Reinforcement Learning (RL) paradigm driven by a unified objective metric and reward system.
Xuerui Su, Liya Guo, Qizhi Pei et al.· 0 citations
An RL method tailored for context management is proposed, which uses context and entropy variation to identify critical editing decisions for branch sampling and estimates action-level advantages from all branched trajectories that pass through the corresponding context editing action.
Zhuo-Shi Pan, Qizhi Pei, Jun-Ru Lu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.