Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients int...
Shi-Cheng Fang, Yi-Wen Zhao, Wen-Bo Tian et al.· 0 citations
A paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context shows that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance...
Zhi-Peng Xu, Jia-Hao Lu, Yi-Ning Zheng et al.· 3 citations
The Embodied Task Agent is introduced, a new paradigm for extending digital agents into the physical world, and OpenETA is released as its open-source implementation, which provides replaceable Planners, composable Tools and Skills, auditable memory, replayable trajectories, and common interfaces for simulation and rea...
Yitong Chen, Zezheng Huai, Si-Xian Li et al.· 3 citations· ⚡1
EvoCUA-1.5 extends self-evolving computer-use agents from offline experience learning to online reinforcement learning, where policies interact with executable sandbox environments and improve from verifiable task outcomes and provides a practical framework for scaling online RL in multi-turn computer-use agents.
Mianqiu Huang, Taofeng Xue, Chong Peng et al.· 1 citation
Applications in materials analysis, molecule design, and protein or antibody screening, together with experiments on scientific reading, idea generation, molecule generation, and antibody screening, show that SCION outperforms existing autonomous research-agent baselines, especially in decomposition, verification, refi...
Y. Zheng, Yuxin Wang, Jiahao Lu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.