Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference learning reduces this cost by selecting informative comparisons, and LLM judges can provide additional scalable feedback. However, the preferences of the judges may deviate fr...
Zhong-Man Du, Hui-Ming Zhang, Hao-Dong Zhu et al.· 0 citations
Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not explicitly capture its r...
Hao-Dong Zhu, Yang-Yang Ren, Chang-Bai Li et al.· 0 citations
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but incurs substantial costs from rollouts and policy updates. Online prompt selection improves efficiency by using per-prompt Bayesian posteriors to predict difficulty and prioritize informative prompts....
Yang-Yang Ren, Hao-Dong Zhu, Sheng Xu et al.· 0 citations
Search agents extend large language models beyond static parametric memory by enabling them to acquire and use external evidence during multi-step reasoning. For knowledge-intensive tasks involving complex or evolving information, their reliability depends not only on retrieving relevant evidence but also on using it t...
Xingyu Guo, Wei Chen, Lin-Lin Yang et al.· 0 citations
A Kalman-Guided Prompt Selection method (KGPS), which reformulates prompt selection as a dynamic state estimation problem rather than static difficulty prediction, and consistently improves both final accuracy and rollout efficiency over strong baselines, establishing state-of-the-art performance among online prompt se...
Hao-Dong Zhu, Yang-Yang Ren, Yanjing Li et al.· arXiv.org· 2 citations· ⚡2
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.