In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been proposed to learn data-collecting policies tailored to reduce online evaluation variance. However, these approaches do not ac...
Claire Chen, S. Liu, Licheng Luo et al.· 1 citation
Verbal feedback can identify errors and prescribe corrections, providing rich supervision for language-model post-training even when reliable programmatic verifiers are unavailable. Such feedback, often generated by a capable model, can be used to condition the teacher in on-policy distillation, which trains the studen...
Amir Moeini, Huai-Jiang Zhu, Daniel Havir et al.· 0 citations
In-context reinforcement learning (ICRL) enables large language model agents to adapt to new environments using their interaction history without updating model parameters. However, repeatedly conditioning on growing histories can lead to substantial token cost. We propose a bounded-history context-management framework...
Rui-Han A. Li, Shang-Tong Zhang, Rohan Chandra· 0 citations
InfRL (Inference-time Reinforcement Learning) offers a practical and domain-agnostic approach to harness reinforcement learning during inference, bridging the gap between static prompting and computationally intensive parameter-level fine-tuning.
Sikun Guo, Amir Hassan Shariatmadari, Jiuqi Wang et al.· Proceedings of the 32nd ACM...· 0 citations
Large language models (LLMs) possess extensive latent knowledge yet remain largely static at inference. Once prompted, their generation policy typically cannot evolve, and post-hoc ''self-reflection'' methods provide no explicit principled learning signals. To address this limitation, we formally model iterative resear...
Sikun Guo, Amir Hassan Shariatmadari, Jiuqi Wang et al.· Proceedings of the 32nd ACM...· 0 citations
This work introduces a novel automata-based approach that provides an efficient memory mechanism and associated Markovian rewards suitable for RL frameworks and empirically demonstrates that this approach learns policies that achieve higher robustness scores and satisfaction rates than those learned by existing approac...
A. Bozkurt, Shang-Tong Zhang, Yuichi Motai· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.