InfRL (Inference-time Reinforcement Learning) offers a practical and domain-agnostic approach to harness reinforcement learning during inference, bridging the gap between static prompting and computationally intensive parameter-level fine-tuning.
Sikun Guo, Amir Hassan Shariatmadari, Jiuqi Wang et al.· Proceedings of the 32nd ACM...· 0 citations
Scientific ideation is driven by curiosity: researchers ask questions that expose knowledge gaps, reveal competing hypotheses, and clarify missing evidence relevant to decision making. Yet, most LLM-based ideation systems optimize the idea text while leaving curiosity under-modeled, resulting in brittle, engine-specifi...
Sikun Guo, Di Wang, Xiaohan Fan et al.· Proceedings of the 32nd ACM...· 0 citations
Large language models (LLMs) possess extensive latent knowledge yet remain largely static at inference. Once prompted, their generation policy typically cannot evolve, and post-hoc ''self-reflection'' methods provide no explicit principled learning signals. To address this limitation, we formally model iterative resear...
Sikun Guo, Amir Hassan Shariatmadari, Jiuqi Wang et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.