Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation ac...
Jia-Hua Yang, Zhiwei Yang, Xian-Peng Zhang et al.· 0 citations
TimeThink is proposed, a reinforcement learning framework that explicitly guides temporal evidence discovery in Video-LLMs and introduces a step-wise temporal process reward that provides localized credit assignment for these clues and a joint process--outcome optimization objective that balances reasoning fidelity wit...
Handong Li, Longteng Guo, Zikang Liu et al.· arXiv.org· 0 citations
This work starts from an empirical observation: when query-relevant visual evidence is explicitly strengthened using the model's own attention, generation becomes more accurate, suggesting that many failures do not arise solely from missing perception, but from an insufficient tendency to trust the evidence the model h...
Xin Zou, Hao Deng, Yibo Yan et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.