Preprint
Jul 2026
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents
Reinforcement learning holds significant potential for training large language models to handle multi-turn interactive tasks, but directly training with outcome rewards often results in slow convergence due to the sparsity of signals and the lack of fine-grained feedback.
Qiang Liu, Taian Guo, Ruizhi Qiao et al.
· 0 citations