Jun 2026
Experience Augmented Policy Optimization for LLM Reasoning
This work proposes Experience-Augmented Policy Optimization (EAPO), which leverages a prior RL-optimized policy as an action-level experience prior and selectively injects experience at critical decision points during rollout to ensure stable and unbiased learning from experience-augmented rollouts.
Jinda Lu, Kexin Huang, Junkang Wu et al.
· arXiv.org · 2 citations