Fine-Tuning on Self-Generated and Reward-Weighted Data: Learning Dynamics, Convergence Rates, and Benefits of Off-Policyness
A unified theory for RE(S) is developed that covers the full spectrum of S, and can be interpreted as a stage-wise optimization process, where each stage takes $S$ gradient steps for minimizing the Kullback-Leibler distance to a fixed reward-weighted rollout distribution.
Zhi-Wei Wang, Yan-Xi Chen, Ya-Liang Li et al.
· 0 citations