Score Centering Stabilizes Off-policy Reinforcement Learning
It is shown that the instability of RL under TIM is primarily caused by drift: a persistent bias between training and inference engines that accumulates with every training step, and an additive score centering term is derived that stabilizes RL under TIM by canceling drift.
Martina Marek, Max Ryabinin
· 2 citations
· ⚡1