Skip to content

Author

Max Ryabinin

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

Score Centering Stabilizes Off-policy Reinforcement Learning

It is shown that the instability of RL under TIM is primarily caused by drift: a persistent bias between training and inference engines that accumulates with every training step, and an additive score centering term is derived that stabilizes RL under TIM by canceling drift.

Martina Marek, Max Ryabinin · 2 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.