Preprint
Jul 2026
Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
This work proposes a game-theoretic framework that gives this reward-retention trade-off an explicit statistical interpretation, and provides a principled method for learning this equilibrium coefficient via reduction to the KL-regularized RL objective, thus allowing for flexible integration into standard fine-tuning pipelines.
Keegan Harris, Brian Lee, Ian Waudby-Smith et al.
· 0 citations