Preprint
Jul 2026
Inverse RL Helps Align AI by Imitating Humans
It is shown that the recovered reward improves a base policy without a supervised loss and yields further gains when optimized after standard supervised fine-tuning and can be used for contextual alignment, in which a single policy can be tailored to the preferences of different audiences.
Michal Wilinski, Liu Leqi, Chirag Nagpal
· 0 citations