WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps
Reward fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically formulate this problem as sampling from a reward-tilted distribution, the solution to a KL-regularized reward-maximization problem. Here, we introduce an opti...