Learning from Unreachable Rewards: Hint-Conditioned Reinforcement Learning for Generative Recommendation
This work proposes Hint-Conditioned Generative Recommendation (HCGRec), a semantic-ID generative recommendation framework that recovers learning signal for such hard training instances and introduces hint-aware credit decomposition, using supervised learning to preserve item-semantic and prefix-structure alignment for hinted tokens and GRPO to optimize the sampled suffix.