Learning from Online User Feedback for Shopping Agents
LOFA combines reinforcement learning over verifiable purchase outcomes with feedback-aware on-policy distillation, which identifies users' in-dialogue directives and converts them into dense token-level supervision, which captures both collaborative behavioral patterns and user-specific preferences.