Large language models are reshaping ecommerce from static recommenders into interactive shopping assistants, yet real-world shopping requires session-level decision support: users reveal and revise constraints, coordinate multiple goals, and expect product-grounded recommendations over a full conversation. Existing ben...
Xin-Wei Yang, Ke-Long Mao, Yu-Dong Guo et al.· 0 citations
Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively s...
Li-Rui Luo, Ke-Long Mao, He-Ming Xia et al.· 0 citations
LOFA combines reinforcement learning over verifiable purchase outcomes with feedback-aware on-policy distillation, which identifies users' in-dialogue directives and converts them into dense token-level supervision, which captures both collaborative behavioral patterns and user-specific preferences.
Haobo Zhang, Kelong Mao, Sulong Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.