LOFA combines reinforcement learning over verifiable purchase outcomes with feedback-aware on-policy distillation, which identifies users' in-dialogue directives and converts them into dense token-level supervision, which captures both collaborative behavioral patterns and user-specific preferences.
Abstract
Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user interaction logs that provide valuable supervision for improving these agents. However, existing approaches primarily rely on offline training signals, such as user-item interactions or synthetic preference data, while largely overlooking the rich supervision contained in users'natural conversational feedback. Moreover, the available online feedback is heterogeneous, sparse, and noisy, making it difficult to transform into reliable learning signals automatically. To address these challenges, we propose LOFA, a framework that enables shopping agents to learn directly from real online interaction logs without human annotation. LOFA combines reinforcement learning over verifiable purchase outcomes with feedback-aware on-policy distillation, which identifies users'in-dialogue directives and converts them into dense token-level supervision. These complementary objectives capture both collaborative behavioral patterns and user-specific preferences. Extensive experiments on real-world e-commerce logs demonstrate that LOFA consistently improves recommendation quality, response helpfulness, and user-satisfaction alignment over strong baselines, highlighting the effectiveness of learning shopping agents from real online user feedback.
The rapid growth of online social platforms has transformed communication and information retrieval, giving rise to social search, where queries-titles are typically expressed in informal, community-specific language. While large language models provide strong general-purpose semantic understanding, their effectiveness...
Tao Su, Jin-Jing Hu, Xiao Wang et al.· arXiv.org· 0 citations
A recommender system extracts user preferences from past interactions to suggest items, widely used in platforms like user-generated content, online shopping, and urban services. These systems aim to provide accurate recommendations, reduce user interaction burden, and enhance user experience while improving socio-econ...
Yiming Cheng, Yitong Ma, Jingyu Wang et al.· Electronics· 1 citation
Agentic recommender systems increasingly employ large language model-based UserAgents to evaluate candidate items through simulated feedback before recommendations are delivered. However, existing UserAgents typically reason in isolation based on limited personal histories, which may lead to perspective narrowing: the...
Zongwei Wang, Min Gao, Guang-Yu Hu et al.· 0 citations
Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.g., clicks, dwell time) rather than explicit, natural language inputs. As a result, users experience a persistent discrepancy between their explicit interests and what passiv...
Zi-Yun Xu, Bo-Sen Ding, Yue Zhang et al.· Proceedings of the 20th ACM...· 1 citation
RecVerse is presented, a GUI-grounded simulation agent that perceives pages through screenshots and produces faithful multi-turn trajectories and significantly outperforms existing baselines in both behavioral fidelity and intent consistency.
DASH is a decision-aware user simulator that jointly generates thinking traces and predicts behavioral actions from heterogeneous cross-domain histories and tailors a rubric-based reward model that evaluates thinking traces along form, content, and logic for RL training.