SODA, a plug-and-play alignment framework that adopts a BPR-style contrastive objective to align recommender representations with target-side distributional representations against negative ones, is developed and demonstrated that SODA consistently strengthens diverse generative recommendation architectures.
Zi-Qiu Xue, Ding-Xian Wang, Yi-Meng Bai et al.· Proceedings of the 20th ACM...· 0 citations
Recent generative recommenders improve scalability by retrieving items through token generation instead of traditional ranking over large candidate sets. Yet their training signals are still dominated by discrete code prediction, which overlooks the soft assignment information naturally produced by the tokenizer. This...
Zi-Qiu Xue, Ding-Xian Wang, Yi-Meng Bai et al.· Proceedings of the 20th ACM...· 0 citations
Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most common...
Bo-Hao Wang, Xiao-Yan Zhao, Yang Zhang et al.· 0 citations
Retrieval-Augmented Generation (RAG) is widely regarded as a novel paradigm born from the limitations of large language models (LLMs)--a mechanism to ground their outputs in external knowledge. This view, however, is incomplete when considered within a broader historical context. In this paper, we argue that the core i...
Xiaoyan Zhao, Yujie Cai, Yang Zhang et al.· 0 citations
Experiments on the LongLaMP dataset show that PrefReward outperforms non-personalized and retrieval-based baselines in both generation quality and personalization interpretability.
Yue Wu, Chengbing Wang, Yimeng Bai et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.