Sep 2026· International Symposium on Networks, Computers and Communications· pp. 1-7· 0 citations· 35 references
Abstract
Large language models (LLMs) deployed as interactive recommendation agents must adapt to user preferences across dialogue rounds without access to ground-truth reward functions. Existing Bayesian-prompting approaches assume a uniform prior over user types, discarding population structure and weakening cold-start performance. We propose an empirical Bayes framework that fits a Dirichlet Compound Negative Multinomial mixture with Feature Saliency (DCNM-FS) to interaction logs without supervision, recovering interpretable user-type archetypes with per-feature saliency scores. At inference time, a closed-form posterior over user types is updated after each choice and injected into the LLM context, while a session-local exponential moving average further personalizes the profile. On a synthetic flight-search dataset and the WebShop e-commerce benchmark, DCNM-FS context improves recommendation accuracy by 30%, with the largest gains in later rounds, where the structured posterior adapts to individual observations.
It is suggested that LLM-derived priors can serve as a practical warm-start mechanism for text-rich bandit recommendation, while also revealing deployment trade-offs.
E. Lee, Oseong Choi, Byungsoo Kang et al.· 0 citations
Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most common...
Bo-Hao Wang, Xiao-Yan Zhao, Yang Zhang et al.· 0 citations
This work proposes a new approach that quantifies the effectiveness of each interaction by the reduction in the assistant's uncertainty, measured via entropy over recommendations, to fine-tune the LLM, enabling strategic interaction generation.
Cedar Site Bai, Zhen-Yu Liao, Duan Li et al.· 0 citations
An extended faithfulness analysis shows that the refined profiles remain largely grounded in the source preferences while preserving task-relevant personalization signals, suggesting that profile-side adaptation serves as a practical complement to universal memory construction for lifelong personalized agents.
Yuting Liu, Wei Wu, Jianzhe Zhao et al.· 0 citations
In subjective tasks, different individuals can have different correct answers for the same input—the ground truth is not fixed but rather determined by each user's personal perspective. Standard foundation models suppress this individual variation by producing population-averaged predictions; conversely, few-shot in-co...
H. Ryu, J. Kang, C. Wallraven· Proceedings of the Thirty-Fi...· 0 citations
Interest changes complicate sequential recommendation when interaction histories are combined with item content. We evaluate MM-DRLSR on four public Amazon and Yelp benchmarks. The model integrates category-overlap drift supervision, history-derived drift representations, lightweight identifier–image–text fusion, candi...
Chang-Cheng Shao, Cheng Zeng, Xiao-Gang Ye et al.· Applied Sciences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.