Skip to content

Instruction-Tuned LLMs via Bayesian Mixture Model Preference Priors

Sep 2026 · International Symposium on Networks, Computers and Communications · pp. 1-7 · 0 citations · 35 references

Abstract

Large language models (LLMs) deployed as interactive recommendation agents must adapt to user preferences across dialogue rounds without access to ground-truth reward functions. Existing Bayesian-prompting approaches assume a uniform prior over user types, discarding population structure and weakening cold-start performance. We propose an empirical Bayes framework that fits a Dirichlet Compound Negative Multinomial mixture with Feature Saliency (DCNM-FS) to interaction logs without supervision, recovering interpretable user-type archetypes with per-feature saliency scores. At inference time, a closed-form posterior over user types is updated after each choice and injected into the LLM context, while a session-local exponential moving average further personalizes the profile. On a synthetic flight-search dataset and the WebShop e-commerce benchmark, DCNM-FS context improves recommendation accuracy by 30%, with the largest gains in later rounds, where the structured posterior adapts to individual observations.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling

Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most common...

Bo-Hao Wang, Xiao-Yan Zhao, Yang Zhang et al. · 0 citations
Preprint Aug 2026

Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation

This work proposes a new approach that quantifies the effectiveness of each interaction by the reduction in the assistant's uncertainty, measured via entropy over recommendations, to fine-tune the LLM, enabling strategic interaction generation.

Cedar Site Bai, Zhen-Yu Liao, Duan Li et al. · 0 citations
Preprint Aug 2026

Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning

An extended faithfulness analysis shows that the refined profiles remain largely grounded in the source preferences while preserving task-relevant personalization signals, suggesting that profile-side adaptation serves as a practical complement to universal memory construction for lifelong personalized agents.

Yuting Liu, Wei Wu, Jianzhe Zhao et al. · 0 citations
Conference Open access Sep 2026

Test-Time User Alignment via Bayesian Population Guidance in Subjective Tasks

In subjective tasks, different individuals can have different correct answers for the same input—the ground truth is not fixed but rather determined by each user's personal perspective. Standard foundation models suppress this individual variation by producing population-averaged predictions; conversely, few-shot in-co...

H. Ryu, J. Kang, C. Wallraven · 0 citations
Open access Aug 2026

Multimodal Interest-Shifting Sequence Recommendation with Offline Policy Optimization and Drift-Aware Representation Learning

Interest changes complicate sequential recommendation when interaction histories are combined with item content. We evaluate MM-DRLSR on four public Amazon and Yelp benchmarks. The model integrates category-overlap drift supervision, history-derived drift representations, lightweight identifier–image–text fusion, candi...

Chang-Cheng Shao, Cheng Zeng, Xiao-Gang Ye et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.