Instruction-Tuned LLMs via Bayesian Mixture Model Preference Priors
Large language models (LLMs) deployed as interactive recommendation agents must adapt to user preferences across dialogue rounds without access to ground-truth reward functions. Existing Bayesian-prompting approaches assume a uniform prior over user types, discarding population structure and weakening cold-start perfor...