Skip to content
Preprint

Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization

Aug 2026 · 0 citations · 49 references
Computer Science

TL;DR

PAC-Bayes-regularized Meta-LoRA is proposed, which uses a meta-learned LoRA initialization as both the adaptation start and prior center, while adjusting update strength according to support-set size and predictive uncertainty to limit overfitting under sparse or ambiguous evidence.

Abstract

Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen conversational domains from only a handful of target-domain interactions. Existing adaptation methods struggle to calibrate update magnitude under sparse evidence and thus overfit, whereas history-transfer methods often entangle user preferences with source-domain artifacts, yielding unreliable personalization priors and negative transfer. To calibrate adaptation to evidence quality, we propose PAC-Bayes-regularized Meta-LoRA, which uses a meta-learned LoRA initialization as both the adaptation start and prior center, while adjusting update strength according to support-set size and predictive uncertainty. This limits overfitting under sparse or ambiguous evidence while permitting stronger personalization as evidence grows. Controlled adaptation alone does not determine which preferences should transfer across domains or how they should be expressed. We therefore functionally decompose personalization priors into user and domain components, using a human-readable prompt for stable preferences and topology-preserving soft tokens for domain-specific hidden-space conditioning. Experiments across multiple benchmarks and personalization tasks show consistent gains over strong baselines. On HiCUPID, our method reduces cross-domain win-rate degradation by 47.9% relative to the best competing baseline and improves win rate by 110.2% under unseen-user cold start.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation

Experiments show that COPE consistently outperforms strong training-free and training-based baselines under sparse feedback, and remains complementary to Retrieval-Augmented Prompting, and further analyses confirm COPE's reliable self-evaluation, meaningful preference patterns, stable general capabilities, and robustne...

Rui-Ke Cao, Fu-Gen Yao, Liang Dong et al. · 0 citations
Preprint Aug 2026

COSMO: Consensus-Driven Shift Modulation for Source-Free Domain Adaptation

COSMO replaces expert-to-expert guidance with co-adaptation through an anchored shared consensus and achieves state-of-the-art performance under matched VLM backbones, indicating that it better balances the retention of valid source-derived evidence with the absorption of complementary VLM evidence.

Bo Li, Junjie Peng, Xiaohua Xie et al. · 0 citations
Sep 2026

Instruction-Tuned LLMs via Bayesian Mixture Model Preference Priors

Large language models (LLMs) deployed as interactive recommendation agents must adapt to user preferences across dialogue rounds without access to ground-truth reward functions. Existing Bayesian-prompting approaches assume a uniform prior over user types, discarding population structure and weakening cold-start perfor...

Ornela Bregu, Nizar Bouguila · 0 citations
Preprint Aug 2026

Cautious Context Steering for Language Model Personalization

Cautious Context Steering (CCS), which adds a lightweight adapter to a frozen backbone LM to decide at each token whether and how strongly user context should affect generation, demonstrating robust generalization to new users and domains.

Gihoon Kim, Jeyoung Lee, S. Woo et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Hypotheses-Guided Self Distillation for Continual Personalization

HypReflect is introduced, a reliable, scalable framework for continual personalization that infers explicit, uncertainty-aware preference hypotheses from diverse user signals, reflectively refines them as new evidence accumulates, and incorporates the resulting user model through hypotheses-guided self-distillation.

Eunjeong Hwang, Kushan Mitra, Dan Zhang et al. · 0 citations
#natural language process... Preprint Sep 2026

HyperTrace: Hypothesis-Based Preference Tracing for Online LLM Personalization

Personalized language models aim to adapt responses to individual users, whose preferences are often latent and revealed gradually through interaction. Existing training-free methods rely on stored histories or retrieved memories, but they often struggle to reconcile long- term preferences with short-term topic-specifi...

Jian-Zhi Shen, Ke-Yu Mao, Ming-Hao Shao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.