Skip to content

COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation

Sep 2026 · 0 citations · 35 references
Computer Science

TL;DR

Experiments show that COPE consistently outperforms strong training-free and training-based baselines under sparse feedback, and remains complementary to Retrieval-Augmented Prompting, and further analyses confirm COPE's reliable self-evaluation, meaningful preference patterns, stable general capabilities, and robustness under shifting preferences and alternative evaluators.

Abstract

While Large Language Models (LLMs) have achieved remarkable results across various benchmarks, their alignment with normative values often results in homogenized responses that fail to address diverse user preferences. Existing training-free methods often occupy valuable context windows through prompt engineering, while training-based methods typically remain static post-training, failing to support the continual optimization required in real-world settings. To address these challenges, we propose COPE (Continual Optimization with Personalized embedding and self-Evaluation), a novel optimization framework tailored for real-world-motivated interaction settings with sparse user feedback. Our framework assigns learnable personalized embeddings to each user and synergistically integrates preference capture, self-evaluation calibration, and personalized response optimization within a single update step. A key innovation of our method is the use of self-evaluation to generate proxy rewards, enabling continuous model updates even when explicit user feedback is unavailable. Experiments show that COPE consistently outperforms strong training-free and training-based baselines under sparse feedback, and remains complementary to Retrieval-Augmented Prompting (RAP). Further analyses confirm COPE's reliable self-evaluation, meaningful preference patterns, stable general capabilities, and robustness under shifting preferences and alternative evaluators.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Hypotheses-Guided Self Distillation for Continual Personalization

HypReflect is introduced, a reliable, scalable framework for continual personalization that infers explicit, uncertainty-aware preference hypotheses from diverse user signals, reflectively refines them as new evidence accumulates, and incorporates the resulting user model through hypotheses-guided self-distillation.

Eunjeong Hwang, Kushan Mitra, Dan Zhang et al. · 0 citations
#natural language process... Preprint Sep 2026

HyperTrace: Hypothesis-Based Preference Tracing for Online LLM Personalization

Personalized language models aim to adapt responses to individual users, whose preferences are often latent and revealed gradually through interaction. Existing training-free methods rely on stored histories or retrieved memories, but they often struggle to reconcile long- term preferences with short-term topic-specifi...

Jian-Zhi Shen, Ke-Yu Mao, Ming-Hao Shao et al. · 0 citations
Preprint Aug 2026

Progressive Content Refinement with Decaying Reward Joint LinUCB

A novel contextual bandit algorithm that explicitly incorporates reward decay modeling that achieves significant performance gains over strong baselines and confirms that the integration of reward decay modeling within the bandit framework is crucial for mitigating over-exploitation and optimizing the iterative refinem...

Shion Ishikawa, Pablo Loyola, Young-joo Chung et al. · 0 citations
Book Open access Aug 2026

Generalizable Multi-Pass Training of Ads Recommendation Models with Foundation Model Guidance

High-capacity Click-Through Rate (CTR) models for ads recommendation often exhibit pronounced one-epoch overfitting: performance peaks after a single training pass (epoch) over the data, while additional epochs degrade generalization as the model memorizes noise in high-variance click labels. To address this challenge,...

Yunzhe Qi, Qinghai Zhou, Bo-Yang Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces

Personalizing large language models (LLMs) is essential for delivering AI assistance that aligns with individual users'styles, intents, and preferences. While per-user fine-tuning can substantially enhance personalization quality, it introduces significant parameter and storage overhead, limiting scalability to large u...

Xin-Yu Li, Hao Zhou, Jian-Feng Zhu et al. · 0 citations
#large language models Review Open access Sep 2026

PALRec: Large Language Model-Based Sequential Recommendation With Parameter-Preserving Augmentation

PALRec is proposed, a parameter-preserving augmentation framework that equips an LLM with recommendation capabilities while keeping its original parameters fixed and consistently outperforms fully fine-tuned counterparts in recommendation accuracy while preserving the LLM’s pre-trained knowledge.

Hyunsoo Na, Minseok Gang, Sang-goo Lee et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.