Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most common...
Bo-Hao Wang, Xiao-Yan Zhao, Yang Zhang et al.· 0 citations
LBR is proposed, a lightweight and model-agnostic framework for mitigating length bias in LLM-based recommendation that substantially alleviates length bias while consistently improving recommendation accuracy and fairness, with negligible additional training and inference overhead.
GALLM constructs a collaborative graph over text tokens and item tokens, and models three types of relations that are transformed into lightweight learnable attention biases and incorporated into the LLM attention mechanism, enabling collaborative-aware token interactions without introducing an additional graph encoder...
Fenglin Yan, Bo-Hao Wang, Jian Zhang et al.· 0 citations
SmartGR is proposed, a novel distillation framework that utilizes Hierarchy-Aware SID Distillation to transfer the teacher's modeling capability across the hierarchy and leverages Beam-Aware Ranking Distillation to distill the teacher's ranking preferences during beam search.
Ziheng Zhang, Yu Cui, Bohao Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.