Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
Experiments show that the proposed Preference Vector framework improves helpfulness without excessive conservatism, allows smooth control over preference trade-offs, and supports scalable multi-preference alignment.
Ren-Wei Liang, Chin-Ting Hsu, Chan-Hung Yu et al.
· Volume 1 · 7 citations