Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
Experiments show that the proposed Preference Vector framework improves helpfulness without excessive conservatism, allows smooth control over preference trade-offs, and supports scalable multi-preference alignment.