Conference
Open access
2026
V-RoLoRA: RLVR-Driven MoE Routing for Steerable Pluralistic Alignment
This work studies value-controllable alignment through discrete condition vectors and proposes Verifiable-reward-Routed LoRA—a parameter-efficient mixture-of-experts LoRA framework enhanced with conditioned gating, which consistently out-performs prompt-based steering and multi-task PEFT baselines.
Jing Wang, Yaomin Wu, Yinglin Wang et al.
· Annual Meeting of the Associ... · 0 citations