To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization
This work argues that, given a sufficiently diverse user population, a curriculum naturally emerges between easy- and hard-to-optimize reward models, and proposes CurriPO, which grows a tree-structured curriculum to accommodate diverse user-specific objectives, covering the population in a single traversal.