Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training
Data-DPO, a target model-oriented SFT data selection method that consistently outperforms existing data selection baselines under multiple data budgets and stably surpasses full data training performance is proposed.
Peng Sun, Yi Yang, Antong Zhang et al.
· 0 citations