LCW-DPO: Direct Preference Optimization with Length Calibration and Confidence Weighting
Abstract
Direct Preference Optimization (DPO) has become an important method for preference alignment of large language models. However, standard DPO still faces two common issues on real-world preference data. First, sequence-level log probability is obtained by accumulating token-level log probabilities, which makes the implicit reward sensitive to response length. Second, standard DPO assigns the same training weight to all preference pairs, making it difficult to handle weak preferences and low-quality samples. In this paper, we propose LCW-DPO (Length-Calibrated and Confidence-Weighted DPO). The method introduces a length-calibrated score that provides a continuous trade-off between sequence-level accumulation and token-level averaging, and softly reweights samples using a confidence score constructed from the reference-model margin and the response-length ratio. Experimental results show that LCW-DPO outperforms DPO and SimPO in preference-ranking accuracy on UltraFeedback Binarized and HH-RLHF, while maintaining more stable performance across multiple length ranges.