Skip to content
Conference

LCW-DPO: Direct Preference Optimization with Length Calibration and Confidence Weighting

Aug 2026 · 2026 12th International Conference on Big Data and Information Analytics (BigDIA) · pp. 128-135 · 0 citations · 25 references

Abstract

Direct Preference Optimization (DPO) has become an important method for preference alignment of large language models. However, standard DPO still faces two common issues on real-world preference data. First, sequence-level log probability is obtained by accumulating token-level log probabilities, which makes the implicit reward sensitive to response length. Second, standard DPO assigns the same training weight to all preference pairs, making it difficult to handle weak preferences and low-quality samples. In this paper, we propose LCW-DPO (Length-Calibrated and Confidence-Weighted DPO). The method introduces a length-calibrated score that provides a continuous trade-off between sequence-level accumulation and token-level averaging, and softly reweights samples using a confidence score constructed from the reference-model margin and the response-length ratio. Experimental results show that LCW-DPO outperforms DPO and SimPO in preference-ranking accuracy on UltraFeedback Binarized and HH-RLHF, while maintaining more stable performance across multiple length ranges.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.