Skip to content

Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization

Jul 2026 · arXiv.org · Vol abs/2607.25136 · 0 citations · 27 references
Computer Science

TL;DR

This method generates candidate responses from the target policy, evaluates helpfulness, factuality, and conciseness with rubric-specialized evaluators, applies a process-critic correction, and retains only high-consensus desirable or undesirable examples.

Abstract

Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-confidence set of on-policy responses can provide a reliable learning signal. Our method, DMAPO (Data-centric Multi-evaluator Agreement for Preference Optimization), generates candidate responses from the target policy, evaluates helpfulness, factuality, and conciseness with rubric-specialized evaluators, applies a process-critic correction, and retains only high-consensus desirable or undesirable examples. This procedure accepts 1,871 of 54,236 Mistral-7B candidates (3.45%). KTO trained on this set reaches 7.50 on MT-Bench, 95.5% length-controlled win rate against a text-davinci-003 reference, and 57.3% IFEval prompt accuracy. Independent pairwise evaluation also favors DMAPO over SimPO: GPT-4o yields a net win rate of 23.3 points on 129 held-out prompts and 24.0 points on 200 out-of-distribution LMSYS-Chat prompts; Claude Opus 4.7 yields 24.1 points on the held-out set. Changing the evaluator model or rubric alters the selected examples but has little effect on downstream performance. A second-backbone study yields a similar 3.41% acceptance rate, although its performance gains are more modest. Across these experiments, consensus filtering offers a data-efficient route to preference optimization for general instructions, at the cost of additional curation compute and dependence on evaluator judgments.

View source

Similar papers

Preprint Aug 2026

MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment

Across cooperative emotional support and adversarial negotiation, min-selection lifts both objectives while sharply cutting their imbalance; on emotional support it raises the weaker axis from 0.37 to 0.64, surpassing human experts and persisting across full multi-turn rollouts.

Tony Tu, Sayan Chakraborty, Ruomeng Xu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection

AlignDiff, a preference data filtering framework driven by intrinsic model signals, first identifies samples with clear preferences using both positive and inverse signals, then prioritizes the more challenging samples based on the average negative log-likelihood gap, encouraging the model to learn richer information f...

Peng Lai, He Zhu, Zhiwen Ruan et al. · 1 citation
Preprint Aug 2026

Private Direct Preference Optimization for LLM Alignment

This paper formalizes preference privacy, a label-DP-style privacy notion for DPO that protects only the relative preference between candidate responses, assuming an adversary who already knows the prompt and responses, and designs PrivDPO, a DPO variant that enforces preference privacy while remaining compatible with...

Yang-Fan Jiang, Fei Wei, Ergute Bao et al. · 0 citations
Conference Aug 2026

LCW-DPO: Direct Preference Optimization with Length Calibration and Confidence Weighting

Direct Preference Optimization (DPO) has become an important method for preference alignment of large language models. However, standard DPO still faces two common issues on real-world preference data. First, sequence-level log probability is obtained by accumulating token-level log probabilities, which makes the impli...

Jia-Yi Li, Jia-Hui Jin, Hai Zhang · 0 citations
Preprint Aug 2026

Preference Data Selection for Mitigating the Alignment Tax in Large Language Models

This paper proposes BALIGN, a balanced data selection strategy that explicitly mitigates catastrophic forgetting while optimizing alignment efficacy, and identifies three key data-centric features that dictate parameter drift: the reference model's log-probability margin, the token length between chosen and rejected re...

Minsu Kim, Jian-Xun Lian, Xing Xie et al. · 0 citations
Jul 2026

Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

Rubric4Setwise is proposed, a training-free method that converts rubric-based evaluation criteria into document set selection signals, achieving the best downstream generation performance with fewer documents and search rounds, validating the effectiveness of closing the loop from evaluation to optimization.

Kai-Lin Jiang, Lei Liu, Jian-Fei Xi et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.