Skip to content
Preprint

Private Direct Preference Optimization for LLM Alignment

Aug 2026 · 0 citations · 77 references
Computer Science

TL;DR

This paper formalizes preference privacy, a label-DP-style privacy notion for DPO that protects only the relative preference between candidate responses, assuming an adversary who already knows the prompt and responses, and designs PrivDPO, a DPO variant that enforces preference privacy while remaining compatible with large-scale LLM training.

Abstract

Direct preference optimization (DPO) is now a standard method for aligning large language models (LLMs) using human preference data. Each DPO example contains a prompt and a pair of candidate model responses. While prompts and responses are often public or model-generated, the relative preference between responses reflects subjective judgments and can reveal sensitive attributes of annotators or end users. Off-the-shelf privacy-preserving approaches are not well matched to this structure, leading to unnecessary noise injection and biased updates in training. In this paper, we formalize preference privacy, a label-DP-style privacy notion for DPO that protects only the relative preference between candidate responses, assuming an adversary who already knows the prompt and responses. We then design PrivDPO, a DPO variant that enforces preference privacy while remaining compatible with large-scale LLM training. Our main observation is that, for neighboring examples differing only in their preference signal, the gradient difference lies on a one-dimensional preference axis determined solely by the text; all preference information flows through this axis. PrivDPO adds calibrated randomness only along this axis via an unbiased randomized rescaling of the DPO objective, avoiding per-example gradient operations. Our experiments on three alignment benchmarks and three LLM families show that PrivDPO consistently achieves strong privacy-utility trade-offs compared with privacy-preserving baselines.

View source

Similar papers

Jun 2026

Data-Efficient Online Training for Direct Alignment in LLMs

In recent years, online Direct Alignment from Preferences (DAP) has emerged as a popular alternative for Reinforcement Learning from Human Feedback (RLHF) due to its training stability and simplicity. In online DAP, training relies on preference data, each composed of a question and a pair of large language model (LLM)...

Chi Zhang, Jia-Chen T. Wang, Kun He et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection

AlignDiff, a preference data filtering framework driven by intrinsic model signals, first identifies samples with clear preferences using both positive and inverse signals, then prioritizes the more challenging samples based on the average negative log-likelihood gap, encouraging the model to learn richer information f...

Peng Lai, He Zhu, Zhiwen Ruan et al. · 1 citation
#machine learning Preprint Aug 2026

PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

PLC-DPO is proposed to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case, which reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples.

Boryeong Cho, Sumyeong Ahn, SeYoung Yun · 0 citations
Preprint Aug 2026

Preference Data Selection for Mitigating the Alignment Tax in Large Language Models

This paper proposes BALIGN, a balanced data selection strategy that explicitly mitigates catastrophic forgetting while optimizing alignment efficacy, and identifies three key data-centric features that dictate parameter drift: the reference model's log-probability margin, the token length between chosen and rejected re...

Minsu Kim, Jian-Xun Lian, Xing Xie et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.