Skip to content

Not All Preferences Deserve Gradients: Understanding Gradient Utility in Offline Reasoning Alignment

Feb 2026 · 0 citations · 21 references
Computer Science

TL;DR

SAGE (Stability-Aware Gradient Efficiency), which maintains difficulty-stratified candidate pools refreshed during training and selects pairs within each pool by a forward-pass signal-to-curvature score, outperforms full-data and size-matched baselines while producing substantially smoother optimization trajectories.

Abstract

Offline preference optimization aligns reasoning models from fixed chosen--rejected pairs, yet standard methods apply gradient updates from every pair regardless of its training value under the current policy. We argue that this uniform treatment is wasteful and potentially harmful. From the perspective of gradient utility, we show that a pair's contribution depends jointly on informativeness and stability. Pair utility drifts as the policy evolves, high-gradient samples can coincide with high-curvature regions, leading to noisy and destabilizing updates, and the most effective supervision comes from stable confident errors where the model is reliably wrong yet curvature remains low. These findings motivate SAGE (Stability-Aware Gradient Efficiency), which maintains difficulty-stratified candidate pools refreshed during training and selects pairs within each pool by a forward-pass signal-to-curvature score. Only pairs with high current utility receive gradient computation; the rest are excluded from backpropagation. On mathematical reasoning benchmarks across multiple model scales, SAGE outperforms full-data and size-matched baselines while producing substantially smoother optimization trajectories.

View source

Similar papers

#machine learning Preprint Sep 2026

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

Gradient-Aligned Reward (GAR), which operates in the policy's own gradient space: truncated backpropagation through the output projection layer extracts a compact gradient vector for each rollout, and cosine similarity with an expert-anchor gradient yields a dense, reasoning-aware reward with less than 9% wall-clock ov...

Le-Qi Zheng, Jin-Bo Su, Fang Niu et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Optimal Design for Active Preference Learning with Biased LLM Judges

Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference learning reduces this cost by selecting informative comparisons, and LLM judges can provide additional scalable feedback. However, the preferences of the judges may deviate fr...

Zhong-Man Du, Hui-Ming Zhang, Hao-Dong Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models

Preference alignment for flow and diffusion models now spans online reinforcement learning and offline preference optimization, but the relation between these methods remains unclear. In particular, existing forward-process alignment methods require fresh samples from the current model, while offline methods based on f...

Yan-Sen Han, Shengyi Liao, Peng Sun et al. · 0 citations
#machine learning Preprint Sep 2026

ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients

Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients into one policy update, achieves the highest average full-budget accuracy and three-budget hypervolume among the compared methods.

Shi-Cheng Fang, Yi-Wen Zhao, Wen-Bo Tian et al. · 0 citations
#machine learning Preprint Aug 2026

PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

PLC-DPO is proposed to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case, which reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples.

Boryeong Cho, Sumyeong Ahn, SeYoung Yun · 0 citations
Preprint Aug 2026

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

Environment-Regularized Policy Optimization (ERPO) replaces the standard Policy-KL regularizer while achieving effective control over query distribution drift, delivering stronger accuracy and substantially more stable behavior under high-temperature decoding and long-horizon training.

Xianlei Zhou, Xiangdi Meng, Yu He et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.