Skip to content

Inference-Time Nash Alignment

Sep 2026 · 0 citations · 49 references
Computer Science

TL;DR

This work forms the problem as obtaining a Nash equilibrium of a two-player zero-sum game between policies, and proposes two algorithms: Best-of-Nash (BoN) and Nash Mirror Descent (NMD), which are proved to achieve a duality gap that matches the problem lower bound.

Abstract

Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large preference datasets. They also need direct access to the model parameters which are not provided by many state-of-the art models. Inference-time alignment offers a cost-effective alternative without updating model parameters. However, existing inference-time methods rely on a scalar reward model derived under a Bradley-Terry assumption, which cannot represent general preferences. Following recent work on fine-tuning with generalized preferences, in this work, we initiate the study of inference-time alignment under general preferences. We formulate the problem as obtaining a Nash equilibrium of a two-player zero-sum game between policies. We propose two algorithms: Best-of-Nash (BoN) and Nash Mirror Descent (NMD). We prove that both algorithms achieve a duality gap that matches the problem lower bound. Empirically, we implement the two methods on three datasets, which shows that our methods substantially outperform the base policy, converging to the performance of the fine-tuned models. Moreover, our results show that NMD remains robust across the regularization parameter.

View source

Similar papers

#machine learning Preprint Sep 2026

NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games

NashDreamer is proposed, a principled MBRL framework for two-player zero-sum IIGs that introduces a centralized Multi-Agent Recurrent State-Space Model (MARSSM) that decouples environment dynamics from the effect of players's strategies on their individual observations.

Tomáš Holeček, Viliam Lisý · 0 citations
Preprint Aug 2026

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

This paper introduces a novel reward-model-free, critic-free, and gradient-based PbRL algorithm compatible with segment preferences named Segment Pairwise Proximal Policy Optimization (SP3O), and provides a theoretical basis for the algorithm and analyze the tradeoff in choosing the segment length.

Evan Assmus, Qi-Ning Zhang, Lei Ying · 0 citations
Preprint Aug 2026

Towards a theory of inference-time alignment with unknown rewards

Generative model alignment has received broad interest, and significant progress has been made in supervised fine-tuning and inference-time computation. Yet, alignment has remained poorly understood from a statistical learning perspective. We formulate inference-time alignment as a weak-to-strong learning problem, wher...

Steve Hanneke, Hongao Wang, Mingyue Xu · 0 citations
#machine learning Preprint Sep 2026

IncentRL: The Trade-Off Between Preference Guidance and Task Performance

Preference-based reward shaping can guide reinforcement learning, but adding preference signals to the reward may unintentionally change the task being optimized. We address this problem with IncentRL, a framework that introduces preference guidance while explicitly characterizing its effect on external-task performanc...

Xue-Ning Wu, Yan-Lan Kang, Shen Yin · 0 citations
#machine learning Preprint Sep 2026

ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment

Best-of-$n$ (BoN) sampling is a simple yet effective inference-time alignment method, but hard maximization provides only coarse control over the trade-off between reward and distribution shift. Soft Best-of-$n$ (Verdun et al. 2025) provides smoother control and converges to the optimal distribution associated with KL-...

Yan-Xiao Liu, Si-Cheng Wan, Deniz Gündüz · 0 citations
#machine learning Preprint Sep 2026

Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation

Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) due to its simplicity, computational efficiency, and implicit modeling of human preferences. Interestingly, iterative extensions of DPO have achieved stronger performance on academic benchmarks, raising two key...

Wen-Bo Zhang, Wen-Zhuo Zhou, Heng-Rui Cai et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.