Skip to content

Inference-Time Nash Alignment

Sep 2026 · 0 citations · 49 references
Computer Science

TL;DR

This work forms the problem as obtaining a Nash equilibrium of a two-player zero-sum game between policies, and proposes two algorithms: Best-of-Nash (BoN) and Nash Mirror Descent (NMD), which are proved to achieve a duality gap that matches the problem lower bound.

Abstract

Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large preference datasets. They also need direct access to the model parameters which are not provided by many state-of-the art models. Inference-time alignment offers a cost-effective alternative without updating model parameters. However, existing inference-time methods rely on a scalar reward model derived under a Bradley-Terry assumption, which cannot represent general preferences. Following recent work on fine-tuning with generalized preferences, in this work, we initiate the study of inference-time alignment under general preferences. We formulate the problem as obtaining a Nash equilibrium of a two-player zero-sum game between policies. We propose two algorithms: Best-of-Nash (BoN) and Nash Mirror Descent (NMD). We prove that both algorithms achieve a duality gap that matches the problem lower bound. Empirically, we implement the two methods on three datasets, which shows that our methods substantially outperform the base policy, converging to the performance of the fine-tuned models. Moreover, our results show that NMD remains robust across the regularization parameter.

View source

Similar papers

#machine learning Preprint Sep 2026

NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games

NashDreamer is proposed, a principled MBRL framework for two-player zero-sum IIGs that introduces a centralized Multi-Agent Recurrent State-Space Model (MARSSM) that decouples environment dynamics from the effect of players's strategies on their individual observations.

Tomáš Holeček, Viliam Lisý · 0 citations
Preprint Aug 2026

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

This paper introduces a novel reward-model-free, critic-free, and gradient-based PbRL algorithm compatible with segment preferences named Segment Pairwise Proximal Policy Optimization (SP3O), and provides a theoretical basis for the algorithm and analyze the tradeoff in choosing the segment length.

Evan Assmus, Qi-Ning Zhang, Lei Ying · 0 citations
Preprint Aug 2026

Towards a theory of inference-time alignment with unknown rewards

This work introduces a novel combinatorial dimension of the reward class which is called the alignment dimension, and shows that it completely characterizes the alignment learnability --- a reward class is alignment learnable if and only if its alignment dimension is finite.

Steve Hanneke, Hongao Wang, Mingyue Xu · 0 citations
#machine learning Preprint Sep 2026

IncentRL: The Trade-Off Between Preference Guidance and Task Performance

Preference-based reward shaping can guide reinforcement learning, but adding preference signals to the reward may unintentionally change the task being optimized. We address this problem with IncentRL, a framework that introduces preference guidance while explicitly characterizing its effect on external-task performanc...

Xue-Ning Wu, Yan-Lan Kang, Shen Yin · 0 citations
#machine learning Preprint Sep 2026

ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment

Best-of-$n$ (BoN) sampling is a simple yet effective inference-time alignment method, but hard maximization provides only coarse control over the trade-off between reward and distribution shift. Soft Best-of-$n$ (Verdun et al. 2025) provides smoother control and converges to the optimal distribution associated with KL-...

Yan-Xiao Liu, Si-Cheng Wan, Deniz Gündüz · 0 citations
#machine learning Preprint Sep 2026

Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation

This work proposes Distilled Preference Probability Policy Optimization (DP3O), an effective and efficient offline alignment algorithm that outperforms state-of-the-art offline methods, achieves performance comparable to iterative DPO, and reduces training time, demonstrating both its effectiveness and efficiency.

Wen-Bo Zhang, Wen-Zhuo Zhou, Heng-Rui Cai et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.