Skip to content
Preprint

Towards a theory of inference-time alignment with unknown rewards

Aug 2026 · 0 citations · 49 references
Computer Science

TL;DR

This work introduces a novel combinatorial dimension of the reward class which is called the alignment dimension, and shows that it completely characterizes the alignment learnability --- a reward class is alignment learnable if and only if its alignment dimension is finite.

Abstract

Generative model alignment has received broad interest, and significant progress has been made in supervised fine-tuning and inference-time computation. Yet, alignment has remained poorly understood from a statistical learning perspective. We formulate inference-time alignment as a weak-to-strong learning problem, where a reference policy (weak model) is assumed to be fairly good and the goal is to produce a strong model that predicts a good response at test time with arbitrarily high probability. Our problem is formulated as learning from scratch --- everything is learned from data rather than assuming access to a good reward estimate, and thus differs from the existing inference-time alignment theory. Our framework shares similarity to the recent work of Joshi et al., (arXiv:2510.15464), where for each prompt, there could be multiple good responses. Our definition of the alignment learnability follows the standard PAC learning principle. We introduce a novel combinatorial dimension of the reward class which we call the alignment dimension, and show that it completely characterizes the alignment learnability --- a reward class is alignment learnable if and only if its alignment dimension is finite. The core of our learning procedure works by learning a pairwise comparator and then running a tournament over candidate responses. We believe that our results might shed light toward establishing a complete theoretical understanding of alignment.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

PACT: From Credit Assignment to Critic Alignment

Policy Aligned Critic Training (PACT), which adopts an Actor-then-Critic update order to apply importance sampling correction to critic training and better align the critic with the updated policy, is developed.

Jia-Yan Fu, Hang Xu, Yong Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Inference-Time Nash Alignment

This work forms the problem as obtaining a Nash equilibrium of a two-player zero-sum game between policies, and proposes two algorithms: Best-of-Nash (BoN) and Nash Mirror Descent (NMD), which are proved to achieve a duality gap that matches the problem lower bound.

Hadi Hosseini, Debmalya Mandal, Duo-Han Zhang · 0 citations
#machine learning Preprint Oct 2026

Metropolis-Hastings Dominates Importance Resampling for Policy Composition

Post-training a large language model (LLM) often requires exploring trade-offs between multiple rewards, but retraining for each trade-off is expensive. Decoding-time policy composition allows these trade-offs to be adjusted by combining reward-specific policies at inference time. This composition targets a weighted pr...

A. Kurennoy, R. Yarullin, Fergal Reid · 0 citations
#reinforcement learning Preprint Sep 2026

Test-Time Weak-to-Strong Alignment: Transferring Implicit Rewards from Weak to Strong Flow Models

The method, AlignGraft, aligns a larger, frozen, never-tuned model by adding the pair's velocity difference during sampling by adding the pair's velocity difference during sampling, and preserves the large model's fidelity at a small constant sampling overhead.

Xin Xie, Fan Zhang, Dong Gong · 0 citations
#artificial intelligence Preprint Sep 2026

MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning

This work proposes a general RL-based framework for Distribution Matching allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution and proposes reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justific...

Sourabh Kulkarni, Ksheeraj Sai Vepuri, B. Demir et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Diffusion Reward Models

Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many way...

Xiang-Yang Wang, Bing-Xiang He, Ze-Yuan Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.