Skip to content
Preprint

VGAS: Variance-Reduced Guidance and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion

Aug 2026 · 1 citation · 52 references
Computer Science Biology Mathematics

TL;DR

Variance-reduced Guidance and Adaptive Selection (VGAS), a simple yet effective inference-time framework that reduces the variance of the guidance estimate for both reward types, applies the reward tilting in the clean-token logits, where the pretrained schedule is preserved, and sets the selection temperature per step.

Abstract

Masked discrete diffusion models perform strongly on text, code, and biological sequences, but their training objective rewards only naturalness, and retraining the generator for every new reward is expensive. Inference-time steering of a frozen model either guides the sampler by the reward gradient or searches over several trajectories, and recent samplers combine the two. Such combinations are assembled as pipelines that leave three choices at their defaults: a guidance estimate resting on one Gumbel draw per sample, a reward tilting placed without reference to the distribution the combination then targets, and a selection temperature held fixed although the spread of per-step rewards drifts. We identify that distribution and settle the three choices against it. We therefore propose Variance-reduced Guidance and Adaptive Selection (VGAS), a simple yet effective inference-time framework that reduces the variance of the guidance estimate for both reward types, applies the reward tilting in the clean-token logits, where the pretrained schedule is preserved, and sets the selection temperature per step. Across regulatory DNA, protein and small-molecule benchmarks, VGAS attains the best training-free reward and matches or surpasses a reward-fine-tuned generator.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Diffusion Reward Models

Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many way...

Xiang-Yang Wang, Bing-Xiang He, Ze-Yuan Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Aligning One-Step Generative Models with Reward-Weighted Transport Distillation

Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge.

Austin S. Wang, Zi-Heng Cheng, Le-Xing Ying · 0 citations
Preprint Aug 2026

Latent Reward Registers for Diffusion Preference Alignment

This work proposes Latent Reward Registers, a mechanism that estimates terminal preference directly from intermediate noisy latents, and achieves significant reward improvement with a favorable reward-quality balance against training-free baselines.

Yuan-Shen Guan, Zipeng Feng, Cheng-Ru Song et al. · 0 citations
Preprint Aug 2026

On-Policy Self-Distillation in Diffusion Models

The results support on-policy self-distillation as an efficient and analyzable approach to diffusion post-training by converting image-level reward guidance into explicit and continually refreshed intermediate supervision, thereby opening a path toward more efficient and diagnosable alignment.

Weina Zhou, Xiongwei Zhu, Ling-Dong Kong et al. · 4 citations
#artificial intelligence Preprint Sep 2026

MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning

This work proposes a general RL-based framework for Distribution Matching allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution and proposes reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justific...

Sourabh Kulkarni, Ksheeraj Sai Vepuri, B. Demir et al. · 0 citations
#machine learning Preprint Sep 2026

When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation

The cross-signal NTK is introduced, a token-level statistic that measures the alignment between reward and distillation gradients at position n and an empirical threshold beyond which naive mixing can lead to persistent training collapse is revealed.

Xin-Ke Jiang, Tao Feng, Zhi-Bang Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.