This paper learns the guidance schedule as a function of diffusion time, conditioning and the current noisy sample, in order to better align sampled images with the text prompt.
Abstract
Modern text-to-image diffusion models rely on classifier-free guidance (CFG) to achieve high image fidelity and text alignment. However, CFG typically applies a static, global scale across all timesteps, samples, and conditions -- a choice that is generally suboptimal and can introduce artifacts, as different states may benefit from different levels of guidance. While time-varying schedules are known to improve quality, designing them by hand is non-trivial and application-dependent. In this paper, we learn the guidance schedule as a function of diffusion time, conditioning and the current noisy sample, in order to better align sampled images with the text prompt. We frame this as a density ratio estimation problem: a discriminator is trained to estimate the time-dependent log-density ratio between the true and guided marginal distributions, while a lightweight generator network predicts the optimal, state-dependent guidance scale. Empirically, our approach outperforms both heuristic CFG schedules and prior methods for learning dynamic guidance on text-to-image generation benchmarks.
This study studies a family of training-free techniques conceptually rooted in Classifier-Free Guidance, most of which were originally proposed on older U-Net diffusion models and validated using metrics that assess image quality in isolation, without accounting for compositional alignment or semantic correspondence.
A. Sergievskii, Artyom Turevich, Sergey Kastryulin· 0 citations
We show that a frozen generic text-to-image diffusion model can perform conditional inpainting across three evaluated natural-image domains with one fixed controller configuration, without inpainting-specific weight training, dataset-specific weight adaptation, or learned inpainting-specific conditioning channels. Step...
PAPT++ is introduced, a risk-aware adversarial generation-training framework for SDG that progressively exposes the classifier to challenging yet semantically consistent variations.
Zhi-Peng Xu, De Cheng, Xinyang Jiang et al.· 0 citations
Pixel diffusion models generate RGB images directly, avoiding the bottleneck of an autoencoder, yet their outputs still systematically underrepresent fine-scale natural-image statistics. We show that adversarial learning provides an effective post-training correction for this deficiency. Starting from a pretrained mode...
Xin Lin, Zhi-Fei Zhang, Yu-Qian Zhou et al.· 0 citations
A Domain-Adversarial Neural Network (DANN) aligns global features but does not control how an added attention block changes the shared representation. We test attention-guided DANN (AG-DANN), which places a Convolutional Block Attention Module (CBAM) before global pooling and scales its residual with a zero-initialized...
Yan-Zhou Qian, Lin-Bo Chen, Shuai Wu et al.· 2026 International Conferenc...· 0 citations
This paper forms few-step generation as a controlled base generative process, and shows that self-consistency loss can be understood through the lens of optimal control, and draws a connection between this approach and reinforcement learning, potentially opening the door to a new set of approaches for few-step generati...
Paribesh Regmi, S. Ghimire, Rui Li· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.