Skip to content

Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling

Sep 2026 · 0 citations · 90 references
Computer Science

TL;DR

The Logit Refiner is introduced, a lightweight autoregressive module that restores intra-scale dependencies by sequentially sampling tokens conditioned on frozen backbone features and generalizes to text-to-image generation, confirming that the mean-field bottleneck persists across VAR variants and is effectively alleviated by the method.

Abstract

Visual Autoregressive Models (VAR) generate images through next-scale prediction, producing all tokens within each scale in parallel. We show that this parallel decoding constitutes a mean-field-style approximation that discards spatial dependencies among same-scale tokens, causing locally incoherent samples regardless of backbone capacity -- a limitation of the decoding rule. Addressing this limitation, we introduce the Logit Refiner, a lightweight autoregressive module that restores intra-scale dependencies by sequentially sampling tokens conditioned on frozen backbone features. Adding only ~10% parameters and less than 5% of the base model's training compute, it plugs into any pretrained VAR checkpoint without retraining. Controlled ablations isolate joint intra-scale sampling -- rather than additional capacity or training -- as the critical ingredient. Across backbones from 310M to 2B parameters on class-conditional ImageNet 256x256, the refiner consistently improves generation quality, enabling a 1.1B-parameter model to surpass one twice its size. The approach further generalizes to text-to-image generation, confirming that the mean-field bottleneck persists across VAR variants and is effectively alleviated by our method. Project page: https://compvis.github.io/logit-refiner/

View source

Similar papers

Preprint Aug 2026

VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation

VISTA is the first gradient-based test-time alignment framework for next-scale autoregressive image generation, and introduces the mechanisms needed to make such optimization stable across scales, together with an extensible objective space that any differentiable constraint on cross-attention can plug into.

Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah · 0 citations
Preprint Aug 2026

Falcon Perception-HD: High Density Perception via Reinforcement Learning

This paper explores post-training reinforcement learning (RL), specifically GRPO, to directly align autoregressive perception models with their evaluation metrics, and designs an RL framework that addresses perception-specific challenges: reward design for set-structured outputs and multi-head sampling control.

Sofian Chaybouti, Yasser Dahou, N. Huynh et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Scores to Samples: Elastic Forcing for Autoregressive Video Generation

Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework mi...

Chi Zhang, Yue-Yi Liu, Hao-Yan Shi et al. · 1 citation
Preprint Sep 2026

UnfoldCRF: Structured Mask Refinement with Image-Conditioned Latent Regions

UnfoldCRF treats refinement as inference in a conditional random field over pixel labels and latent region variables, and zeroing individual messages shows where the corrections come from: the pairwise messages mostly fix boundaries, the region messages mostly fix non-boundary errors.

Chun-Ming He, Rihan Zhang, Lei Xu et al. · 0 citations
Preprint Aug 2026

Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective

This work takes a source-centric perspective on VAR editing, in which the encoded source image tokens serve as the primary visual state and the editing process focuses on condition-induced changes, and proposes EditMod, which compares source- and target-conditioned predictions under a shared autoregressive context.

Hongyi Fang, Chu-Wen Xie, Ben-Jia Zhou et al. · 0 citations
#machine learning Preprint Sep 2026

Unifying Distributional Training for One-Step Visual Generation

Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce a unified theoretical framework that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature...

Chi Zhang, Hao-Yan Shi, Yue-Yi Liu et al. · 2 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.