Skip to content
Preprint

LoopVAE: Recurrent Depth Across Scales for Visual Tokenization

Sep 2026 · 0 citations · 31 references
Computer Science

TL;DR

With convolutional and Transformer operators and single- or multi-resolution latent interfaces, LoopVAE establishes recurrent depth across scales as a parameter-sharing design axis for visual tokenization.

Abstract

Hierarchical visual tokenizers typically allocate different processing blocks to different spatial scales. We ask how much of this computation can use the same parameters. LoopVAE reuses a scale- and loop-conditioned core within and across scales, while keeping resolution-changing transitions independent. A four-block core executes 28 block applications per encoder or decoder. On ImageNet-256, the 29M-parameter convolutional model reaches 0.28 rFID and 32.54 dB PSNR under an approximately 30-epoch two-stage training budget, using approximately 65% fewer parameters than the 84M reference VAEs. A non-adversarial Transformer ablation with the same execution graph finds competitive PSNR and SSIM under global sharing, although unshared blocks improve LPIPS. Targeted loop interventions show that completing the trained recurrence improves reconstruction and that even small feature updates can have substantial downstream effects. Truncation also exposes output-range errors, distinguishing useful recurrent computation from reliable early exit. Runtime profiling reveals the execution tradeoff: fewer stored weights require more arithmetic and longer runtime in the tested configurations. With convolutional and Transformer operators and single- or multi-resolution latent interfaces, LoopVAE establishes recurrent depth across scales as a parameter-sharing design axis for visual tokenization.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference

AdaVSkip is proposed, which equips each layer with two lightweight routers that independently determine whether visual tokens pass through by or skip the self-attention and MLP modules, and maintains strong task performance with substantially less computation.

Yu-Yao Sun, Tao Deng, Shuang-Hua Li et al. · 0 citations
#machine learning Preprint Sep 2026

Looped Diffusion Transformer

Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the par...

Yong Xien Chng, Tian-Yi Chen, Wen-Wen Tong et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Training-Free Hidden-State Refinement for Flow-Matching Image Generators

A training-free looping framework that repeatedly applies selected transformer layers inside each denoising call is introduced, which improves primary and auxiliary quality metrics with competitive quality--efficiency trade-offs across two Scale-RAE model scales.

Yuan-Yi Yan, Xin-Zhe Rao, Can-Yu Shen et al. · 0 citations
Preprint Aug 2026

Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding

Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature compression critical under bandwidth and computation constraints. Existing ViT feature codecs typically flatten heterogeneous global and pa...

Donghui Feng, Feng-Xi Zhang, Changsheng Gao et al. · 0 citations
#machine learning Preprint Sep 2026

Beyond Selection: Token Parameterization for Extreme Visual Token Compression

Braco is a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling that forms the favorable empirical accuracy-efficiency frontier under compression.

Rui-Lian Zhong, Yu Li, Zhe-Yu Yan et al. · 0 citations
#machine learning Preprint Sep 2026

QuantForge: Discovering Residual Decompositions for MXFP4 Post-Training Quantization

Four-bit post-training quantization can reduce the memory demands of large language models, but preserving accuracy under strict MXFP4 W4A4 requires coordinating several design choices. Coordinate transforms change block-encoding errors, which in turn affect the residuals propagated through the network. The useful algo...

Qiu-Lin Shang, Zhou-Tong Wu, Jie Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.