Skip to content
Preprint

Information-Guided Frontier Decoding: Contextual Utility-Driven Commitment in dMLLMs

Aug 2026 · 0 citations · 48 references
Computer Science

TL;DR

IGFD is proposed, a training-free decoding strategy that ranks candidates using token confidence, neighborhood uncertainty, and structural commitment risk and consistently outperforms existing decoding strategies across the majority of benchmarks and diffusion MLLM backbones under identical decoding budgets.

Abstract

Decoding quality in diffusion multimodal language models (dMLLMs) depends heavily on the order in which masked tokens are committed. Existing confidence-based strategies prioritize locally easy tokens, but confidence does not necessarily reflect contextual usefulness. As a result, structurally easy tokens such as punctuation may be committed before informative semantic anchors, weakening context propagation and increasing error accumulation. We propose Information-Guided Frontier Decoding (IGFD), a training-free decoding strategy that ranks candidates using token confidence, neighborhood uncertainty, and structural commitment risk. IGFD encourages early commitment of reliable semantic anchors while delaying fragile structural tokens, improving contextual support during decoding. A dynamic candidate frontier further constrains token selection to locally expandable regions under the same decoding budget. The method requires no additional training, auxiliary models, or extra forward passes. Experiments across multimodal understanding, reasoning, grounding, and hallucination benchmarks show that IGFD consistently outperforms existing decoding strategies across the majority of benchmarks and diffusion MLLM backbones under identical decoding budgets.

View source

Similar papers

#natural language process... Preprint Sep 2026

SFAD: Speculative Factuality-Aware Decoding

SFAD is presented, a speculative decoding framework that enhances contextual faithfulness without inference degradation and substantially improves faithfulness while achieving $2.48\times$ speedup, offering a practical solution for efficient LLMs.

Guan-Qiao Chen, Di Wang, Lijie Hu · 0 citations
Preprint Aug 2026

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

The Visual Information-Guided Sampler (VIG-Sampler), which prioritizes tokens based on their attention to image tokens, and imposes a constraint that penalizes candidate tokens whose image-attention distributions are similar to those of previously selected tokens, thereby increasing the information gain of the decoded...

Insu Lee, Woo-Soon Park, Wonseok Shin et al. · 1 citation
Preprint Aug 2026

Context-Aware Cluster Decoding: Semantic Anchor-Driven Coherence in dMLLMs

Diffusion multimodal large language models (dMLLMs) frequently produce long-form outputs marred by semantic drift and repetition, with quality generally degrading as output length increases. We identify two structural deficiencies in existing decoding methods as primary drivers of these failures: confidence-based scori...

Yikai Zhao, Qiyan Zhao, Jia-Quan Zhang et al. · 0 citations
Preprint Aug 2026

Dependency-Aware Revocable Decoding for Efficient Diffusion Large Language Model Inference

DARD is proposed, a training-free framework that separates tokens into masked, candidate, and unmasked states and adaptively regulates their influence on subsequent decoding, and consistently improves the speed-quality Pareto frontier over recent revocable decoding methods.

Woo-Soon Park, Insu Lee, Minyoung Noh et al. · 1 citation
#natural language process... Preprint Aug 2026

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

RPS is proposed, a novel training-free decoding method that seeks mid-entropy positions as promising candidate pivots (where to decode), and determines their token assignment that yields the greatest downstream benefit via lookahead evaluation (what to decode).

Yu-Shi Ye, Xu Chen, Hao-Yun Jiang et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models

This paper measures dLLM safety behavior under harmful prompts by tracing intermediate token distributions and commitment decisions throughout denoising, and proposes Refusal-Aware Early Commitment (RAEC), a simple training-free decoding method that commits persistent refusal signals from early steps.

Guoli Wang, Haonan Shi, Tu Ouyang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.