Skip to content
Preprint

Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors

Aug 2026 · 0 citations · 46 references
Computer Science

TL;DR

This paper proposes Dual-Stream Cross-Anchor Correction (DSCC), which injects object-level visual anchors into the language model itself during fine-tuning and reaches the long-caption, low-hallucination region.

Abstract

Object hallucination in multimodal large language models arises when language priors and corpus co-occurrence bias outweigh the visual evidence, with nothing tying an object mention to the image. Most remedies intervene at decoding time, yet under a unified protocol their benefit is confined to short captions; supervised fine-tuning (SFT) on a detail-rich corpus lengthens captions, but over forty percent still name absent objects. This paper proposes Dual-Stream Cross-Anchor Correction (DSCC). Unlike work that post-processes decoding, DSCC injects object-level visual anchors into the language model itself during fine-tuning: a perception stream aligns object-level hidden states at an intermediate layer to frozen text anchors by a bidirectional contrastive objective; a cognition stream lets deeper layers query those anchors by cross-attention at every generation step; and a two-stage curriculum gate couples them, making evidence retrieval a structural constraint on generation. Under one backbone and one scoring protocol, experiments span long-caption hallucination, object-existence discrimination and cross-domain generalisation, with vanilla SFT on the same corpus and schedule as a length- and density-matched control separating the data effect from the architectural gain. DSCC alone reaches the long-caption, low-hallucination region: captions roughly 1.9 times the baseline length at 88.19% precision per object mention, the highest under a density-independent criterion. Ablations expose a synergy: the perception stream alone degrades precision yet reverses sign when stacked on the cognition stream. No universal superiority is claimed: three out-of-domain benchmarks yield a predictable, falsifiable domain-conditionality, the synergy being bound to the anchors'semantic domain and breaking on charts and illusions.

View source

Similar papers

Open access Aug 2026

Mitigating Hallucination in Long Referring Expressions via Training-Free, Anchor-Preserved Visual Grounding

Long referring expressions create two coupled sources of hallucination in visual grounding. A detector can select an object that matches only part of the instruction, while a structured vision–language model (VLM) branch can hallucinate a target head or an attribute–object binding. We propose DeRecG, a training-free, a...

Hao-Xuan Song, Li-Huan Shao · 0 citations
Conference Aug 2026

Beyond CLIP: A Critical Analysis of Representational Misalignment When LLMs Replace Text Encoders in Diffusion-based Image Generation

Recent text-to-image diffusion systems have begun replacing CLIP and T5 text encoders with decoder-only large language models (LLMs), motivated by their stronger language understanding. This substitution, however, does not straightforwardly improve image-text alignment: naively using an LLM as the prompt encoder can su...

Yogesh Kakde, Jitendra Jaiswal, B. Sahoo et al. · 0 citations
#machine learning Preprint Sep 2026

The Alignment Illusion in Multimodal Large Language Models

Internal visual-text alignment in MLLMs is best read as a geometric diagnostic of the visual stream inside the language model rather than a direct proxy for content-level cross-modal interaction, and is most informative when calibrated by controlled task evidence.

Hong-Han Wang, Yun-Tao Wang, Hui-Chao Ding · 0 citations
#artificial intelligence Preprint Aug 2026

RelCheck: Dual-Evidence Spatial Grounding for VLM Hallucination Correction

Multimodal large language models (MLLMs) fre- quently generate text that is inconsistent with the input image. While object- and attribute-level hallucinations have received considerable attention, relational hallucinations (incorrect de- scriptions of spatial or interactive relationships between objects) remain largel...

Siddhi Patil, N. Saxena, William B. Andreopoulos · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.