Skip to content
Open access

Mitigating Hallucination in Long Referring Expressions via Training-Free, Anchor-Preserved Visual Grounding

Unknown authors
Aug 2026 · Information · 0 citations · 20 references

Abstract

Long referring expressions create two coupled sources of hallucination in visual grounding. A detector can select an object that matches only part of the instruction, while a structured vision–language model (VLM) branch can hallucinate a target head or an attribute–object binding. We propose DeRecG, a training-free, anchor-preserved framework that addresses both sources. A frozen Grounding DINO detector supplies an anchor, and a frozen Qwen2-VL model decomposes the expression for candidate verification. Confidence-Constrained Disagreement Arbitration (CCDA) replaces the anchor only when a candidate passes head-validity, score-floor, score-margin, and spatial-disagreement gates. Lexical–Visual Candidate Expansion (LVCE) repairs parser-induced head errors before arbitration. Across six RefCOCO-family splits (46,842 expressions), LVCE+CCDA improves weighted accuracy at an intersection-over-union (IoU) threshold of 0.5 from 53.40% to 55.00%, while weighted mean IoU rises from 53.28% to 54.66%. With the same frozen thresholds, Ref-L4 validation accuracy rises from 31.02% to 34.73% and mean IoU from 33.43% to 36.72% over 13,420 long expressions; 646 anchor errors are corrected and 148 are induced (p<10−74), with no degradation on short-phrase Flickr30k Entities grounding. The complete pipeline requires no training, uses 6.10 GiB peak allocated GPU memory, and processes 1.15 images/s on an RTX 5090 D.

Read PDF