Preprint
Aug 2026
Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors
This paper proposes Dual-Stream Cross-Anchor Correction (DSCC), which injects object-level visual anchors into the language model itself during fine-tuning and reaches the long-caption, low-hallucination region.
LingKai Bu, Qian Gao, Jun Fan et al.
· 0 citations