Counterfactual Anatomy-guided Spatial-Temporal Decoding for Annotation-Free Hallucination Mitigation in Medical VLMs
Experiments on the SLAKE and MIMIC-CXR datasets demonstrate that CAST consistently outperforms strong baselines and surpasses decoding strategies reliant on ground truth, and indicates that compact, automatically selected regions provide highly effective contrastive guidance without expert annotations.