Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separ...
Zi-Heng Wang, Ming-Xuan Xie, Yi-Lin Liu et al.· 0 citations
MULVEC is proposed, a role-aware method whose compiler produces a structured query record that is mapped to four retrieval roles: Global describes the full target, Desired states what should appear, Preserve states what should remain, and Forbidden states what should disappear.
Zihao Zhang, Da-Yan Wu, Xin-Ze Liu et al.· 0 citations
Focused On-demand Visual Evidence Adaptation is proposed, a cache-friendly approach that builds a reusable visual memory and dynamically retrieves a bounded subset for a draft state and demonstrates that state-conditioned evidence retrieval is an effective alternative to reusing a fixed visual representation throughout...
He Zhu, Da-Yan Wu, Zihao Zhang et al.· 0 citations
M UL V EC is proposed, a role-aware method whose compiler produces a structured query record that is mapped to four retrieval roles: Global describes the full target, Desired states what should appear, Preserve states what should remain, and Forbidden states what should disappear.
Zihao Zhang, Da-Yan Wu, Xin-Ze Liu et al.· 0 citations
Verifier-Guided Decoding (VGD), a decoding framework in which a lightweight verifier examines each emerging object mention, rolls back the KV cache when the mention is identified as high risk, suppresses the object and its synonyms, and regenerates the affected continuation, achieves state-of-the-art object hallucinati...
Lei Yang, Xinze Liu, Dayan Wu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.