Skip to content

Author

Yidong Wang

We have 2 of 2 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Pay More Attention To Text In High-Resolution MLLMs

Failures of high-resolution MLLMs are commonly attributed to a visual problem, motivating zooming, cropping, and related visual interventions to recover fine-grained evidence or suppress interference. Yet recent studies suggest that relevant visual evidence is already encoded in intermediate representations, indicating...

Z. Mao, Wen-Zhuo Zhao, Xian-Jie Liu et al. · 0 citations
Jul 2026

Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA

Thinking-Once is proposed, a training-free, single-visual-pass evidence-routing method that reconstructs question-conditioned attention at this window, preserves core entity tokens and compact background context, and routes this evidence to later layers without extra visual encoding.

Z. Mao, Xian-Jie Liu, Tian-Yu Meng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.