Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, while isolated crops can lose the context needed to interpre...
Yao-Xin Niu, Zhangquan Chen, Yang Zhang et al.· 0 citations
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer caus...
Haojie Huang, Xin-Lei Yu, Cheng-Ming Xu et al.· 2 citations
While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. These limitations often stem from"Rollout Silencing"and low-quality gradient signals in standard sampling procedures. In this work, we propose...
Yi-Meng Ye, Shuang Chen, Wen-Xuan Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.